Search papers, labs, and topics across Lattice.
5
0
6
2
Class-name descriptors can mislead vision-language models, but selecting attributes directly from images boosts accuracy and interpretability significantly.
Multimodal unlearning could revolutionize how we handle sensitive data in AI, enabling targeted removal without sacrificing model performance.
PRISM achieves up to 15% performance gains in visuomotor tasks by leveraging short-term memory, challenging the notion that immediate sensory input suffices for complex decision-making.
Joint image-depth generation can be achieved with a single model trained on sparse data, outperforming existing methods by a significant margin.
Forget salient cues – now you can *steer* visual representations in ViTs with language, focusing on any object you want without hurting overall performance.