Search papers, labs, and topics across Lattice.
5
0
7
2
Latent objects can revolutionize how Video-LLMs manage memory, achieving a 10-point performance boost while slashing memory usage by 50%.
Latent visual reflection allows models to achieve superior fine-grained perception in long visual contexts while slashing inference time by nearly half.
Optimizing for runtime in multimodal training can be energy-inefficient, as data movement and overlap on Grace Hopper chips dominate energy consumption, not raw compute.
Forget brute-force context windows: a small vision-language model can compress hour-long videos below theoretical limits by intelligently prioritizing relevant content.
Scale up offline policy training for diffusion LLMs without breaking the bank: dTRPO slashes trajectory computation costs while boosting performance up to 9.6% on STEM tasks.