Search papers, labs, and topics across Lattice.
9
0
9
3
Adapting supervision weights based on the evolution of divergence histories boosts reasoning performance in language models without extra computational overhead.
The shift from parameter-centric to system-level adaptation in continual learning could redefine how we build and interact with AI models.
Retaining nearly all of a model's capability while slashing visual token usage by over 80% reveals a transformative approach to VLM compression.
Merging RL experts effectively requires balancing sharp, informative signals with stable, dispersed components, a challenge that ResMerge addresses with innovative spectral techniques.
Achieve near-lossless performance in autonomous driving VLMs with 90% token reduction – without any training.
LVLMs can run 2.3x faster with only a 2% accuracy drop, thanks to a new pruning method that understands which visual tokens are most relevant to the text.
Decoupling masked reconstruction and contrastive alignment in audio-visual representation learning yields surprisingly large gains in zero-shot retrieval, outperforming SOTA by a significant margin.
Ditch slow, verbose chain-of-thought reasoning: PLUME's latent reasoning slashes inference time by 30x while boosting multimodal embedding performance.
Training-free zero-shot image retrieval just got a whole lot better: WISER's "retrieve-verify-refine" pipeline achieves state-of-the-art results by intelligently fusing text-to-image and image-to-image retrieval.