Search papers, labs, and topics across Lattice.
Shanghai Jiao Tong University
6
0
12
7
Visual tool-use in multimodal LLMs may create an illusion of effectiveness, with many models achieving gains that are not causally justified.
Multimodal context learning remains critically underdeveloped, with top models barely scratching the surface of effective performance.
Forget flat, lifeless speech: this model uses self-critique to generate expressive speech rivaling GPT-4o-Audio, even with significantly less training data.
LLMs can now scale depth more effectively: a new attention mechanism recovers diluted features in deeper layers, boosting performance with negligible overhead.
Multi-robot coverage can now handle multiple sensory demands simultaneously, with provable guarantees on performance even when those demands are initially unknown.
Ditch the slow, iterative zooming during MLLM inference: Region-to-Image Distillation lets you bake those agentic zooming benefits directly into a single forward pass.