Search papers, labs, and topics across Lattice.
Shanghai Jiao Tong University
3
0
7
17
Achieving up to 17脳 faster inference for long-context LLMs without compromising output quality could redefine efficiency standards in AI applications.
Forget static pipelines: SLMs can learn to dynamically seek help from LLMs, leading to better performance and transferability.
Decoupling temporal and spatial reasoning in video grounding unlocks significant performance gains, outperforming existing MLLM-based methods by a large margin.