Search papers, labs, and topics across Lattice.
13
0
14
7
ReCache achieves a staggering 92.43% reduction in KV-tensor memory while maintaining nearly identical performance in tool-augmented language models.
WIDE achieves a remarkable 55.1% performance boost at 50% sparsity, revolutionizing how LLMs can efficiently allocate computation at the token level.
Confidence in reasoning models can be dramatically improved by strategically leveraging position-aware signals, leading to better performance in challenging tasks.
Rigid geometric compression can collapse reasoning space, while generative reconstruction preserves semantic integrity and enhances reasoning accuracy.
Achieving 134 FPS with under 50 ms latency, ViCoStream redefines the capabilities of VideoLLMs for real-time streaming applications.
Achieving robot navigation policy training in under 20 seconds could revolutionize the deployment of DRL in robotics.
AdaSR achieves superior reasoning performance in dynamic environments by enabling models to adaptively allocate computation during streaming input.
Achieving nearly 10 times faster reranking without sacrificing performance, CompRank revolutionizes the efficiency of LLMs in retrieval tasks.
FPGAs can beat GPUs at dynamically allocating computation for LLM inference, thanks to a new architecture that fuses operations, uses mixed precision, and caches KV values on-chip.
Untangling the mess of "streaming LLMs," this paper delivers a clear taxonomy that distinguishes between streaming generation, streaming inputs, and interactive architectures.
LVLMs can reason about video streams *much* faster and better by thinking concurrently with the incoming data, not in batches.
Forget simple image search: MCMR reveals how current multimodal models struggle with the complex, interdependent constraints of real-world product search.
LLMs actually *do* improve time series forecasting, especially for cross-domain generalization, overturning prior doubts with a massive 8-billion observation study.