Search papers, labs, and topics across Lattice.
UC San Diego 2 Zhejiang University
2
0
4
0
JetSpec shatters the speed ceiling of speculative decoding, achieving up to 9.64x acceleration on complex tasks while maintaining high acceptance rates.
Current LLM memory systems falter when faced with the continuous, machine-generated interaction streams typical of real-world agentic applications, highlighting a critical need for causality-aware and tool-augmented memory architectures.