Search papers, labs, and topics across Lattice.
Affiliation:
3
0
5
Long-horizon reasoning traces repeatedly backtrack to early planning steps via predictable query clusters, revealing that a handful of representative "beacon" vectors can guide 5.8脳 KV cache compression with near-zero quality loss.
Diminishing returns in inference-time scaling reveal that more computation doesn't always equate to better performance in local computer-use agents.
Tangram transforms multi-turn LLM serving by achieving up to 2.6 times throughput improvement without sacrificing accuracy, all while sidestepping the pitfalls of memory fragmentation.