Search papers, labs, and topics across Lattice.
7
0
10
21
OasisKV achieves up to 2.1x throughput gains in LLM inference while using significantly less memory, challenging the limits of current HBM constraints.
Task-vector subtraction in VLA policies can lead to unexpected control failures, with edits harming unrelated tasks and revealing the fragility of behavioral locality.
State history can dramatically enhance VLA model performance, but the interface choice is crucial for optimizing its impact.
Repurposing retired GPUs can cut costs dramatically, but without clean energy, it risks quadrupling carbon emissions for LLM inference.
Dynamic quantization, a widely adopted optimization for efficient ML serving, can leak your data to adversaries sharing the same batch.
Forget static, homogeneous multi-agent systems: Team-of-Thoughts unlocks superior performance by dynamically orchestrating heterogeneous agents based on calibrated coordination and self-assessed domain expertise.
Fusing kernels in SwiGLU MLP blocks slashes memory bandwidth bottlenecks, yielding up to 13.2% speedups on H100 GPUs during agentic LLM inference.