Search papers, labs, and topics across Lattice.
Affiliation:
5
0
7
9
Identical nodes can exhibit up to 5% performance variation, challenging assumptions about uniformity in data center hardware.
Halving multi-hop accuracy loss while maintaining single-hop recall, Kamera redefines how multimodal agents can efficiently reuse cached information without retraining.
Leyline transforms agentic LLM performance by enabling efficient, policy-driven cache edits that drastically reduce latency and improve solve rates.
Routing queries across GPUs can be cheaper than moving the cache, challenging conventional wisdom in LLM architecture design.
Turns out, the latest and greatest GPU isn't always the most energy-efficient: NVIDIA's H100 surprisingly beats the H200 for compute-bound workloads under power constraints.