Search papers, labs, and topics across Lattice.
ByteDance
2
0
2
LLMVisor achieves up to 4.4x improvement in latency attribution accuracy for multi-tenant LLMs, revealing hidden inefficiencies in GPU resource usage.
Lodestar slashes time-to-first-token by up to 4.42x in heterogeneous GPU clusters through intelligent, adaptive request routing.