Search papers, labs, and topics across Lattice.
University of Michigan
2
0
5
LLMVisor achieves up to 4.4x improvement in latency attribution accuracy for multi-tenant LLMs, revealing hidden inefficiencies in GPU resource usage.
Multi-turn RL agents can learn far more effectively by explicitly monitoring and controlling uncertainty at both the token and turn levels, leading to more stable training and higher performance.