Search papers, labs, and topics across Lattice.
School of Electrical and Computer Engineering, The University of Sydney
4
0
6
2
HorizonServe achieves up to 4.9x better SLO attainment and drastically lowers latency for omni-model serving by intelligently coordinating GPU resource allocation.
MIM's superior robustness against non-IID data could redefine the benchmarks for distributed self-supervised learning frameworks.
AgentServe achieves up to 2.8x improvement in time-to-first-token and 2.7x in tokens-per-output-token for agentic workloads on a single GPU by strategically isolating prefills and decodes.
Stop leaving performance on the table: jointly optimizing resource allocation and request batching with reinforcement learning can yield up to 24x speedups for multi-tenant GPU inference.