Search papers, labs, and topics across Lattice.
5
0
6
27
DeepShare is a scheduler that uses a continuous tenant-assurance signal to coordinate these decisions at runtime to achieve a more advantageous utilization-QoS trade-off than optimizing quotas, scheduling, and resource sharing independently.
Optimizing GPU resource allocation can cut multi-agent workflow completion times by nearly 37% while saving substantial GPU usage.
SpecBox slashes end-to-end latency by up to 2.9x while cutting memory usage by nearly 46%, revolutionizing how LLM agents manage sandbox resources.
CrossPool achieves a staggering 10.4x reduction in tail latency for bursty long-context requests by decoupling model weights from KV-cache in GPU memory.
Maestro slashes memory usage by over 67% while boosting service level objectives for LLM-based multi-agent systems, revolutionizing resource management in cloud settings.