Search papers, labs, and topics across Lattice.
4
0
7
1
DeepShare is a scheduler that uses a continuous tenant-assurance signal to coordinate these decisions at runtime to achieve a more advantageous utilization-QoS trade-off than optimizing quotas, scheduling, and resource sharing independently.
Optimizing GPU resource allocation can cut multi-agent workflow completion times by nearly 37% while saving substantial GPU usage.
EcoVLA boosts energy efficiency for VLA models by up to 236% while ensuring real-time performance, revolutionizing how robotic systems manage inference costs.
Maestro slashes memory usage by over 67% while boosting service level objectives for LLM-based multi-agent systems, revolutionizing resource management in cloud settings.