Search papers, labs, and topics across Lattice.
Affiliation:
2
0
3
2
Operator-level scaling can reduce GPU usage by over a third while maintaining strict service level objectives for LLMs.
A fault in one GPU process no longer needs to crash them all: this paper introduces mechanisms for fault-resilient NVIDIA MPS, enabling more robust multi-tenant GPU clusters.