Search papers, labs, and topics across Lattice.
Affiliation:
1
0
2
5
Operator-level scaling can reduce GPU usage by over a third while maintaining strict service level objectives for LLMs.