Search papers, labs, and topics across Lattice.
Telecom Cloud Computing Research Institute
2
0
4
Energy consumption in LLM serving can be cut by nearly half without sacrificing performance, thanks to a new framework that intelligently manages GPU frequency scaling.
Forget fixed-precision quantization: STQuant slashes optimizer memory by 84% in large model training by dynamically adapting bit-widths across layers and training steps.