Search papers, labs, and topics across Lattice.
2
0
3
8
Energy consumption in LLM serving can be cut by nearly half without sacrificing performance, thanks to a new framework that intelligently manages GPU frequency scaling.
Extending speculation horizons can backfire, but SparseSpec-L smartly navigates this trade-off to achieve substantial speedups in long-context inference without sacrificing output quality.