Search papers, labs, and topics across Lattice.
3
0
3
0
Energy consumption in LLM serving can be cut by nearly half without sacrificing performance, thanks to a new framework that intelligently manages GPU frequency scaling.
LUMEN slashes recovery times in distributed LLM serving by intelligently coordinating load-aware recovery strategies during worker failures.
RDMA failover can be made significantly more efficient and correct by selectively retransmitting only the requests that were actually lost during a link failure, avoiding redundant retransmissions and semantic violations.