Search papers, labs, and topics across Lattice.
EPFL EcoCloud, Google DeepMind, EcoCloud, EPFL Tsinghua University, EPFL MangoBoost Inc, EPFL Korea University
Google DeepMind2
0
4
MXSens redefines LLM quantization by achieving state-of-the-art accuracy with mixed precision, leveraging sensitivity-aware bitwidth allocation.
Splitting attention and feedforward networks onto separate GPUs can unlock 4x higher MoE LLM throughput, but only if you carefully tune the GPU partitioning strategy based on the workload.