Search papers, labs, and topics across Lattice.
4
0
3
24
Achieving a staggering reduction in hardware footprint while maintaining near-perfect accuracy, Uni-SFU redefines the efficiency of activation function implementations in neural networks.
APEX achieves over 99% overlap accuracy in expert prefetching, slashing per-token latency by up to 26% while enhancing energy efficiency for edge MoE inference.
Hybrid Mamba-Transformer models can get 4x faster time to first token and 1.4x higher throughput by disaggregating prefill and decode phases onto specialized accelerator packages.
LLMs can run up to 35% faster on chiplet architectures thanks to a new lossless exponent compression technique that slashes inter-chiplet communication overhead.