Search papers, labs, and topics across Lattice.
3
0
8
Unused AI computation at the edge can be harnessed to boost performance in traditional tasks without sacrificing the efficiency of primary workloads.
Achieving a 90% boost in prefill throughput for MoE models could redefine the efficiency of large-scale language model serving.
Flash-WAM achieves real-time inference for world-action models by reducing latency from 8.1 seconds to 348 milliseconds without sacrificing performance.