Search papers, labs, and topics across Lattice.
3
0
7
7
FlashRT achieves up to 70x latency reduction in multimodal applications by intelligently guiding coding agents through a multi-pass optimization process.
Vortex achieves up to 4.7 times higher throughput for large language models, revolutionizing how researchers can prototype and evaluate sparse attention algorithms.
Real-time video generation just got a whole lot faster: Monarch-RT achieves up to 95% attention sparsity without quality loss and outperforms FlashAttention, finally enabling 16 FPS video generation on a single GPU.