Search papers, labs, and topics across Lattice.
6
0
6
9
FIBER achieves up to 2.25x speedup in LLM serving by decoupling tensor computation from register ownership, revolutionizing GPU efficiency.
Deltoris achieves a staggering 34.2脳 speedup for real-time VLA inference, revolutionizing how embodied AI can operate on edge devices.
DeGS redefines the performance landscape for 3D Gaussian Splatting, achieving up to 7.25x throughput improvements while maintaining high processing element utilization.
Kaleido achieves a remarkable 5.9x speedup in video diffusion transformers by intelligently reusing computations based on latent space correlations.
Modern LLM performance hinges on dependency structures rather than individual instruction latencies, revealing a critical insight for GPU optimization.
GF-DiT achieves up to 6.01脳 throughput improvement and 95% latency reduction by dynamically adapting GPU parallelism in response to workload demands.