Search papers, labs, and topics across Lattice.
5
0
3
PLoRA slashes decode latency for multi-LoRA serving by over 6x while using pooled memory and near-data processing, revolutionizing how we deploy specialized AI models.
Achieving a 7.6脳 speedup in distributed 3D scene reconstruction without sacrificing quality could redefine efficiency benchmarks in computer graphics.
AoiZora cuts video diffusion denoising latency by 1.42x on TPU v5e sub-slices by smartly aligning logical sharding with physical hardware topology.
FlashCP achieves up to 1.63x faster training for large language models by eliminating redundant communication and optimizing workload balance.
A fault in one GPU process no longer needs to crash them all: this paper introduces mechanisms for fault-resilient NVIDIA MPS, enabling more robust multi-tenant GPU clusters.