Search papers, labs, and topics across Lattice.
3
0
3
4
FIBER achieves up to 2.25x speedup in LLM serving by decoupling tensor computation from register ownership, revolutionizing GPU efficiency.
Modern LLM performance hinges on dependency structures rather than individual instruction latencies, revealing a critical insight for GPU optimization.
GF-DiT achieves up to 6.01脳 throughput improvement and 95% latency reduction by dynamically adapting GPU parallelism in response to workload demands.