Search papers, labs, and topics across Lattice.
UC Santa Cruz
5
0
4
4
fabric_ext revolutionizes dataflow management in GPU-CXL fabrics, optimizing performance for large language model tasks by executing extensible policies across diverse hardware hooks.
Kops enables significant performance improvements in eBPF by allowing new operations to be added without compromising kernel safety or increasing the trusted computing base.
Concordia achieves fault tolerance for LLM inference by seamlessly integrating persistent kernel checkpointing, enabling rapid recovery without CPU bottlenecks.
Kernel launch overhead is a bigger bottleneck than you think: GPUOS achieves up to 15.3x speedup by fusing operations at runtime.
Diagnose performance bottlenecks in large-scale AI training 100x faster with a new observability system that adds almost no overhead.