Search papers, labs, and topics across Lattice.
This paper introduces CommBench, a benchmark designed to evaluate the ability of large language models (LLMs) to generate efficient GPU communication code, addressing the complexity of GPU architectures and networking. The study reveals that even the most advanced model, GPT-5.5, only achieves correct and competitive performance on 30.7% of the tasks, highlighting a significant gap between LLM-generated code and expert-written implementations. By providing a robust evaluation framework and a unified metric for correctness and performance, CommBench sets a new standard for assessing AI's capabilities in systems programming.
LLMs struggle to match expert-level performance in GPU communication tasks, with top models achieving only 30.7% success in generating efficient code.
Training and serving large language models (LLMs) rely heavily on high-performance GPU communication, yet implementing efficient GPU communication primitives requires deep expertise in GPU architectures, networking hardware, and distributed communication patterns, making them particularly challenging for code generation models. We present CommBench, a comprehensive benchmark for GPU communication programming, consisting of over 100 expert-curated tasks spanning point-to-point communication, collective operations, expert-parallel communication, compute--communication fusion, and communication utility functions, with reference implementations either written by GPU communication experts or distilled from production codebases. We further introduce a cheat-resistant evaluation framework that automatically compiles, executes, and validates generated code on multi-GPU systems, and a unified metric that jointly measures functional correctness and communication performance. Evaluating leading frontier and open-source code generation models on both intra-node NVLink and inter-node RDMA platforms reveals that even the strongest model, GPT-5.5, correctly implements and achieves competitive performance on only 30.7\% of the benchmark tasks. Our results expose a substantial gap between current LLMs and expert-written GPU communication code, establishing CommBench as a challenging benchmark for advancing AI-assisted systems programming.