Search papers, labs, and topics across Lattice.
KernelArc introduces a multi-agent framework designed for autonomous optimization of GPU kernels across diverse workloads, leveraging strategy-specialized agents that operate in parallel. The framework employs a unique coordination mechanism through shared memory and deterministic benchmarks, resulting in significant performance improvements on NVIDIA H100 and B200 GPUs. Evaluations demonstrate that KernelArc achieved top rankings on the SOL-ExecBench leaderboard for various tasks, illustrating the effectiveness of shared multi-agent search in enhancing kernel optimization within a constrained candidate budget.
Shared multi-agent search in KernelArc outperforms traditional methods, achieving top rankings in GPU kernel optimization tasks while maintaining a fixed candidate budget.
We present KernelArc, a multi-agent framework for autonomous GPU kernel optimization across heterogeneous workloads. Strategy-specialized agents run in parallel and coordinate through conclusions-only shared memory, a deterministic benchmark guard, and read-only cross-agent state with plateau-triggered drafting. We evaluate \kernelarc{} on NVIDIA H100 and B200 GPUs using category-representative SOL-ExecBench workloads. The resulting implementations span custom BF16 GEMM, static cuBLASLt Expert-API configuration tables, fused mixture-of-experts backward, shape-gated decoder-layer fusion, native NVFP4 grouped-query attention, and paged prefill attention. At the public SOL-ExecBench leaderboard snapshot recorded on July~30, 2026, these submissions ranked first on representative L1, L2, Quantization, and FlashInfer tasks. The trajectories support the paper's central motivation: shared multi-agent search can broaden exploration and reach stronger incumbents within a fixed candidate budget, while the value of individual coordination features depends on the kernel and optimization stage.