Search papers, labs, and topics across Lattice.
This paper introduces Redwood, an AI accelerator designed and deployed entirely by an AI system in just two weeks, addressing the mismatch between rapidly evolving AI workloads and traditional hardware design cycles. By integrating hardware and software into a single optimization loop, Redwood achieves significant performance improvements, delivering 1.75x throughput and 3.4x performance-per-watt compared to existing solutions. Notably, this marks the first instance of a production-ready AI accelerator autonomously designed by AI, showcasing the potential for recursive self-improvement in hardware development.
Redwood achieves a groundbreaking 3.4x performance-per-watt gain, demonstrating that AI can autonomously design and deploy cutting-edge hardware in record time.
Modern AI workloads and the hardware that runs them evolve on different timescales: architectural definition precedes volume silicon by years, while target workloads shift in months. Design decisions are therefore committed under deep uncertainty and paid for twice, once in the generality added as a hedge, and again when new workloads map poorly onto frozen silicon. As Moore's Law stagnates, specialization is the main remaining source of performance-per-watt and demands a design cycle that runs at the cadence of the workloads. We present an end-to-end AI system that collapses the software-to-silicon stack into a single optimization loop, where hardware and software are co-designed and verified under one objective. Its first demonstration is Redwood, a frontier AI accelerator built for single-batch, low-power, ultra-low-latency inference for physical AI. From a high-level specification by two human architects, the system autonomously generated the performance model, RTL design, UVM environments, formal proofs, firmware, and kernels in under two weeks with no human intervention below the specification. Every block reached 95% coverage via commercial EDA tools, our proprietary formal engine, and hardware-in-the-loop validation. Specification changes were reverified and redeployed to hardware in under 48 hours. Redwood Nano, its ultra-low-power FPGA variant, runs multi-billion-parameter models like Llama and Qwen. Projected onto Samsung 8 nm, the Jetson Orin Nano's process class, Redwood delivers 1.75x the throughput at 1.9x lower power, a 3.4x performance-per-watt gain against a measured Jetson baseline on the same models. Qwen running on Redwood also helped design next-generation Redwood, an early step toward recursive self-improvement. To our knowledge, this is the first production-worthy AI accelerator designed end-to-end by an AI system and running a modern AI model.