Search papers, labs, and topics across Lattice.
This paper reviews the evolution of Google's Tensor Processing Units (TPUs) from version v2 to Ironwood, emphasizing their architectural stability and adaptability to diverse deep learning workloads. Key advancements include a tenfold increase in high-bandwidth memory capacity, a hundredfold increase in peak node performance, and a staggering 3600-fold improvement in overall supercomputer performance over eight years. The study also highlights innovations in resilience and sustainability, such as optical circuit switches and reduced carbon emissions per floating point operation, which are critical for the future of AI training infrastructure.
A staggering 3600-fold increase in supercomputer performance over eight years underscores the transformative evolution of Google's TPUs in AI training.
This paper (to appear in the July/August 2026 issue of IEEE Micro magazine) summarizes five generations of Google s TPUs, from TPU v2 to Ironwood, highlighting their evolution as scalable, resilient, power-efficient, sustainable supercomputers for AI training. It details the TPU s stable architecture, which has surprisingly easily accommodated the rapidly changing deep neural network workloads, such as the rise of Transformers. Key advancements over eight years include 10x increase in HBM capacity and bandwidth per node, a 100x increase in peak node performance, and a 3600x increase in supercomputer performance. The paper also discusses the role of optical circuit switches, built-in self test, and hardware replay in enhancing resilience and how TPU's environmental impact is reduced with substantial improvements in performance per Watt and in carbon emissions per floating point operation. It concludes by identifying six features that may well characterize successful training accelerators of this decade.