StoryHardware
Google announces Ironwood, its seventh-generation TPU built for inference
Google DeepMindTensor Processing Unit (TPU)
Google introduced Ironwood, its seventh-generation TPU designed specifically for inference workloads. It scales to pods of 9,216 liquid-cooled chips delivering 42.5 exaflops of compute—over 24 times the world's largest supercomputer. Ironwood achieves 2x improved performance-per-watt over the previous generation and 192GB of HBM per chip.
- Scales to 9,216 liquid-cooled chips delivering 42.5 exaflops; 24x more compute than El Capitan supercomputer
- 192GB HBM per chip (6x prior generation); 7.37 TB/s memory bandwidth per chip (4.5x prior)
- 2x performance-per-watt improvement; 1.2 TBps inter-chip bandwidth; built for inference not training
- Enables what Google calls 'age of inference': AI agents proactively generating insights rather than responding to queries