StoryHardware

Google announces Ironwood, its seventh-generation TPU built for inference

Google DeepMindTensor Processing Unit (TPU)

Google introduced Ironwood, its seventh-generation TPU designed specifically for inference workloads. It scales to pods of 9,216 liquid-cooled chips delivering 42.5 exaflops of compute—over 24 times the world's largest supercomputer. Ironwood achieves 2x improved performance-per-watt over the previous generation and 192GB of HBM per chip.

  • Scales to 9,216 liquid-cooled chips delivering 42.5 exaflops; 24x more compute than El Capitan supercomputer
  • 192GB HBM per chip (6x prior generation); 7.37 TB/s memory bandwidth per chip (4.5x prior)
  • 2x performance-per-watt improvement; 1.2 TBps inter-chip bandwidth; built for inference not training
  • Enables what Google calls 'age of inference': AI agents proactively generating insights rather than responding to queries
Read the original · Blog

More on Tensor Processing Unit (TPU)

Primer