StoryHardware

Google announces Trillium, its sixth-generation TPU, with 4.7x the peak compute per chip of TPU v5e

Google DeepMindTensor Processing Unit (TPU)

Google announced Trillium, its sixth-generation TPU, delivering 4.7x peak compute, 2x memory capacity and bandwidth, and 2x interchip interconnect bandwidth versus TPU v5e. The chips are 67% more energy-efficient; models including Gemini 1.5 Flash announced at I/O were trained and served on TPUs.

  • Supports up to 256 TPUs in a single pod; hundreds via multislice; tens of thousands via Google Jupiter network.
  • Third-generation SparseCore accelerator handles ultra-large embeddings for ranking and recommendation workloads.
  • 2x HBM capacity and bandwidth enable larger models with more weights and faster memory access.
  • Available later in 2024; represents the most sustainable TPU generation to date.
Read the original · Blog

More on Tensor Processing Unit (TPU)

Primer