StoryTest

Tenstorrent claims Blackhole Galaxy clusters run language models about 3x faster than GPUs, at TT-Deploy JP

TenstorrentBlackhole

At TT-Deploy JP, Tenstorrent reported Galaxy Blackhole superclusters achieving 900 tokens per second per user on Kimi K2.6 (claimed 3x faster than GPUs) and 400+ tokens per second on DeepSeek-R1 671B. Tenstorrent also launched the TT-Ascalon S RISC-V CPU and announced 120+ Galaxy systems deployed with ai& for sovereign AI in Japan.

  • 900 tokens/second/user on Kimi K2.6, claimed 3x faster than GPUs; 400+ tokens/second on DeepSeek-R1 671B.
  • TT-Ascalon S IP claims ~140% performance per mm² in footprint roughly 50% smaller than its predecessor.
  • Largest deployment: 120+ Galaxy systems with ai& for sovereign AI infrastructure in Japan.
Read the original · Blog

More on Blackhole

Primer