Primer
AI accelerator
A processor built for the massively parallel arithmetic of neural networks, such as a GPU or Google's TPU.
- Neural networks are dominated by multiplying large grids of numbers, so accelerators pack thousands of simple multiply-add units working in parallel.
- They gain speed with low-precision numbers, 8 or even 4 bits instead of 32, trading tiny accuracy losses for several times more throughput.
- Big models span thousands of chips, so memory bandwidth and chip-to-chip links matter as much as raw compute; NVIDIA's lead rests on these and its software.
- A top chip now draws roughly a kilowatt, pushing data centres to liquid cooling, and nearly all are manufactured by TSMC in Taiwan.