Primer

Transformer

The neural-network architecture behind today's large AI models, built around a mechanism called attention.

  • Input is split into tokens (word pieces, image patches, amino acids); attention lets every token weigh every other token when working out its meaning.
  • Unlike earlier networks that read one step at a time, it processes all tokens in parallel, which suits GPUs and let training scale to trillions of tokens.
  • Introduced by Google researchers in 2017, the same design now handles text, images, audio, proteins and robot actions.
  • Attention cost grows with the square of input length, so doubling the context roughly quadruples that work; much research aims to cut this.
Primer