Primer
Transformer
The neural-network architecture behind today's large AI models, built around a mechanism called attention.
- Input is split into tokens (word pieces, image patches, amino acids); attention lets every token weigh every other token when working out its meaning.
- Unlike earlier networks that read one step at a time, it processes all tokens in parallel, which suits GPUs and let training scale to trillions of tokens.
- Introduced by Google researchers in 2017, the same design now handles text, images, audio, proteins and robot actions.
- Attention cost grows with the square of input length, so doubling the context roughly quadruples that work; much research aims to cut this.