Primer
Large language model
A transformer trained on trillions of words to predict the next token, which yields a general ability to write, answer and reason.
- Predicting text well forces the model to absorb grammar, facts and patterns of argument, because those are what make the next word predictable.
- Frontier models have hundreds of billions to trillions of parameters, the numbers tuned in training, and a run can cost tens to hundreds of millions of dollars.
- Pre-training yields a raw text predictor; further training on example dialogues and human or AI feedback turns it into a usable assistant.
- It can state falsehoods fluently ('hallucination') because it is optimised to produce plausible text, not verified truth.