Primer
Reasoning model
A language model trained to produce a long chain of intermediate reasoning before it gives a final answer.
- It drafts steps, checks them and backtracks, much as a person works on scratch paper; this text is often hidden from the user.
- The behaviour is trained mostly with reinforcement learning on problems whose answers can be checked automatically, such as maths and code.
- It opened a second way to scale: accuracy rises with the computation spent thinking at answer time, not only with model size.
- Costs are answers that take seconds to minutes and many more tokens; gains are smaller where there is no checkable right answer.