Primer

Reasoning model

A language model trained to produce a long chain of intermediate reasoning before it gives a final answer.

  • It drafts steps, checks them and backtracks, much as a person works on scratch paper; this text is often hidden from the user.
  • The behaviour is trained mostly with reinforcement learning on problems whose answers can be checked automatically, such as maths and code.
  • It opened a second way to scale: accuracy rises with the computation spent thinking at answer time, not only with model size.
  • Costs are answers that take seconds to minutes and many more tokens; gains are smaller where there is no checkable right answer.
Primer