Primer
Reinforcement learning
Training an AI by letting it act, scoring the outcomes with a reward, and strengthening the choices that led to high scores.
- No one supplies the right answer; the system discovers strategies by trial and error, which lets it exceed the humans it might otherwise copy.
- DeepMind's AlphaGo used it to beat top professional Lee Sedol at Go in 2016; robots learn to walk from millions of simulated attempts.
- Rewarding verifiably correct maths answers and code that passes tests is the core recipe behind today's reasoning models.
- The main risk is 'reward hacking': the system maximises the score in unintended ways, such as editing a test instead of fixing the code.