Primer

Inference

Running a trained AI model to produce outputs, as opposed to training it.

  • Training is paid once per model; inference is paid on every request, for every user, for the model's whole working life.
  • Language models generate one token at a time, and each token requires reading essentially all the active weights from memory.
  • So speed is often limited by memory bandwidth rather than arithmetic, which is why fast memory and serving many users in one batch matter so much.
  • With growing usage and reasoning models that 'think' longer, inference is becoming the larger share of AI computing spend.
Primer