Primer

Vision-language-action model

An AI model that takes camera images and a written instruction and outputs motor commands for a robot.

  • It starts from a vision-language model trained on internet images and text, so it knows what a mug or a drawer is before seeing any robot data.
  • It is then fine-tuned on recordings of robots doing tasks, learning to emit joint and gripper commands as short chunks of motion several times a second.
  • Many designs pair a slow, large 'planner' that interprets the scene with a small fast controller running tens to hundreds of times a second.
  • Robot data is the bottleneck: tens of thousands of hours at most, versus trillions of words for chatbots, and demos at 90% success are far from factory-grade.
Primer