Primer
Vision-language-action model
An AI model that takes camera images and a written instruction and outputs motor commands for a robot.
- It starts from a vision-language model trained on internet images and text, so it knows what a mug or a drawer is before seeing any robot data.
- It is then fine-tuned on recordings of robots doing tasks, learning to emit joint and gripper commands as short chunks of motion several times a second.
- Many designs pair a slow, large 'planner' that interprets the scene with a small fast controller running tens to hundreds of times a second.
- Robot data is the bottleneck: tens of thousands of hours at most, versus trillions of words for chatbots, and demos at 90% success are far from factory-grade.