BaseStarts from a vision-language model
WALL-OSS builds on a 3-billion-parameter model that already understands images and text, then is taught to output robot motions.
ReasoningThink, plan, then move
One network can go from an instruction to a written reasoning step, a list of subtasks and then smooth joint motions, or skip straight to motion.
ActionsCoarse tokens, then smooth motion
Training first predicts compressed action tokens, then a flow-matching head that generates continuous, smooth arm trajectories.
PhysicsWALL-B predicts the world
WALL-B learns vision, language and action together from real-home data, giving the robot an intuitive sense of how objects react to its moves.