StorySoftware
Figure unveils Helix, a vision-language-action model that runs full upper-body control onboard
Figure introduced Helix, a vision-language-action model for humanoid upper-body control. A slow 7–9 Hz planner (7B parameters) reasons about scenes and language, while a fast 200 Hz policy (80M parameters) drives the motors. Trained on ~500 hours of teleoperated data and runs entirely onboard on low-power embedded GPUs.
- Two-system architecture: 7–9 Hz planner (7B parameters) + 200 Hz policy (80M parameters).
- Controls 35 degrees of freedom across the humanoid upper body, torso, head, wrists and fingers.
- Trained on only ~500 hours of teleoperated data—far less than prior vision-language-action datasets.
- Runs entirely onboard on low-power embedded GPUs; single unified weights for all tasks with no task-specific fine-tuning.