StoryResearch

π*0.6 learns from its own deployments with the Recap reinforcement-learning method, more than doubling throughput on some hard tasks

Physical Intelligenceπ model family

Physical Intelligence introduced π*0.6, trained with Recap (reinforcement learning combining human demonstrations, expert corrections during deployment, and autonomous trial-and-error). On hard tasks like espresso-making, throughput more than doubled and failures dropped by half or more.

  • Recap method integrates human demo data, real-time expert correction, and RL into one training pipeline.
  • Espresso task: throughput >2x baseline; failures reduced by 50%+; sustained 18-hour continuous operation.
  • Laundry folding: ~3 minutes per item in novel homes; box assembly: ~2.5 minutes per unit in factory.
  • Autonomous trial-and-error after deployment closes sim-to-real gap faster than pre-deployment training alone.
Read the original · Blog

More on π model family

Primer