Policy training pipeline
A task in plain English and a model of your robot in; a trained, verified policy out. No one hand-tunes a reward.
- Now: trained policies per task, verified in simulation, for parallel-jaw grippers
- Ranked on held-out success, never on reward
- Delivered with its success rate per condition, and footage
- Next: the pipeline itself, licensed to automation teams
We want the pipeline in companies' hands as soon as possible. What stands between here and there is proof on real hardware. If you run a cell, that is where you come in.