// SOFTWARE · AUTONOMOUS POLICY TRAINING
Describe the task. Get a verified policy.
Our pipeline turns a task written in plain English and a model of your robot into a trained manipulation policy. Nobody hand-tunes a reward. Every candidate is graded on its measured success under conditions it never trained on, never on its own reward.
01 — Results so far
Measured in simulation, on conditions held out from training.
// Simulation results in NVIDIA Isaac Lab, with domain randomisation on the second gripper. 95% is the deterministic policy (no sampling noise) on approach distances it never trained at. Real-hardware validation is the next milestone.
02 — How it works
A search over rewards, judged by measurement.
Our loop writes, tests and revises rewards on its own, taking over the slow, expert part of training a robot policy. A person sets the task and the bar. The system does the iterations.
-
Task in English
"Grip the block squarely, carry it low over the bin, set it down, then let go." That sentence plus your robot's model is the whole input.
Input -
Generate rewards
A language model writes candidate rewards from a vocabulary of tested terms. It is told every failure the system has already seen.
LLM · Eureka method -
Gate before training
Checks that cost milliseconds to minutes reject a broken reward before it can waste an hour of GPU time.
Cheapest check first -
Train in parallel
Thousands of simulated robots per GPU, several candidates and seeds at once. Runs that stall or collapse are stopped automatically.
NVIDIA Isaac Lab -
Grade on held-out conditions
Each policy is re-tested at distances and settings it never trained on. That number, not the reward, decides the ranking.
Measured success -
Diagnose the failure
Every failed episode gets a category: never gripped, dropped on the way, stopped short, released too high. The dominant failure picks the next fix.
Failure catalogue -
Advance the curriculum
Pick, then carry, then place. Each stage starts from the last stage's winner, and moves on only after it clears its own bar.
Staged training -
Verified policy out
You get the policy, its success rate under each held-out condition, and footage. The measurement is part of what you receive.
Deliverable
Steps 02–06 repeat until the policy clears the bar or the system reports it is stuck, with the reason. No one watches the dashboards.
- PPick
- TCarry
- APlace
- B–CWider approaches
- NNoise down: deterministic policy
03 — Gates
Every reward is checked before it reaches a GPU.
Each check runs before the next, more expensive one. The first gate tests the task itself: a scripted reference policy has to succeed in every condition, and the success test has to reject drops, throws and a gripper that never lets go. A task has to be solvable before any reward search begins.
Have a pick, kitting or machine-tending task?Phase 1 is ready now: we can train and verify a policy for your gripper and part in simulation.
Tell us the task →04 — Why held-out grading
We grade every policy on conditions it never trained on.
Held-out grading measures what a policy will do on the parts and positions it meets next. A policy scored on its own training conditions is partly scored on memory: in one of our own placing runs, the top policy succeeded 91% of the time under training conditions and 43% under held-out ones. Ranking on the first number would have shipped the wrong policy.
// Placing stage, prototype gripper, simulation. The system then trained on until the deterministic policy reached 95% held-out.
05 — It learns from its mistakes
Every failure becomes a rule the next search starts with.
When a run fails in a new way, the cause is written into a failure catalogue with its evidence and the safeguard that now prevents it. Every future reward search reads that catalogue. A fix made once applies to every policy after it, on every robot.
One recent example: a carry reward paid an off-centre grip almost as well as a good one, so the policy traded grip quality for speed. The fix was not a new weight. It was a rule: carry pay starts only after the grip passes the success test's own check. That rule now applies to every task the system trains.
06 — Compared
What changes for an automation team.
| Hand-tuned RL | MK2 pipeline | |
|---|---|---|
| Reward | Written and re-tuned by an expert, per task | Generated, gated and revised automatically |
| Ranking | Reward curves and a person's judgement | Measured success on held-out conditions |
| Failures | Rediscovered on the next project | Catalogued, and checked on every future run |
| Supervision | Someone watches every run | Runs stop, advance or report stuck on their own |
| What you receive | A checkpoint | A policy, its numbers per condition, and footage |
| New gripper | Start over | About a day of setup; same task, same checks |
07 — Phases
From trained policies to the pipeline in your hands.
Every number on this page is from simulation, and the bridge to real hardware is already going in: domain randomisation on friction, drive stiffness and sensor noise, so a policy cannot rely on one perfect simulated world. Real hardware is what comes next.
-
Ready now
Phase 1 — Trained policies per task
We run the pipeline for a given gripper and task and deliver the policy with its held-out success rates and footage. Verified in simulation; parallel-jaw grippers today.
-
Critical next step
Phase 2 — Real-hardware validation
A real gripper and arm in a real cell, with real success rates reported next to the simulated ones.
This is the step that lets us sell the pipeline to companies, and we want to get there as fast as possible. What we need is hardware: a gripper, an arm and a concrete pick, kitting or tending task.
-
As soon as Phase 2 proves it
Phase 3 — The pipeline, licensed to companies
Automation teams run the pipeline on their own robots and tasks. The same pipeline then brings up the Dex Hand: the declarations a new gripper needs are exactly the ones a multi-finger hand needs.
Can your cell be our first real-hardware test?A gripper, an arm and a concrete task are all it takes.
Offer a real-hardware cell →08 — Read more
How we grade a grasp policy.
The blog covers why we rank on held-out conditions, and what the failure catalogue caught on its way to a 95% pick-and-place policy.
Contact
Help us get it on real hardware.
Tell us the gripper, the part and what success looks like. We will reply with how we would train it and how we would measure it, in simulation now and in your cell next.
Email us →