// SOFTWARE · AUTONOMOUS POLICY TRAINING

Describe the task. Get a verified policy.

Our pipeline turns a task written in plain English and a model of your robot into a trained manipulation policy. Nobody hand-tunes a reward. Every candidate is graded on its measured success under conditions it never trained on, never on its own reward.

Measured in simulation, on conditions held out from training.

95% Held-out pick-and-place, parallel-jaw gripper
0 Reward weights tuned by hand
~1 day To bring up a new parallel gripper
2 Grippers through the same pipeline

// Simulation results in NVIDIA Isaac Lab, with domain randomisation on the second gripper. 95% is the deterministic policy (no sampling noise) on approach distances it never trained at. Real-hardware validation is the next milestone.

A search over rewards, judged by measurement.

Our loop writes, tests and revises rewards on its own, taking over the slow, expert part of training a robot policy. A person sets the task and the bar. The system does the iterations.

  1. Task in English

    "Grip the block squarely, carry it low over the bin, set it down, then let go." That sentence plus your robot's model is the whole input.

    Input
  2. Generate rewards

    A language model writes candidate rewards from a vocabulary of tested terms. It is told every failure the system has already seen.

    LLM · Eureka method
  3. Gate before training

    Checks that cost milliseconds to minutes reject a broken reward before it can waste an hour of GPU time.

    Cheapest check first
  4. Train in parallel

    Thousands of simulated robots per GPU, several candidates and seeds at once. Runs that stall or collapse are stopped automatically.

    NVIDIA Isaac Lab
  5. Grade on held-out conditions

    Each policy is re-tested at distances and settings it never trained on. That number, not the reward, decides the ranking.

    Measured success
  6. Diagnose the failure

    Every failed episode gets a category: never gripped, dropped on the way, stopped short, released too high. The dominant failure picks the next fix.

    Failure catalogue
  7. Advance the curriculum

    Pick, then carry, then place. Each stage starts from the last stage's winner, and moves on only after it clears its own bar.

    Staged training
  8. Verified policy out

    You get the policy, its success rate under each held-out condition, and footage. The measurement is part of what you receive.

    Deliverable

Steps 02–06 repeat until the policy clears the bar or the system reports it is stuck, with the reason. No one watches the dashboards.

  • PPick
  • TCarry
  • APlace
  • B–CWider approaches
  • NNoise down: deterministic policy

Every reward is checked before it reaches a GPU.

Each check runs before the next, more expensive one. The first gate tests the task itself: a scripted reference policy has to succeed in every condition, and the success test has to reject drops, throws and a gripper that never lets go. A task has to be solvable before any reward search begins.

~5 min, onceEnvironment gateThe task is solvable, and the success test cannot be gamed
MillisecondsReward spec checkKnown terms, valid parameters, no zero weights
~1 secondReward regression suiteEvery term pays the right sign, and a wrong grip pays nothing
SecondsScripted-attempt rankingA correct attempt must outscore a drop, a swat or never releasing
~1 hour GPUTraining plus held-out gradingOnly what passed everything above

Have a pick, kitting or machine-tending task?Phase 1 is ready now: we can train and verify a policy for your gripper and part in simulation.

Tell us the task →

We grade every policy on conditions it never trained on.

Held-out grading measures what a policy will do on the parts and positions it meets next. A policy scored on its own training conditions is partly scored on memory: in one of our own placing runs, the top policy succeeded 91% of the time under training conditions and 43% under held-out ones. Ranking on the first number would have shipped the wrong policy.

91% Same policy, training conditions
43% Same policy, held-out conditions: the number we rank on

// Placing stage, prototype gripper, simulation. The system then trained on until the deterministic policy reached 95% held-out.

Every failure becomes a rule the next search starts with.

When a run fails in a new way, the cause is written into a failure catalogue with its evidence and the safeguard that now prevents it. Every future reward search reads that catalogue. A fix made once applies to every policy after it, on every robot.

One recent example: a carry reward paid an off-centre grip almost as well as a good one, so the policy traded grip quality for speed. The fix was not a new weight. It was a rule: carry pay starts only after the grip passes the success test's own check. That rule now applies to every task the system trains.

What changes for an automation team.

Hand-tuned RLMK2 pipeline
RewardWritten and re-tuned by an expert, per taskGenerated, gated and revised automatically
RankingReward curves and a person's judgementMeasured success on held-out conditions
FailuresRediscovered on the next projectCatalogued, and checked on every future run
SupervisionSomeone watches every runRuns stop, advance or report stuck on their own
What you receiveA checkpointA policy, its numbers per condition, and footage
New gripperStart overAbout a day of setup; same task, same checks

From trained policies to the pipeline in your hands.

Every number on this page is from simulation, and the bridge to real hardware is already going in: domain randomisation on friction, drive stiffness and sensor noise, so a policy cannot rely on one perfect simulated world. Real hardware is what comes next.

Can your cell be our first real-hardware test?A gripper, an arm and a concrete task are all it takes.

Offer a real-hardware cell →

How we grade a grasp policy.

The blog covers why we rank on held-out conditions, and what the failure catalogue caught on its way to a 95% pick-and-place policy.

Help us get it on real hardware.

Tell us the gripper, the part and what success looks like. We will reply with how we would train it and how we would measure it, in simulation now and in your cell next.

Email us →