Graded on conditions it never saw
We built a pipeline that trains robot grasping policies with no hand-written reward. The part we put the most care into is knowing when a result is real. Here is how we grade a policy, and five failures that grading caught.
A search instead of hand-tuning
Our pipeline searches for the reward instead of hand-writing it. A reward is the formula that scores every moment of every attempt when a robot learns to pick something up with reinforcement learning, and writing one by hand is expert, slow work. Reward the hand for being near the part and it learns to hover. Reward lifting and it learns to flick the part into the air. Every task, part and gripper needs another round of tuning.
In our search, a language model writes candidate rewards from a vocabulary of tested terms, following NVIDIA's Eureka method. The system trains each candidate in simulation, measures what it actually did, and revises. No one edits a weight by hand.
A search is only as good as its judge, though. If the judge can be fooled, the search finds the way to fool it.
Rule one: rank on measured success
A reward that hands out big numbers while doing nothing useful would win any ranking based on reward. So we rank on measured task success: did the block end up resting in the bin, released gently, after a clean grip?
Then we go one step further and measure it on conditions the policy never trained on. If training used approach distances of 2 to 7 cm, grading uses 7 to 10 cm.
That gap is real. One placing policy succeeded 91% of the time under training conditions and 43% under held-out ones. Ranked on the first number, the system would have stopped there and called it solved.
Five things the grading caught
Each one looked like progress on some dashboard. Each one left behind a permanent rule that every later search starts with.
1. A grip that held by luck. Pick success read 99%. Watching the simulation showed the gripper closing near an edge of the block, where it pivots and gets set down crooked. The success test checked contact and angle but not where the grip was. We added a grip-position check, with the ideal grip point defined for each robot rather than assumed to be the centre.
2. Friction that went missing. Blocks slid out of a firm grip while the arm held still. The simulated wrist was being teleported to its target pose many times a second. That reset the contact state the physics engine builds static friction from, so a grip set to friction 0.75 held like friction 0.08. Every one of the heavier blocks slid out. Driving the wrist by velocity instead took slip to 0% at every block weight. No reward could have fixed it, and the first gate in our pipeline now tests the task itself before any reward search runs.
3. A policy that needed its own noise. During training, policies act with random noise added, which helps them explore. One policy placed 84% of the time with that noise and 38% without it, and a deployed robot runs without it. The pipeline now ends every task with staged "noise down" training and ranks the final stage on the noise-free policy. That is where our 95% held-out pick-and-place figure comes from.
4. Paid to carry a sloppy grip. On a second, domain-randomised gripper, the carry reward paid any two-finger grip nearly as well as a good one. Training traded grip quality for speed, and the pipeline's diagnosis showed it plainly: most failed episodes carried the block, just badly. The fix was not a bigger weight. It was a rule: carry pay starts only after the grip passes the success test's own check.
5. Stuck, and saying so. When every method the system knows has been tried against a failure and nothing beats the current best, it stops spending GPU hours and reports stuck, naming the failure. That is how a person learns the task, not the reward, needs a new idea.
Why this matters on a production line
An automation team needs to know how often the policy will succeed on the parts and positions it will actually meet. So the deliverable from our pipeline is the policy and its success rate under each held-out condition, with footage.
- No hand-tuning: rewards are generated, gated and revised automatically.
- Graded honestly: ranked on held-out success, never on the reward.
- Learns from mistakes: every failure becomes a check on every future run, on every robot.
Where we are
Everything above is simulation, in NVIDIA Isaac Lab. A box-jaw prototype gripper is solved end to end at 95% held-out pick-and-place. An industry-standard adaptive gripper with domain randomisation is working through the same stages now. The next milestone is real hardware: a real cell, with real success rates next to the simulated ones.
If you run a pick, kitting or machine-tending cell and want it to be one of those first real-world tests, we would like to hear from you.