A field guide to reward hacking in tool-using agents
The shortcuts agents find when a check only looks at the surface — and the tests that close each one.
Our goal is narrow on purpose: specialist open models that match a frontier model on the tasks a company repeats every day — trained on outcomes that were checked, not on answers that merely sounded right.
How do you know a verifier can't be gamed before you train on it?
A separate attacker searches for answers that score high while failing the task. Every shortcut it finds becomes a permanent test the verifier must pass.
Can a small model match a frontier model on one task family without copying the whole model?
Distill the task, not the model: a warm start on verified examples, then reinforcement learning where the same verifiers are the reward.
How do you keep training every week without the model forgetting what it already does well?
Replay from the golden set, per-task regression gates, and automatic rollback when a task's score slips below its bar.
What should happen to old examples when a verifier gets better?
Deduplicate, resolve contradictions, re-grade the whole set against each new verifier version, and keep provenance for every example.
When is a model good enough to take real traffic — and when should it hand the work back?
Side-by-side scoring against the frontier model on live tasks, shadow mode before routing, and thresholds the customer sets.
A preference score says an answer looks better. For an agent that moves money, edits code or books a shipment, that isn't the same as the work being done.
When a model is rewarded for what a grader likes, it gets very good at pleasing the grader. The score keeps climbing while real task success flattens — the gap in the figure.
When the reward comes from checking the outcome in your own systems — did the refund post, did the tests pass — there is much less room to look right without being right.
The shortcuts agents find when a check only looks at the surface — and the tests that close each one.
Why a narrow specialist trained on verified examples can match a generalist on the work it repeats.
Replay, regression gates and rollback for models that never stop learning.
Tracking where every example came from, so a set can be filtered, audited or rebuilt at any time.
Side-by-side scoring, shadow mode, and the thresholds that decide a hand-off — or a hand-back.
Keeping a golden set honest as the checks that built it get stricter.
How a number was produced matters as much as the number. Write-ups include setup, verifiers and failure cases.
Fixed seeds, versioned configs and the exact verifier revision behind every run.
What a method can't do goes on the first page, not in a footnote.
Research on customer tasks runs only with permission, inside the customer's own account.
You get a first golden set, the verifiers that built it, and an honest read on which tasks your own model can take over.
Noted — we'll be in touch about scoping the pilot.