Sign in Book demo →
katyar/Research
Research

Teaching small models to do real work.

Our goal is narrow on purpose: specialist open models that match a frontier model on the tasks a company repeats every day — trained on outcomes that were checked, not on answers that merely sounded right.

Focus areas5
Write-ups6 in progress
Dataopen benchmarks, or yours in your cloud

The research agenda

01 · five open questions
01

Verifier robustness & reward hacking

How do you know a verifier can't be gamed before you train on it?

Our approach

A separate attacker searches for answers that score high while failing the task. Every shortcut it finds becomes a permanent test the verifier must pass.

active
02

Specialists under verifiable reward

Can a small model match a frontier model on one task family without copying the whole model?

Our approach

Distill the task, not the model: a warm start on verified examples, then reinforcement learning where the same verifiers are the reward.

active
03

Continuous training without drift

How do you keep training every week without the model forgetting what it already does well?

Our approach

Replay from the golden set, per-task regression gates, and automatic rollback when a task's score slips below its bar.

in progress
04

Golden-data curation

What should happen to old examples when a verifier gets better?

Our approach

Deduplicate, resolve contradictions, re-grade the whole set against each new verifier version, and keep provenance for every example.

in progress
05

Measuring parity

When is a model good enough to take real traffic — and when should it hand the work back?

Our approach

Side-by-side scoring against the frontier model on live tasks, shadow mode before routing, and thresholds the customer sets.

open question

Write-ups

02 · drafts and upcoming
6 write-ups
Verifiers

A field guide to reward hacking in tool-using agents

The shortcuts agents find when a check only looks at the surface — and the tests that close each one.

draft
Training

Distill the task, not the model

Why a narrow specialist trained on verified examples can match a generalist on the work it repeats.

draft
Training

Training every week without forgetting

Replay, regression gates and rollback for models that never stop learning.

coming soon
Data

Provenance for golden data

Tracking where every example came from, so a set can be filtered, audited or rebuilt at any time.

draft
Evaluation

What parity should mean before a model takes traffic

Side-by-side scoring, shadow mode, and the thresholds that decide a hand-off — or a hand-back.

draft
Data

Re-grading: what happens to old data when a verifier improves

Keeping a golden set honest as the checks that built it get stricter.

coming soon

How we do research

03 · four commitments
01

Publish methods, not just results

How a number was produced matters as much as the number. Write-ups include setup, verifiers and failure cases.

02

Reproducible by default

Fixed seeds, versioned configs and the exact verifier revision behind every run.

03

Limits stated up front

What a method can't do goes on the first page, not in a footnote.

04

Your data stays in your cloud

Research on customer tasks runs only with permission, inside the customer's own account.

Start with one agent

Point us at one agent. We'll show you its parity curve.

You get a first golden set, the verifiers that built it, and an honest read on which tasks your own model can take over.

01One agent, one task family, fixed length.
02Connector in your VPC. Redaction before anything moves.
03You keep the golden set, the verifiers and the weights.
04Your frontier model stays in place until parity is proven.