Every call your agents make to a frontier model is a lesson. We collect it, verify it, refine it into a golden dataset — and train an open model on it until it does the job just as well. On your data. In your walls.
Right now, the value of an AI call ends the moment the answer arrives. katyar checks each call against your own systems and keeps the ones that were right — as data you own and a model you run.
A small open model starts out knowing nothing about your business. Train it on verified examples of your own work, cycle after cycle, and it climbs to the model it learned from. Scroll — watch the gap close.
Transcripts, tool calls, the human who corrected the answer, what happened next. Today it expires in a log bucket.
A frontier model is brilliant and generic. Your context goes in with every prompt and is forgotten when the window closes.
Every token buys an answer, not an asset. When pricing, policy or the model changes, it's someone else's roadmap.
Continuous collection captures what your agents do. Continuous training turns it into golden data and better weights. Each figure below runs live.
One connector beside your AI gateway records every conversation, tool call and response — then joins it to what actually happened in your systems. Personal data is redacted before anything leaves your VPC.
// captured trace — redacted in your VPCuser "Order #4471 never arrived, I'm [EMAIL]"tool orders.lookup(4471) → shipped 12d agotool carrier.track(…) → lost in transittool refund.issue(48.20) → okassistant "Refunded $48.20 — sorry about that."outcome ledger −48.20 · no reopen · CSAT 5// queued for verification
Every trace runs through a stack of verifiers — from our library and ones your team writes. A trace is admitted only if every check passes. One failure and it's out, with the reason logged.
Duplicates collapse. Contradictions get resolved. When a verifier improves, every older example is re-graded against it. What survives is versioned — a dataset where nothing unchecked gets in.
A fine-tuning warm start on the golden set, then RL where the same verifiers become the reward. The open model is rewarded only for doing the task the way your systems say is correct — on your compute, into your weights.
Each task is scored side by side against the frontier model. Past the parity line, traffic moves to your model. Below it, the frontier keeps the work — and keeps generating the next lesson.
The model you deploy generates new traces. New traces make the golden set better. A better set trains a better model. The dot is one interaction going all the way round.
Every request, response and tool call from agents already in production.
Joined to outcomes and human edits. Personal data removed inside your VPC.
Library and custom verifiers decide what counts as correct. Failures are dropped.
Deduplicated, re-graded, versioned. A proprietary asset that grows every cycle.
Verifiers become the reward. An open model learns your tasks on your compute.
Tasks that match the frontier move to your model — and start producing new traces.
A verifier decides what counts as a correct answer, and becomes the reward the model trains on. Use ours, write your own, or both. A bio lab's definition of a valid protocol should come from the bio lab.
Tests pass in a sandbox; the diff builds.
Query runs; result set matches the reference.
Right tool, valid schema, arguments grounded in context.
Ledger, policy and ticket state all agree.
Numbers reconcile to the book of record.
No personal data in outputs or training rows.
Reagents in stock, temperatures within SOP, samples exist in LIMS.
Coverage rules applied the way your adjusters apply them.
They get good at specific tasks by doing them next to someone better, with a lead checking the work. After enough reps, they own the queue. We train open models exactly the same way.
Watches the senior handle every ticket.
Learns from frontier traces on your tasks. Serves no traffic.
Takes easy tickets. A lead reviews every answer.
Answers in shadow. Verifiers grade every output against the frontier.
Owns the routine queue. Escalates the strange ones.
Tasks past parity route to it. Everything else falls back to frontier.
Knows your systems better than any new hire.
Runs the domain. The frontier stays on call for the long tail.
A model that understands your company, and the verified data that made it. Both are yours, and both get better every cycle.
Open weights trained on your work, running where you choose. It knows your tools, your policies and your customers' vocabulary — and it doesn't change because someone else shipped a new version.
Every example checked by verifiers before it's admitted. When a better open model ships, you retrain on it instead of starting over. In the next era, this is the balance-sheet item.
Collection, training and inference inside your boundary. Nothing trains anyone else's model.
One failed check and a trace is out. The reason is logged, not hidden.
No one-off fine-tune that goes stale. Every week of production makes the next model better.
Tasks below parity keep routing to the frontier model. No big-bang cutover, ever.
Each task handed off moves spend from per-token rent to infrastructure you run.
Better open weights next quarter? Retrain on the same golden set and move on.
Every company will end up running models trained on its own work. The ones collecting verified data today will be the ones who get there — task by task, with a model they can point to.
Your agents run on the best model available. Every trace they produce starts building your golden set.
High-volume, verifiable tasks pass parity and move to your model. The frontier handles what's left.
A concrete, sovereign model that knows your company — with the frontier as a specialist you call, not a dependency.
Nothing about how your agents work today has to change on day one. What changes is what you're left holding.
| Frontier API only | With katyar | |
|---|---|---|
| Who owns the weights | the provider | you |
| Where company context lives | in the prompt, every call | in the model |
| What happens to agent traces | expire in logs | become golden data |
| Model and price changes | someone else's roadmap | your release schedule |
| Where the data goes | out, on every request | stays in your VPC |
| Cost over time | rises with usage | falls as tasks hand off |
We'll look at one of your agents with you and sketch its parity curve, live.
At everything, yes. At the twenty tasks your agents repeat every day, not necessarily. A specialist trained on your work can match a generalist on those tasks, and Claude still handles everything else.
No. Your engineers keep building agents the way they do today. katyar handles collecting, checking, training and promoting.
Your model can't take over a task until it clears the bar you set, and it keeps being checked after that. If a score slips, the task moves back to Claude automatically. You can also move it back yourself with one click.
No. Collection, checking and training run inside your own cloud account. Sensitive fields are redacted before anything is stored, and you approve the rules first.
The loop never stops: new work keeps coming in, getting checked and improving your model. New and harder work stays on the frontier model, and your verified library lets you test any new release against yours within an hour.
Your model learns from outcomes checked in your own systems: refunds that went through, tests that passed, fixes your people made. The checks do the teaching. At the start of every pilot we go through your AI providers' terms with your legal team.
You do. The model, the verified library and the checks all live in your cloud, in open formats. You keep all of it.
You get a first golden set, the verifiers that built it, and an honest read on which tasks your own model can take over.
Noted — we'll be in touch about scoping the pilot.