Quickstart
Go from an agent in production to verified examples and a parity score. You'll deploy the connector inside your own cloud, approve what gets redacted, pick one task, attach verifiers, and promote your model once it matches the frontier.
Before you start
- An agent in production that calls a model through an AI gateway or SDK.
- Admin access to the cloud account (AWS, GCP or Azure) where the connector will run.
- A way to see outcomes for the task you'll pick — ticket state, ledger entries, test results.
- Someone who can approve redaction rules, usually your security or data owner.
1Deploy the connector in your cloud
Sign in with the CLI, then deploy the connector next to your gateway. In mirror mode it reads a copy of traffic, so your agent's requests never wait on it.
# sign in and choose the account the connector runs in
$ katyar login
$ katyar connector deploy \
--cloud aws --region eu-west-1 \
--gateway https://llm-gateway.internal.example.com \
--mode mirror
✓ connector running in your account (aws · eu-west-1)
✓ mirroring traffic from llm-gateway.internal.example.com
What the connector captures
- Prompts and responses, including multi-turn conversations.
- Tool calls, their arguments and their results.
- The model and route that served each request.
- Joins to the outcome events you connect in step 3.
2Approve redaction rules
Before anything is stored, the connector proposes redaction rules from a sample of live traffic. Nothing is written until someone with the data-owner role approves them.
# katyar/redaction.yaml — proposed from 500 sampled traces
rules:
- { field: "*.email", action: mask } # seen in 38% of samples
- { field: "*.phone", action: mask }
- { field: "customer.address", action: drop }
- { pattern: card_number, action: drop }
- { field: "order.id", action: keep } # needed to join outcomes
$ katyar redaction review
$ katyar redaction approve --by data-owner@yourco.com
3Choose a task
Start with one task that repeats often and has an outcome you can check. katyar groups your traffic into candidate tasks and shows how much volume each has and how checkable it is.
$ katyar tasks suggest --last 7d
TASK TRACES/WK OUTCOME SOURCE CHECKABLE
refund-triage 28,140 ledger + ticket state high
order-lookup 41,002 orders API high
sql-reports 9,318 warehouse high
contract-review 3,044 human edits medium
escalations 6,210 — low
$ katyar tasks create refund-triage --outcomes ledger,tickets
Skip tasks marked low. If you can't say what a correct answer looks like, a verifier can't either — those tasks stay on your frontier model.
4Add verifiers
Verifiers decide which traces count as correct. A trace joins the golden set only if every verifier on its task passes. Most teams combine checks from the library with one written by the people who own the process.
From the library
$ katyar verifiers add refund-triage \
tools.schema_valid \
support.ledger_agrees --policy refunds@v7 \
privacy.no_pii
Write your own
# verifiers/refund_window.py
from katyar import verifier, Reward
@verifier("yourco.refund_window", version="1.0")
def refund_window(trace) -> Reward:
order = trace.outcome("orders").get(trace.args["order_id"])
refund = trace.outcome("ledger").refund_for(order.id)
return Reward.all(
refund is not None, # money actually moved
order.age_days <= policy.window("refunds"), # inside the window
abstain_if=order.disputed, # unsure → not golden
)
$ katyar verifiers test verifiers/refund_window.py --sample 200
✓ 200 traces · 171 pass · 22 fail · 7 abstain
$ katyar verifiers publish verifiers/refund_window.py --task refund-triage
5Watch parity and promote
As verified examples build up, katyar trains your open model on them and scores it side by side with the frontier model on the same task. A task can only move once your model clears the threshold you set.
$ katyar parity refund-triage
MODEL SUCCESS PARITY STATUS
frontier (current) 0.92 100% serving
yours · v14 0.89 97% ready to promote
$ katyar promote refund-triage --threshold 95% --canary 10%
✓ 10% of refund-triage traffic now served by yours · v14
✓ rollback armed: parity below 95% for 30 min → back to frontier
Raise the canary as confidence grows, or move the task back yourself at any time:
$ katyar promote refund-triage --canary 100%
$ katyar rollback refund-triage # immediate; the frontier takes the task back
What happens next
The loop keeps running after the first promotion:
- New traffic keeps flowing through your verifiers, so the golden set grows every week.
- Each training cycle produces a new model version, scored against the frontier before it can serve.
- Add your next task with
katyar tasks create, starting from the suggestions list. - Export the golden set whenever you like, in open formats. It's yours.