Sign in Book demo →
katyar/Docs
Docs

Build with katyar, one task at a time.

Get started / Quickstart

Quickstart

Go from an agent in production to verified examples and a parity score. You'll deploy the connector inside your own cloud, approve what gets redacted, pick one task, attach verifiers, and promote your model once it matches the frontier.

SetupAbout 30 minutes
Runs inYour cloud account
You needAdmin on one agent

Before you start

  • An agent in production that calls a model through an AI gateway or SDK.
  • Admin access to the cloud account (AWS, GCP or Azure) where the connector will run.
  • A way to see outcomes for the task you'll pick — ticket state, ledger entries, test results.
  • Someone who can approve redaction rules, usually your security or data owner.
Nothing leaves your cloud. The connector, verifiers, golden sets and training jobs all run inside your account. katyar never receives raw traces.

1Deploy the connector in your cloud

Sign in with the CLI, then deploy the connector next to your gateway. In mirror mode it reads a copy of traffic, so your agent's requests never wait on it.

# sign in and choose the account the connector runs in
$ katyar login
$ katyar connector deploy \
    --cloud aws --region eu-west-1 \
    --gateway https://llm-gateway.internal.example.com \
    --mode mirror

✓ connector running in your account (aws · eu-west-1)
✓ mirroring traffic from llm-gateway.internal.example.com

What the connector captures

  • Prompts and responses, including multi-turn conversations.
  • Tool calls, their arguments and their results.
  • The model and route that served each request.
  • Joins to the outcome events you connect in step 3.

2Approve redaction rules

Before anything is stored, the connector proposes redaction rules from a sample of live traffic. Nothing is written until someone with the data-owner role approves them.

# katyar/redaction.yaml — proposed from 500 sampled traces
rules:
  - { field: "*.email",          action: mask }  # seen in 38% of samples
  - { field: "*.phone",          action: mask }
  - { field: "customer.address", action: drop }
  - { pattern: card_number,      action: drop }
  - { field: "order.id",         action: keep }  # needed to join outcomes

$ katyar redaction review
$ katyar redaction approve --by data-owner@yourco.com
Redaction runs before storage, not after. Masked and dropped fields never reach disk, so they can't end up in a golden set or a training run.

3Choose a task

Start with one task that repeats often and has an outcome you can check. katyar groups your traffic into candidate tasks and shows how much volume each has and how checkable it is.

$ katyar tasks suggest --last 7d

TASK              TRACES/WK   OUTCOME SOURCE          CHECKABLE
refund-triage        28,140   ledger + ticket state   high
order-lookup         41,002   orders API              high
sql-reports           9,318   warehouse               high
contract-review       3,044   human edits             medium
escalations           6,210   —                       low

$ katyar tasks create refund-triage --outcomes ledger,tickets

Skip tasks marked low. If you can't say what a correct answer looks like, a verifier can't either — those tasks stay on your frontier model.

4Add verifiers

Verifiers decide which traces count as correct. A trace joins the golden set only if every verifier on its task passes. Most teams combine checks from the library with one written by the people who own the process.

From the library

$ katyar verifiers add refund-triage \
    tools.schema_valid \
    support.ledger_agrees --policy refunds@v7 \
    privacy.no_pii

Write your own

# verifiers/refund_window.py
from katyar import verifier, Reward

@verifier("yourco.refund_window", version="1.0")
def refund_window(trace) -> Reward:
    order = trace.outcome("orders").get(trace.args["order_id"])
    refund = trace.outcome("ledger").refund_for(order.id)
    return Reward.all(
        refund is not None,                           # money actually moved
        order.age_days <= policy.window("refunds"),  # inside the window
        abstain_if=order.disputed,                    # unsure → not golden
    )
$ katyar verifiers test verifiers/refund_window.py --sample 200
✓ 200 traces · 171 pass · 22 fail · 7 abstain

$ katyar verifiers publish verifiers/refund_window.py --task refund-triage
Abstaining is a feature. When a verifier can't be sure, it abstains and the trace stays out of the golden set. Fewer, cleaner examples train a better model than more, noisier ones.

5Watch parity and promote

As verified examples build up, katyar trains your open model on them and scores it side by side with the frontier model on the same task. A task can only move once your model clears the threshold you set.

$ katyar parity refund-triage

MODEL                SUCCESS   PARITY   STATUS
frontier (current)      0.92     100%   serving
yours · v14             0.89      97%   ready to promote

$ katyar promote refund-triage --threshold 95% --canary 10%
✓ 10% of refund-triage traffic now served by yours · v14
✓ rollback armed: parity below 95% for 30 min → back to frontier

Raise the canary as confidence grows, or move the task back yourself at any time:

$ katyar promote refund-triage --canary 100%
$ katyar rollback refund-triage   # immediate; the frontier takes the task back
Rollback is automatic. Promoted tasks keep being checked. If the score slips below your threshold, traffic moves back to the frontier model without anyone having to step in.

What happens next

The loop keeps running after the first promotion:

  • New traffic keeps flowing through your verifiers, so the golden set grows every week.
  • Each training cycle produces a new model version, scored against the frontier before it can serve.
  • Add your next task with katyar tasks create, starting from the suggestions list.
  • Export the golden set whenever you like, in open formats. It's yours.
Was this page helpful? Suggest an edit
Need a hand?

We'll set up your first task with you.

Book a setup call →