Checks whether the AI that grades your AI can be trusted.

Many teams use one AI model to grade the answers of another. The grading model is called a judge. judgekeeper compares the judge with answers a person already graded, and tells you how often the judge is right.

pip install judgekeeper

Free and open source (MIT license). It runs on your own computer. There is no account and no server.

The problem

Nobody checks the judge

Your AI product writes answers. To know whether they are good, you could read them all yourself, but that is slow. So teams ask another AI model to grade them: the judge. A judge is fast and cheap. It can also be wrong, in three ways.

Question: Which planet is the largest? Answer: Saturn is the largest planet. HUMAN: FAIL JUDGE: PASS

It disagrees with people

A judge can pass an answer a person would fail. A judge that passes everything makes every test look fine.

The same question, judged three times Run 1 PASS Run 2 PASS Run 3 PASS FAIL changed its mind

It is inconsistent

Ask a judge the same question three times and you can get two different answers.

Scores for the same product model update

It changes silently

The company behind the judge model can update it. Then every score moves, although your product did not change.

How it works

Five steps, from labels to a report

The idea is simple. A person grades a small set of answers once. judgekeeper then measures the judge against those grades, and keeps measuring.

  1. Pick about 100 answers from your product. A person marks each one pass or fail. That mark is called a human label. You can label in a spreadsheet, or in a small page judgekeeper opens on your own computer:

    judgekeeper label items.jsonl --out labels.csv
  2. judgekeeper saves the labeled examples as an anchor set: the fixed answer key every later check uses. It records a hash, a short fingerprint of the file such as sha256 2e75fe64157b…. If anyone edits a label later, the fingerprint no longer matches and judgekeeper stops.

    judgekeeper import-labels labels.csv -o anchors.jsonl
  3. Your judge grades every example. Each grade is a verdict. Each full pass over the examples is a run. Running three times shows how often the judge changes its mind.

    judgekeeper judge anchors.jsonl --callable my_judges:judge --runs 3 --out runs/mine

    Here my_judges:judge is your own Python function. judgekeeper can also call Anthropic or OpenAI models directly, or any program. See Works with what you already use.

  4. judgekeeper lines up each verdict with the human label for the same answer. It counts where they agree and where they do not, run by run.

    judgekeeper validate anchors.jsonl runs/mine --out reports/mine
  5. You get one page. It says how often the judge is right, how noisy it is, where it is weak, and lists every disagreement with the judge's own reason. It ends with a verdict: usable as a gate, usable with care or not trustworthy as a gate. Open reports/mine/report.html in your browser, or see a real one.

Already have the judge's verdicts and human labels in one spreadsheet? Skip to one command. It does steps 2 to 5 for you:

judgekeeper check results.csv --judge verdict --human label --out reports/my-judge

What the numbers mean

Try it: how good is this judge?

Illustration: you choose the numbers

Move the sliders. The table counts the four ways a judge and a person can line up. The numbers below it are the ones judgekeeper reports, computed the same way.

Each answer lands in one box
Judge passedJudge failed
People passed18agree2judge too harsh
People failed3judge too lenient17agree
TPR (true positive rate)
0.90

Of the answers people passed, the judge passed 18 of 20.

TNR (true negative rate)
0.85

Of the answers people failed, the judge failed 17 of 20.

Kappa
0.75

Agreement after taking away what lucky guessing would give. 0 is no better than chance, 1 is perfect.

Raw agreement (for comparison only)
88%

The share of all answers where judge and people agree. It can look high for a useless judge.

Usable with care: TNR is between 0.80 and 0.90.

Likely ranges with this few answers (95 percent Wilson intervals): TPR 0.70 to 0.97, TNR 0.64 to 0.95.

The report's verdict uses these rules: not trustworthy as a gate if TPR or TNR is below 0.80 or kappa is below 0.6; usable with care if TPR or TNR is below 0.90; otherwise usable as a gate. A gate is an automatic check that can stop a release.

Two more things it measures

Noise and position bias

Illustration

Noise: does the judge change its mind?

judgekeeper runs the judge several times on the same answers and counts how often a verdict changes. The report calls this the noise floor. Here, 1 of 5 answers flipped. With only one run the noise is reported as unknown, never as zero.

Illustration

Position bias: does order matter?

Order A, then B A B judge: A is better Order B, then A B A judge: A is better judge: B is better

When a judge compares two answers, it can favour whichever it reads first. judgekeeper asks both ways, A then B and B then A. If the verdict changes with the order, that is position bias.

In the real result below, 3.6% of items changed verdict in at least one of three runs, and 9.9% of verdicts changed when the two answers swapped places.

After the first check

It keeps checking

A judge that is right today can drift tomorrow. Three commands keep watch.

Gate: stop a release when the judge slips

CI (continuous integration) is the set of automatic checks that run on every code change. judgekeeper gate reads a report and returns a status. Each status has an exit code, a number a program returns when it ends. Anything but 0 stops the build.

judgekeeper gate reports/my-judge/report.json
judge gate: PASS (exit 0)
StatusMeansExit
PASSThe judge clears every check.0
FAILThe judge fell below a threshold, or dropped by more than its own noise.1
ANCHORS_CHANGEDSomeone edited the frozen examples.3
FLAKYToo few runs, or the judge flips too often to tell.4
JUDGE_CHANGEDA different model, prompt or setting: scores are not comparable.5

It never fails on a change smaller than the judge's own run-to-run noise.

Migrate: switch judge models safely

Illustration

Moving to a new judge model? Run both on the same frozen examples. migrate compares them answer by answer.

judgekeeper migrate anchors.jsonl runs/old/ runs/new/ --out reports/migration/
AnswerPeopleOld judgeNew judge
#1passpasspass
#2failpassfail (fixed)
#3failfailfail
#4passpassfail (broken)

BETTER: the new judge agrees with people more, by more than the noise. Switch, and expect scores to move because the judge improved.

Attribute: was it your app or the judge?

Your score dropped. judgekeeper grades the frozen examples again. Their answers never change, so if the judge's verdicts on them moved, the judge changed. If not, your app did.

judgekeeper attribute reports/now/report.json --app-score-before 0.82 --app-score-after 0.74

SYSTEM_CHANGE The judge still grades the frozen examples the same way, so the drop comes from your app.

StatusMeansExit
STABLENothing moved beyond the noise.0
JUDGE_DRIFTThe judge's verdicts on the frozen examples moved.6
SYSTEM_CHANGEThe judge held still; your app's score moved.7

Integrations

Works with what you already use

judgekeeper does not replace your tools. It reads what they already save. Open a tile for the exact command.

A spreadsheetCSV or JSONL file

One row per verdict, with a column for the judge's verdict and one for the human label. Columns named id, run, input, output and reason are used when present.

judgekeeper check results.csv --judge verdict --human label --out reports/my-judge/

Verdicts can be pass/fail, true/false, yes/no, correct/incorrect or 1/0. Scores need a rule such as --pass-if "score>=0.5". judgekeeper never guesses.

Your own Python function--callable

Write a function that takes one item (its id, input and output, never the human label) and returns a verdict:

# mypkg/judges.py
def my_judge(item):
    return call_my_model(item)   # "pass", "fail", True, False, a score, or {"verdict": ..., "reason": ...}
judgekeeper judge anchors.jsonl --callable mypkg.judges:my_judge --runs 3 --out runs/mine/
Any language--exec

Any program works. judgekeeper sends one item as JSON on standard input; the program prints a verdict.

judgekeeper judge anchors.jsonl --exec "node judge.js" --runs 3 --out runs/mine/
promptfooresults.json

Run your eval three times, then point judgekeeper at the results file and your human labels.

promptfoo eval -o results.json --repeat 3
judgekeeper import promptfoo results.json --metric helpfulness --labels labels.csv --out reports/helpfulness/

Full promptfoo guide

DeepEvaltest run files
export DEEPEVAL_RESULTS_FOLDER=deepeval-results   # then run your DeepEval tests 3 times
judgekeeper import deepeval deepeval-results/ --metric "Correctness [GEval]" --labels labels.csv --out reports/correctness/

Full DeepEval guide

Inspect AIeval logs
inspect eval task.py --epochs 3 --log-format json
judgekeeper import inspect logs/ --metric model_graded_qa --labels labels.csv --out reports/qa/

Full Inspect AI guide

MLflowassessments on traces
pip install "judgekeeper[mlflow]"
judgekeeper import mlflow --experiment my-app-eval --metric correctness --out reports/correctness/

Full MLflow guide

Langfusescores

Set LANGFUSE_PUBLIC_KEY and LANGFUSE_SECRET_KEY in your environment (see Setup and keys), then:

judgekeeper import langfuse --judge-score helpfulness --human-score helpfulness_human --from 2026-09-01 --pass-if "score>=0.5" --out reports/helpfulness/

Full Langfuse guide

Anthropic models--runner anthropic

judgekeeper calls the model for you with your own key, read from the environment. The prompt file holds your grading instructions; an example is prompts/single.md.

judgekeeper judge anchors.jsonl --runner anthropic --model claude-haiku-4-5-20251001 --prompt prompts/single.md --runs 3 --out runs/haiku/

How to set the key safely

OpenAI models--runner openai
judgekeeper judge anchors.jsonl --runner openai --model "$JUDGE_MODEL" --prompt prompts/single.md --runs 3 --out runs/openai/

$JUDGE_MODEL is the model name you use. How to set the key safely

Any OpenAI-compatible server--base-url

Azure OpenAI, OpenRouter, Together, a LiteLLM gateway (for AWS Bedrock or Google Vertex), Ollama or vLLM. For a local Ollama, no key is needed:

judgekeeper judge anchors.jsonl --runner openai --model "$JUDGE_MODEL" --base-url http://localhost:11434/v1 --prompt prompts/single.md --runs 3 --out runs/local/

More endpoints on the setup page

GitHub Actiona check on every change

The action runs the judge, writes the report and gates on it. Keys come from repository secrets.

- uses: judgekeeper/judgekeeper@main   # pin to a release tag
  env:
    ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
  with:
    anchors: evals/anchors.jsonl
    runner: anthropic
    model: claude-haiku-4-5-20251001
    prompt: prompts/judge.md
    runs: 3

The full workflow

pytestpytest-judgekeeper

Gate a report your pipeline already wrote, from your test suite. It never calls a judge.

pytest --judgekeeper-report reports/my-judge/report.json
def test_judge_still_agrees_with_humans(judgekeeper_gate):
    judgekeeper_gate("reports/my-judge/report.json")
A coding agentskill file

Coding agents such as Claude Code can set judgekeeper up for you. Give them skills/judgekeeper/SKILL.md. It explains each step, and tells the agent never to invent human labels.

A real result

Claude Haiku 4.5 as a judge, on LLMBar

LLMBar is a public set of questions, each with two answers, where people already marked the better one. We ran claude-haiku-4-5-20251001 as the judge on 419 of them, 3 times, in both answer orders.

TPR
0.95
TNR
0.88
Kappa
0.83

Verdict: usable as a gate. But the average hides something. LLMBar groups its questions into slices: Natural questions are ordinary ones, and the Adversarial slices were built to be hard for judges. Kappa by slice:

Natural0.94
Adversarial/GPTInst0.92
Adversarial/Neighbor0.84
Adversarial/Manual0.64
Adversarial/GPTOut0.56

Strong on ordinary questions (kappa 0.94), much weaker on the hardest adversarial ones (0.56 and 0.64). A judge that looks great on average can be weak exactly where it matters. The report shows it; an average hides it.

Open the full report

Privacy

Your data and keys stay with you

An API key is the password your code uses to call a model provider. With judgekeeper, your key goes from your machine straight to your provider, and nowhere else.

Your computer or your CI judgekeeper your examples and labels your key, in an environment variable reports, written to your disk Your model provider Anthropic, OpenAI or your own server a judgekeeper server: none exists
  • No server. judgekeeper runs where you run it. Nothing is sent to its makers.
  • No key flag. No option takes a key. --api-key-env takes the name of a variable, never the key.
  • Keys come only from environment variables, settings your shell or CI passes to programs, such as ANTHROPIC_API_KEY.
  • A test checks it. judgekeeper's own test suite checks that no key appears in any file or message it writes.
  • Error messages are scrubbed. Anything that looks like a key is replaced with [REDACTED] before it is printed.

Step-by-step key setup for your laptop, GitHub Actions and cloud sessions

Get started

Try it in a minute

pip install judgekeeper
judgekeeper demo

Before the first release on PyPI, install from GitHub instead:

pip install "judgekeeper @ git+https://github.com/judgekeeper/judgekeeper"

uv can run judgekeeper without installing it:

uvx judgekeeper demo

Before the first release on PyPI:

uvx --from git+https://github.com/judgekeeper/judgekeeper judgekeeper demo
git clone https://github.com/judgekeeper/judgekeeper
cd judgekeeper
pip install -e .
judgekeeper demo

judgekeeper demo checks a recorded judge on 100 made-up examples. It needs no key and no internet, takes a few seconds, and writes a report you can open in your browser. Then try the tutorial.

Questions

Short, honest answers

How many human labels do I need?

About 100, roughly half pass and half fail. With fewer than 60 the report warns that the error bars are wide. It also warns when the split is more lopsided than 80/20.

Why TPR and TNR first, and not kappa?

They are easier to act on. TNR in particular catches a judge that passes everything, which raw agreement hides. Kappa is in the report too, next to them.

Does it cost anything?

judgekeeper is free. The only cost is your judge's own API calls, which you pay to your model provider. The demo, check and the framework imports make no calls at all.

Does it replace my eval framework?

No. It works alongside promptfoo, DeepEval, Inspect AI, MLflow, Langfuse or your own code, and reads the files they already save.

Does my data leave my machine?

No. judgekeeper has no server. The only network calls are the ones your judge makes to your model provider, and reads from a platform you import from.

What if the judge model is updated?

Every verdict records the judge's identity: provider, model, prompt and settings. The gate notices a different judge and says JUDGE_CHANGED instead of comparing scores that are not comparable. A weekly scheduled check also catches silent updates.