It disagrees with people
A judge can pass an answer a person would fail. A judge that passes everything makes every test look fine.
Many teams use one AI model to grade the answers of another. The grading model is called a judge. judgekeeper compares the judge with answers a person already graded, and tells you how often the judge is right.
pip install judgekeeperFree and open source (MIT license). It runs on your own computer. There is no account and no server.
The problem
Your AI product writes answers. To know whether they are good, you could read them all yourself, but that is slow. So teams ask another AI model to grade them: the judge. A judge is fast and cheap. It can also be wrong, in three ways.
A judge can pass an answer a person would fail. A judge that passes everything makes every test look fine.
Ask a judge the same question three times and you can get two different answers.
The company behind the judge model can update it. Then every score moves, although your product did not change.
How it works
The idea is simple. A person grades a small set of answers once. judgekeeper then measures the judge against those grades, and keeps measuring.
Pick about 100 answers from your product. A person marks each one pass or fail. That mark is called a human label. You can label in a spreadsheet, or in a small page judgekeeper opens on your own computer:
judgekeeper label items.jsonl --out labels.csvjudgekeeper saves the labeled examples as an anchor set: the fixed answer key every later check uses. It records a hash, a short fingerprint of the file such as sha256 2e75fe64157b…. If anyone edits a label later, the fingerprint no longer matches and judgekeeper stops.
judgekeeper import-labels labels.csv -o anchors.jsonlYour judge grades every example. Each grade is a verdict. Each full pass over the examples is a run. Running three times shows how often the judge changes its mind.
judgekeeper judge anchors.jsonl --callable my_judges:judge --runs 3 --out runs/mineHere my_judges:judge is your own Python function. judgekeeper can also call Anthropic or OpenAI models directly, or any program. See Works with what you already use.
judgekeeper lines up each verdict with the human label for the same answer. It counts where they agree and where they do not, run by run.
judgekeeper validate anchors.jsonl runs/mine --out reports/mineYou get one page. It says how often the judge is right, how noisy it is, where it is weak, and lists every disagreement with the judge's own reason. It ends with a verdict: usable as a gate, usable with care or not trustworthy as a gate. Open reports/mine/report.html in your browser, or see a real one.
Already have the judge's verdicts and human labels in one spreadsheet? Skip to one command. It does steps 2 to 5 for you:
judgekeeper check results.csv --judge verdict --human label --out reports/my-judgeWhat the numbers mean
Illustration: you choose the numbers
Move the sliders. The table counts the four ways a judge and a person can line up. The numbers below it are the ones judgekeeper reports, computed the same way.
| Judge passed | Judge failed | |
|---|---|---|
| People passed | 18agree | 2judge too harsh |
| People failed | 3judge too lenient | 17agree |
Of the answers people passed, the judge passed 18 of 20.
Of the answers people failed, the judge failed 17 of 20.
Agreement after taking away what lucky guessing would give. 0 is no better than chance, 1 is perfect.
The share of all answers where judge and people agree. It can look high for a useless judge.
Usable with care: TNR is between 0.80 and 0.90.
Likely ranges with this few answers (95 percent Wilson intervals): TPR 0.70 to 0.97, TNR 0.64 to 0.95.
The report's verdict uses these rules: not trustworthy as a gate if TPR or TNR is below 0.80 or kappa is below 0.6; usable with care if TPR or TNR is below 0.90; otherwise usable as a gate. A gate is an automatic check that can stop a release.
Two more things it measures
Illustration
judgekeeper runs the judge several times on the same answers and counts how often a verdict changes. The report calls this the noise floor. Here, 1 of 5 answers flipped. With only one run the noise is reported as unknown, never as zero.
Illustration
When a judge compares two answers, it can favour whichever it reads first. judgekeeper asks both ways, A then B and B then A. If the verdict changes with the order, that is position bias.
In the real result below, 3.6% of items changed verdict in at least one of three runs, and 9.9% of verdicts changed when the two answers swapped places.
After the first check
A judge that is right today can drift tomorrow. Three commands keep watch.
CI (continuous integration) is the set of automatic checks that run on every code change. judgekeeper gate reads a report and returns a status. Each status has an exit code, a number a program returns when it ends. Anything but 0 stops the build.
judgekeeper gate reports/my-judge/report.json| Status | Means | Exit |
|---|---|---|
| PASS | The judge clears every check. | 0 |
| FAIL | The judge fell below a threshold, or dropped by more than its own noise. | 1 |
| ANCHORS_CHANGED | Someone edited the frozen examples. | 3 |
| FLAKY | Too few runs, or the judge flips too often to tell. | 4 |
| JUDGE_CHANGED | A different model, prompt or setting: scores are not comparable. | 5 |
It never fails on a change smaller than the judge's own run-to-run noise.
Illustration
Moving to a new judge model? Run both on the same frozen examples. migrate compares them answer by answer.
judgekeeper migrate anchors.jsonl runs/old/ runs/new/ --out reports/migration/| Answer | People | Old judge | New judge |
|---|---|---|---|
| #1 | pass | pass | pass |
| #2 | fail | pass | fail (fixed) |
| #3 | fail | fail | fail |
| #4 | pass | pass | fail (broken) |
BETTER: the new judge agrees with people more, by more than the noise. Switch, and expect scores to move because the judge improved.
Your score dropped. judgekeeper grades the frozen examples again. Their answers never change, so if the judge's verdicts on them moved, the judge changed. If not, your app did.
judgekeeper attribute reports/now/report.json --app-score-before 0.82 --app-score-after 0.74SYSTEM_CHANGE The judge still grades the frozen examples the same way, so the drop comes from your app.
| Status | Means | Exit |
|---|---|---|
| STABLE | Nothing moved beyond the noise. | 0 |
| JUDGE_DRIFT | The judge's verdicts on the frozen examples moved. | 6 |
| SYSTEM_CHANGE | The judge held still; your app's score moved. | 7 |
Integrations
judgekeeper does not replace your tools. It reads what they already save. Open a tile for the exact command.
One row per verdict, with a column for the judge's verdict and one for the human label. Columns named id, run, input, output and reason are used when present.
judgekeeper check results.csv --judge verdict --human label --out reports/my-judge/Verdicts can be pass/fail, true/false, yes/no, correct/incorrect or 1/0. Scores need a rule such as --pass-if "score>=0.5". judgekeeper never guesses.
Write a function that takes one item (its id, input and output, never the human label) and returns a verdict:
# mypkg/judges.py
def my_judge(item):
return call_my_model(item) # "pass", "fail", True, False, a score, or {"verdict": ..., "reason": ...}judgekeeper judge anchors.jsonl --callable mypkg.judges:my_judge --runs 3 --out runs/mine/Any program works. judgekeeper sends one item as JSON on standard input; the program prints a verdict.
judgekeeper judge anchors.jsonl --exec "node judge.js" --runs 3 --out runs/mine/Run your eval three times, then point judgekeeper at the results file and your human labels.
promptfoo eval -o results.json --repeat 3
judgekeeper import promptfoo results.json --metric helpfulness --labels labels.csv --out reports/helpfulness/export DEEPEVAL_RESULTS_FOLDER=deepeval-results # then run your DeepEval tests 3 times
judgekeeper import deepeval deepeval-results/ --metric "Correctness [GEval]" --labels labels.csv --out reports/correctness/inspect eval task.py --epochs 3 --log-format json
judgekeeper import inspect logs/ --metric model_graded_qa --labels labels.csv --out reports/qa/pip install "judgekeeper[mlflow]"
judgekeeper import mlflow --experiment my-app-eval --metric correctness --out reports/correctness/Set LANGFUSE_PUBLIC_KEY and LANGFUSE_SECRET_KEY in your environment (see Setup and keys), then:
judgekeeper import langfuse --judge-score helpfulness --human-score helpfulness_human --from 2026-09-01 --pass-if "score>=0.5" --out reports/helpfulness/judgekeeper calls the model for you with your own key, read from the environment. The prompt file holds your grading instructions; an example is prompts/single.md.
judgekeeper judge anchors.jsonl --runner anthropic --model claude-haiku-4-5-20251001 --prompt prompts/single.md --runs 3 --out runs/haiku/judgekeeper judge anchors.jsonl --runner openai --model "$JUDGE_MODEL" --prompt prompts/single.md --runs 3 --out runs/openai/$JUDGE_MODEL is the model name you use. How to set the key safely
Azure OpenAI, OpenRouter, Together, a LiteLLM gateway (for AWS Bedrock or Google Vertex), Ollama or vLLM. For a local Ollama, no key is needed:
judgekeeper judge anchors.jsonl --runner openai --model "$JUDGE_MODEL" --base-url http://localhost:11434/v1 --prompt prompts/single.md --runs 3 --out runs/local/The action runs the judge, writes the report and gates on it. Keys come from repository secrets.
- uses: judgekeeper/judgekeeper@main # pin to a release tag
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
with:
anchors: evals/anchors.jsonl
runner: anthropic
model: claude-haiku-4-5-20251001
prompt: prompts/judge.md
runs: 3Gate a report your pipeline already wrote, from your test suite. It never calls a judge.
pytest --judgekeeper-report reports/my-judge/report.jsondef test_judge_still_agrees_with_humans(judgekeeper_gate):
judgekeeper_gate("reports/my-judge/report.json")Coding agents such as Claude Code can set judgekeeper up for you. Give them skills/judgekeeper/SKILL.md. It explains each step, and tells the agent never to invent human labels.
A real result
LLMBar is a public set of questions, each with two answers, where people already marked the better one. We ran claude-haiku-4-5-20251001 as the judge on 419 of them, 3 times, in both answer orders.
Verdict: usable as a gate. But the average hides something. LLMBar groups its questions into slices: Natural questions are ordinary ones, and the Adversarial slices were built to be hard for judges. Kappa by slice:
Strong on ordinary questions (kappa 0.94), much weaker on the hardest adversarial ones (0.56 and 0.64). A judge that looks great on average can be weak exactly where it matters. The report shows it; an average hides it.
Privacy
An API key is the password your code uses to call a model provider. With judgekeeper, your key goes from your machine straight to your provider, and nowhere else.
--api-key-env takes the name of a variable, never the key.ANTHROPIC_API_KEY.[REDACTED] before it is printed.Step-by-step key setup for your laptop, GitHub Actions and cloud sessions
Get started
pip install judgekeeper
judgekeeper demoBefore the first release on PyPI, install from GitHub instead:
pip install "judgekeeper @ git+https://github.com/judgekeeper/judgekeeper"uv can run judgekeeper without installing it:
uvx judgekeeper demoBefore the first release on PyPI:
uvx --from git+https://github.com/judgekeeper/judgekeeper judgekeeper demogit clone https://github.com/judgekeeper/judgekeeper
cd judgekeeper
pip install -e .
judgekeeper demojudgekeeper demo checks a recorded judge on 100 made-up examples. It needs no key and no internet, takes a few seconds, and writes a report you can open in your browser. Then try the tutorial.
Questions
About 100, roughly half pass and half fail. With fewer than 60 the report warns that the error bars are wide. It also warns when the split is more lopsided than 80/20.
They are easier to act on. TNR in particular catches a judge that passes everything, which raw agreement hides. Kappa is in the report too, next to them.
judgekeeper is free. The only cost is your judge's own API calls, which you pay to your model provider. The demo, check and the framework imports make no calls at all.
No. It works alongside promptfoo, DeepEval, Inspect AI, MLflow, Langfuse or your own code, and reads the files they already save.
No. judgekeeper has no server. The only network calls are the ones your judge makes to your model provider, and reads from a platform you import from.
Every verdict records the judge's identity: provider, model, prompt and settings. The gate notices a different judge and says JUDGE_CHANGED instead of comparing scores that are not comparable. A weekly scheduled check also catches silent updates.