This page is generated from docs/reference.md
by scripts/render_reference.py. New to judgekeeper? Start with the
home page or the tutorial.
judgekeeper reference
Every command, file format, flag, exit code and config key. The README has the short version.
- Anchor sets, judging and the report
- check: a table in, a report out
- How verdicts are read
- Bring your own judge
- Labels from a spreadsheet
- label: a local labeling page
- import: promptfoo, DeepEval and Inspect AI results
- import mlflow and import langfuse
- ScoreRecords: import records and export records
- demo
- Unknown judge fields
- Your API keys
- Gate CI on the judge
- pytest plugin
- Migrate to a new judge
- Attribute a score change
- Exit codes
- GitHub Action
Anchor sets, judging and the report
judgekeeper freeze anchors.jsonl # writes anchors.manifest.json (counts + sha256)
judgekeeper judge anchors.jsonl --runner anthropic --model claude-haiku-4-5-20251001 \
--prompt prompts/pairwise.md --runs 3 --out runs/my-judge/
judgekeeper validate anchors.jsonl runs/my-judge/ --out reports/my-judge/
An anchor set is JSONL with id, input and human_label on every line. Pairwise items add output_a and output_b with labels A/B; single-output items add output with labels pass/fail. slice and notes are optional. After freeze, every command checks the anchor set against its manifest and exits with code 3 if it changed.
judge writes one run-NN.jsonl per run. Line 1 is a header with the judge fingerprint and a source object ({kind: judgekeeper | table | callable | exec | promptfoo | deepeval | inspect | records | mlflow | langfuse, file, metric}, plus version, notes and warnings for imports; imported and custom judges also record the --pass-if rule and label map used). Every judgment line carries the judge fingerprint (provider, model, served snapshot, endpoint, prompt hash, rubric version, temperature, timestamp). Pairwise items are judged in both AB and BA order.
validate writes report.json and a self-contained report.html. The headline is TPR, TNR (each with a 95% Wilson interval) and Cohen's kappa, mean over runs. Kappa is chance-corrected agreement with the human labels; TPR and TNR are the numbers to act on. The interval's n is the number of human-positive (or negative) items, not items × runs, because runs re-judge the same items. The report also has: the fingerprint and source, per-run numbers, a confusion matrix on majority verdicts, the noise floor across runs (per-item flip rate, run-vs-run kappa, majority verdict; with one run it reads "unknown: one run supplied", never zero), AB/BA position bias, a per-slice table, and every disagreement with the judge's rationale.
The verdict:
| condition | says |
|---|---|
| TPR or TNR below 0.80, or kappa below 0.6 | "not trustworthy as a gate" |
| TPR and TNR between 0.80 and 0.90 (and kappa at least 0.6) | "usable with care" |
| otherwise | "usable as a gate" |
Flags: position bias above 0.10, more than 10% of items flipping between runs ("noisy, use majority of runs"), one run only (noise floor unknown), judge errors above 2% of judgments, an incomplete judge fingerprint ("judge identity incomplete: <fields>"), fewer than 60 labeled items ("error bars are wide; aim for about 100") and a class split more lopsided than 80/20. report.json keys read by gate are stable; session 4 added source, normaliser, errors, label_quality, notes, fingerprint_unknown, verdict.level, headline.tpr_ci/tnr_ci and noise_floor.status.
check: a table in, a report out
judgekeeper check results.csv --judge verdict --human label [--id id] [--run run] \
[--input input --output output --reason reason] [--pass-if RULE] [--label-map MAP] --out reports/x/
- CSV, TSV or JSONL, one row per judgment. Column flags default to columns of those names when present (
--judge verdict,--human label,--id id,--run run,--input input,--output output,--reason reason). Columnsoutput_aandoutput_bmake the items pairwise (verdictsA/B). - Several rows for one id with different
--runvalues are repeat runs. With no run column, or one run, the noise floor is "unknown: one run supplied" andgateon the report returnsFLAKYfor that reason. An item missing from a run is a judge error in that run. - Without an id column, an item's id is the sha256 of its canonical input and output (both outputs for pairwise items), first 16 hex characters, and the report says the ids were derived. Duplicate ids within a run are a usage error that lists them.
- A row with an empty human label is skipped and counted (
source.n_unlabeled). Different human labels for one id are a usage error. --out(defaultjudgekeeper-report) getsanchors.jsonlwith its manifest,runs/run-NN.jsonlandreport.json/report.html, built by the same pipeline asvalidate. Imported data rarely says which model judged: the fingerprint is recorded with unknown fields (see Unknown judge fields).
In Python (pandas is optional; anything with to_dict(orient="records") works):
import judgekeeper
report = judgekeeper.check_table(df_or_rows_or_path, judge="verdict", human="label",
pass_if=None, label_map=None, fingerprint={"model": "my-judge"},
out="reports/x") # out=None writes to a temporary directory
How verdicts are read
One normaliser serves check, check_judge, --callable, --exec and import-labels. It turns a raw judge output into pass/fail (or A/B) or error, plus a rationale:
- a bool:
Trueis pass; - a string, case-insensitively, through the label map. Defaults:
pass/fail,true/false,yes/no,correct/incorrect,1/0, and a leadingPASSorFAILtoken (PASS: looks right). Pairwise:A/B.--label-map "good=pass,bad=fail"adds entries; - a number, through a required
--pass-ifrule:score>=0.5,>,<=,<,==. The name is free text, except that for a dict it names the key to read (relevance>=3). Integers1/0follow the label map when no rule is given; - a
(verdict, reason)tuple or two-element list; - a dict with the verdict under
verdict,pass,passed,labelorscore(first found) and the reason underreason,rationale,explanationorcomment.
A value that looks like a verdict but is not in the map (good), or a number with no rule, is a usage error that lists every such value seen. judgekeeper never guesses. A judge that raised, returned nothing (None, an empty string, NaN) or returned something unreadable (a dict with no verdict key, a list of three) is recorded as error, never as fail: errors are excluded from every agreement metric, counted in the report (errors) and flagged above 2% of judgments. Human labels use the same label map but never --pass-if.
Bring your own judge
Python:
judgekeeper.check_judge(judge, anchors, runs=3, pass_if=None, label_map=None, fingerprint=None,
out=None, yes=False) # returns the report dict
judge(item: dict) returns a bool, str, float, tuple or dict; it may be async def (this also works inside a running event loop, as in a notebook). item is the anchor item without human_label and notes. Pairwise items are judged twice, with output_a and output_b swapped the second time; the second verdict is mapped back to the original labels. anchors is a frozen anchor JSONL path, or a list of items (written under out and frozen). fingerprint takes what you know: provider, model, snapshot, endpoint, prompt (text, hashed) or prompt_hash, rubric_version, temperature; everything else is recorded as unknown. Above 1,000 judge calls it needs yes=True.
Command line:
judgekeeper judge anchors.jsonl --callable mypkg.judges:my_judge --runs 3 --out runs/x/
judgekeeper judge anchors.jsonl --exec "node judge.js" --runs 3 --out runs/x/
Both print the number of judge calls (items × runs, × 2 for pairwise) before starting and need --yes above 1,000. --pass-if and --label-map work as above; --model, --temperature and --prompt (any file; hashed, with rubric_version taken from its frontmatter if it has one) fill in the fingerprint. --callable imports from the working directory and calls the function one item at a time. --exec runs the command once per item, --workers at a time (default 4), with --timeout seconds each (default 300).
The --exec contract: one item as JSON on stdin per invocation. Stdout is either a bare verdict (pass, FAIL: wrong total, 0.83) or a JSON object the normaliser reads ({"verdict": "pass", "reason": "..."}). A non-zero exit, a timeout or empty stdout is an error judgment; the last lines of stderr go into its error text (scrubbed of keys).
Node (judge.js):
// judge.js: run with --exec "node judge.js"
let data = "";
process.stdin.on("data", (chunk) => (data += chunk));
process.stdin.on("end", async () => {
const item = JSON.parse(data); // { id, input, output } or { id, input, output_a, output_b }
// Call your model here. This placeholder passes any non-empty output.
const pass = item.output.trim().length > 0;
const reason = pass ? "output is not empty" : "empty output";
console.log(JSON.stringify({ verdict: pass ? "pass" : "fail", reason }));
});
Python (judge.py):
# judge.py: run with --exec "python judge.py"
import json
import sys
item = json.load(sys.stdin) # {"id", "input", "output"} or {"id", "input", "output_a", "output_b"}
# Call your model here. This placeholder passes any non-empty output.
ok = bool(item["output"].strip())
reason = "output is not empty" if ok else "empty output"
print(json.dumps({"verdict": "pass" if ok else "fail", "reason": reason}))
sys.exit(0) # a non-zero exit is recorded as an error, not a fail
Labels from a spreadsheet
judgekeeper template items.jsonl -o labels.csv # or items.csv
judgekeeper import-labels labels.csv -o anchors.jsonl [--label-map "good=pass,bad=fail"]
template writes id,input,output,human_label,notes (pairwise: output_a,output_b) with human_label and notes empty, ready for Excel or Google Sheets. Ids come from an id column or are derived as in check. import-labels reads the sheet back, checks every label through the normaliser (an unmapped label is a usage error listing them), skips and lists unlabeled rows, keeps non-empty notes and slice, writes the anchor file and freezes it. It prints the label-quality warnings that also appear in every report.
label: a local labeling page
judgekeeper label items.jsonl [--out labels.csv] [--port 8765] [--no-browser] # or items.csv
Opens a page in your browser that shows one item at a time: the input and the output (pairwise: A and B side by side). Keys: 1 pass (pairwise: A, also a), 2 fail (pairwise: B, also b), d defer, u undo, n note, arrow keys to move; a button mirrors every key. A judge verdict in the items (a judge_verdict, verdict or judge column, with judge_reason, reason or rationale) sits behind a "Show judge" button, hidden by default so it does not anchor the labeler.
Every change is written to --out at once (to a temporary file, then renamed over it) as id,input,output,human_label,notes (pairwise output_a,output_b), the shape template writes, so import-labels and import --labels read it unchanged. Single items get pass/fail, pairwise items A/B; deferred items stay unlabeled (deferral is not saved to the file). Reopening with the same --out resumes where it left off; an --out with ids that are not in the items is refused. items may itself be a sheet from template, partly filled in. When every item is labeled or deferred, the page shows the counts, the split, the label-quality warnings (fewer than 60 labels, worse than 80/20) and the import-labels command to run next.
The server uses only the standard library and binds 127.0.0.1, never 0.0.0.0. The URL carries a random token that the page and every request must present; a request whose Host header is not 127.0.0.1:<port> is refused (DNS rebinding); the page loads nothing from the network (a Content-Security-Policy enforces it) and shows every string as text. Ctrl-C stops it, and it stops by itself after 2 hours without a request.
import: promptfoo, DeepEval and Inspect AI results
judgekeeper import <tool> <path>... [--metric NAME] [--labels labels.csv] [--pass-if RULE] \
[--label-map MAP] [--runs-by-order] [--id-var NAME] [--map MAP] --out reports/x/
<tool> is promptfoo, deepeval, inspect or records. A path is a file, a directory or a quoted glob; files are read in the order given (a directory or glob in name order). The output is what check writes: anchors.jsonl with its manifest, runs/run-NN.jsonl and report.json / report.html, with source.kind set to the tool and source.version to the tool's own format version when the file states one (promptfoo results.version, Inspect version). Exit 0 on success, 2 on a usage error. Per-tool pages: promptfoo, DeepEval, Inspect AI.
--metric NAME: the judge to validate (promptfoo assertionmetricor type, DeepEval metricname, Inspect scorer name). Optional when the files hold one; with several, leaving it out is a usage error that lists the names.--labels TABLE: human labels, CSV, TSV or JSONL with anidcolumn and ahuman_label(orlabel) column, read through the same normaliser ascheck. It wins over human labels in the files (promptfoo web-UI ratings, Inspect score edits) and the report notes how many it replaced and how many differed. Rows with an empty label are skipped.--pass-if,--label-map: as incheck.--pass-ifis applied to the record'sscore(DeepEvalscore, promptfoo componentscore, numeric Inspect values); without it the verdict is the tool's own pass flag or label. Human labels never use--pass-if.--runs-by-order: number each item's verdicts in a file 1, 2, 3 in order of appearance, in place of any run index the file carries, instead of stopping when one run holds several verdicts for one item. Works with every tool. The promptfoo reader always numbers repeats this way, because promptfoo strips the repeat index.--id-var NAME(promptfoo only): the test var holding the item id.--map MAP(records only): see below.
How records become a report: human records become anchor labels, judge records become judgments and code records (promptfoo's contains, javascript, ...) are ignored with a note. Several files, or several run indices in one file (Inspect epochs, promptfoo repeats), become separate runs; an item a run does not judge is an error judgment in that run. With one run the noise floor is "unknown: one run supplied". An item with a verdict and no human label, or a label and no verdict, is dropped, counted in source.n_judged_unlabeled / source.n_labeled_unjudged and noted in the report. Ids that had to be derived (a hash of input and output, as in check) are noted too.
Fingerprint: each judgment keeps the judge identity its record carries (model, prompt hash, temperature, timestamp). The run header records the fields every judgment agrees on; a field they disagree on is unknown in the header, with a note. A raw prompt is hashed into prompt_hash and never written. Tool warnings (promptfoo's unrecorded default grader, DeepEval's positional names) are report flags.
In Python:
import judgekeeper
report = judgekeeper.import_results("deepeval", ["deepeval-results/"], metric="Correctness [GEval]",
labels="labels.csv", pass_if=None, label_map=None,
runs_by_order=False, out="reports/x")
# also id_var="qid" (promptfoo) and column_map="target_id=trace_id,..." (records)
import mlflow and import langfuse
judgekeeper import mlflow --experiment NAME_OR_ID [--run-id ID]... [--metric NAME] \
[--tracking-uri URI] [--id-from KEY] [--temperature T] [--labels labels.csv] \
[--pass-if RULE] [--label-map MAP] [--anchors-out anchors.jsonl] --out reports/x/
judgekeeper import langfuse --judge-score NAME --human-score NAME \
(--from DATE [--to DATE] | --to DATE | --max-items N) [--rate PER_MINUTE] \
[--labels labels.csv] [--pass-if RULE] [--label-map MAP] [--anchors-out anchors.jsonl] --out reports/x/
These read a platform instead of files: they take no paths, and their flags are usage errors with any other tool. The output, --labels, --pass-if, --label-map and --runs-by-order are as for the file readers above. Per-platform pages: MLflow, Langfuse.
MLflow (needs pip install "judgekeeper[mlflow]"; without it, a usage error saying so):
--experiment NAME_OR_ID: required.--metric NAME: the assessment name. Optional with one judge assessment name; with several, leaving it out is a usage error that lists them.--run-id ID(repeatable): only those runs' judge assessments; an id not in the experiment is a usage error. Default: every run, one judgekeeper run each, in start order (source.fileis the MLflow run id).--tracking-uri URI: defaultMLFLOW_TRACKING_URI, then MLflow's own default. Databricks readsDATABRICKS_HOSTandDATABRICKS_TOKENthrough MLflow.--id-from KEY: item ids from this trace tag or request input key; a trace without it is a usage error. Default: a hash of the trace's request input, noted in the report.--temperature T: recorded in the fingerprint; MLflow does not store it.LLM_JUDGEassessments are judgments,HUMANassessments labels,CODEignored. A human assessment that overrides a judge assessment is the label; the overridden value stays the judge's verdict.feedback.erroris anerrorverdict. Fingerprint:source.source_idas model (provider before:/),scorerName@scorerVersionmetadata as rubric version, prompt unknown.
Langfuse (standard library only):
--judge-score NAME,--human-score NAME: required; the same name splits by source (ANNOTATIONis human).--from DATE,--to DATE: a date (2026-09-01) or ISO time, UTC when no offset is given;--frominclusive,--toexclusive.--max-items N: stop after N scores. One of the three is required.--rate PER_MINUTE: request cap, default 30. HTTP 429 waits forRetry-Afterand retries up to 5 times, then exits 1.--metricdefaults to--judge-score. Verdicts bydataType:BOOLEAN,NUMERIC(with--pass-if),CATEGORICAL;TEXTandCORRECTIONare skipped with a note. Item id issubject.id; rationale iscomment. One run.- Fingerprint: one
GET /api/public/v2/evaluatorscall, matched by score name, formodelConfig.provider,modelConfig.model,versionand a hash ofprompt; unknown with a note when no evaluator matches ormodelConfigis null. Temperature is always unknown. - A failed request (HTTP error, unreachable host, refused cross-host redirect) exits 1 with the status and path only.
--anchors-out PATH (both): also write every item with a human label as an anchor set (id, input, output, human_label; input and output only when the platform returned them, else ""; no rationales, no annotator ids) and freeze it, ready for judgekeeper judge. Usage error with the file readers.
In Python:
report = judgekeeper.import_results("mlflow", metric="correctness", out="reports/x",
source={"experiment": "my-app-eval", "run_ids": None,
"tracking_uri": None, "id_from": None,
"temperature": None},
anchors_out="anchors.jsonl")
report = judgekeeper.import_results("langfuse", pass_if="score>=0.5", out="reports/y",
source={"judge_score": "helpfulness",
"human_score": "helpfulness_human",
"from_": "2026-09-01", "to": None,
"max_items": None, "rate": 30})
ScoreRecords: import records and export records
A ScoreRecord is one verdict. Field names follow OpenInference annotations, so other tools' exports map onto it by renaming:
| field | meaning |
|---|---|
target_id | the item. Missing: derived from input and output as in check, and the report says so |
name | the metric or scorer (default judge); choose one with --metric |
annotator_kind | LLM (a judge verdict), HUMAN (a label) or CODE (ignored). Case-insensitive; LLM_JUDGE reads as LLM. Default LLM |
label | the verdict or label: pass/fail, true/false, a bool, or anything --label-map maps |
score | a number; read as the verdict with --pass-if, or when label is empty |
explanation | the judge's rationale |
run | repeat index, or empty. Without one, each file is one run (see --runs-by-order) |
input, output | the item's text (or any JSON) |
evaluator | {provider, model, prompt, temperature, version}; version is the rubric version, prompt is hashed. Also prompt_hash, snapshot and endpoint, which export records writes |
created_at | when the verdict was made |
HUMAN records label the metric they are named after, or any metric when their name is not a judge metric in the file (promptfoo's ratings are named human).
judgekeeper import records records.jsonl --out reports/x/
judgekeeper import records export.csv --map "target_id=trace_id,name=metric,label=value,explanation=comment,annotator_kind=source" --out reports/x/
judgekeeper export records reports/x/ -o records.jsonl
import records reads JSONL, CSV or TSV. --map takes field=column pairs; evaluator fields are evaluator.model=judge_model and so on (a CSV can also have evaluator.model columns, or an evaluator column holding JSON). A mapped column that does not exist is a usage error listing the columns.
export records <dir> -o records.jsonl [--anchors anchors.jsonl] writes judgekeeper's runs and anchors as ScoreRecords: one HUMAN record per anchor item and one LLM record per judgment, with the judgment's fingerprint as evaluator (the prompt as prompt_hash) and error judgments as an empty label. <dir> is a directory check or import wrote (or its report.json), or a runs directory with --anchors. import records on the result gives the same report. Single-output anchor sets only; slice and notes are not carried.
demo
judgekeeper demo [--out judgekeeper-demo] validates a recorded judge on 100 synthetic single-output items (arithmetic questions, 3 recorded runs) shipped inside the package, writes report.html and report.json, and prints the path and the TPR, TNR and kappa. No key, no network. The data is made by scripts/make_demo_data.py; see src/judgekeeper/demo_data/NOTICE.
Unknown judge fields
Every fingerprint field except created_at may be unknown: provider, model, snapshot, prompt hash, rubric version and temperature are null, and an unknown endpoint is the string "unknown" (because endpoint: null already means the provider's default endpoint). Reports show such fields as "unknown" and flag "judge identity incomplete: <fields>". Files from earlier versions still load: a missing field reads as unknown, except a missing endpoint, which reads as the provider default.
gate compares the baseline and the report field by field. A field known on both sides that differs is JUDGE_CHANGED, as always. A field unknown on either side adds a warning and does not block, unless --require-fingerprint is passed, in which case the status is JUDGE_CHANGED with the reason "cannot prove same judge". The default is to warn because imported data rarely carries temperature or snapshot, and blocking on that would make the gate unusable for it; pass --require-fingerprint once your judge records its full identity.
Your API keys
- judgekeeper has no server. Your key stays in your environment, and requests go from your machine (or your CI runner) straight to the provider or the endpoint you name.
- Nothing judgekeeper writes contains a key: judgment files, reports, gate, migration and attribution files are all scrubbed. Before any text reaches disk or your terminal, the value of
ANTHROPIC_API_KEY,OPENAI_API_KEY,LANGFUSE_PUBLIC_KEY,LANGFUSE_SECRET_KEY,MLFLOW_TRACKING_PASSWORD, the variable named with--api-key-envand any variable ending in_API_KEY,_TOKENor_SECRET(DATABRICKS_TOKEN,MLFLOW_TRACKING_TOKEN) is replaced with[REDACTED], as is anything shaped like a provider key (sk-ant-…,sk-…,Bearer …,Basic …) and theuser:password@part of a URL. - Platform readers take credentials from the environment only:
LANGFUSE_PUBLIC_KEYandLANGFUSE_SECRET_KEY(host fromLANGFUSE_BASE_URLorLANGFUSE_HOST, defaulthttps://cloud.langfuse.com) for Langfuse; whatever MLflow reads (MLFLOW_TRACKING_URI,MLFLOW_TRACKING_TOKEN,DATABRICKS_HOST,DATABRICKS_TOKEN, ...) for MLflow. The Langfuse reader only issues GET requests, sends the keys only to the configured host, refuses redirects to another host and never prints the Authorization header, a response body or the keys. - Error text is scrubbed too. A failed provider call prints one line and exits 1;
--debugadds the traceback, still scrubbed. A failed call is never scored. - There is no flag that takes a key, and there never will be.
--api-key-envtakes the name of a variable. - In CI, keep keys in repository secrets and pass them to the step as
env, never as inputs.
A custom endpoint: any OpenAI-compatible server (Azure OpenAI's v1 endpoint, OpenRouter, Together, a LiteLLM proxy, Ollama, vLLM) with --runner openai, or a gateway in front of Anthropic with --runner anthropic. OPENAI_BASE_URL and ANTHROPIC_BASE_URL also work; the flag wins when both are set. With --base-url and no key set, judgekeeper sends a placeholder key, since local servers need none. A URL with user:password@ in it is rejected.
judgekeeper judge anchors.jsonl --runner openai --model "$JUDGE_MODEL" \
--base-url http://localhost:11434/v1 --prompt prompts/pairwise.md --runs 3 --out runs/local/
A custom key variable:
export OPENROUTER_API_KEY=... # in your shell or CI secrets, never in a file
judgekeeper judge anchors.jsonl --runner openai --model "$JUDGE_MODEL" \
--base-url https://openrouter.ai/api/v1 --api-key-env OPENROUTER_API_KEY \
--prompt prompts/pairwise.md --runs 3 --out runs/openrouter/
The endpoint's host goes into the fingerprint as endpoint (null for the provider default): the same model name behind a different endpoint is a different judge, and gate treats a changed endpoint as JUDGE_CHANGED. AWS Bedrock and Google Vertex have no native client; put a gateway such as LiteLLM in front of them and use --base-url.
Gate CI on the judge
judgekeeper baseline set reports/my-judge/report.json # copies to .judgekeeper/baseline.json; commit it
judgekeeper baseline show # fingerprint, anchors hash, kappa / TPR / TNR
judgekeeper gate reports/my-judge/report.json # writes gate.json and gate.md next to the report
gate compares a report with fixed thresholds and, if there is one, the baseline (--baseline, default .judgekeeper/baseline.json when it exists). It decides one status, checking in this order:
ANCHORS_CHANGED: the anchor set hash differs from the baseline's.JUDGE_CHANGED: provider, model, snapshot, endpoint, prompt hash, rubric version or temperature differ from the baseline. Scores are not compared across judges.--allow-judge-changeturns this into a warning and gates on absolute thresholds only. A field unknown on either side is a warning, orJUDGE_CHANGED("cannot prove same judge") with--require-fingerprint; see Unknown judge fields.FLAKY: fewer than 3 runs, so the noise floor is unknown (with one run: "noise floor unknown: one run supplied").- Absolute thresholds: kappa mean >= 0.6, TPR mean >= 0.8, TNR mean >= 0.8, AB/BA disagreement <= 0.10.
- Against the baseline: kappa, TPR and TNR may not drop by more than the noise band, which is the larger of the baseline's and this report's run-to-run spread, and at least 0.02.
- A check that fails on the mean but passes on the best run, or more than 10% of items flipping between runs, is
FLAKY("use the majority of more runs") instead ofFAIL. - Otherwise
PASS.
JUDGE_CHANGED points at migrate, below: compare the old and new judge, then --rebase to make the new one the baseline.
--flaky-as pass or --flaky-as fail maps FLAKY to exit 0 or 1; the status in gate.json and gate.md stays FLAKY. Without the flag FLAKY exits 4. gate.md is a short summary for a PR comment or $GITHUB_STEP_SUMMARY: the status, why, a metric / baseline / now / delta / noise band table and the judge fingerprint.
Thresholds live in an optional judgekeeper.toml (--config, default ./judgekeeper.toml when it exists), one table per command: [gate] here, [migrate] and [attribute] below. Unknown tables and keys are a usage error.
[gate]
kappa_min = 0.6
tpr_min = 0.8
tnr_min = 0.8
ab_ba_disagreement_max = 0.10
min_band = 0.02 # smallest noise band, for judges that never vary between runs
flip_rate_max = 0.10 # share of items that may flip between runs before a failure is FLAKY
pytest plugin
Installed with judgekeeper (pytest11 entry point judgekeeper.pytest_plugin); pip install pytest-judgekeeper installs the same thing under the name the pytest plugin list uses. It runs the gate logic on a report.json your pipeline already wrote. It never calls a judge or an API and writes nothing.
def test_judge_still_agrees_with_humans(judgekeeper_gate):
judgekeeper_gate("reports/my-judge/report.json") # baseline=None, config=None, allow=("PASS",), flaky_as=None
@pytest.mark.judgekeeper(report="reports/my-judge/report.json", flaky_as="pass")
def test_with_the_marker():
... # runs only if the gate allows the report
pytest --judgekeeper-report reports/my-judge/report.json [--judgekeeper-baseline B] \
[--judgekeeper-config judgekeeper.toml] [--judgekeeper-flaky-as pass|fail]
- A status not in
allowfails the test with thegate.mdsummary as the message.flaky_as="pass"addsFLAKYtoallow;"fail"removes it. A missing or unreadable report fails the test with the reason. - The marker checks the report before the test body runs; a marked test whose gate fails never runs its body.
--judgekeeper-reportadds one test,judgekeeper-gate, so a CI job can gate without writing a test.- Fixture and marker paths are relative to the pytest rootdir; command-line paths to the directory pytest was started in. With no baseline or config,
.judgekeeper/baseline.jsonandjudgekeeper.tomlunder the rootdir are used when they exist, asjudgekeeper gatedoes.
Migrate to a new judge
When a judge model is deprecated, or you want a cheaper or better one, judge the same frozen anchor set with both and compare:
judgekeeper judge anchors.jsonl --runner anthropic --model "$OLD_MODEL" --prompt prompts/pairwise.md --runs 3 --out runs/old/
judgekeeper judge anchors.jsonl --runner anthropic --model "$NEW_MODEL" --prompt prompts/pairwise.md --runs 3 --out runs/new/
judgekeeper migrate anchors.jsonl runs/old/ runs/new/ --out reports/migration/ [--rebase] [--fail-on worse|different]
Both run directories must have been judged against the same frozen anchor set (exit 3 otherwise). migrate writes migration.json and a self-contained migration.html with:
- Who changed: both fingerprints side by side, changed fields highlighted, served snapshots.
- Each judge against humans: kappa, TPR and TNR, mean and range over runs.
- Old judge against new judge: kappa between the two judges' majority verdicts, the share of items whose verdict changed, and for each change whether both judges were stable on it (unanimous across their own runs) or either was flipping anyway ("within noise"). Changed-and-stable items are fixed (the new judge now agrees with the human label) or broken, with the net.
- Per slice: old kappa, new kappa, delta, changed-and-stable count.
- Changed items: id, slice, human label, old and new verdict, fixed or broken, both rationales.
- Pass-rate bridge: a = P(new says pass | old said pass) and b = P(new says pass | old said fail) on the anchor set, with Wilson 95% intervals, overall and per slice, so that
new_rate ≈ old_rate × a + (1 − old_rate) × b. It assumes your production traffic resembles the anchor set. Pairwise sets use "prefers A" for "pass". - Status, with a sentence on what to do. The noise band is the larger of each judge's run-to-run kappa spread and
min_band.
| status | meaning | what to do |
|---|---|---|
EQUIVALENT | kappa vs humans moved within the noise band, and at most 2% of items changed while both judges were stable | keep comparing scores across the switch; rebase |
BETTER | kappa vs humans rose by more than the noise band | switch and rebase; scores move because the judge improved |
WORSE | kappa vs humans fell by more than the noise band | do not migrate yet |
DIFFERENT | kappa within the band, but more than 2% of items changed while both judges were stable | as good, on different items: old and new scores are not comparable item by item; rebase if you switch |
migrate exits 0 when the analysis completes, whatever the status. --fail-on worse exits 1 on WORSE; --fail-on different exits 1 on WORSE or DIFFERENT. --rebase copies the new judge's report to the baseline (--baseline, default .judgekeeper/baseline.json) and keeps migration.json as .judgekeeper/migrations/<UTC date>-<old model>-to-<new model>.json, the audit trail. Commit both; after a rebase gate no longer reports JUDGE_CHANGED.
[migrate]
min_band = 0.02 # smallest kappa noise band
max_changed_share = 0.02 # changed-and-stable share above which equal kappa is DIFFERENT
scripts/migration_demo.sh runs this on LLMBar: RUNNER, OLD_MODEL and NEW_MODEL come from the environment, and it has no default model ids. Take them from the provider's current models and deprecations pages.
> TODO: the live migration demo is pending. No API key was available in the build environment, so no migration report is committed yet and no migration numbers appear here. When one runs, its report goes under docs/examples/.
Attribute a score change
Your app's eval score moved. Did your system change, or did the judge? The anchor set is frozen, so the outputs being judged are identical every time: any movement in verdicts on it comes from the judge. Re-judge the anchor set, validate, and compare with the baseline:
judgekeeper attribute reports/now/report.json [--baseline .judgekeeper/baseline.json] \
[--app-score-before X --app-score-after Y]
It compares per-item majority verdicts between the baseline and the current report. An item counts as moved only if it was stable (unanimous across runs) in both. Judge drift is declared when more than 2% of items moved, or kappa vs humans moved outside the noise band. This works when the declared fingerprint is identical, which is what a silent provider-side update looks like, and it reports whether the served snapshot changed.
STABLE: the judge did not move on the anchor set.JUDGE_DRIFT: it did. Re-validate the judge; to keep the new behaviour,migrateand rebase.SYSTEM_CHANGE: the judge is stable on the anchor set and the app scores you supplied (pass rates between 0 and 1) differ by more than the judge's own run-to-run pass-rate spread on the anchor set (at leastmin_band). The score change comes from your system.
Both reports need per-item verdicts (items, written by validate from this version on); an older baseline is a usage error that tells you to regenerate it. attribute writes attribution.json and attribution.md (for $GITHUB_STEP_SUMMARY) next to the report, or under --out. The example workflow in docs/examples/workflows/judge-gate.yml runs it weekly after the gate job re-judges the anchor set.
[attribute]
min_band = 0.02 # smallest noise band, for kappa and for app scores
max_moved_share = 0.02 # share of stable items that may move before it is JUDGE_DRIFT
Exit codes
| exit code | meaning | commands |
|---|---|---|
| 0 | success; PASS; STABLE; migrate finished | all |
| 1 | FAIL; a judge call failed; migrate --fail-on matched; a Langfuse request failed | gate, judge, migrate, import langfuse |
| 2 | usage error (bad arguments, report, baseline or config; unmapped verdicts or labels; duplicate ids; several metrics without --metric; more than 1,000 judge calls without --yes; a missing extra or platform key; a Langfuse import with no time window or --max-items) | all |
| 3 | anchor set changed: hash mismatch, ANCHORS_CHANGED, runs or reports from a different anchor set | all that read anchors or reports |
| 4 | FLAKY (--flaky-as maps it to 0 or 1) | gate |
| 5 | JUDGE_CHANGED | gate |
| 6 | JUDGE_DRIFT | attribute |
| 7 | SYSTEM_CHANGE | attribute |
GitHub Action
action.yml at the repo root runs judge, validate and gate, appends gate.md to the job summary, uploads report.html, report.json and gate.json as an artifact, and exits with the gate's exit code. API keys come from the caller's env, never from inputs; api-key-env names the variable when it is not the provider's standard one. A copy-paste workflow that runs weekly and on PRs touching the judge prompt is in docs/examples/workflows/judge-gate.yml:
name: judge gate
on:
schedule:
- cron: "17 6 * * 1" # weekly: catches provider-side judge drift between PRs
pull_request:
paths: ["prompts/judge.md", "evals/anchors.jsonl", ".judgekeeper/baseline.json", "judgekeeper.toml"]
jobs:
gate:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: judgekeeper/judgekeeper@main # pin to a release tag or commit sha
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
with:
anchors: evals/anchors.jsonl
runner: anthropic
model: claude-haiku-4-5-20251001
prompt: prompts/judge.md
runs: 3
flaky-as: pass
| input | default | |
|---|---|---|
anchors | required | frozen anchor set JSONL, manifest next to it |
runner | required | anthropic, openai or replay |
model, prompt | judge model id and prompt (anthropic, openai) | |
fixture | recorded judgments JSONL for replay (no API key) | |
base-url | endpoint instead of the provider default (anthropic, openai) | |
api-key-env | name of the variable holding the key, never the key itself | |
runs | 3 | judge runs; the gate needs at least 3 |
baseline | .judgekeeper/baseline.json if present | baseline report |
config | judgekeeper.toml if present | gate thresholds |
flaky-as | pass | pass, fail, or empty to keep exit code 4 |
out-dir | judgekeeper-out | where runs, report and gate files go |
artifact-name | judgekeeper-gate | uploaded artifact name |
python-version | 3.12 | Python used to run judgekeeper |
Outputs: status and exit-code. This repo's own CI runs the action on replay fixtures on every push.