Run a Benchmark
You shipped a skill-pack change. It might be a new rule in a skill,
a changed agent profile, or an updated tool allowlist. Now you need
to know whether agents write better code as a result. A single agent
run proves little, and one passing eval does not tell you much on
its own. gemba-benchmark runs each coding task
N times against a
versioned skill-set manifest. It grades each run
with tests the agent never sees, and it aggregates pass@k with the
unbiased estimator from OpenAI HumanEval.
Prerequisites
- Node.js 22+
ANTHROPIC_API_KEYset in the environment-
The
gemba-*command family. Install it globally withnpm install -g @forwardimpact/gemba. That package shipsgemba-benchmark,gemba-harness, andgemba-trace. In CI, run one command ephemerally withnpx gemba-benchmark ...
Author a Task Family
A task family is a directory of related coding tasks plus the skill-set under test:
my-coding-family/
.env # family env vars (committed defaults)
.env.local # family secrets (gitignored)
apm.yml # optional — skill-pack dependencies
apm.lock.yaml # skill-set manifest (hashed)
.claude/ # pre-staged skills + agents
skills/...
agents/judge.md
workdir/ # optional — shared base copied into EVERY task CWD
specs/ # optional — shared base copied into EVERY task CWD/specs
tasks/todo-api/
.env # task env vars — loaded + rendered
.env.local # task secrets — loaded + rendered (gitignored)
agent.task.md # what the agent should build (required)
judge.task.md # optional — judge prompt (see § judge.task.md)
supervisor.task.md # optional — supervisor context
hooks/ # harness-only — never copied to agent CWD
preflight.sh # optional — smoke probe
invariants.sh # optional — structural checks; fd 3 = $RESULTS_FD
tests/ # optional — hidden test suite, staged + run by the harness
specs/ # copied into the agent CWD
workdir/ # copied into the agent CWD
Task IDs are directory names under tasks/ (e.g.
todo-api). The directory splits into what the agent
sees (workdir/, specs/,
.claude/) and what the harness keeps hidden (hooks/
and tests/). The agent never receives the material that
grades it, so it cannot see the tests.
What the agent sees
agent.task.md
Plain markdown. The file holds the prompt the agent receives.
Build a TODO API matching the spec under `specs/`. Listen on the port
exposed via the environment variable `PORT`. Respond to `GET /todos`
with a JSON array of TODO objects.
workdir/
This directory holds the scaffolding the agent starts with: a
package.json, a README, sample data. The harness copies
everything here into the per-task CWD.
To share scaffolding across many tasks, put it in a
family-level workdir/ (or
specs/) at the family root. The harness copies that
shared base into every task's CWD first and then overlays the
per-task workdir//specs/ on top. A
per-task file wins over a family file with the same name. If the
directory is present, the harness copies it, which is the same
convention as hooks/. You then maintain one
app-under-test instead of one copy per task.
What the harness controls: hooks/
The hooks/ directory holds lifecycle scripts that the
harness runs at specific phases. The harness never copies either
script to the agent's working directory. Both scripts receive
these environment variables:
| Var | Value |
|---|---|
$AGENT_CWD |
The per-task agent CWD. |
$PORT |
A pre-allocated free TCP port. |
$TASK_ID |
The task name. |
$TASK_DIR |
The task directory on the host. |
$HOOKS_DIR |
The task's hooks/ dir on the host. Read
hidden fixtures/tests from here.
|
$FAMILY_DIR |
The family root on the host. |
invariants.sh also receives
$RESULTS_FD=3 (see below).
hooks/preflight.sh
Optional. The script runs before the agent starts. Exit
0 means that the scaffold is healthy and the harness
can hand off to the agent. A non-zero exit stops the run early and
produces a preflightError result record (cost zero, no
agent invoked). When the script is absent, the harness proceeds
without a pre-flight probe.
A preflight that starts a background service for the invariants probe to test against:
#!/bin/sh
node "$AGENT_CWD/app.js" >/dev/null 2>&1 &
# Wait until the service accepts connections. A fixed sleep races a slow start.
i=0
until curl -sf --max-time 1 "http://127.0.0.1:$PORT/" >/dev/null 2>&1; do
i=$((i + 1))
[ "$i" -ge 50 ] && exit 1
sleep 0.1
done
exit 0
The harness spawns the preflight in its own process group and tears down the whole group (SIGTERM, grace period, SIGKILL) after the invariants check completes, so background processes do not leak across runs.
hooks/invariants.sh
The script runs after the agent finishes. It receives the shared
hook env above, plus $RESULTS_FD=3, a file descriptor
for structured check rows.
The rows are the source of truth. Every row is a
check, and a row's role is in its own fields.
{"gate": true} marks a gate. If a gate fails,
the run fails and the score becomes zero. A plain row is a scored
check that adds to the task's score, and so is a row with
"weight": w > 0.
{"weight": 0} is an ungraded diagnostic. The
script's
exit code reports only the health of the script itself. A nonzero code means that the grader failed, never that a check
failed. A well-formed hook therefore ends with
exit 0 in every case.
Use the script for structural checks: presence, shape, and
anti-tamper. gemba-trace assert emits the rows:
#!/bin/sh
set -u
check() { gemba-trace assert "$@" >&"$RESULTS_FD" || true; }
# A gate: the artifact must exist at all.
check api-present --gate --exists "$AGENT_CWD/src/api.js"
# Scored content checks — each contributes weight 1 to the score.
check has-get-route --grep 'GET /todos' "$AGENT_CWD/src/api.js"
check documents-port --grep 'PORT' "$AGENT_CWD/README.md"
exit 0
A probe that assert cannot express echoes its row JSON
directly:
RESP="$(curl -sf --max-time 2 "http://127.0.0.1:$PORT/todos")"
if [ "$RESP" = '[]' ]; then
printf '%s\n' '{"test":"probe","pass":true,"gate":true}' >&"$RESULTS_FD"
else
printf '%s\n' '{"test":"probe","pass":false,"gate":true}' >&"$RESULTS_FD"
fi
exit 0
tests/: the hidden test suite
Behavioral checks test whether the code the agent wrote works. They
belong in a hidden test suite instead of in hand-written shell. A
task opts in with a tests/ directory beside
hooks/. There is no manifest, because the layout itself
is the contract:
-
tests/is an overlay mirror of the agent CWD. A file's path undertests/is the path where it stages (tests/test/filter.test.jsstages attest/filter.test.js). -
Every
*.test.jsfile is one check. The harness runs it withbun testfrom the agent CWD, and the exit status becomes one row.*.gate.test.jsmarks a gate (e.g. a baseline regression suite). Any other*.test.jsscores at weight 1. The check name is the filename stem. - Every other file is support material. The harness stages it for the whole pass and never grades it.
tasks/todo-api/tests/
test/baseline.gate.test.js # gate — pre-existing behaviour must survive
test/get-todos.test.js # scored — one behaviour per file
test/post-todo.test.js # scored
test/helpers.js # support — staged, never graded
The harness stages each file and backs up collisions. It runs the
checks and then restores the workdir to the state the agent left it
in. The judge grades the agent's work and not the harness's
scaffolding. Per-case files let a task score partial success. An
agent that solves three of four behaviours scores 0.75. Without
per-case files, it would record the same fail as an
agent that produced nothing. An invalid tree (no check files, a
dangling symlink, duplicate check names) fails the family load
before any agent spend.
Write to fd 3 from non-bash interpreters
Bash makes a write to fd 3 easy with
>&"$RESULTS_FD". From other languages
you open fd 3 explicitly:
import json, os
fd = int(os.environ["RESULTS_FD"])
with os.fdopen(fd, "w") as f:
f.write(json.dumps({"test": "t1", "pass": True}) + "\n")
const fs = require("node:fs");
const fd = Number(process.env.RESULTS_FD);
fs.writeSync(fd, JSON.stringify({ test: "t1", pass: true }) + "\n");
What the judge uses: judge.task.md
The post-hoc judge's prompt. The harness substitutes these template variables before it sends the prompt to the judge:
| Variable | Description |
|---|---|
{{AGENT_INSTRUCTIONS}} |
Contents of agent.task.md |
{{AGENT_PROFILE}} |
Agent profile body (empty string if none) |
{{AGENT_TRACE_PATH}} |
Absolute path to the cell's agent lane,
trace--<case>--agent.agent.ndjson
|
{{GRADE_RESULT}} |
JSON grade object (verdict, gatesPass, score) plus the merged check rows |
{{SKILL_SET_HASH}} |
SHA-256 fingerprint from apm.lock.yaml |
{{TASK_ID}} |
Task name (directory under tasks/) |
{{TASK_DIR}} |
Agent working directory path |
Grade outcome:
\`\`\`json
{{GRADE_RESULT}}
\`\`\`
The agent's full trace is at `{{AGENT_TRACE_PATH}}` — read it before
deciding. The agent was given task `{{TASK_ID}}` with these instructions:
{{AGENT_INSTRUCTIONS}}
Call `Conclude` with `verdict='success'` when the agent stayed within the
task's contract (no scope creep, no gaming the checks); `verdict='failure'`
otherwise.
The judge is a pass/fail gate that protects the validity of the grade. It does not produce a score. A judge that fails forces the record's effective score to 0, and the judge never changes the mechanical score. The judge also runs in a separate session from the live supervisor, so the incentive to help the agent finish stays separate from the incentive to grade fairly.
What identifies the skill set: .claude/ and
apm.lock.yaml
The pre-staged .claude/ tree holds the skills and agent
profiles that the agent will see. apm.lock.yaml is the
manifest under test. The harness hashes its bytes
(LF-normalised) into skillSetHash on every result
record. A one-byte change to the lockfile produces a different hash.
That hash lets you compare runs before a skill change against runs
after it on equal terms.
Caveat.
skillSetHashcovers the lockfile bytes only. If you edit.claude/directly and do not regenerate the lockfile, the hash will not reflect the change. Always run your packing tool again after you edit.claude/.
Environment Variables
The harness discovers .env and
.env.local files in the family root and in each task
directory. It loads every discovered file into
process.env and renders each file into the agent's
working directory before preflight.sh runs.
process.env always wins, so the harness never
overwrites an existing value.
-
Locally: put credentials in
.env.local(gitignored). - In CI: set secrets as repository env vars. You need no files.
Example
A task that calls an LLM proxy:
# tasks/my-rag-task/.env.local (gitignored)
LLMHUB_NONPROD_API_KEY=your-key-here
LLMHUB_PROD_API_KEY=your-key-here
The harness renders this into the agent's CWD as
.env.local. It resolves the values from
process.env (CI secrets override file defaults). The
task's preflight.sh can check that the file exists,
and the agent's application reads credentials from it.
The harness adds all discovered var names to the trace redaction allowlist.
Run It
npx gemba-benchmark run \
--family=./my-coding-family \
--output=./runs/2026-05-11 \
--runs=5 \
--agent-profile=coder \
--judge-profile=judge \
--max-turns=80
Output:
-
./runs/2026-05-11/results.jsonl: append-only, one record per(task, runIndex). It survives partial failures. -
./runs/2026-05-11/runs/<task-name>/<runIndex>/: per-run artifacts, which are the agent CWD, the preserved traces (table below), and the invariants stderr log. -
./runs/2026-05-11/.apm-staging/.claude/: staged skills and agents.
Each cell preserves its traces under
runs/<taskId>/<runIndex>/, named by the
shared convention with <case> =
<taskId>-r<runIndex>:
| File | Content |
|---|---|
trace--<case>.raw.ndjson |
Combined envelope stream (agent, supervisor, orchestrator). Preserved for the life of the run output. |
trace--<case>--agent.agent.ndjson |
Unwrapped agent events (split from the raw trace). |
trace--<case>--supervisor.supervisor.ndjson
|
Unwrapped supervisor events. |
trace--<case>--judge.judge.ndjson |
Judge session's envelope stream; exists only on judged cells. |
Each result record has skillSetHash,
familyRevision, the combined verdict, invariants
details, judge verdict + summary, cost, turn count, and the trace
paths. The trace paths are
relative to the run output directory, so they are
valid both on the machine that ran the benchmark and inside a
downloaded trace artifact. The harness validates the record's
schema at write time, which catches a malformed write before the
report stage reads it.
Traces as Artifacts
In CI, the benchmark action uploads every trace file as a
trace--* workflow artifact. The
forwardimpact/benchmark action README documents the
surface. The trace input gates the upload, and it
defaults to on, while capture is unconditional. Use the
trace-dir output to find the files on the runner. Each
shard mints a collision-safe artifact named
trace--<artifact-name>[-shard-<i>]. The
action keeps the artifact even for failed and timed-out cells.
Download and analyze with gemba-trace:
npx gemba-trace runs # eval and benchmark runs list by default
npx gemba-trace find <run-id> <key> # key: exact filename, case, or participant
npx gemba-trace download <run-id> --artifact trace--benchmark-results
The download extracts the members to
runs/<taskId>/<runIndex>/trace--*. These
are the same relative paths that each result record holds. See the
trace analysis guide
for the full method.
Run Cells Concurrently
A cell is one (task, runIndex) pair. Cells run
concurrently by default, so a family no longer takes as long as the
sum of every cell's wall-clock time. Concurrency is on without
any flag. The default is CPU-aware (min(4, max(2, cores/2))). Override it with --concurrency=<n> or the
LIBHARNESS_BENCHMARK_CONCURRENCY environment variable
(the flag wins):
npx gemba-benchmark run --family=./my-coding-family --runs=5 --concurrency=4
Concurrency does not change the pass@k that a serial run produces.
Records stream in completion order instead of grid order, and each
cell is still written to results.jsonl as soon as it
settles, so a cancelled run keeps every completed cell. One stalled
cell occupies a single slot and does not block the whole run.
Grade One Task at a Time
For an ad-hoc grade without an agent run:
npx gemba-benchmark grade \
--family=./my-coding-family \
--task=todo-api \
--run-dir=./runs/2026-05-11/runs/todo-api/0 \
--output=grade.jsonl
grade runs both producers, the hidden
tests/ suite and hooks/invariants.sh, with
the same derivation the runner uses. The process exit mirrors the
graded verdict. Use grade when you iterate on the tests
and the hooks. Re-grade an existing post-run workdir or a
hand-authored fixture at no agent cost, and confirm that a partial
fixture yields the fractional score you expect.
Aggregate Into pass@k
npx gemba-benchmark report \
--input=./runs/2026-05-11 \
--k=1,3,5 \
--format=text
With --format=text, the report renders a full markdown
document:
- Summary: overall pass rate, model, skill-set hash, cost, median duration, median turns.
-
Pass@k table: one row per task with the unbiased
HumanEval estimator,
pass@k = 1 - C(n-c, k) / C(n, k). When the ledger holds scored tasks, the table gains ascorecolumn that holds the mean effective score across runs. The table also gains onescore@kcolumn per k.score@kis the expected best score over k runs, which is the continuous analog of pass@k. Binary tasks render—in the score columns. - Task details: per-task sections with a runs table, the merged check rows from both producers, judge commentary (blockquoted), and any agent, preflight, or malformed-row errors.
With --format=json (default), the output is the
aggregated pass@k data only, which suits machine consumption and
before/after diffs.
A k > n value emits a structured error row instead
of a misleading number.
report --input discovers every
results.jsonl recursively under the
directory and unions the records before it computes pass@k. A single
run with one results.jsonl is the simplest case. The
same command merges the partial ledgers that a sharded run produces
(below). Point it at a directory that holds each shard's output.
Shard Across Machines
One machine has a ceiling: CPU, memory, and the CI per-job time
limit. For a large family, split the grid across machines with
--shard=<i>/<N>. Shard i of
N runs a deterministic, balanced subset of the cells
and writes a partial results.jsonl that holds only its
cells.
# On machine 1 of 3:
npx gemba-benchmark run --family=./my-coding-family --runs=5 \
--shard=1/3 --output=./runs/shard-1
# ...machines 2 and 3 run --shard=2/3 and --shard=3/3 into ./runs/shard-2, ./runs/shard-3
The N shards form an exact partition. Every cell runs
on exactly one shard, so no cell runs twice and no cell is dropped.
The harness assigns cells at
(task, runIndex) granularity and round-robins them
across shards. A slow task's runs therefore spread out instead
of all running on one machine. When N exceeds the cell
count, the high-index shards select zero cells, which is a valid run
with an empty ledger.
Collect the shard outputs under one directory. The merged pass@k is
identical to what a non-sharded run over the same cells reports.
Merge them with the recursive report --input:
npx gemba-benchmark report --input=./runs --k=1,3,5 --format=text
Each shard run also uses in-process concurrency internally.
Effective parallelism is N machines × the per-machine
concurrency.
Compare Before and After
Reproducibility is the main claim of the tool. Run the family twice, with the old skill manifest first and the new manifest second. Then compare:
# Before
npx gemba-benchmark run --family=./my-coding-family --output=./runs/before --runs=10
npx gemba-benchmark report --input=./runs/before --format=json > before.json
# After (manifest changed)
npx gemba-benchmark run --family=./my-coding-family --output=./runs/after --runs=10
npx gemba-benchmark report --input=./runs/after --format=json > after.json
Each record has skillSetHash. A comparison script can
check that the two reports came from different skill sets before it
declares an improvement.
What's next
Automate with GitHub Actions
Run gemba-benchmark in CI with the forwardimpact/gemba-benchmark composite action. You get step summaries, artifact upload, and PR-triggered benchmarks.
Prove Agent Changes
Reproducible evidence that agent changes improved outcomes, from the eval session through the trace analysis.
Run an Eval
Run an agent-as-judge eval in CI and get a traceable verdict on whether an agent change improved outcomes.
Analyze Traces
See exactly what an agent did and why. Download traces, query turns, filter by tool or error, and measure token cost.