Analyze Traces
You need to see exactly what the agent did, so that you can debug
failures and verify improvements. gemba-trace reads the
NDJSON traces that gemba-harness produces. It gives you
structured queries over every turn, tool call, and result.
Prerequisites
- Node.js 22+
-
The
gemba-*commands. Install them withnpm install -g @forwardimpact/gemba, or run each one throughnpx gemba-trace ... -
A trace file, either the
--outputfrom agemba-harnessrun or a file you download from CI withgemba-trace download
Get the trace
Local runs already produce a trace at the
--output path. For CI runs, list recent workflow runs
and download:
npx gemba-trace runs # list recent workflow runs
npx gemba-trace download 24497273755 # downloads to /tmp/trace-24497273755/
The default runs pattern covers kata, agent, eval, and
benchmark workflow names, so benchmark-driven eval runs list with no
flags. The download extracts the artifact zip's
.ndjson members. These are the
trace--<case>--<participant>.<role>.ndjson
lane files plus the combined
trace--<case>.raw.ndjson. Every one of them is
direct input to every query command below. The download produces a
structured.json only when the artifact has a single
.ndjson member, and the common bundles have several.
When you know the run but not the file, find resolves
one lane by key. The key may be a participant name, a case id, or an
exact member filename. A key that matches several members fails and
lists the candidates so that you can narrow it:
npx gemba-trace find 24497273755 agent # participant key
npx gemba-trace find 24497273755 fix-bug-r0 # case key (eval runs)
npx gemba-trace find 24497273755 trace--fix-bug-r0--agent.agent.ndjson
Orient with the overview
Start with the overview before you look at individual turns.
Analysis verbs take their trace files through
--file and print human-readable text by default. Add
--format json for the machine-parseable envelope:
npx gemba-trace overview --file /tmp/trace-24497273755/trace--default--agent.agent.ndjson --format json
{
"summary": { "result": "success", "totalCostUsd": 0.42, "numTurns": 18 },
"verdict": null,
"turnCount": 34,
"tools": [{ "tool": "Bash", "count": 12 }, { "tool": "Read", "count": 8 }],
"taskPrompt": "Refactor src/utils/format.js so that formatDate and formatCurrency share..."
}
verdict carries the lead's terminal verdict on the
combined trace of a coordinated run. A split lane and a single-agent
run read null.
The timeline command shows the shape of the session in
a few lines. It prints one line per assistant turn, with the tools
used and the token counts:
npx gemba-trace timeline --file /tmp/trace-24497273755/trace--default--agent.agent.ndjson
[1] Read in:12.3K out:0.8K Let me read the current implementation...
[3] Bash in:13.1K out:1.2K Running the existing tests first...
[5] Edit in:14.0K out:2.1K I'll extract the shared locale helper...
[7] Bash in:15.2K out:0.4K Running tests to verify the refactor...
Find errors
List every tool result where the agent's tool call failed:
npx gemba-trace errors --file /tmp/trace-24497273755/trace--default--agent.agent.ndjson
Each result includes the turn index, the toolUseId that
links it back to the assistant turn that made the call, and the
error content.
Filter by tool or role
See every turn where the agent used a specific tool. The output
holds both the tool_use request and its
tool_result response:
npx gemba-trace tool /tmp/trace-24497273755/trace--default--agent.agent.ndjson Bash
tool takes the trace file as a positional argument,
because it works on a single trace plus a tool name. Or use
filter for structural queries by role, tool name, or
error status:
npx gemba-trace filter --file /tmp/trace-24497273755/trace--default--agent.agent.ndjson --tool Edit
npx gemba-trace filter --file /tmp/trace-24497273755/trace--default--agent.agent.ndjson --error
npx gemba-trace filter --file /tmp/trace-24497273755/trace--default--agent.agent.ndjson --role user
Search across the trace
Search all turn content with a regex pattern (search is
single-file, so the file is a positional argument):
npx gemba-trace search /tmp/trace-24497273755/trace--default--agent.agent.ndjson 'permission denied' --context 1
--context 1 includes one turn on each side of every
match. --limit 10 caps the number of results.
--full emits the complete content block instead of a
short excerpt.
Read the agent's reasoning
The text blocks in assistant turns show what the agent said it would do, and the tool calls show what it did. Extract only the text blocks:
npx gemba-trace reasoning --file /tmp/trace-24497273755/trace--default--agent.agent.ndjson --from 5 --to 15
[
{ "index": 5, "text": "I'll extract the shared locale helper..." },
{ "index": 9, "text": "Tests pass. Now adding coverage for de-DE..." }
]
Compare reasoning output to the tool calls
to find mismatches between intent and execution.
Measure token usage and cost
npx gemba-trace stats --file /tmp/trace-24497273755/trace--default--agent.agent.ndjson --format json
{
"totals": {
"inputTokens": 142800, "outputTokens": 18400,
"totalCostUsd": 0.42, "durationMs": 94200,
"durationLabel": "cumulative invocation time",
"resultEventTurns": 18, "population": "result-event-sum"
},
"perTurn": [{ "messageId": "msg_01", "inputTokens": 12300, "outputTokens": 800, "population": "api-message", ... }]
}
The totals are the sum over all result events in
the trace. A supervised or facilitated session has one result event
per invocation, so if you read only the last one, you undercount the
session cost. The perTurn breakdown is one row per API
message. Its outputTokens comes from a snapshot of the
stream, so it is a lower bound and not the final count. Every figure
states its population. A trace with no result event still reports
per-message totals, and it marks cost and duration unavailable
instead of a misleading 0.
stats --by-tool attributes token usage and a cost-share
fraction to each tool. The fractions sum to 1.0. Turns that made no
tool call go into the (no-tool) bucket.
stats --summary prints the totals block only. Both
views report the same result-event totals, so their per-bucket token
sums match the un-flagged stats totals.
Track these numbers across runs over time. A single trace is a snapshot, and a series of traces shows whether your changes had an effect.
Split multi-agent traces
For supervised or facilitated runs, split the combined trace into per-source files, so that you can see what each agent saw on its own:
npx gemba-trace split /tmp/trace-24497273755/trace--default.raw.ndjson --mode=facilitate --case=demo
This produces files in the same directory. The names follow the
trace--<case>--<participant>.<role>.ndjson
convention:
trace--demo--facilitator.facilitator.ndjson and one
trace--demo--<participant>.agent.ndjson per
participant. Each file works as input to every query command above.
For supervised runs, use --mode=supervise to get
trace--<case>--agent.agent.ndjson and
trace--<case>--supervisor.supervisor.ndjson.
--case defaults to default. Matrix
workflows pass the case id, so per-shard artifacts stay separate.
Eval traces
Benchmark-driven eval runs emit the same convention, and the case id
identifies the cell. <case> is
<taskId>-r<runIndex>, so every cell in the
grid has its own lanes
(trace--fix-bug-r0--agent.agent.ndjson). The judge gets
its own lane, trace--<case>--judge.judge.ndjson.
Members extract nested per cell
(runs/<taskId>/<runIndex>/trace--*). Raw
and judge files are enveloped
{source, seq, event} streams, and split lanes carry
unwrapped events. Every file-consuming verb takes both shapes as
they are, with no eval-specific flags.
Navigate individual turns
When you need to inspect a specific moment in the trace:
npx gemba-trace turn /tmp/trace-24497273755/trace--default--agent.agent.ndjson 8
npx gemba-trace batch /tmp/trace-24497273755/trace--default--agent.agent.ndjson 5 10
npx gemba-trace head --file /tmp/trace-24497273755/trace--default--agent.agent.ndjson --lines 5
npx gemba-trace tail --file /tmp/trace-24497273755/trace--default--agent.agent.ndjson --lines 5
turn and batch are single-file
(positional). batch returns turns in the half-open
range [from, to). head and
tail are cross-trace (--file). They take
their count through --lines, which defaults to 10.
Aggregate without writing wrappers
These verbs answer questions that used to need a script.
tool-calls emits one record per
tool_use block and pairs each block with its
tool_result by toolUseId. Orphaned calls
show (no result), and tool-calls never
drops them:
npx gemba-trace tool-calls --file /tmp/trace-24497273755/trace--default--agent.agent.ndjson
commands lists every Bash command (filter with
--match <regex>). paths gives a
frequency-sorted list of the distinct
Read/Edit/Write file paths
(filter with --prefix):
npx gemba-trace commands --file /tmp/trace-24497273755/trace--default--agent.agent.ndjson --match '^git'
npx gemba-trace paths --file /tmp/trace-24497273755/trace--default--agent.agent.ndjson --prefix /app
These verbs complement tool (every turn for one tool)
and tools (frequency across all tools). Use
tool-calls when you want one record that holds both the
use and the result.
Compare two traces
compare puts two traces side by side. It reports turn
count, distinct tools, paths touched, cost, and a per-tool delta.
The header shows the case name and the participant for each side:
npx gemba-trace compare trace--demo--agent.agent.ndjson trace--demo--supervisor.supervisor.ndjson
Identical traces emit zero deltas. An empty trace emits zeroed
counters with an (empty) marker, and it does not error.
compare takes its two files as positional arguments and
does not take --file.
Analyze several traces at once
Cross-trace verbs accept more than one trace. Repeat
--file, or pass a quoted glob. The verb expands the
glob itself:
npx gemba-trace paths --file 'traces/*.ndjson' --prefix /app
npx gemba-trace tool-calls --file run-a.ndjson --file run-b.ndjson
With more than one resolved file, every record includes its source,
so you can tell the traces apart. Per-record verbs prefix each line
with <basename>: (grep -H
convention). The aggregators (paths,
tools) carry a sources array in
--format json. A single resolved file has no source
prefix, and a glob that matches exactly one file counts as a single
file. Source attribution uses the file's
basename, so two traces with the same basename in
different directories collide. Rename them, or run from inside one
directory to keep them distinct.
What to look for
When you debug a failure, use this sequence:
-
overview: see whether the run succeeded or failed, and how many turns it took. errors: see which tool calls failed.-
tool <name>on the tool that failed: see what input the agent sent. -
reasoningaround those turns: see whether the agent understood the error. -
searchfor the error message: see whether it appeared earlier than you expected.
When you verify an improvement, compare stats across
before-and-after runs. Fewer retries, lower token usage, and shorter
duration are the signals that a profile or prompt change improved
outcomes.
What's next
Prove Agent Changes
Reproducible evidence that agent changes improved outcomes, from the eval session through the trace analysis.
Run an Eval
Run an agent-as-judge eval in CI and get a traceable verdict on whether an agent change improved outcomes.
Run a Benchmark
Prove a skill-pack change improved coding outcomes. Run a task family across N runs, grade with hidden tests, and report pass@k.