Report
uv run tdb report \
jobs/<job> \
[more dirs...] \
[--out report.json] \
[--format table|json|md] \
[--threshold 0.8] \
[--task-ids ...]
Inputs
- Harbor job directories. Each trial's
result.jsonandverifier/grade.jsonare read (falling back toverifier/reward.json). The job'sconfig.jsondefines the expected tasks. - Standalone directories of
<task_id>/grade.jsonproduced bytdb gradewith no Harbor involved. The directory name is used as the model label, so external submissions sit on the same leaderboard.
Aggregation
Trials are grouped by agent × model × effort × task. The effort level is
read from each trial's recorded agent kwargs (reasoning_effort, or
variant for opencode); trials that ran at the harness default, and all
standalone directories, are labelled default.
- Per task: mean, min, and max score over attempts, and "all attempts pass"
(every attempt scores at least
--threshold, a pass^k style metric). - Per agent × model × effort:
mean_scoreis the mean over all expected tasks; tasks never attempted count as 0 and are listed as missing. Trials that errored (agent crash, no verifier output, grade statuserror) score 0 and are listed. - Breakdowns by question type, design element, and rubric dimension are the mean of the per-trial rubric breakdowns.
- Token and cost totals come from each trial's trajectory
final_metrics, falling back to Harbor'sagent_result.
Every trial with a non-graded status or network violations is listed with
its reasons.
Canary
If a job contains a network canary trial that did not pass, tdb report
refuses to report and exits with an error. Canary trials are never scored.