Report
trialdesignbench.report
Aggregate graded trials into a benchmark report and leaderboard.
Accepts Harbor job directories (trial dirs with result.json and
verifier/grade.json) and plain directories of standalone
<task_id>/grade.json files, so external submissions are comparable.
Unattempted or errored tasks count as 0 and are listed explicitly.
aggregate(trials, task_ids, threshold)
Per agent x model x effort aggregates; unattempted tasks count as 0.
build_report(sources, *, threshold=DEFAULT_THRESHOLD, task_ids=None, allow_canary_failure=False)
Build a ReportSummary; refuses when a network canary failed.
effort_label(effort)
Display label for an effort level (default for the harness default).
is_job_dir(path)
Whether path looks like a Harbor job directory.
leaderboard(report)
Ranked rows by mean score, then pass rate.
load_job(job_dir)
Trials, expected task ids, and canary failures for one Harbor job.
load_standalone(root)
Trials from a directory of <task_id>/grade.json files.
render_markdown(report)
Markdown report with leaderboard, per-task tables, and flagged trials.