Skip to content

Report

trialdesignbench.report

Aggregate graded trials into a benchmark report and leaderboard.

Accepts Harbor job directories (trial dirs with result.json and verifier/grade.json) and plain directories of standalone <task_id>/grade.json files, so external submissions are comparable.

Unattempted or errored tasks count as 0 and are listed explicitly.

aggregate(trials, task_ids, threshold)

Per agent x model x effort aggregates; unattempted tasks count as 0.

build_report(sources, *, threshold=DEFAULT_THRESHOLD, task_ids=None, allow_canary_failure=False)

Build a ReportSummary; refuses when a network canary failed.

effort_label(effort)

Display label for an effort level (default for the harness default).

is_job_dir(path)

Whether path looks like a Harbor job directory.

leaderboard(report)

Ranked rows by mean score, then pass rate.

load_job(job_dir)

Trials, expected task ids, and canary failures for one Harbor job.

load_standalone(root)

Trials from a directory of <task_id>/grade.json files.

render_markdown(report)

Markdown report with leaderboard, per-task tables, and flagged trials.