Overview
trialdesignbench is a thin evaluation framework. It owns the task schema,
task materialization, the grader, scoring rules, aggregation, and provenance.
Harbor is the execution backend:
it runs first-party agent harnesses (Claude Code, Codex CLI, Grok Build,
OpenCode; see Agents) inside Docker with concurrency, retries,
trajectories, and token accounting.
The two tools meet only through files:
- the Harbor task directory format (
task.toml,instruction.md,environment/,tests/), - a generated
job.yamlpassed toharbor run -c, - the job directory Harbor writes (
result.json,verifier/reward.json,verifier/reward-details.json,agent/trajectory.json).
trialdesignbench never imports Harbor, so the core package stays light.
Harbor is an optional extra. Both require Python 3.12+.
Pipeline
flowchart TD
intake["Curated intake JSON"] -->|tdb dataset import| dataset["Canonical dataset"]
dataset -->|tdb build| tasks["Harbor task directories"]
tasks -->|"tdb run<br/>(harbor run -c job.yaml)"| job["Harbor job directory"]
job -->|tdb report| report["report.json"]
job -->|"tdb regrade<br/>(harbor job regrade)"| job
Grading is decoupled from inference. tdb grade is a pure function of the
submission artifacts, the hidden rubrics, and the judge configuration, so it
runs identically inside Harbor's verifier container, standalone on any
directory with output.json and output.R, and in tests.
Quick start
# 1. Canonical dataset from curated submissions (+ protocol/SAP Markdown)
uv run tdb dataset import \
data/json/*.json \
--out tmp/dataset \
--documents path/to/docs
uv run tdb dataset check tmp/dataset
# 2. Shared environment image (R, pinned CRAN snapshot, agent CLIs, skills)
uv run tdb env build
uv run tdb env check --canary
# 3. Harbor tasks
uv run tdb build \
tmp/dataset \
--out tmp/tasks
# 4. Run agents (needs `trialdesignbench[harbor]`)
export ANTHROPIC_API_KEY=... # the agent's model API key and the judge's key
uv run tdb run \
--tasks tmp/tasks \
--agent claude-code \
--model anthropic/claude-opus-5-5 \
--effort high \
--n-attempts 3 \
--canary
# 5. Aggregate
uv run tdb report \
"jobs/<job-name>" \
--format md
Each step has its own article:
- Dataset: intake import, the canonical format, documents.
- Environment: the shared image and the network policy.
- Build: Harbor task materialization.
- Agents: supported agents, credentials, reasoning effort, closed-book settings.
- Closed book: how egress control works, which agent web tools run provider-side, and why some agents are refused.
- Run:
job.yaml, network allowlists, matrices, regrading. - Grade: deterministic checks, rubric judging, scoring.
- Judge: judge model, credentials, and network access.
- Report: aggregation and leaderboards.
- Reproducibility: what is pinned and recorded.