Skip to content

Overview

trialdesignbench is a thin evaluation framework. It owns the task schema, task materialization, the grader, scoring rules, aggregation, and provenance. Harbor is the execution backend: it runs first-party agent harnesses (Claude Code, Codex CLI, Grok Build, OpenCode; see Agents) inside Docker with concurrency, retries, trajectories, and token accounting.

The two tools meet only through files:

  • the Harbor task directory format (task.toml, instruction.md, environment/, tests/),
  • a generated job.yaml passed to harbor run -c,
  • the job directory Harbor writes (result.json, verifier/reward.json, verifier/reward-details.json, agent/trajectory.json).

trialdesignbench never imports Harbor, so the core package stays light. Harbor is an optional extra. Both require Python 3.12+.

Pipeline

flowchart TD
  intake["Curated intake JSON"] -->|tdb dataset import| dataset["Canonical dataset"]
  dataset -->|tdb build| tasks["Harbor task directories"]
  tasks -->|"tdb run<br/>(harbor run -c job.yaml)"| job["Harbor job directory"]
  job -->|tdb report| report["report.json"]
  job -->|"tdb regrade<br/>(harbor job regrade)"| job

Grading is decoupled from inference. tdb grade is a pure function of the submission artifacts, the hidden rubrics, and the judge configuration, so it runs identically inside Harbor's verifier container, standalone on any directory with output.json and output.R, and in tests.

Quick start

# 1. Canonical dataset from curated submissions (+ protocol/SAP Markdown)
uv run tdb dataset import \
    data/json/*.json \
    --out tmp/dataset \
    --documents path/to/docs
uv run tdb dataset check tmp/dataset

# 2. Shared environment image (R, pinned CRAN snapshot, agent CLIs, skills)
uv run tdb env build
uv run tdb env check --canary

# 3. Harbor tasks
uv run tdb build \
    tmp/dataset \
    --out tmp/tasks

# 4. Run agents (needs `trialdesignbench[harbor]`)
export ANTHROPIC_API_KEY=... # the agent's model API key and the judge's key
uv run tdb run \
    --tasks tmp/tasks \
    --agent claude-code \
    --model anthropic/claude-opus-5-5 \
    --effort high \
    --n-attempts 3 \
    --canary

# 5. Aggregate
uv run tdb report \
    "jobs/<job-name>" \
    --format md

Each step has its own article:

  • Dataset: intake import, the canonical format, documents.
  • Environment: the shared image and the network policy.
  • Build: Harbor task materialization.
  • Agents: supported agents, credentials, reasoning effort, closed-book settings.
  • Closed book: how egress control works, which agent web tools run provider-side, and why some agents are refused.
  • Run: job.yaml, network allowlists, matrices, regrading.
  • Grade: deterministic checks, rubric judging, scoring.
  • Judge: judge model, credentials, and network access.
  • Report: aggregation and leaderboards.
  • Reproducibility: what is pinned and recorded.