Agents
tdb run launches agents through Harbor's
pre-integrated agents. Harbor
can run many more agents than TrialDesignBench accepts: an agent is supported
only when it can run closed book, meaning that its web tools can be turned
off, its network egress can be limited to the model API, and Harbor records
an ATIF trajectory that the grader can scan.
The tables below mirror the single list in the code,
trialdesignbench.agents.AGENTS (a test keeps them in sync); tdb run
refuses any other agent.
Supported agents
| Agent | Harness | Pinned version | Installed | --model providers |
--auth |
|---|---|---|---|---|---|
claude-code |
Claude Code | 2.1.285 | in the image | anthropic |
api, subscription |
codex |
Codex CLI | 0.159.1 | in the image | openai |
api, subscription |
grok-build |
Grok Build | 1.0.44 | by Harbor at setup | xai |
api |
opencode |
OpenCode | 1.18.33 | by Harbor at setup | anthropic, openai, xai |
api |
--modelis always<provider>/<model>, for exampleanthropic/claude-opus-5-5,openai/gpt-6-astra, orxai/grok-4.7. The provider decides which API key is required and which API host is allowlisted, so a model from any other provider is refused.- The pinned version is passed to Harbor as the agent
version;--agent-versionmay only repeat it. Pins live intrialdesignbench.environment.PINSand are recorded as image labels. - In the image: the CLI is preinstalled in the shared image, so Harbor skips its install step. By Harbor at setup: Harbor's adapter reinstalls the CLI at every trial setup no matter what the image contains, so a few install hosts are reachable during setup only (see Network).
Run one agent, or a matrix by repeating the pair:
uv run tdb run \
--tasks tmp/tasks \
--agent opencode \
--model anthropic/claude-opus-5-5
uv run tdb run \
--tasks tmp/tasks \
--agent claude-code \
--model anthropic/claude-opus-5-5 \
--agent grok-build \
--model xai/grok-4.7
Reasoning effort
Reasoning effort changes a model's capability as much as the model choice
does, so tdb run treats it as a first-class setting next to --agent and
--model. --effort <level> is checked before launch, becomes the
agent-specific Harbor kwarg below, and is recorded in tdb-run.json
(agents[].effort, and in agents[].kwargs) and in each trial's
result.json. tdb report groups results by agent × model × effort, so the
same model at two levels never shares a leaderboard row.
| Agent | Harbor kwarg | Levels (lowest to highest) | Becomes |
|---|---|---|---|
claude-code |
reasoning_effort |
low, medium, high, xhigh, max |
claude --effort <level> |
codex |
reasoning_effort |
none, minimal, low, medium, high, xhigh, max |
codex -c model_reasoning_effort=<level> |
grok-build |
reasoning_effort |
none, minimal, low, medium, high, xhigh, max |
grok --reasoning-effort <level> |
opencode |
variant |
none, minimal, low, medium, high, xhigh, max |
opencode --variant <level> |
The levels are those Harbor 0.23.0 accepts for the kwarg (harbor agent
schema <name>), confirmed against the pinned CLIs. Each model supports a
subset, which the provider documents and which changes with new models;
tdb run cannot check that on the host. What happens with a level the model
lacks differs by agent:
claude-code: the CLI runs the highest supported level at or below the requested one (for examplexhighruns ashighon Opus 4.6). A value the CLI does not know is dropped with a warning and the default applies, which is whytdb runrefusesnoneandminimalhere.codex: the level is sent to the OpenAI API asreasoning.effort, which enumerates exactly these seven values. A level the model does not support fails the request, so the trial errors (scored 0 and listed bytdb report).grok-build: a model accepts only the levels its menu advertises (grok-4.7:low,medium,high,xhigh); reasoning cannot be turned off ongrok-4.x.opencode: variant names are the model'sreasoning_optionseffort values in the catalog the image ships at/opt/tdb/opencode-models.json(for exampleclaude-opus-5-5:lowtomax;gpt-6-astra:none,low,medium,high,xhigh;grok-4.7:lowtoxhigh). OpenCode silently ignores a variant the model does not define, so check the catalog before a large run.
Without --effort (or with --effort default) no kwarg is set and each
harness applies its own default: Claude Code the model's default (high for
most models, medium for Opus 5.5), Codex medium, Grok Build the model's
default (high on grok-4.x), and OpenCode sends no effort, so the provider
default applies. These defaults depend on the model and change between CLI
versions, so tdb run warns when an agent runs without an explicit level,
and tdb report labels such runs default.
--effort is given once for every agent, or once per --agent; default
is the placeholder for an agent that keeps the harness default. The same
agent and model may appear at several levels in one job:
uv run tdb run \
--tasks tmp/tasks \
--agent codex \
--model openai/gpt-6-astra \
--effort xhigh
uv run tdb run \
--tasks tmp/tasks \
--agent claude-code \
--model anthropic/claude-opus-5-5 \
--effort medium \
--agent claude-code \
--model anthropic/claude-opus-5-5 \
--effort max \
--agent grok-build \
--model xai/grok-4.7 \
--effort default
Levels are not comparable across agents or providers: high is Claude
Code's default but two steps above Codex's. Compare levels within one agent
and model, and report the level with every result.
Authentication
Agent credentials and the judge credential are separate
concerns. Agents authenticate to their own model API as described here; the
verifier's judge always needs ANTHROPIC_API_KEY, whichever agent runs.
Credentials are read from the environment of the shell that runs tdb run.
Harbor's adapters pick them up from there, so nothing is passed on the
command line. tdb run checks that the required variables or files exist
before launching and fails otherwise. Secret values are never written to
job.yaml or tdb-run.json; only variable names are recorded.
API keys (--auth api, the default)
Export the key for the provider of each --model:
| Provider | Variable | API host | Keys from |
|---|---|---|---|
anthropic |
ANTHROPIC_API_KEY |
api.anthropic.com |
Claude Console |
openai |
OPENAI_API_KEY |
api.openai.com |
OpenAI Platform |
xai |
XAI_API_KEY |
api.x.ai |
xAI Console |
export ANTHROPIC_API_KEY=... # agent (anthropic/ models) and the judge
export OPENAI_API_KEY=... # codex, or opencode with openai/ models
export XAI_API_KEY=... # grok-build, or opencode with xai/ models
uv run tdb run \
--tasks tmp/tasks \
--agent codex \
--model openai/gpt-6-astra
With claude-code, or opencode with an anthropic/ model, the same
ANTHROPIC_API_KEY serves the agent and the judge.
Subscriptions (--auth subscription)
Only claude-code and codex support subscription login through Harbor.
Subscription runs default to one concurrent trial (--n-concurrent 1).
| Agent | Plans | Requires | tdb run sets |
|---|---|---|---|
claude-code |
Pro, Max, Team, Enterprise | CLAUDE_CODE_OAUTH_TOKEN |
CLAUDE_FORCE_OAUTH=1 |
codex |
ChatGPT Plus, Pro, Business, Enterprise | ~/.codex/auth.json, or a file named by CODEX_AUTH_JSON_PATH |
CODEX_FORCE_AUTH_JSON=1 |
Claude Code: create a long-lived token with claude setup-token and export
it. CLAUDE_FORCE_OAUTH=1 makes Harbor drop ANTHROPIC_API_KEY from the
agent, so the key can stay exported for the judge.
claude setup-token # prints a one-year OAuth token
export CLAUDE_CODE_OAUTH_TOKEN=...
export ANTHROPIC_API_KEY=... # still required by the judge
uv run tdb run \
--tasks tmp/tasks \
--agent claude-code \
--model anthropic/claude-opus-5-5 \
--auth subscription
Codex: sign in once with codex login (or codex login --device-auth on a
headless machine). Harbor uploads that auth.json into each trial and does
not set OPENAI_API_KEY for the agent.
codex login # writes ~/.codex/auth.json
export ANTHROPIC_API_KEY=... # required by the judge
uv run tdb run \
--tasks tmp/tasks \
--agent codex \
--model openai/gpt-6-astra \
--auth subscription
Closed-book settings
tdb run adds these settings to every agent it launches. They are
recorded in tdb-run.json (agents[].kwargs, env_keys, and
network_policy.disabled_tools).
| Agent | Web tools disabled | How |
|---|---|---|
claude-code |
WebSearch, WebFetch |
disallowed_tools plus a native settings file with permissions.deny; DISABLE_TELEMETRY, DISABLE_ERROR_REPORTING, CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC |
codex |
web_search |
web_search=disabled |
grok-build |
web_search, web_fetch |
disable_web_search=true and [features] web_fetch = false in the generated config, both also pinned in the image's /etc/grok/requirements.toml, which user config cannot override |
opencode |
webfetch, websearch |
permission denies in opencode_config and OPENCODE_PERMISSION; auto-update, sharing, LSP, formatters, and snapshots off; model catalog read from the image (OPENCODE_MODELS_PATH) instead of fetched |
No MCP servers are ever added. Provider-side web tools never touch the container network, which is why disabling them in the harness matters; the grader additionally scans each trajectory for web tool calls (see Grade). Which of each agent's tools run provider-side, and the probe runs that established it, are documented in Closed book.
Network
Every task has two allowlists (see Environment):
[agent], duringagent.run(): only the model API hosts of the agent and auth mode.[environment], during agent setup and the healthcheck: the same hosts plus the install hosts of agents that Harbor installs at setup.
| Agent | --auth |
Agent phase | Setup adds | Status |
|---|---|---|---|---|
claude-code |
api |
api.anthropic.com |
nothing | verified by canary and a smoke run |
claude-code |
subscription |
api.anthropic.com |
nothing | unverified: add any OAuth refresh host a smoke run shows |
codex |
api |
api.openai.com |
nothing | verified by table only |
codex |
subscription |
chatgpt.com, auth.openai.com |
nothing | unverified |
grok-build |
api |
api.x.ai |
archive.ubuntu.com, security.ubuntu.com, ports.ubuntu.com, x.ai |
canary and Harbor install-only run pass; needs a smoke run |
opencode |
api |
the provider's API host | raw.githubusercontent.com, github.com, nodejs.org, registry.npmjs.org |
canary and Harbor install-only run pass; needs a smoke run |
Entries not verified by a smoke run make tdb run print a warning. Confirm
them with tdb env check --canary --agent <name> [--provider <provider>] and
a smoke run before relying on them. For grok-build and opencode, the canary
(both phases) and a Harbor --install-only run (the real install under the
setup allowlist, amd64) pass; a run with a real API key would confirm that
the CLIs need no other host during agent.run(). grok-build's setup reaches
the Ubuntu mirrors because Harbor's adapter always runs apt-get install
ca-certificates.
Refused agents
These Harbor agents were evaluated and are refused by tdb run with the
reason below. Any other Harbor agent is refused as unsupported.
| Agent | Reason |
|---|---|
antigravity-sdk |
Harbor's runner enables every Antigravity SDK tool, including the Google-side search_web and read_url_content, and offers no option to disable them. |
antigravity-cli |
Harbor's adapter rewrites the Antigravity CLI settings on every run and offers no option to disable its web search tool. |
kimi-code |
Harbor records no ATIF trajectory for it, so the grader cannot verify closed book and every trial would score 0. |
muse-code |
No ATIF trajectory (every trial would score 0), no documented way to disable its web tools, and Harbor cannot pin its version. |
The evidence behind these reasons and what would have to change to accept one of these agents are in Closed book.
Adding or updating an agent
Everything lives in src/trialdesignbench/agents.py:
- Read the agent's Harbor adapter (
harbor agent schema <name>lists its kwargs). Check how it installs (does it skip when the pinned version is present?), which credentials it reads, which web tools it exposes and how to disable them, and that it writes an ATIF trajectory. Run the probe to learn which of its web tools run on the provider's servers; those are the ones that need a switch. - Add a pin to
ImagePinsand a matchingARGand label to the Dockerfile (a test enforces this), plus any closed-book settings the image must carry. - Add an
AgentProfiletoAGENTS: providers, pin,preinstalledorsetup_hosts, anEffort(the Harbor kwarg that carries the reasoning effort and the levels it accepts, fromharbor agent schema <name>and the CLI's help), closed-bookkwargsandenv,disabled_tools, and an optionalSubscription. Add a new provider toPROVIDERSif needed. - Add the agent to the tables on this page (a test checks the supported,
reasoning effort, and refused lists), and make sure the grader's
WEB_TOOL_NAMEScovers its disabled tools (also tested). - Run
tdb env check --canary --agent <name>and a smoke run, then setverified=True.