Build
trialdesignbench.build
Materialize canonical dataset tasks as Harbor task directories.
Each task gets:
<tasks_dir>/<task_id>/
task.toml Harbor config (network allowlist placeholder)
instruction.md prompt template + question skeleton + source document
environment/ empty (prebuilt image) or our Dockerfile context
tests/
Dockerfile verifier image: shared image + /tests files
test.sh runs `tdb grade`
rubrics.json hidden grading spec
Network policy has two phases. [environment] is the baseline during agent
setup (and the healthcheck); [agent] applies during agent.run(). Both
allowed_hosts lists are written empty (deny all) with marker comments;
tdb run fills them per agent and auth mode: the model API hosts for the
agent phase, plus Harbor's install hosts for the setup baseline when the
agent is not preinstalled in the image.
build_task(record, rubrics, task_src, out, *, options, dataset_version, dataset_digest, template)
Write one Harbor task directory.
build_tasks(dataset_dir, out, *, options, task_ids=None, template_path=None)
Build Harbor tasks for a dataset and write tdb-build.json.
default_template()
The packaged prompt template (copy of the intake system prompt).
render_instruction(record, document, template)
Agent instruction: template, sandbox note, question block, outputs, document.
render_task_toml(record, *, options, dataset_version, dataset_digest, task_digest, template_sha256)
Harbor task.toml for one task.
render_test_sh(grader_source, version)
Verifier test.sh for the chosen grader source.
render_tests_dockerfile(options)
Verifier image definition: shared image plus /tests files.