Skip to content

Schema

trialdesignbench.schema

Versioned data models shared by every TrialDesignBench stage.

All models are frozen and reject unknown fields, so a file written by one version of the package is either read back exactly or rejected loudly. Each model carries a schema_version so on-disk artifacts are self-describing.

TaskType = Literal['reproduction'] module-attribute

Kind of benchmark task. Only reproduction exists today; design generation (Task 2) will extend this literal.

CheckResult

Bases: _Model

One deterministic check. error means the check could not run.

DatasetManifest

Bases: _Model

Canonical dataset index (dataset.json).

DocumentRef

Bases: _Model

Reference to the agent-visible source document (protocol or SAP).

Question

Bases: _Model

One evaluation question. The agent sees skeleton() only.

skeleton()

Agent-visible prompt entry with every answer field nulled.

Mirrors extract_prompts.py from the intake tooling so the agent contract matches what reviewers curated.

ReportSummary

Bases: _Model

Aggregated benchmark results (report.json).

RubricSet

Bases: _Model

Hidden grading spec (rubrics.json). Never shown to the agent.

RunManifest

Bases: _Model

Provenance for one tdb run invocation (tdb-run.json).

Submission

Bases: _Model

Artifacts found in a submission directory.

SubmissionSource

Bases: _Model

Which curated intake submission a task was imported from.

TaskGrade

Bases: _Model

Full grading result (grade.json).

TaskRecord

Bases: _Model

Agent-visible task definition (<task_id>/task.json).