Schema
trialdesignbench.schema
Versioned data models shared by every TrialDesignBench stage.
All models are frozen and reject unknown fields, so a file written by one
version of the package is either read back exactly or rejected loudly.
Each model carries a schema_version so on-disk artifacts are
self-describing.
TaskType = Literal['reproduction']
module-attribute
Kind of benchmark task. Only reproduction exists today; design generation (Task 2) will extend this literal.
CheckResult
Bases: _Model
One deterministic check. error means the check could not run.
DatasetManifest
Bases: _Model
Canonical dataset index (dataset.json).
DocumentRef
Bases: _Model
Reference to the agent-visible source document (protocol or SAP).
Question
Bases: _Model
One evaluation question. The agent sees skeleton() only.
skeleton()
Agent-visible prompt entry with every answer field nulled.
Mirrors extract_prompts.py from the intake tooling so the agent
contract matches what reviewers curated.
ReportSummary
Bases: _Model
Aggregated benchmark results (report.json).
RubricSet
Bases: _Model
Hidden grading spec (rubrics.json). Never shown to the agent.
RunManifest
Bases: _Model
Provenance for one tdb run invocation (tdb-run.json).
Submission
Bases: _Model
Artifacts found in a submission directory.
SubmissionSource
Bases: _Model
Which curated intake submission a task was imported from.
TaskGrade
Bases: _Model
Full grading result (grade.json).
TaskRecord
Bases: _Model
Agent-visible task definition (<task_id>/task.json).