Judge
trialdesignbench.judge
Rubric judges.
A judge maps (question, criteria, agent artifacts) to one verdict per
criterion. Judges never see other questions' rubrics. AnthropicJudge is the
default; FakeJudge gives deterministic verdicts for tests and dry runs.
The anthropic SDK is imported lazily so the core package and conversion-only
workflows never need it.
AnthropicJudge
Judge backed by the Anthropic Messages API with structured JSON output.
All criteria of one question are batched into a single call. With
votes > 1, the call is repeated and verdicts are majority-voted. Any
call that still fails after retries marks every criterion of that
question error.
FakeJudge
Deterministic judge for tests.
verdicts maps criterion ids to verdicts; rule computes a verdict from
(criterion, artifacts) for anything not in the mapping; otherwise
default is used.
JudgeArtifacts
dataclass
What the judge may see for one question.
build_user_message(question, criteria, artifacts)
Judge input for one question. Never includes other questions.
error_results(question, criteria, message)
Mark every criterion of a question as error.
judge_prompt_sha256()
Hash identifying the judge prompt and response contract.
supports_temperature(model)
Whether the judge should send an explicit temperature for this model.