Skip to content

Judge

trialdesignbench.judge

Rubric judges.

A judge maps (question, criteria, agent artifacts) to one verdict per criterion. Judges never see other questions' rubrics. AnthropicJudge is the default; FakeJudge gives deterministic verdicts for tests and dry runs.

The anthropic SDK is imported lazily so the core package and conversion-only workflows never need it.

AnthropicJudge

Judge backed by the Anthropic Messages API with structured JSON output.

All criteria of one question are batched into a single call. With votes > 1, the call is repeated and verdicts are majority-voted. Any call that still fails after retries marks every criterion of that question error.

FakeJudge

Deterministic judge for tests.

verdicts maps criterion ids to verdicts; rule computes a verdict from (criterion, artifacts) for anything not in the mapping; otherwise default is used.

JudgeArtifacts dataclass

What the judge may see for one question.

build_user_message(question, criteria, artifacts)

Judge input for one question. Never includes other questions.

error_results(question, criteria, message)

Mark every criterion of a question as error.

judge_prompt_sha256()

Hash identifying the judge prompt and response contract.

supports_temperature(model)

Whether the judge should send an explicit temperature for this model.