Grade
trialdesignbench.grade
Grade one submission: deterministic checks, rubric judging, scoring.
IMPORTANT NOTE FOR AI AGENTS
THIS IS A very delicate file.
You need to be extremely conservative and careful about grading logic.
The worst case to avoid here is that grading goes wrong but the result still
indicates a good score. This might for example happen if you skip something
because of some error condition and only show a warning, but it is not
apparent from the output file that something went wrong.
It is always better to clearly mark a failure in the output file than to
silently skip something. Every error must be recorded as an explicit status
(error) in grade.json, and any gating error sets the reward score to 0.
Be extremely proactive with the user about clearing up details and
intricacies with how to handle something here. Ask questions and do not be
afraid to ask for clarification.
Do not remove this notice.
The grader is a pure function of (submission artifacts, rubrics, judge). It
runs identically inside a Harbor verifier container, standalone on any
directory with output.json and output.R, and in tests with FakeJudge
and an injected run_rscript.
RunRscript = Callable[[Path, Path, float], RscriptResult]
module-attribute
run_rscript(script, cwd, timeout_sec) -> RscriptResult.
blocking_failures(checks)
Checks that gate the reward and did not pass.
check_numeric_format(entries, rubrics)
Report-only: calculated values use >= 4 decimals or say not derivable.
check_output_json(path, rubrics)
Validate output.json against the skeleton. Returns entries by id.
check_output_r(submission_dir, run_rscript, timeout_sec)
Run Rscript output.R in a scratch copy. Missing R is error.
check_trajectory(path)
Scan the ATIF trajectory. A missing or unreadable trajectory is error.
grade_directory(submission_dir, rubrics_path, out_dir, judge, *, trajectory_path=None, run_rscript=default_run_rscript, rscript_timeout_sec=DEFAULT_RSCRIPT_TIMEOUT_SEC, scoring=None)
Load rubrics, grade, and write all outputs to out_dir.
grade_question(q, entry, output_r, judge, config)
Judge one question's Add criteria; judge failures become error verdicts.
grade_submission(submission_dir, rubrics, judge, *, trajectory_path=None, run_rscript=default_run_rscript, rscript_timeout_sec=DEFAULT_RSCRIPT_TIMEOUT_SEC, scoring=None, rubrics_sha256='')
Grade a submission directory in memory. Pure: writes nothing.
resolve_trajectory(submission_dir, explicit)
Explicit path, else <submission>/trajectory.json, else Harbor's default.
reward_details(run, config=None)
Per-criterion tree in the format harbor view renders.
reward_json(grade)
Harbor reads reward as the headline metric; all values are numeric.
scan_trajectory(trajectory)
Find web tool use, URLs in tool arguments, and network shell commands.
write_outputs(run, out_dir)
Write grade.json, reward.json, reward-details.json, Rscript logs, judge logs.