Runs

One registered. None completed yet.

  • RewardBench 2 — independent re-judging

    Status: pre-registered. Three validators, blind to the published label and to each other, re-judge a stratified sample of 210 items from the Allen Institute's RewardBench 2 test split. Registered 2026-08-19, amended before any judgment, published 2026-08-29. The run has not started; no judgment exists. Read the pre-registration →

What a proof run has to contain

Every run on this page is held to the same five parts, fixed before it begins. The dataset, pinned to an exact revision and licence. The sample, with its size and its stratification stated and the size labelled honestly as either derived from a power analysis or fixed as a design constant. The hypotheses, each paired with the result that would falsify it. The statistics, named in advance — which agreement coefficient, which interval method, which bootstrap seed — so that the analysis shown afterwards is checkably the analysis that was planned. And the stopping rule: the run ends at the registered sample size, not when the number looks good.

The registered text carries a content hash you can recompute yourself, and the finished result is published next to it with the hash still matching. If a design has to change after registration, the change is dated and logged in the registration itself rather than applied quietly.

This is the same discipline as our sample verification bundle, applied to a public benchmark instead of a review batch: the raw records, the named methods, and a script that recomputes every figure from scratch. How we report agreement, and why the interval matters →