163 tasks scored against the babysitter-vs-vanilla rubric ·
tb2.1 ffccbe05 · tb3 d7ff2f36 ·
run 01KZNWNZGBF2DJXHKG8XCC790C + 01KZP295QJFKTZ7MPZ4J656B19 · 2026-08-10 14:52 UTC
Every task is verifiable, sandboxed and headless, so of the
rubric dimensions are constant across this corpus and are pinned rather than
scored. net_live renormalizes over the that vary, preserving the
base rubric's scale so its pre-registered thresholds (babysitter ≥ 20, vanilla ≤ -15) still
apply.
Each dimension is scored 0–3 with cited evidence. Benefit dimensions push toward babysitter, cost dimensions push toward vanilla; the weight is that dimension's share of its side of the scale.
Diverging stacked bar, centred on the neutral “borderline” band.
Each bar is a 10-point bin, stacked by the verdict its tasks received. Bins holding more
than one colour are the panel at work: net_live here is the first judge's score, while the
verdict is the majority of three — so a panelled task can land on the other side of its own bin.
net_live against the task's expert time estimate (log scale). Spearman ρ = . If this were ≈1.0 the rubric would be an expensive proxy for a number already in task.toml.
Score distribution per live dimension. A dimension whose mass sits in one bucket adds a constant, not information.
Share of each category's tasks routed to babysitter. Categories with n < 3 are grouped
as “Other”. The two benchmarks ship different category vocabularies (security vs
Security, machine-learning vs ML); they are left unmerged rather
than mapped by guesswork.
All 163 judgments. Click a task to see the verbatim instruction the judge scored from, its per-dimension scores with the evidence cited for each, the counterfactual failure mode, and the panel votes where a panel ran. Search covers instruction text as well as evidence.