Routing evaluation · Terminal-Bench 2.1 + 3

Which Terminal-Bench tasks
benefit from Babysitter?

163 tasks scored against the babysitter-vs-vanilla rubric · tb2.1 ffccbe05 · tb3 d7ff2f36 · run 01KZNWNZGBF2DJXHKG8XCC790C + 01KZP295QJFKTZ7MPZ4J656B19 · 2026-08-10 14:52 UTC

01 · Summary

How the corpus splits

Every task is verifiable, sandboxed and headless, so of the rubric dimensions are constant across this corpus and are pinned rather than scored. net_live renormalizes over the that vary, preserving the base rubric's scale so its pre-registered thresholds (babysitter ≥ 20, vanilla ≤ -15) still apply.

What the criteria mean

Each dimension is scored 0–3 with cited evidence. Benefit dimensions push toward babysitter, cost dimensions push toward vanilla; the weight is that dimension's share of its side of the scale.

Verdict split by benchmark

Diverging stacked bar, centred on the neutral “borderline” band.

Show data table

Distribution of net_live

Each bar is a 10-point bin, stacked by the verdict its tasks received. Bins holding more than one colour are the panel at work: net_live here is the first judge's score, while the verdict is the majority of three — so a panelled task can land on the other side of its own bin.

Show data table

Does net_live just re-measure task length?

net_live against the task's expert time estimate (log scale). Spearman ρ = . If this were ≈1.0 the rubric would be an expensive proxy for a number already in task.toml.

Which dimensions actually discriminate?

Score distribution per live dimension. A dimension whose mass sits in one bucket adds a constant, not information.

Show data table

Babysitter share by category

Share of each category's tasks routed to babysitter. Categories with n < 3 are grouped as “Other”. The two benchmarks ship different category vocabularies (security vs Security, machine-learning vs ML); they are left unmerged rather than mapped by guesswork.

Show data table
02 · Details

Every judgment

All 163 judgments. Click a task to see the verbatim instruction the judge scored from, its per-dimension scores with the evidence cited for each, the counterfactual failure mode, and the panel votes where a panel ran. Search covers instruction text as well as evidence.