Calibration record
Every number below is computed weekly from live resolutions only, with bootstrap confidence intervals (1000 resamples). When an interval crosses zero we say so; nothing here is a claim until it separates.
snapshot 2026-09-27 05:05 UTC · 205 resolved · 429 scored forecasts · 755 shadow scores · 65 abstained
Brier by verdict class
the core experiment: does the verdict class carry information about error?| Verdict | n | Brier | Bootstrap CI |
|---|---|---|---|
| CAUTION | 89 | 0.228 | 95% CI [0.187, 0.267] |
| CONSTRAINED | 340 | 0.280 | 95% CI [0.256, 0.305] |
ABSTAIN rows never publish a probability, so they cannot appear here; their value is measured below, through the shadow.
Value of abstention
an ungated shadow model forecasts every question; we compare its error on the questions we abstained fromshadow on published set
0.253
n=498
shadow on abstained set
0.244
n=257
abstention value (diff)
-0.009
95% CI [-0.042, 0.025]
Positive means the questions our gate refused were genuinely harder, even for an ungated model. The interval currently crosses zero: direction only, not yet a claim.
Layer vs shadow on the published set
same questions, same evidence pool; the only difference is the decision layerRisk-coverage
verdict classes are the natural thresholds of a selective forecaster| Publish up to | Coverage | Risk (Brier) | n |
|---|---|---|---|
| CONSTRAINED | 68.8% | 0.280 | 340 |
| CAUTION | 86.8% | 0.270 | 429 |
naive AURC 0.242 · with few points this is descriptive, not comparative