Anteproof › Calibration

Calibration record

How much to trust a number before you use it. Every forecast we have ever issued is scored against what actually happened; this page is that scorecard, by question family, with nothing left out.

Read from the record at 20 September 2026, 19:12 UTC.

The launch gate

Four statistical criteria, all required, evaluated on point-in-time backtests ("pastcasts") of the forecaster version that is live. Passing the gate is what lets us market a family; failing it is published just the same. A verdict covers one horizon and one domain, both named on the card: when the horizon selected below has had no gate run on it, this says so rather than showing another horizon's verdict.

GATE PASSED evaluated Sep 6, 2026 · forecaster m5-v5 · domain rollup_threshold · horizon 7 days
CriterionValueRule
Resolved forecastsn = 900≥ 150met
Brier skill over 50%-baseline0.112≥ 0.100met
Skill significant (bootstrap upper 95% CI)-0.014< 0met
Bias, calibration-in-the-large+0.50|z| < 2met
Bias, conditional (Spiegelhalter)+0.15|Z| < 2met
Brier 0.222 vs baseline 0.250 pool base rate 43%

Reliability

When we say 30%, does it happen about 30% of the time? Each point is one probability bin: what we said on the horizontal axis, what happened on the vertical. Perfect calibration sits on the diagonal. Point size shows how many forecasts are in the bin.

Horizon
Pastcast pool · up to 7 days (the gate's horizon)n 900Brier 0.222skill 0.112vs base rate 0.109
0%0%25%25%50%50%75%75%100%100%what we saidwhat happenedsaid 4% · happened 3% · n=60said 14% · happened 6% · n=17said 26% · happened 27% · n=33said 36% · happened 31% · n=119said 45% · happened 50% · n=254said 54% · happened 58% · n=249said 64% · happened 56% · n=127said 73% · happened 72% · n=36said 83% · happened 100% · n=4said 94% · happened 100% · n=1
Live track · up to 7 days (the gate's horizon)n 27Brier 0.134skill 0.466vs base rate 0.447TOO FEW TO JUDGE
0%0%25%25%50%50%75%75%100%100%what we saidwhat happenedsaid 5% · happened 0% · n=11said 11% · happened 50% · n=2said 22% · happened 0% · n=1said 36% · happened 100% · n=3said 47% · happened 0% · n=1said 53% · happened 50% · n=2said 60% · happened 100% · n=1said 72% · happened 50% · n=2said 87% · happened 100% · n=2said 92% · happened 100% · n=2

By question family

Brier score is the mean squared error of a probability: 0 is perfect. Two baselines: skill is against always saying 50% (Brier 0.25); vs base rate is against always saying the family's own observed frequency, a harder bar and the one that matters when a family's outcomes are lopsided. Families with fewer than 30 resolved forecasts are shown but not judged.

Pastcast pool forecaster m5-v5, scheme forecast_referenced · up to 7 days (the gate's horizon)

FamilynBase rateMean forecastBrierSkillvs base rateBias z
GDELT news tone (daily) 864 47% 46% 0.223 0.109 0.106 +0.56
Hacker News points (daily)n < 30 28 57% 52% 0.232 0.071 0.052 +0.60
Wikipedia page views (daily)n < 30 8 13% 40% 0.105 0.581 0.042 -1.79
All families 900 47% 46% 0.222 0.112 0.109 +0.50

Live track forecasts issued in real time, all versions

FamilynBase rateMean forecastBrierSkillvs base rateBias z
Prediction-market questionsn < 30 25 36% 32% 0.129 0.485 0.442 +0.65
manifoldn < 30 2 100% 65% 0.195 0.218 +1.26
All familiesn < 30 27 41% 34% 0.134 0.466 0.447 +1.00

Against prediction markets

A separate test from the launch gate above, which is about our own question families. Here we forecast questions that a prediction market was already pricing, and score ourselves against the crowd's price at the same moment. The model was shown that price when it forecast.

Manifold Markets. On 315 resolved Manifold Markets questions, the crowd's Brier score at the same moment was 0.172; ours was 0.186. On questions a liquid market already prices, we do not beat the market.

Polymarket. On 16 resolved Polymarket questions, the crowd's Brier score at the same moment was 0.037; ours was 0.050. Too few to conclude anything; it is here because we said we would publish it either way.

manifold (earlier wording, retired). On 383 resolved manifold (earlier wording, retired) questions, the crowd's Brier score at the same moment was 0.155; ours was 0.199. On questions a liquid market already prices, we do not beat the market.

Our forecasts are for the questions markets don't cover. Brier score: mean squared error of a probability, 0 is perfect. Forecasts issued with the market price visible to the model, scored on the market's price at the forecast's own cutoff.

Every gate evaluation

The complete history, verbatim. Sample size restarts whenever the question scheme or the forecaster version changes, because pooling regimes was measured to flatter the numbers.

How to read this

Pastcast. A forecast issued as if on a past date, with retrieval, thresholds and the model's own knowledge clamped to what was available then, then scored on the real outcome. It is how a young forecaster earns a sample size before it has lived long enough to accumulate one. Every pastcast is stored with the same append-only discipline as a live forecast.

Live track. Forecasts issued in real time and timestamped before the outcome was knowable. Small today; it grows every week and will eventually replace the pastcast pool as the record that matters.

The criteria. At least 150 resolved forecasts; a Brier skill score of at least 0.10 over the always-50% baseline; that skill statistically significant under a block bootstrap (blocks by metric and date, so clustered questions do not count as independent evidence); and two bias tests, calibration-in-the-large and Spiegelhalter's conditional test, both inside |z| < 2.

Two baselines. Always-50% is the classical Brier baseline and the one the gate uses. Always-base-rate ("climatology") is stricter: it credits nothing for knowing that a family resolves YES 46% of the time, only for telling questions apart. We compute it on the pool's own observed base rate, which flatters the baseline and understates our skill, the conservative direction. A family that beats 50% but not its base rate has learned the base rate and little else.

Horizon. Days from a forecast's as-of date (its issue date, for a live forecast) to resolution, bucketed. Skill at seven days says nothing about skill at ninety, so the selector above scopes every table on this page; buckets with no resolved forecasts are greyed. The page opens on the horizon the launch gate judges, because the archive also holds pastcasts at other horizons and pooling them was measured to flatter the headline; "all" pools them on purpose.

Against prediction markets. Questions that a market was already pricing when we forecast them, scored against the crowd's price at that same moment. The model sees that price in its context, so this measures whether we add anything to it. On liquid markets we do not, and we say so; the record is here because we said we would publish it either way. It is not part of the launch gate, which judges our own question families.

What this does not claim. Calibration on one family does not transfer to another, and a passing gate on backtests is not a passing gate on the live track. That is why both are shown, side by side, and why the live numbers carry a flag until they are large enough to mean something.