Judgment & behaviour
Calibration — being right about how right you are
A forecaster is calibrated when the things they call 70% likely happen about 70% of the time. It is measurable, it is trainable, and it is a different skill from being smart.
What calibration means
Take every forecast you have ever made at "70% likely". If about 70 out of 100 of them happened, you are calibrated at that level. If 40 happened, you are overconfident. If 95 happened, you are underconfident and should have said 95.
Notice what this does not ask. It does not ask whether you were right about any particular case. It asks whether your probabilities mean what they say.
That makes it one of the few things in investing you can actually score, and it is why weather forecasters — who make thousands of probabilistic statements that resolve quickly — are among the best-calibrated professionals measured.
The Brier score
Brier's rule is the standard measure and it is simple: for each forecast, take the difference between what you said and what happened (1 or 0), and square it. Average over all forecasts. Lower is better; 0 is perfect.
Squaring is what makes it work. It punishes confident errors far harder than timid ones, so you cannot improve your score by shouting, and it removes the incentive to hedge everything to 50% — a forecaster who says 50% to everything scores 0.25 and never beats anyone who knows anything.
A Brier score is only meaningful against a benchmark. The one that matters is the base rate: what you would have scored by predicting the historical frequency every time. Beating that is skill; failing to is not.
Two things the score contains
A Brier score decomposes, and the two parts are different skills.
Reliability is calibration proper — do your 70%s happen 70% of the time? This is trainable, and it is mostly a matter of feedback.
Resolution is whether you say anything useful — do you ever move away from the base rate? A perfectly calibrated forecaster who says the base rate to every question is honest and useless.
You want both, and improving one can cost the other. The forecaster who discovers they are overconfident and retreats to 55% on everything has improved reliability by destroying resolution.
How this connects to Pythia
Both halves of the product are scored this way, deliberately.
When you attach a numeric forecast to a thesis, it resolves at its horizon and your Brier score is computed over your own predictions. Once enough of them have graded, the decomposition appears too, so you can see whether you are miscalibrated or simply not saying much — it waits for a sample rather than computing reliability off four forecasts, for the reason in the limits below.
A resolved forecast cannot be deleted: a score over the ones you chose to keep would measure nothing. An unresolved one can be withdrawn, which is a different act — changing your mind before the fact is legitimate, and it is the only version of deleting a prediction the database permits.
And we hold our own scores to it. The validation surface publishes cohort outcomes against a benchmark with a walk-forward Brier and its Murphy decomposition, on the same rule, including the periods where it does not flatter us.
The limits of this idea
Calibration needs volume. Ten forecasts tell you almost nothing; the noise swamps the signal. Do not read a Brier score off a handful of resolved theses and conclude anything about yourself.
It needs resolution criteria fixed in advance. "The company does well" is not scoreable, and a claim whose meaning can drift will always resolve in your favour.
And it is not the same as making money. A well-calibrated investor who is calibrated about unimportant things, or who sizes positions badly, will lose to a sloppier forecaster with better judgement about what matters. Calibration is a check on your confidence, not a substitute for having something to be confident about.
Check yourself
5 questions. Nothing is recorded unless you are signed in, and nothing here affects anything else.
Sources
- Tetlock & Gardner, Superforecasting (opens in a new tab) — The Good Judgment tournament — what accurate forecasters do differently, and that it improves with feedback.
- Brier (1950), Verification of Forecasts Expressed in Terms of Probability (opens in a new tab) — The scoring rule, from weather forecasting, that makes calibration measurable.
- Kahneman, Thinking, Fast and Slow (opens in a new tab) — Why confidence and accuracy come apart, and why feedback is what closes the gap.