Skip to content
Convexly
Learn

What is a Brier score, and what is calibration?

Calibration asks whether the probabilities you stated match how often the events happened. The Brier score is the number that measures it: the mean squared error between what you said would happen and what happened. Lower is better, 0 is perfect, a coin-flip forecaster scores 0.25. The formula, a worked example, the baseline-adjusted variant that makes it comparable across traders, and why it barely predicts profit on Polymarket.

The answer first

The Brier score measures probability-forecast accuracy as the mean squared error between the stated probability and the realized outcome, coded 0 or 1:

BS = (1/N) · ∑ (pi − oi)2

The range is [0, 1] for binary outcomes. A perfect forecast contributes 0; the worst possible forecast contributes 1; a forecaster who always says 50% scores exactly 0.25 on any binary sequence. It was introduced by Glenn Brier in 1950 for weather forecasting and remains the standard scalar accuracy metric for probability forecasts, including on prediction markets where every trade is an implied probability statement.

What calibration means

Calibration is the property of a probability forecaster whose stated probabilities match the long-run frequency of the predicted outcome. A forecaster who says 70% on a long series of events is well-calibrated if the event actually resolves yes about 70% of the time. On a prediction market every entry price is an implicit forecast (paying 70 cents says 70%), so a wallet's calibration is checked band by band: group its entries by price, then compare the average price paid in each band with the share that actually resolved true. The Brier score compresses that comparison into one number. It became the standard scalar measure of forecasting accuracy after the Murphy (1973) decomposition split it into reliability, resolution, and uncertainty components, and Tetlock's Good Judgment Project used it as the primary outcome measure in the IARPA forecasting tournament that identified the superforecaster cohort.

Worked example

Take three resolved forecasts. You said 80% on an event that resolved yes, 60% on an event that resolved no, and 90% on an event that resolved yes:

(0.8 − 1)2 = 0.04
(0.6 − 0)2 = 0.36
(0.9 − 1)2 = 0.01
BS = (0.04 + 0.36 + 0.01) / 3 = 0.137

Two things to notice. The one miss (60% on an event that did not happen) contributes nine times as much error as the 80% hit, because the error is squared. And 0.137 is meaningless on its own: it only becomes interpretable against a baseline, which is what the next section fixes.

Reference values

On the Good Judgment Project's IARPA geopolitical tournament, the median forecaster scored around 0.20 and the top-2% superforecaster cohort around 0.16. US National Weather Service next-day precipitation probabilities run around 0.10. Prediction markets sit between those extremes, but the absolute number depends heavily on the category mix a trader bets on: a wallet that lives in 50/50 sports toss-ups faces a structurally higher Brier floor than one betting extreme base-rate markets. Raw cross-trader Brier comparisons are therefore not meaningful without an adjustment.

Calibration vs resolution (the Murphy decomposition)

The Brier score decomposes into three components (Murphy 1973):

BS = Reliability − Resolution + Uncertainty

Reliability measures how close the forecaster's stated probabilities are to the realized base rates within each forecast bin. Lower reliability is better. It is what most non-technical readers mean by calibration. Resolution measures how varied the forecaster's predictions are across outcome categories: how well the forecaster discriminates events that did happen from events that did not. Higher resolution is better. Uncertainty is the irreducible variance of the outcome itself; it does not depend on the forecaster.

A forecaster who always predicts 50% on a 50/50 domain is perfectly calibrated (zero reliability error) but has zero resolution. A forecaster who always predicts the empirical base rate is perfectly calibrated on the marginal but adds no information. The forecasters who look good on a Brier ranking are those with good reliability AND high resolution.

Skill-Brier: the baseline-adjusted version

Convexly computes skill-Brier: observed Brier minus the Brier a trivial always-predict-the-base-rate forecaster would have scored on the same set of events. Negative means the wallet beats the trivial baseline; positive means it is worse. This aligns with how skill scores are constructed in the weather forecasting literature (the Brier Skill Score), and it is the input to the posture pillar of Edge Score.

Posture: why the Edge Score pillar is not called calibration

In the original V3b fit the posture pillar was the standardized negation of skill-Brier (z(-skill_brier)) with an OLS coefficient of +0.79; the constants were refit on 2026-07-13, so the live coefficients differ (dated method-change note at /methodology). That sign rewards higher standardized values of negative skill-Brier, which corresponds to worse calibration relative to the marginal-frequency baseline. Labeling that pillar "calibration" would have been misleading: the pillar does not reward forecasting accuracy in the traditional sense. Renaming it posture in the V1 paper aligned the label with the direction of the effect, without overclaiming what the pillar tracks. The reasoning is set out at /research/why-posture-not-calibration.

What a Brier score does NOT tell you

A good Brier score does not imply a profitable trader. Across the 8,656-wallet Polymarket cohort in the V1 study, the Spearman rank correlation between Brier score and realized PnL is only +0.148. In Cohen's effect-size convention that is a small effect: real, but not a dominant signal.

The reason is the shape of the profit distribution. Polymarket realized PnL is fat-tailed: the Hill tail index estimated on the cohort is 1.28, below the alpha = 2 threshold above which sample variance is well-behaved. In a distribution like this the realized rank is dominated by a small number of very large positions; the per-trade accuracy of the forecaster matters less than the accuracy-weighted-by-stake of the few trades where the forecaster concentrated. A wallet that concentrated and was right once can outrank a wallet that was accurate across hundreds of small positions, and the profit figure does not distinguish the two. That is the empirical motivation for Edge Score having three pillars instead of one.

Three more things calibration does not do. It does not bound expected PnL: a perfectly calibrated forecaster who takes no position makes nothing. It does not separate skill from luck on a single wallet's history; that is the job of the realized entry edge and its bootstrap interval, the statistic behind the four-state verdict. And a score fitted on it does not port across venues: per the V1-M cross-venue paper, the fitted coefficients did not transfer between Polymarket and Manifold. The lesson-6 worked example, with the calibration bands and per-category splits the analyzer renders, is /learn/calibration-and-brier.

Where the methodology lives

The V1 methodology paper (full validation suite, Fama-French bootstrap null at 10,000 permutations, derivation of skill-Brier and the posture pillar coefficient) is at /research/edge-score-methodology-v1. The cross-venue extension is at /research/edge-score-methodology-v1m. The cohort study behind the +0.148 figure is at /research/polymarket-10k-wallet-study. An earlier top-100 calibration audit was withdrawn on 9 August 2026 and its findings should not be cited; the withdrawal notice is at /research/polymarket-whale-audit. Code and reproduction scripts are held in the Convexly repository, which is private, so they are available on request rather than by download.

Check a wallet's Brier score

Paste any Polymarket wallet address at the analyzer to see its raw Brier, skill-Brier, and where that places it against the 8,656-wallet reference cohort. First check free, no signup, Polymarket public data only.

Convexly publishes new methodology research roughly every 6-8 weeks plus the /learn series on a rolling cadence. Get the next paper in your inbox when it ships:

Frequently asked

What is calibration in forecasting?
Calibration is the property of a probability forecaster whose stated probabilities match the long-run frequency of the predicted outcome. A perfectly calibrated forecaster who says 70% across many forecasts is correct on 70% of those forecasts. The Brier score is the canonical scalar measure of how close a forecaster's stated probabilities come to the realized outcomes; lower is better.
What is the Brier score formula?
For N binary forecasts, the Brier score is the average of (p - o) squared, where p is the stated probability and o is the realized outcome coded 0 or 1. A perfect forecast contributes 0, the worst possible forecast contributes 1, and a forecaster who always says 0.5 scores exactly 0.25 on any binary sequence.
What is a good Brier score?
It depends on the domain. On the Good Judgment Project's geopolitical tournament the median forecaster scored around 0.20 and the top-2% superforecaster cohort around 0.16. US National Weather Service next-day rain probabilities run around 0.10. On prediction markets the absolute number depends heavily on the category mix, which is why Convexly reports the baseline-adjusted version (skill-Brier) rather than comparing raw Brier scores across wallets.
Is a lower Brier score better?
Yes. The Brier score is an error measure: 0 is a perfect record, 1 is the worst possible record, and 0.25 is the always-say-50% baseline on binary outcomes. Lower is always better.
Does a good Brier score mean a profitable trader?
Not by itself. Across the 8,656-wallet Polymarket cohort in the Convexly V1 study, the Spearman rank correlation between Brier score and realized PnL is only +0.148. Profit on prediction markets is fat-tailed and dominated by a few large concentrated positions, so per-forecast accuracy is one input among several, not the whole story.
What is skill-Brier?
Skill-Brier is observed Brier minus the wallet's own marginal-frequency Brier: the score a trivial forecaster who always predicts the base rate of resolution would have earned on the same events. Negative skill-Brier means the wallet beats that trivial baseline. The adjustment makes scores comparable across traders who bet on markets with very different base rates, and it is the input to the posture pillar of Edge Score.
Why is calibration a weak predictor of profit on Polymarket?
Across the 8,656-wallet Polymarket cohort, Spearman rank correlation between raw Brier score and realized PnL is only +0.148. The reason is that Polymarket PnL is fat-tailed (Hill tail index = 1.28; below the alpha = 2 threshold above which OLS variance is well-behaved), so a few large concentrated positions dominate realized profit. A trader who is well-calibrated but spreads tiny bets across many markets does not capture much of the available edge, and a trader who concentrates can post a large profit figure without the record saying anything about forecasting accuracy.
What is the difference between calibration and resolution?
Calibration is the question: when the forecaster says 70%, does the event happen 70% of the time on average? Resolution is the question: does the forecaster's set of probabilities discriminate between events that did happen and events that did not? A forecaster who always says 50% is perfectly calibrated on a 50% base-rate domain but has zero resolution (no information). Brier score combines both via the Murphy decomposition: BS = Reliability - Resolution + Uncertainty.
Why does Convexly call the pillar posture instead of calibration?
Because on the Polymarket training cohort, the OLS coefficient on z(-skill_brier) in the original V3b fit was +0.79 (constants refit 2026-07-13, so the live coefficients differ; dated method-change note at /methodology), meaning the composite rewarded higher standardized negative skill-Brier, which is worse calibration relative to the marginal-frequency baseline. Labeling a pillar 'calibration' while assigning it a coefficient that rewards worse calibration would have been misleading. Posture is the term the V1 paper adopted for the sign-aligned component of skill-Brier in the composite, without overclaiming that the pillar tracks forecasting accuracy in the traditional sense.
Where can I see the Brier score of a Polymarket wallet?
Paste the wallet address at /tools/polymarket-wallet-analyzer. The free analyzer reports the wallet's raw Brier score, its skill-Brier (baseline-adjusted), and where that places it against an 8,656-wallet reference cohort. It reads Polymarket public data only. Your first check needs no signup or signature; more wallets are free with an account.

Related explainers

Related reading