What is a Brier score, and what is calibration?
Calibration asks whether the probabilities you stated match how often the events happened. The Brier score is the number that measures it: the mean squared error between what you said would happen and what happened. Lower is better, 0 is perfect, a coin-flip forecaster scores 0.25. The formula, a worked example, the baseline-adjusted variant that makes it comparable across traders, and why it barely predicts profit on Polymarket.
The answer first
The Brier score measures probability-forecast accuracy as the mean squared error between the stated probability and the realized outcome, coded 0 or 1:
BS = (1/N) · ∑ (pi − oi)2
The range is [0, 1] for binary outcomes. A perfect forecast contributes 0; the worst possible forecast contributes 1; a forecaster who always says 50% scores exactly 0.25 on any binary sequence. It was introduced by Glenn Brier in 1950 for weather forecasting and remains the standard scalar accuracy metric for probability forecasts, including on prediction markets where every trade is an implied probability statement.
What calibration means
Calibration is the property of a probability forecaster whose stated probabilities match the long-run frequency of the predicted outcome. A forecaster who says 70% on a long series of events is well-calibrated if the event actually resolves yes about 70% of the time. On a prediction market every entry price is an implicit forecast (paying 70 cents says 70%), so a wallet's calibration is checked band by band: group its entries by price, then compare the average price paid in each band with the share that actually resolved true. The Brier score compresses that comparison into one number. It became the standard scalar measure of forecasting accuracy after the Murphy (1973) decomposition split it into reliability, resolution, and uncertainty components, and Tetlock's Good Judgment Project used it as the primary outcome measure in the IARPA forecasting tournament that identified the superforecaster cohort.
Worked example
Take three resolved forecasts. You said 80% on an event that resolved yes, 60% on an event that resolved no, and 90% on an event that resolved yes:
(0.8 − 1)2 = 0.04
(0.6 − 0)2 = 0.36
(0.9 − 1)2 = 0.01
BS = (0.04 + 0.36 + 0.01) / 3 = 0.137
Two things to notice. The one miss (60% on an event that did not happen) contributes nine times as much error as the 80% hit, because the error is squared. And 0.137 is meaningless on its own: it only becomes interpretable against a baseline, which is what the next section fixes.
Reference values
On the Good Judgment Project's IARPA geopolitical tournament, the median forecaster scored around 0.20 and the top-2% superforecaster cohort around 0.16. US National Weather Service next-day precipitation probabilities run around 0.10. Prediction markets sit between those extremes, but the absolute number depends heavily on the category mix a trader bets on: a wallet that lives in 50/50 sports toss-ups faces a structurally higher Brier floor than one betting extreme base-rate markets. Raw cross-trader Brier comparisons are therefore not meaningful without an adjustment.
Calibration vs resolution (the Murphy decomposition)
The Brier score decomposes into three components (Murphy 1973):
BS = Reliability − Resolution + Uncertainty
Reliability measures how close the forecaster's stated probabilities are to the realized base rates within each forecast bin. Lower reliability is better. It is what most non-technical readers mean by calibration. Resolution measures how varied the forecaster's predictions are across outcome categories: how well the forecaster discriminates events that did happen from events that did not. Higher resolution is better. Uncertainty is the irreducible variance of the outcome itself; it does not depend on the forecaster.
A forecaster who always predicts 50% on a 50/50 domain is perfectly calibrated (zero reliability error) but has zero resolution. A forecaster who always predicts the empirical base rate is perfectly calibrated on the marginal but adds no information. The forecasters who look good on a Brier ranking are those with good reliability AND high resolution.
Skill-Brier: the baseline-adjusted version
Convexly computes skill-Brier: observed Brier minus the Brier a trivial always-predict-the-base-rate forecaster would have scored on the same set of events. Negative means the wallet beats the trivial baseline; positive means it is worse. This aligns with how skill scores are constructed in the weather forecasting literature (the Brier Skill Score), and it is the input to the posture pillar of Edge Score.
Posture: why the Edge Score pillar is not called calibration
In the original V3b fit the posture pillar was the standardized negation of skill-Brier (z(-skill_brier)) with an OLS coefficient of +0.79; the constants were refit on 2026-07-13, so the live coefficients differ (dated method-change note at /methodology). That sign rewards higher standardized values of negative skill-Brier, which corresponds to worse calibration relative to the marginal-frequency baseline. Labeling that pillar "calibration" would have been misleading: the pillar does not reward forecasting accuracy in the traditional sense. Renaming it posture in the V1 paper aligned the label with the direction of the effect, without overclaiming what the pillar tracks. The reasoning is set out at /research/why-posture-not-calibration.
What a Brier score does NOT tell you
A good Brier score does not imply a profitable trader. Across the 8,656-wallet Polymarket cohort in the V1 study, the Spearman rank correlation between Brier score and realized PnL is only +0.148. In Cohen's effect-size convention that is a small effect: real, but not a dominant signal.
The reason is the shape of the profit distribution. Polymarket realized PnL is fat-tailed: the Hill tail index estimated on the cohort is 1.28, below the alpha = 2 threshold above which sample variance is well-behaved. In a distribution like this the realized rank is dominated by a small number of very large positions; the per-trade accuracy of the forecaster matters less than the accuracy-weighted-by-stake of the few trades where the forecaster concentrated. A wallet that concentrated and was right once can outrank a wallet that was accurate across hundreds of small positions, and the profit figure does not distinguish the two. That is the empirical motivation for Edge Score having three pillars instead of one.
Three more things calibration does not do. It does not bound expected PnL: a perfectly calibrated forecaster who takes no position makes nothing. It does not separate skill from luck on a single wallet's history; that is the job of the realized entry edge and its bootstrap interval, the statistic behind the four-state verdict. And a score fitted on it does not port across venues: per the V1-M cross-venue paper, the fitted coefficients did not transfer between Polymarket and Manifold. The lesson-6 worked example, with the calibration bands and per-category splits the analyzer renders, is /learn/calibration-and-brier.
Where the methodology lives
The V1 methodology paper (full validation suite, Fama-French bootstrap null at 10,000 permutations, derivation of skill-Brier and the posture pillar coefficient) is at /research/edge-score-methodology-v1. The cross-venue extension is at /research/edge-score-methodology-v1m. The cohort study behind the +0.148 figure is at /research/polymarket-10k-wallet-study. An earlier top-100 calibration audit was withdrawn on 9 August 2026 and its findings should not be cited; the withdrawal notice is at /research/polymarket-whale-audit. Code and reproduction scripts are held in the Convexly repository, which is private, so they are available on request rather than by download.
Check a wallet's Brier score
Paste any Polymarket wallet address at the analyzer to see its raw Brier, skill-Brier, and where that places it against the 8,656-wallet reference cohort. First check free, no signup, Polymarket public data only.
Convexly publishes new methodology research roughly every 6-8 weeks plus the /learn series on a rolling cadence. Get the next paper in your inbox when it ships:
Frequently asked
What is calibration in forecasting?
What is the Brier score formula?
What is a good Brier score?
Is a lower Brier score better?
Does a good Brier score mean a profitable trader?
What is skill-Brier?
Why is calibration a weak predictor of profit on Polymarket?
What is the difference between calibration and resolution?
Why does Convexly call the pillar posture instead of calibration?
Where can I see the Brier score of a Polymarket wallet?
Related explainers
- /learn/calibration-and-brier: lesson 6, the worked example with calibration bands and per-category splits
- /learn/realized-edge: the entry-price-based skill read Convexly pairs with a bootstrap interval
- /learn/edge-score: the composite that uses skill-Brier as one of three pillars