What is Edge Score?
A descriptive composite that ranks a Polymarket wallet's past behavior against a fixed reference cohort. It is not a measure of skill, and it failed its own forward test. Open methodology, public data, scorable on any wallet for free.
The one-paragraph version
Edge Score is a descriptive composite that ranks a Polymarket wallet's past behavior against a fixed reference cohort. It is not a measure of skill: it failed its own forward test, so a high score says the record looks like the profitable part of that cohort (8,656 wallets) looked, not that the wallet will profit next. Three behaviors go into it: posture (the calibration pillar: how the wallet's entry prices compared with a trivial base-rate forecast), conviction (how much of the result came from the single largest event), and discipline (how many positions the wallet resolved, with fewer scoring higher). The raw composite is mapped to a 0-100 percentile rank against the cohort. The skill-or-luck read on Convexly is a different statistic: the realized entry edge with its interval, covered in the four-state verdict.
The numbers behind it
The three pillars are z-scored: posture is the standardized negation of baseline-adjusted Brier score, conviction the standardized PnL concentration in the wallet's single largest event, and discipline the standardized log of resolved position count with a negative sign in the composite. The coefficients were originally fit by ordinary least squares against signed log realized PnL on the reference cohort of 8,656 Polymarket wallets, sampled April 15-16 2026 (the V3b fit), and are never refit at inference time; on 2026-07-13 the constants were refit after a bounded-input correction, so the V3b fit statistics in this paragraph describe that original fit, not current scores (the dated method-change note is at /methodology). In the V3b fit, out-of-fold Spearman rank correlation with signed log realized PnL (in-sample) was +0.514, against +0.148 for a Brier-only baseline. That +0.514 was a cross-sectional, in-sample association, not established forward skill: the per-wallet temporal holdout we committed to in advance did not clear the +0.30 threshold it was filed against. The frozen-coefficient leg held an out-of-sample Spearman of +0.111 (95% CI 0.05 to 0.18) with forward PnL, and the refit leg landed at -0.082 (95% CI -0.149 to -0.006), below zero. Both legs are below the filed bar.
Why three pillars and not just one
If calibration were the dominant skill on prediction markets, a trader's Brier score alone would predict their profit rank cleanly. It does not. Across the full 8,656-wallet Polymarket cohort, the Spearman rank correlation between raw Brier score and realized PnL is only +0.148. The empirical story is that prediction-market profit is fat-tailed (Hill tail index = 1.28; below the alpha = 2 threshold above which OLS variance is well-behaved), and a few large concentrated positions dominate realized PnL.
Adding conviction (concentration) and discipline (position count) lifted the in-sample out-of-fold Spearman rank correlation with PnL to +0.514 in the V3b fit; per the 2026-07-13 constants refit, that figure describes the original fit, not current scores. The intuition: a trader who is well-calibrated but spreads tiny bets across many markets does not capture much of the available edge; a trader who concentrates the right way at the right times tends to outperform. Edge Score does not tell you HOW to find the right concentration; it tells you what the historical pattern of the profitable cohort looks like, and how a given wallet ranks against that pattern.
The three pillars in plain English
Posture (calibration)
Posture is the standardized negation of baseline-adjusted Brier. Baseline-adjusted Brier is observed Brier minus the wallet's own marginal-frequency Brier (the trivial always-predict-the-base-rate baseline). The coefficient on this pillar in the original V3b fit was +0.79 (constants refit 2026-07-13; dated note at /methodology). Higher posture means worse calibration relative to the baseline-trivial alternative, which on this cohort empirically aligns with higher realized profit. The pillar does not measure forecasting accuracy in the traditional sense; it measures the sign-aligned contribution of calibration to PnL on this specific cohort. The renaming from "calibration" to "posture" in the V1 paper preserves the measured effect without overclaiming what the pillar tracks.
Conviction (concentration)
Conviction is the standardized share of total realized PnL attributable to the wallet's single largest event. The coefficient on this pillar in the original V3b fit was +2.72, the largest of the three. Higher conviction means a more barbell-concentrated profit profile: most of the wallet's return comes from one event. On the training cohort, the wallets that compound the most are not the ones that distribute risk evenly across the catalog; they are the ones that concentrate when the conviction trade appears.
Discipline (position count)
Discipline is the standardized log1p of resolved position count, with a negative sign in the composite. The coefficient on this pillar in the original V3b fit was -1.15. Higher discipline (in the composite-contribution sense) corresponds to fewer resolved positions: the most profitable wallets on the training cohort hold fewer, larger positions. A trader who places hundreds of small bets across the catalog tends to score lower on Edge Score even with comparable calibration.
What the score number actually means
Edge Score is on a 0-100 percentile scale by construction. A score of 50 is exactly the cohort median. A score of 90 means the wallet is in the top decile of the frozen reference cohort by composite ranking. The published board's top score moves with each daily rescore, so read it from the board rather than from this page. Below 30 means the composite places the wallet in the bottom third by skill ranking, regardless of where their realized PnL lands.
Two important caveats. First, the percentile is computed against the frozen training cohort, not against any new wallet population the analyzer encounters. A new wallet that is materially different from the training cohort (e.g., a wallet that only bets on one category, or one that has very few resolved positions) is being scored by extrapolation. Second, the composite ranks wallets cross-sectionally; it does not bound expected returns for any individual wallet. Realized PnL on Polymarket is fat-tailed and individual outcomes vary widely.
What Edge Score does NOT do
Edge Score does not predict whether a particular wallet will profit on the next market. It does not separate skill from luck on a single wallet's realized PnL history (the per-wallet temporal holdout that addresses this is covered in the V1.5 follow-up paper). It does not transfer cleanly across venues: per the V1-M paper, the fitted coefficients diverge materially between Polymarket and Manifold, with the discipline pillar flipping sign at permutation p = 0.0001. And it does not substitute for category-specific calibration analysis, time-period analysis, or position-sizing diagnostics that depend on bankroll context.
Where the methodology lives
The V1 methodology paper (frozen coefficients, full validation suite, Fama-French bootstrap null at 10,000 permutations) is at /research/edge-score-methodology-v1. The V1-M cross-venue extension (15,106-user Manifold cohort, sweepcash within-user paired comparison plus the 2026-05-04 politics-excluded sensitivity update) is at /research/edge-score-methodology-v1m. The V1.5 deferred experiments (per-wallet temporal holdout + per-quarter Information Coefficient stability) were filed ex-ante before any analysis ran. The data bundle for that paper's Manifold half, which held 15,106 aggregated user records, was withheld on 2026-09-04 pending Manifold's written position on data rights; the stdlib-only Python script that re-runs the Manifold analysis is still published. Corrected 2026-08-10: that bundle was Manifold only. The Polymarket cohort Edge Score is fitted on has never been published, so no number on this page is currently reproducible by you from a published file.
Score a wallet
Paste any Polymarket wallet address at the analyzer. Your first check needs no signup, no signature, and no private key; more wallets are free with an account. Reads Polymarket public data only.
Convexly publishes new methodology research roughly every 6-8 weeks plus the /learn series on a rolling cadence. Get the next paper in your inbox when it ships:
Frequently asked
What is Edge Score?
How is Edge Score calculated?
What is a good Edge Score?
Why three pillars instead of just calibration?
Is Edge Score the same as PnL or rank by realized profit?
Where is the methodology published?
Can I score my own wallet?
What does Edge Score NOT do?
Related explainers
- /learn/brier-score: what a Brier score actually measures, calibration, baseline-adjusted Brier (skill-Brier), and why calibration alone barely predicts profit on Polymarket