PnL vs Edge Score
Same cohort, two rankings. Left: top 20 by realized PnL (what Polymarket shows you). Right: top 20 by Edge Score, a descriptive behavioral profile, not a validated skill ranking: in its one forward test the frozen composite held an out-of-sample Spearman of +0.11 (95% CI [0.05, 0.18]) with forward PnL, below the +0.30 bar set in advance. The honest read of any single record is its realized-edge 95% interval against zero, which neither column shows.
Tap any wallet to expand its specific analysis: behavioral pillars, sample quality, and its realized-edge interval against zero where a frozen skill-versus-luck read exists.
Top 20 by realized PnL
What the Polymarket leaderboard shows
rank . wallet . edge . pnl . edge rank . tap to expand
Top 20 by Edge Score
Descriptive behavioral composite of the resolved record
rank . wallet . edge . pnl . pnl rank . tap to expand
What this artifact is saying
Across the full 8,656-wallet Polymarket cohort in our V1 methodology paper, the Spearman rank correlation between a wallet's Brier score (calibration) and its realized PnL is +0.148 (in-sample, descriptive). The top 1 percent of wallets by absolute PnL captured 36.2 percent of all signed profit on the platform. In a market that concentrated, a PnL leaderboard is heavily shaped by a small number of large outcomes.
Convexly's Edge Score is a composite of three behavioral pillars fitted against signed log PnL: posture, conviction, and discipline. The original V3b fit's in-sample out-of-fold Spearman with signed log PnL was +0.514, but that figure shares positions with the PnL it is scored against and its dominant input is PnL-derived, so it overstates predictive skill; the constants were also refit on 2026-07-13, so that fit statistic does not describe the current column (dated method-change note at /methodology). In the composite's one forward test, a per-wallet temporal holdout with the threshold set before the analysis ran, the frozen score held an out-of-sample Spearman of +0.11 (95% CI [0.05, 0.18]) with forward PnL. That did not clear the +0.30 bar. Treat the right-hand column as a descriptive behavioral profile, not a validated skill ranking.
The divergence you see above therefore shows that the two orderings disagree, not which one is right. Wallets in the left column but not the right have a PnL rank ahead of their composite rank; wallets in the right column but not the left have a composite rank ahead of their PnL rank. The honest read of any single record is whether its realized-edge 95% interval clears zero, which neither column displays. A wallet-by-wallet realized-edge scan of the published top-50 cohort is being prepared as a research note; see the top-50 skill scan in the research index.
The V1-M extension paper published 2026-04-22 adds a 15,106-user Manifold cohort and a paired-window sweepcash sensitivity analysis. The original public bundle reported median concentration 8.9 percentage points lower in the sweepcash window on 1,647 paired users, with concentration delta defined on n=333. A 2026-05-04 recovered-cohort rerun excluding political and election markets reports -9.2pp, 95% CI [-12.8, -3.6], Wilcoxon p=0.0021 on n=515 of 1,208 paired users with defined concentration delta. The concentration shift survives the sensitivity check, but causal attribution to sweepcash alone remains limited by the catalog confound.
Score any Polymarket wallet against this cohort
Paste any 0x address into the analyzer and see its three- pillar Edge Score in under 30 seconds. Your first wallet is free with no signup; more are free with an account. No wallet signature, ever.
Methodology: Edge Score (Convexly Research 2026), a frozen-coefficient behavioral composite; the dated method-change note is at /methodology. The composite is a descriptive behavioral profile, not a validated skill ranking. Cohort refreshes daily at 06:00 UTC via a GitHub Actions cron. Use this as a comparison view, not a wallet recommendation list.