Skip to content
Convexly

Data hooks for journalists · April 2026

Three Polymarket findings

Three empirical findings from the Convexly V1 (8,656 Polymarket wallets) and V1-M (15,106 Manifold users plus a 1,647-paired sub-cohort) reference datasets. Each is reported with the underlying number, the cohort, the test, the confidence interval where applicable, and a link to the methodology source. Corrected 2026-08-10: this paragraph said all three are reproducible against the public data bundle. Only Finding 02 was. Updated 2026-09-04: that bundle, the 15,106-user Manifold cohort, was withheld pending Manifold's written position on data rights, so no finding on this page is currently reproducible by a reader from a published file. The aggregate results are unchanged. Findings 01 and 03 are computed on the 8,656-wallet Polymarket cohort, which has never been published in any form; they are recomputable by us and not currently by a reader. Retention status (registry): Findings 01 and 03 are recorded as lost for reader reproduction on exactly those grounds. A journalist who needs the Polymarket inputs should email research@convexly.app rather than assume a download exists. Use of these findings does not require attribution beyond a Convexly link, but a citation to the V1 paper for any methodology claim is requested.

Finding 01

PnL leaderboards are not skill rankings

On a Hill tail-index distribution with alpha = 1.28, variance is formally infinite and ranking by realized PnL is dominated by tail-event winners. The top 1% of ranked wallets in the Convexly V1 cohort captures 36.2% of signed profit. Across the same cohort, calibration (Brier score) correlates with signed log PnL at Spearman r = +0.148 (95% CI +0.128 to +0.169), a small effect: calibration barely orders the leaderboard. PnL leaderboards on Polymarket measure unit of luck on a fat-tailed distribution, not unit of skill. The Truth Leaderboard at convexly.app/truth-leaderboard shows the two orderings side by side on Convexly's closed 50-address board, not on this cohort.

Cohort8,656 Polymarket wallets, V1 cohort fixed 2026-04-15
TestHill tail-index alpha = 1.28; cohort Spearman (Brier vs signed log PnL) = +0.148 [+0.128, +0.169]
Visual referencePnL vs Edge Score

Quotable

On Polymarket, PnL leaderboards order wallets by unit of luck on a Hill alpha = 1.28 distribution, not unit of skill. The top 1% of 8,656 ranked wallets captures 36.2% of signed profit, and forecasting accuracy barely orders the board at all: Brier score against signed log PnL is Spearman +0.148.

Finding 02

Concentration differs across the sweepcash window

The original public bundle reported median per-trader concentration 8.9 percentage points lower in the sweepcash window than the pre-sweepcash bridge window (95% bootstrap CI [-17.0pp, -1.1pp]; Wilcoxon signed-rank p = 0.0137 on n=333 of 1,647 paired users with concentration delta defined). The 2026-05-04 recovered-cohort politics-excluded rerun reports -9.2pp, 95% CI [-12.8, -3.6], p=0.0021 on n=515 of 1,208 paired users with concentration delta defined. The concentration conclusion survives; causal attribution remains guarded because the sweepcash window overlapped the 2024-US-election catalog. The same Edge Score V3b methodology that ranks Polymarket wallets has its discipline-pillar coefficient flip sign on Manifold (permutation p = 0.0001).

Cohort1,647 paired Manifold traders (n=333 with concentration delta defined)
TestWithin-user paired bootstrap (1,000 reps) + Wilcoxon signed-rank
Effect sizeOriginal -8.9pp (95% CI [-17.0, -1.1], p = 0.0137); politics-excluded rerun -9.2pp (95% CI [-12.8, -3.6], p = 0.0021)

Quotable

In the original V1-M bundle, median portfolio concentration was 8.9 percentage points lower in the sweepcash window than in the pre-sweepcash bridge window (95% CI [-17.0, -1.1], Wilcoxon signed-rank p = 0.0137 on n=333 of 1,647 paired users with concentration delta defined). A 2026-05-04 politics-excluded recovered-cohort rerun reports -9.2pp, 95% CI [-12.8, -3.6], p=0.0021 on n=515 of 1,208 paired users. Treat this as a paired-window result, not a clean causal estimate of sweepcash alone.

Finding 03

Calibration explains about 2% of who profits

Across 8,656 Polymarket wallets, Brier score (lower is better-calibrated) correlates with signed log realized PnL at Spearman r = +0.148. That is, calibration explains roughly 2% of the rank variance in who profits on Polymarket. The Convexly V1 frozen-coefficient composite (Edge Score V3b) brings the out-of-fold Spearman to +0.514 by weighting calibration with two other pillars (conviction and discipline) at frozen weights of +0.7876, +2.7220, and -1.1508. The Fama-French (2010) bootstrap null on 10,000 PnL permutations rejects the null of no association at p < 0.0001. Calibration alone is necessary but not sufficient to rank wallets by realized PnL on Polymarket in-sample (cross-wallet, not a forward test); conviction and discipline measurably matter, and both are observable in the public position tape.

Cohort8,656 Polymarket wallets (full V1 cohort)
TestSpearman rank correlation; Fama-French (2010) bootstrap null
Calibration onlySpearman r = +0.148
Edge Score V3b (in-sample, cross-wallet)Out-of-fold (in-sample CV) Spearman = +0.514, p < 0.0001
In-sample rolling diagnostic (same-window PnL, leaked)Mean Spearman = +0.391 (in-sample, NOT forward)
V1.5 per-wallet temporal holdoutSpearman = +0.111 [+0.046, +0.175], FAILED +0.30 threshold

Quotable

Across 8,656 Polymarket wallets, calibration alone (Brier score) correlates with profit at Spearman r = +0.148, explaining roughly 2% of the rank variance (rho squared 0.022). A three-pillar composite (calibration, conviction, discipline) with frozen coefficients reaches out-of- fold Spearman +0.514, p < 0.0001 against a Fama- French bootstrap null. Calibration is necessary but not sufficient; conviction and discipline measurably matter.

How to use these

Each finding has been checked against the audit rules in the Convexly codebase. Two figures on this page do not meet them and are marked in place: the top-1-percent profit share and the per-user tail index carry no interval, and the first is computed from a realized-PnL column with a known open defect. Every sub-sample (such as the n=333 with concentration delta) is disclosed in the same sentence as the headline N, and no claim of pre-registration is made about the V1 cohort itself. The V1.5 deferred experiments (per-wallet temporal holdout, IC temporal stability) are externally pre-registered at AsPredicted #287368.

For sourcing in print, "Convexly Research" is the byline. Numbers come from the V1 paper at /research/edge-score-methodology-v1 and the V1-M paper at /research/edge-score-methodology-v1m. The comparison (the two results diverge; see that page's dated correction) with the Gomez-Cram, Guo, Jensen, Kung SSRN paper (April 20, 2026, revised April 25, 2026) is documented at /research/convexly-v1-vs-gomez-cram-comparison.