Scan data as of 2026-06-09; frozen edition published 2026-06-10.
Research note ยท scan date 2026-06-09
We ran our skill read on every wallet in our own published top-50 cohort. 35 of 50 had enough resolved positions to test. The number that cleared the corrected statistical bar: zero.
Convexly publishes a top-50 Polymarket cohort ranked by a behavioral composite, Edge Score (the leaderboard page renders the cohort's positive-PnL subset). We turned our own skill test on that published cohort: for each wallet with at least 30 resolved positions, we computed the realized entry edge, meaning whether entries resolved true more often than the price paid implied, with a bootstrap 95% interval, plus a concentration screen.
Method note (2026-07-16): this scan was computed under the pre-2026-07-15 machine, where high single-event concentration demoted a read to flagged; concentration is now a descriptive style label and the sufficiency function moved to an independent-event floor. The frozen artifact below is unchanged.
Read basis (2026-07-29): the per-wallet records behind this scan were taken from Polymarket's recent-trade window, which omits redeemed positions and is therefore not a complete resolved record. We are recomputing the cohort against the on-chain Polygon settlement record and will publish the revised scan alongside this one with the per-wallet deltas, in whichever direction they move. The frozen artifact below is unchanged.
Results: 19 of 35 read clean as not separable from chance, with an interval that includes zero and no concentration flag. Another 15 are too concentrated to read as a clean edge, with one event carrying most of the net result; most of those intervals include zero too. Exactly 1 record's interval clears zero on the positive side, which is roughly what chance alone predicts: testing 35 records at a 2.5% threshold is expected to produce about 0.9 false positives. That record does not survive a multiple-comparisons correction, and the same wallet is net-negative on PnL. A wallet with a very large public P&L that people cite as proof of skill shows an edge of +0.4 probability points, 95% interval [-16.0, +15.6], across 38 resolved positions.
The honest takeaway is NOT "these wallets have no skill." It is that the resolved records, at these sample sizes, mostly cannot distinguish skill from chance either way, and a P&L screenshot tells you even less.
Per-wallet table
Labels are on-chain Polymarket usernames or addresses only. The one positive-interval row is annotated in place: it is an uncorrected single test selected out of 35, statistically consistent with chance at the cohort level, it is not in the FDR-corrected cleared set, and its net PnL is negative. Dollar PnL figures are withheld pending a cross-pipeline reconciliation; the sign column reflects the published cohort artifact. Diagnostics on public on-chain records, not investment advice; a past read is not a forecast.
| Read | Label | Edge Score | n resolved | Realized edge (pp) | 95% BCa CI | Concentration | Net PnL sign |
|---|---|---|---|---|---|---|---|
| interval clears zero on the positive side (uncorrected) | ShouShouKKos 0xc2fb28Uncorrected single test of 35; NOT in the FDR-corrected cleared set; net PnL negative. Not a skill verdict. For scale, 35 one-sided tests at a 2.5 percent threshold work out to 0.875 positive intervals if no wallet had any edge and the 35 had been drawn at random, and these 35 were not drawn at random, so read 0.875 as a floor, not a forecast. | 12.7 | 137 | +5.6 | [+1.2, +9.5] | n/a | negative |
| not separable from chance | Zarvantis 0x30c92f | 77.1 | 47 | -2.7 | [-8.9, +3.2] | n/a | positive |
| not separable from chance | beachboy4 0xc2e780 |
Computed from a frozen 2026-06-09 snapshot of the published top-50 Edge Score cohort. That snapshot carries no proof that the wallet's record was read in full, so these figures describe the snapshot and not the whole record. Positions outside it can move them.
Reads: 1 interval-clears-zero on the positive side (uncorrected), 19 not separable from chance, 15 too concentrated to read, 15 with fewer than 30 resolved positions. Zero of the 50 are in the FDR-corrected cleared set.
Method
- Universe: the 50 addresses in our published cohort artifact. They were screened once, on a 2026-04-15 snapshot, for having at least 5 resolved positions and at least $5,000 realized PnL, then ranked by the then-current Edge Score and cut to the top 100. No address has entered the set since 2026-04-19 and the set has only shrunk, to 50. The daily job rescores these same 50 addresses and republishes them; it does not screen Polymarket. Every row currently clears a floor of 20 scored resolved positions, meaning resolved positions for which this read recovered an entry price, and the rows are ordered by that count, not by Edge Score. Membership was conditioned partly on having made money, which is selection on the dependent variable, so no count over these rows describes Polymarket wallets. The rendered leaderboard page shows only the positive-PnL rows, so that surface is conditioned on profit twice.
- Realized entry edge per wallet: outcome minus dollar-weighted entry price (fill prices weighted by dollars filled; terminology corrected 2026-08-14, computation unchanged), equal-weighted across resolved unique positions; 95% interval via BCa bootstrap. Readability floor: 30 scored resolved positions (35 of 50 qualify; the other 15 are reported as insufficient, not tested). The April screen is part of why: it admitted a set of which roughly half falls below this floor, so the untestable 15 are an artifact of how the set was chosen, not a finding about the wallets.
- Concentration screen: the single biggest winning event's PnL divided by net PnL; at or above 0.6 the record is reported as too concentrated to read. The ratio is not bounded by 1 and exceeds it when one event is larger than the net result. A non-computable ratio (analyzer-side net PnL at or below zero) cannot trigger the concentrated read; the record is then classified by its interval alone. This applies to the one positive-interval row: its net-negative PnL makes the ratio non-computable, so that record was not concentration-screened. The sign column comes from the cohort artifact, which can disagree with the analyzer-side figure.
- Multiple comparisons: 35 one-sided tests at a 2.5% threshold work out to 0.875 positive intervals if no wallet had any edge and the 35 had been drawn at random. These 35 were not drawn at random, so read 0.875 as a floor, not a forecast. One was observed. The zero also does not rest on the concentration screen or on an independence assumption: dropping the concentration screen and correcting across all 35 tests still clears none, and neither does a correction that stays valid when wallets trade the same events.
- Dollar PnL figures are withheld: the cohort artifact and the live analyzer disagree on realized PnL for some wallets, and we do not publish numbers our own pipelines dispute. The reconciliation is tracked internally; the sign column is the cohort artifact's.
Diagnostics on public on-chain records, not investment advice. A past read is not a forecast, and no corrected evidence of skill is not evidence of no skill: several intervals are wide and include large positive values, so the record is underpowered, not settled. If you think the math is wrong, the interval is the thing to argue with; we publish corrections when receipts say so.