Learn

What is a prediction-market wallet skill audit?

An independent, method-frozen read of whether a Polymarket wallet's resolved record reflects demonstrated skill or luck, built on realized entry edge with a 95 percent interval and its resolved-position count, never a headline number on its own.

The answer first

A prediction-market wallet skill audit asks one narrow, testable question of a Polymarket wallet: does its resolved record reflect demonstrated skill, or is it consistent with luck? It is independent, because it runs on the venue's public record rather than a self-reported screenshot, and method-frozen, because the scoring method is version-controlled and does not move under you. It is not reproducible in the third-party sense: the repository is private, so re-running it is something Convexly can do and a reader cannot. The audit does not endorse a wallet, take custody, or route investment advice. It reads the record and returns a verdict.

The engine is realized entry edge: how much more often the wallet's entries resolved true than the prices it paid implied, on the probability scale. A reading of +5 probability points would mean entries resolved true about five points more often than the paid prices implied. Convexly reads that statistic with a bias-corrected and accelerated (BCa) bootstrap 95 percent interval and always states the resolved-position count in the same breath, because at realistic sample sizes the point estimate on its own is mostly noise.

What the audit reads

A skill audit is not a single metric. It combines four inputs, and the floors cap the read on their own:

  • Realized entry edge, with its 95 percent interval and count: the entry-price test, reported as probability points with a BCa 95 percent interval and the number of resolved positions behind it. The number never travels alone.
  • Calibration: whether the wallet's implied probabilities line up with how often those positions actually resolve true, complementing the entry-price read.
  • Concentration: a style label attached when a single event drives 60 percent or more of the net result. It never changes the verdict; it says the dollar outcome rode on a few events, and the caveat travels with it.
  • Resolved-position and independent-event counts: a floor of 30 resolved positions AND 20 independent resolved events before a skill read is even attempted, because a thinner or more clustered record cannot tell an edge from chance either way. The event floor fails closed when the count is missing.

Worked example: the bigger number is the weaker read

Two real rows from the frozen 2026-06-09 scan of our own published top-50 cohort (full table at /research/top50-skill-scan):

  • Wallet 0xaaaf7f: realized entry edge +12.2pp across 32 resolved positions, 95 percent interval [-0.7, +23.7]. The interval includes zero, so the record is not distinguishable from chance despite the large point estimate.
  • Wallet 0xc2fb28: realized entry edge +5.6pp across 137 resolved positions, 95 percent interval [+1.2, +9.5]. Less than half the point estimate, but the interval clears zero: the stronger read. As one uncorrected test among 35 it is still consistent with chance at the cohort level and is not in the false-discovery-rate-controlled cleared set.

Read this example with what came next. On 2026-07-29 this same wallet was reconciled against the on-chain settlement record, and over its complete resolved record the realized edge is -19.8pp, 95 percent interval [-34.2, -6.9]: an interval sitting entirely below zero, where the windowed read had one sitting entirely above it. The complete read is refused a verdict anyway, because an entry price was recoverable for only 24.6% of the record, under the 50 percent floor. Nothing about the wallet changed between the two readings; what changed is which record was read. The lesson above still holds, and this is the harder half of it: an interval is only as trustworthy as the record it was computed over, and a recent-trade window is not the record.

That inversion is the whole point of running an audit rather than trusting a headline. A +12.2pp reading across 32 resolved positions tells you less than a +5.6pp reading across 137, because the interval width scales with sample size. Anyone quoting a wallet's edge without its interval and its resolved-position count in the same breath is quoting noise.

The verdict: distinguishable from chance, or not

The inputs feed a deterministic read, evaluated top-down with the sufficiency gates first, so a flattering number can never out-vote a floor:

  • Insufficient: fewer than 30 resolved positions, fewer than 20 independent resolved events (fail-closed when the count is missing), a coverage floor missed, or no usable interval. Too thin or too clustered to tell an edge from chance either way.
  • Not distinguishable from chance: the BCa 95 percent interval includes zero. This is not a claim of no skill; it is a claim that the record cannot establish one at this sample size.
  • Skill distinguishable from chance: renders only when both sufficiency floors and the coverage checks are met and the BCa 95 percent lower bound is strictly above zero. Style labels such as concentrated, or holding both outcome legs of the same markets to resolution, attach descriptively and keep their caveats. Retrospective and in-sample; a past read is not a forecast.
  • Flagged: reserved for a wallet-integrity screen that is not yet live; the current machine never produces it. A held record would read as held for an integrity review, not as a verdict either way.

When a whole cohort is screened at once, one positive interval does not survive being one test among many. That is the job of the false-discovery-rate correction, which sets the frozen bar for the cleared set. The frozen definitions live in the lexicon, and the interval construction is documented on the methodology page.

What a skill audit does NOT do

It does not measure dollar PnL, which on a fat-tailed market is dominated by a few large positions and by luck at these sample sizes. It does not forecast future trades: the per-wallet forward holdout did not clear the filed threshold, so a skill audit is not a prediction of a single wallet's next result. And it is not a recommendation to follow, mirror, or copy any trader or wallet. Convexly does not take custody, broker trades, or route investment advice. The audit reads a record; the decision stays with you.

Run a skill audit

Paste any Polymarket wallet address at the analyzer to get its realized entry edge, the BCa 95 percent interval, the resolved-position count, calibration and concentration, and the verdict with its style labels. Free, no signup, Polymarket public data only.

Convexly publishes new methodology research roughly every 6-8 weeks plus the /learn series on a rolling cadence. Get the next paper in your inbox when it ships:

Frequently asked

What is a prediction-market wallet skill audit?

An independent, method-frozen read of whether a Polymarket wallet's public record reflects demonstrated skill or luck. It computes realized entry edge, how much more often the wallet's entries resolved true than the prices it paid implied, reads that figure with a 95 percent confidence interval and the count of resolved positions behind it, checks calibration and the sufficiency floors (30 resolved positions and 20 independent resolved events), and returns one of the states: skill distinguishable from chance, not distinguishable from chance, or insufficient sample (a fourth, flagged, is reserved for a not-yet-live integrity screen). Concentration travels as a descriptive style label on the verdict. Polymarket public data only. Convexly does not take custody, broker trades, or route investment advice.

How does a skill audit tell skill from luck?

It never reads a single number on its own. The core statistic, realized entry edge, always travels with its 95 percent confidence interval and the number of resolved positions behind it. If the interval sits entirely above zero the record is distinguishable from chance at the frozen bar; if it includes zero it is not, no matter how large the point estimate. In the frozen 2026-06-09 scan of our own published cohort, one wallet showed +12.2 probability points of edge across 32 resolved positions with a 95 percent interval of [-0.7, +23.7], which includes zero, so it is not separable from chance. Another showed a smaller +5.6 points across 137 resolved positions with a 95 percent interval of [+1.2, +9.5] that clears zero.

Why does the confidence interval matter more than the headline number?

Because at realistic sample sizes the point estimate is mostly noise, and the interval width scales with how many positions have resolved. A +12.2 probability-point reading across 32 resolved positions (95 percent interval [-0.7, +23.7]) tells you less than a +5.6-point reading across 137 resolved positions (95 percent interval [+1.2, +9.5]), because the smaller number has the tighter interval. Any wallet edge quoted without its 95 percent interval and its resolved-position count in the same breath is quoting noise.

What does a skill audit check besides realized entry edge?

Calibration, whether the wallet's implied probabilities line up with how often those positions actually resolve true; concentration, whether a single event drives 60 percent or more of the net result, which attaches the concentrated style label and its caveat; the resolved-position count, with a floor of 30; and the independent-event count, with a fail-closed floor of 20, because the interval is an event-clustered bootstrap and independent events are its effective sample size. A skill verdict renders only when both floors and the coverage checks are met and the 95 percent lower bound of realized entry edge is strictly above zero. The floors win: a flattering point estimate on its own never reaches the skilled state.

Does a skill audit forecast a wallet's future or recommend a wallet to follow?

No. The audit is retrospective and in-sample: it reads whether a resolved record is distinguishable from chance, not whether the next trade will win. The per-wallet forward holdout did not clear the filed threshold, so a skill audit does not predict a single wallet's future performance. Convexly does not take custody, broker trades, route investment advice, or recommend copying any trader or wallet.

Where can I run a skill audit on a Polymarket wallet?

Paste any address at the Polymarket wallet analyzer. The free tool returns realized entry edge in probability points, its 95 percent confidence interval, the resolved-position count, calibration and concentration, the verdict, and its style labels. No signup, Polymarket public data only.

Related explainers

Related reading

LearnCohort audit

AnswersHow to audit a cohort of Polymarket wallets

AnswersPrediction market market quality audit

AnswersSkill vs luck trading PnL