The four-state verdict
Every wallet read lands in exactly one state. The gates run in a fixed order, the sufficiency floors always come first, and a flashy point estimate can never outrank them.
The answer first
When Convexly reads a wallet, the output is one of four states: skilled, not separable from luck, thin sample, or a reserved integrity hold (not yet produced; it waits on a wash/Sybil screen that has not shipped). Style labels such as concentrated or holding both outcome legs of the same markets to resolution describe how an edge was earned; they travel with the verdict and never change it. One deterministic rule set computes the state from the same inputs on every surface: the analyzer, the board, the badge, the API, and the enterprise deliverable. Same wallet, same fields, same verdict, everywhere.
The intuition, with one worked example
Imagine a wallet showing a +9 point realized edge whose single biggest event carries 72% of its net result. If it clears the floors (30 resolved positions, 20 independent events, coverage) and its interval clears zero, the verdict is skilled with the concentrated style attached and a caution accent: the per-position edge is real, but the dollar outcome rode on a few events, so a copier would experience the lumpy path, not the average. Now imagine those same positions collapse into 14 independent events. The verdict is thin sample, whatever the point estimate says, because the interval is an event-clustered bootstrap and 14 independent events is too few to trust it. The checks that can invalidate a number always run before the number is allowed to impress you.
The actual method
The four API states, quoted verbatim from the published definitions:
Skilled (skilled)
"n_resolved >= 30, at least 20 independent resolved events, and the BCa bootstrap 95% lower bound on realized edge is above zero, all on the resolved record. Retrospective and in-sample, uncorrected for multiple comparisons. See styles for how the edge is earned and its payoff shape."
Not separable from luck (luck)
"Not separable from chance on the resolved record: the BCa bootstrap 95% interval includes zero, so this record does not distinguish the realized edge from chance either way."
Thin sample (insufficient)
"Too few resolved positions (n_resolved < 30), too few independent resolved events (fewer than 20), no computable edge interval, or an order-book entry price recoverable for less than 50% of the resolved record, so the resolved record cannot be read either way. The last cause is NOT a sample-size statement: the record can be large and read in full while only part of it carries a recoverable entry price to measure an edge against, and n_resolved reports the priced subset the interval was computed on."
Reserved integrity hold (flagged)
"Reserved for a record held pending an integrity review. The current verdict machine does not produce this state: single-event concentration and multi-leg (complete-set) structure are reported as styles, not as a demotion."
Four API states, more labels on the page. The four above are the API's enum values, and they are what a response carries. The wallet surfaces render a longer vocabulary derived from the same read, because two of the cases below were being filed under a state whose words contradicted the number printed beside them. Nothing here is a second machine: one read produces both.
Ahead of its entry prices
A 95% interval sitting entirely above zero, but read alongside many others and not corrected for that. It is the honest reading of a positive record that has not cleared a family correction, and it is what the published board rows show today instead of skilled.
Behind its entry prices
The mirror: a 95% interval sitting entirely below zero. Filed under the luck state by the frozen four-state machine, which is why a row could print 'not separable from chance' next to a visibly negative interval until this was split out.
Recent-window record only
The read covers the venue's recent-trade window rather than the complete resolved record, so no verdict is issued on it at all.
The gates evaluate top-down, first match wins: (0) an outcome-basis check first, so an edge scored on anything other than resolution outcomes caps the state at thin sample; (1) the sufficiency floors: 30 resolved positions AND 20 independent resolved event clusters (fail-closed: a read missing the cluster count can never render skilled), plus coverage floors and interval availability; (2) the interval-versus-zero read; and only then (3) skilled, when the BCa lower bound is strictly above zero. Style labels are computed alongside and attach to whatever state the gates produce. The interval used is always the BCa bootstrap interval, never a looser approximation.
Where you see this on the site
The analyzer prints the state as its headline verdict. The wallet board annotates every row with the same read, the embeddable skill badge renders it with the caveat baked in, and the API returns it as the state field with the definition attached. The full check is described end to end at /learn/skill-audit.
What this does NOT mean
Not separable from luck does not mean bad wallet; it means the interval includes zero, so the record does not settle the question either way. The rule runs in both directions: where the interval sits entirely below zero the record is behind the prices it paid, and saying it was not separable from luck would be just as false, so that read is labelled for what it is. It still is not a verdict on the wallet. Fees, spreads and fills push realized edge below zero with no judgment failure anywhere, and the read stays retrospective, in-sample, and uncorrected for multiple comparisons. A concentrated style label is a note on how the money was made, not a verdict on the wallet or a suggestion of wrongdoing; it keeps the caveat that the dollar outcome rode on a few events and that concentrated records can mask gamed activity. Thin sample is not an insult; it is the refusal to read 12 positions, or 14 independent events, as if they were 300. And skilled is a property of the resolved past record: retrospective, in-sample, and uncorrected for multiple comparisons, which is exactly why the cohort-level work in lesson 9 exists. None of the states is advice, and none is a forecast.
Convexly publishes new methodology research roughly every 6-8 weeks plus the /learn series on a rolling cadence. Get the next paper in your inbox when it ships:
Frequently asked
What are the four states?
What does it take to earn the skilled state?
What happened to the too-concentrated state?
Related explainers
- /learn/skill-audit: the full wallet check these states summarize
- /learn/concentration-flag: the concentrated style label and its caveat
- /learn/resolved-position-count: why the floor is 30 resolved positions
Related reading
LearnBrier score
LearnCalibration
LearnCohort audit