Lesson 4 of 12

The four-state verdict

Every wallet read lands in exactly one state. The gates run in a fixed order, the sufficiency floors always come first, and a flashy point estimate can never outrank them.

The answer first

When Convexly reads a wallet, the output is one of four states: skilled, not separable from luck, thin sample, or a reserved integrity hold (not yet produced; it waits on a wash/Sybil screen that has not shipped). Style labels such as concentrated or holding both outcome legs of the same markets to resolution describe how an edge was earned; they travel with the verdict and never change it. One deterministic rule set computes the state from the same inputs on every surface: the analyzer, the board, the badge, the API, and the enterprise deliverable. Same wallet, same fields, same verdict, everywhere.

The intuition, with one worked example

Imagine a wallet showing a +9 point realized edge whose single biggest event carries 72% of its net result. If it clears the floors (30 resolved positions, 20 independent events, coverage) and its interval clears zero, the verdict is skilled with the concentrated style attached and a caution accent: the per-position edge is real, but the dollar outcome rode on a few events, so a copier would experience the lumpy path, not the average. Now imagine those same positions collapse into 14 independent events. The verdict is thin sample, whatever the point estimate says, because the interval is an event-clustered bootstrap and 14 independent events is too few to trust it. The checks that can invalidate a number always run before the number is allowed to impress you.

The actual method

The four API states, quoted verbatim from the published definitions:

  • Skilled (skilled)

    "n_resolved >= 30, at least 20 independent resolved events, and the BCa bootstrap 95% lower bound on realized edge is above zero, all on the resolved record. Retrospective and in-sample, uncorrected for multiple comparisons. See styles for how the edge is earned and its payoff shape."

  • Not separable from luck (luck)

    "Not separable from chance on the resolved record: the BCa bootstrap 95% interval includes zero, so this record does not distinguish the realized edge from chance either way."

  • Thin sample (insufficient)

    "Too few resolved positions (n_resolved < 30), too few independent resolved events (fewer than 20), no computable edge interval, or an order-book entry price recoverable for less than 50% of the resolved record, so the resolved record cannot be read either way. The last cause is NOT a sample-size statement: the record can be large and read in full while only part of it carries a recoverable entry price to measure an edge against, and n_resolved reports the priced subset the interval was computed on."

  • Reserved integrity hold (flagged)

    "Reserved for a record held pending an integrity review. The current verdict machine does not produce this state: single-event concentration and multi-leg (complete-set) structure are reported as styles, not as a demotion."

Four API states, more labels on the page. The four above are the API's enum values, and they are what a response carries. The wallet surfaces render a longer vocabulary derived from the same read, because two of the cases below were being filed under a state whose words contradicted the number printed beside them. Nothing here is a second machine: one read produces both.

  • Ahead of its entry prices

    A 95% interval sitting entirely above zero, but read alongside many others and not corrected for that. It is the honest reading of a positive record that has not cleared a family correction, and it is what the published board rows show today instead of skilled.

  • Behind its entry prices

    The mirror: a 95% interval sitting entirely below zero. Filed under the luck state by the frozen four-state machine, which is why a row could print 'not separable from chance' next to a visibly negative interval until this was split out.

  • Recent-window record only

    The read covers the venue's recent-trade window rather than the complete resolved record, so no verdict is issued on it at all.

The gates evaluate top-down, first match wins: (0) an outcome-basis check first, so an edge scored on anything other than resolution outcomes caps the state at thin sample; (1) the sufficiency floors: 30 resolved positions AND 20 independent resolved event clusters (fail-closed: a read missing the cluster count can never render skilled), plus coverage floors and interval availability; (2) the interval-versus-zero read; and only then (3) skilled, when the BCa lower bound is strictly above zero. Style labels are computed alongside and attach to whatever state the gates produce. The interval used is always the BCa bootstrap interval, never a looser approximation.

Where you see this on the site

The analyzer prints the state as its headline verdict. The wallet board annotates every row with the same read, the embeddable skill badge renders it with the caveat baked in, and the API returns it as the state field with the definition attached. The full check is described end to end at /learn/skill-audit.

What this does NOT mean

Not separable from luck does not mean bad wallet; it means the interval includes zero, so the record does not settle the question either way. The rule runs in both directions: where the interval sits entirely below zero the record is behind the prices it paid, and saying it was not separable from luck would be just as false, so that read is labelled for what it is. It still is not a verdict on the wallet. Fees, spreads and fills push realized edge below zero with no judgment failure anywhere, and the read stays retrospective, in-sample, and uncorrected for multiple comparisons. A concentrated style label is a note on how the money was made, not a verdict on the wallet or a suggestion of wrongdoing; it keeps the caveat that the dollar outcome rode on a few events and that concentrated records can mask gamed activity. Thin sample is not an insult; it is the refusal to read 12 positions, or 14 independent events, as if they were 300. And skilled is a property of the resolved past record: retrospective, in-sample, and uncorrected for multiple comparisons, which is exactly why the cohort-level work in lesson 9 exists. None of the states is advice, and none is a forecast.

Convexly publishes new methodology research roughly every 6-8 weeks plus the /learn series on a rolling cadence. Get the next paper in your inbox when it ships:

Frequently asked

What are the four states?

Skilled, not separable from luck, thin sample, and a reserved integrity hold (in the API: skilled, luck, insufficient, flagged). The flagged state is reserved for a not-yet-live wallet-integrity screen, so live reads land in the first three. Every wallet read on every Convexly surface is derived by one deterministic state machine, so two surfaces fed the SAME FIELDS can never disagree about a wallet. That qualifier is load-bearing: surfaces reading different records of the same wallet, such as a windowed venue read and a complete on-chain read, can and do differ, and the difference is a change in what was read rather than a change in the wallet.

What does it take to earn the skilled state?

The sufficiency gates and the interval at once: at least 30 resolved positions, at least 20 independent resolved event clusters (the interval is an event-clustered bootstrap, so independent events are its effective sample size), coverage floors met, and a BCa bootstrap 95% interval whose lower bound sits strictly above zero. The read is retrospective and in-sample, and it is uncorrected for multiple comparisons, which is why cohort-level results apply a further correction (lesson 9).

What happened to the too-concentrated state?

It became a style label. Under the 2026-07-15 realized-edge reframe, high single-event concentration (60% or more of the net result) no longer demotes a read; it attaches the concentrated style to whatever verdict the gates produce, with the caveat that the per-position edge can be clean while the dollar outcome rode on a few events, and that concentrated records can mask gamed activity. The integrity function moved to the independent-event floor: positions that collapse into fewer than 20 independent resolved events cap the read at thin sample, and that floor fails closed.

Related explainers

Related reading

LearnBrier score

LearnCalibration

LearnCalibration and Brier

LearnCohort audit