Lesson 4 of 12

The four-state verdict

Every wallet read lands in exactly one state. The gates run in a fixed order, the sufficiency floors always come first, and a flashy point estimate can never outrank them.

The answer first

When Convexly reads a wallet, the output is one of four states: skilled, not separable from luck, thin sample, or a reserved integrity hold (not yet produced; it waits on a wash/Sybil screen that has not shipped). Style labels such as concentrated or holding both outcome legs of the same markets to resolution describe how an edge was earned; they travel with the verdict and never change it. One deterministic rule set computes the state from the same inputs on every surface: the analyzer, the board, the badge, the API, and the enterprise deliverable. Same wallet, same fields, same verdict, everywhere.

The intuition, with one worked example

Imagine a wallet showing a +9 point realized edge whose single biggest event carries 72% of its net result. If it clears the floors (30 resolved positions, 20 independent events, coverage) and its interval clears zero, the verdict is skilled with the concentrated style attached and a caution accent: the per-position edge is real, but the dollar outcome rode on a few events, so a copier would experience the lumpy path, not the average. Now imagine those same positions collapse into 14 independent events. The verdict is thin sample, whatever the point estimate says, because the interval is an event-clustered bootstrap and 14 independent events is too few to trust it. The checks that can invalidate a number always run before the number is allowed to impress you.

The actual method

The four states, quoted from the published API definitions:

  • Skilled (skilled)

    "n_resolved >= 30, at least 20 independent resolved event clusters, coverage floors met, and the BCa bootstrap 95% lower bound on realized edge is strictly above zero, all on the resolved record. Retrospective and in-sample, uncorrected for multiple comparisons. Style labels (for example concentrated, or holds both outcome legs of the same markets to resolution) travel with the verdict to say how the edge was earned; they never change the state."

  • Not separable from luck (luck)

    "Not separable from chance on the resolved record: the BCa bootstrap 95% interval includes zero, so this record does not distinguish the realized edge from chance either way."

  • Thin sample (insufficient)

    "Too few resolved positions (n_resolved < 30), fewer than 20 independent resolved event clusters, a coverage floor missed, or no computable edge interval, so the resolved record cannot be read either way. The event-cluster floor fails closed: a read missing the cluster count can never render skilled."

  • Reserved integrity hold (flagged)

    "Reserved for a wallet-integrity screen (wash and Sybil activity) that is not yet live. The current machine never produces this state; when it ships, a held record reads as held for an integrity review, not as a skill verdict either way."

The gates evaluate top-down, first match wins: (0) an outcome-basis check first, so an edge scored on anything other than resolution outcomes caps the state at thin sample; (1) the sufficiency floors: 30 resolved positions AND 20 independent resolved event clusters (fail-closed: a read missing the cluster count can never render skilled), plus coverage floors and interval availability; (2) the interval-versus-zero read; and only then (3) skilled, when the BCa lower bound is strictly above zero. Style labels are computed alongside and attach to whatever state the gates produce. The interval used is always the BCa bootstrap interval, never a looser approximation.

Where you see this on the site

The analyzer prints the state as its headline verdict. The wallet board annotates every row with the same read, the embeddable skill badge renders it with the caveat baked in, and the API returns it as the state field with the definition attached. The full check is described end to end at /learn/skill-audit.

What this does NOT mean

Not separable from luck does not mean bad wallet; it means the interval includes zero, so the record does not settle the question either way. A concentrated style label is a note on how the money was made, not a verdict on the wallet or a suggestion of wrongdoing; it keeps the caveat that the dollar outcome rode on a few events and that concentrated records can mask gamed activity. Thin sample is not an insult; it is the refusal to read 12 positions, or 14 independent events, as if they were 300. And skilled is a property of the resolved past record: retrospective, in-sample, and uncorrected for multiple comparisons, which is exactly why the cohort-level work in lesson 9 exists. None of the states is advice, and none is a forecast.

Convexly publishes new methodology research roughly every 6-8 weeks plus the /learn series on a rolling cadence. Get the next paper in your inbox when it ships:

Frequently asked

What are the four states?

Skilled, not separable from luck, thin sample, and a reserved integrity hold (in the API: skilled, luck, insufficient, flagged). The flagged state is reserved for a not-yet-live wallet-integrity screen, so live reads land in the first three. Every wallet read on every Convexly surface is derived by one deterministic state machine from the same inputs, so two surfaces can never disagree about the same wallet.

What does it take to earn the skilled state?

The sufficiency gates and the interval at once: at least 30 resolved positions, at least 20 independent resolved event clusters (the interval is an event-clustered bootstrap, so independent events are its effective sample size), coverage floors met, and a BCa bootstrap 95% interval whose lower bound sits strictly above zero. The read is retrospective and in-sample, and it is uncorrected for multiple comparisons, which is why cohort-level results apply a further correction (lesson 9).

What happened to the too-concentrated state?

It became a style label. Under the 2026-07-15 realized-edge reframe, high single-event concentration (60% or more of the net result) no longer demotes a read; it attaches the concentrated style to whatever verdict the gates produce, with the caveat that the per-position edge can be clean while the dollar outcome rode on a few events, and that concentrated records can mask gamed activity. The integrity function moved to the independent-event floor: positions that collapse into fewer than 20 independent resolved events cap the read at thin sample, and that floor fails closed.

Related explainers

Related reading

LearnBrier score

LearnCalibration

LearnCalibration and Brier

LearnCohort audit