Skip to content
Convexly
Lesson 10 of 12

Negative controls and evidence grades

Before trusting a result, we run the same test on records that should show nothing. Then the grades compress what the evidence supports, not what anyone hopes.

The answer first

A negative control is a placebo test for a method: run the identical pipeline on random records that have no reason to carry skill, and see what it finds. If the placebo run finds as much as the real run, the method is measuring noise. The grades (A/B/C/D) then fold each wallet's evidence into one letter with frozen, quoted gates, so a reader can triage a cohort without re-deriving the statistics.

The intuition, with one worked example

Suppose a client hands us 40 wallets and asks which show skill. The pipeline says 3 survive the corrected test. Is 3 impressive? On its own, unanswerable. So we draw 500 random cohorts of 40 wallets each from outside the client's list, size-matched, and push every one through the identical pipeline. If only a few percent of random cohorts produce 3 or more survivors, the client's list beats chance. If a third of random cohorts do, the honest report says the list looks like chance, however the list was assembled. This matters because wallet lists are usually assembled from winners someone noticed, and noticing winners is itself a luck-harvesting filter.

The actual method

The negative control: 500 seeded, size-matched placebo cohorts, drawn without replacement per draw from wallets outside the set under review, run through the identical false-discovery pipeline. The report states the empirical probability that chance matches or beats the real cohort's survivor count, and it stays prominent in the deliverable. The seeding makes the draws reproducible.

One disclosure travels with the number. The random pool is restricted to data-rich, scoreable wallets (at least 30 resolved positions with a usable interval), because a wallet with too few resolved positions cannot be put through the same test. The baseline is therefore a baseline over scoreable wallets, not the full address space, and that restriction is conservative: it makes the random draws clear the test at least as often as a full-address-space draw would, so the reviewed cohort's separation gets harder to claim, never easier.

The grade legend, quoted from the frozen specification:

  • A, Strong

    "Skilled read that survives correction at q=0.05 on a complete, diversified record with a tight interval."

  • B, Moderate

    "Skilled read with one weakness: survives only at q=0.10, or capped coverage, or high single-event share, or a wide interval."

  • C, Weak

    "A not-separable-from-chance read, a skilled read capped by a structural style (concentrated, or holding both outcome legs to resolution), or a skilled read with two or more weaknesses."

  • D, Insufficient

    "Too few resolved positions to read an edge either way; no record evidence to grade."

Grade A specifically requires survival at q = 0.05 with no coverage cap on the record, a diversified book, and an interval no wider than 12 probability points. Each graded wallet also carries one binding caveat naming its largest structural weakness (capped coverage, one event driven, or single-category dependence), or none.

The chance arithmetic on our own cohort

The simplest negative-control logic needs no simulation at all. In the frozen 2026-06-09 scan of our own published top-50 cohort, 35 wallets were testable at a 2.5 percent one-sided threshold, so under a null of zero skill the expected number of uncorrected positives is 35 × 0.025 = 0.875, call it about 0.9. Observed: exactly 1 of 35, which is what noise predicts, and it did not survive the correction at q = 0.10. The cohort's result matches its own chance baseline, and that is exactly what we published at /research/top50-skill-scan. The 500-draw control generalizes the same question to cohorts where the answer is less clean.

The honest-null rule

When the control cannot be computed (no random pool was supplied to a run, or the pool is smaller than the cohort), it is reported as null with the reason stated, never fabricated. A synthesized baseline would quietly poison every number that leans on it, and a reader who cannot see which runs had a real control cannot trust any of them. The same rule governs the rest of the pipeline: an inconclusive result is a publishable result.

Where you see this on the site

Both live in the enterprise cohort audit deliverable, delivered per period with the survivor counts, the negative-control baseline, and the grade column side by side. The per-wallet statistic being controlled and the correction each draw runs through are covered in the realized-edge and false-discovery-rate explainers, and /enterprise describes the engagement itself.

What this does NOT mean

A grade is not a rating of a person and not a forecast; it compresses the strength of the resolved record evidence, nothing more. D is not an insult: it says the record is too thin to grade either way. A null cohort result, where nobody clears and the placebo band explains everything, is a valid, reportable deliverable, not a failed engagement. And the negative control does not certify a method as correct; it only shows whether this result, on this cohort, beats what chance produces on same-sized random lists.

Convexly publishes new methodology research roughly every 6-8 weeks plus the /learn series on a rolling cadence. Get the next paper in your inbox when it ships:

Frequently asked

What is a negative control in a wallet cohort audit?

A placebo run: 500 size-matched cohorts of wallets drawn at random from outside the set under review, pushed through the identical statistical pipeline. If random cohorts produce as many survivors as the real cohort, the apparent result is what chance looks like, and the report says so plainly. The draws are seeded, so anyone re-running the pipeline gets the same placebo cohorts.

What do the A/B/C/D grades mean?

They compress the strength of the record evidence per wallet, with frozen gates. A (Strong) is a skilled read surviving correction at q=0.05 on a complete, diversified record with a tight interval. B (Moderate) is a skilled read with exactly one weakness. C (Weak) is a not-separable-from-chance read, a skilled read capped by a structural style (concentrated, or holding both outcome legs to resolution), or a skilled read with two or more weaknesses. D (Insufficient) means too few resolved positions to read an edge either way.

Why run placebo cohorts at all?

Because how a wallet list was picked can manufacture a result. A list assembled from wallets someone noticed because they won is pre-filtered for luck, and it will look better than random even if nobody in it has skill. The negative control anchors every cohort finding to what chance produces on same-sized random lists, which is the only fair baseline.

What is the pool restriction, and why is it disclosed?

The random pool contains only data-rich, scoreable wallets (at least 30 resolved positions with a usable bootstrap interval), because a wallet with too few resolved positions cannot be put through the same test. The chance baseline is therefore a baseline over scoreable wallets, not the full address space. That restriction biases the baseline conservatively: the random draws clear the test at least as often as a draw from the full address space would, making the reviewed cohort's separation harder to claim, never easier.

What happens when the control cannot be computed?

It is reported as null with the reason (for example, no random pool was supplied to the run), never fabricated. An honest null is a valid result; a synthesized baseline would poison every number that leans on it.

Related explainers

Related reading