Negative controls and evidence grades
Before trusting a result, we run the same test on records that should show nothing. Then the grades compress what the evidence supports, not what anyone hopes.
The answer first
A negative control is a placebo test for a method: run the identical pipeline on random records that have no reason to carry skill, and see what it finds. If the placebo run finds as much as the real run, the method is measuring noise. The grades (A/B/C/D) then fold each wallet's evidence into one letter with frozen, quoted gates, so a reader can triage a cohort without re-deriving the statistics.
The intuition, with one worked example
Suppose a client hands us 40 wallets and asks which show skill. The pipeline says 3 survive the corrected test. Is 3 impressive? On its own, unanswerable. So we draw 500 random cohorts of 40 wallets each from outside the client's list, size-matched, and push every one through the identical pipeline. If only a few percent of random cohorts produce 3 or more survivors, the client's list beats chance. If a third of random cohorts do, the honest report says the list looks like chance, however the list was assembled. This matters because wallet lists are usually assembled from winners someone noticed, and noticing winners is itself a luck-harvesting filter.
The actual method
The negative control: 500 seeded, size-matched placebo cohorts, drawn without replacement per draw from wallets outside the set under review, run through the identical false-discovery pipeline. The report states the empirical probability that chance matches or beats the real cohort's survivor count, and it stays prominent in the deliverable. The seeding makes the draws reproducible.
One disclosure travels with the number. The random pool is restricted to data-rich, scoreable wallets (at least 30 resolved positions with a usable interval), because a wallet with too few resolved positions cannot be put through the same test. The baseline is therefore a baseline over scoreable wallets, not the full address space, and that restriction is conservative: it makes the random draws clear the test at least as often as a full-address-space draw would, so the reviewed cohort's separation gets harder to claim, never easier.
The grade legend, quoted from the frozen specification:
A, Strong
"Skilled read that survives correction at q=0.05 on a complete, diversified record with a tight interval."
B, Moderate
"Skilled read with one weakness: survives only at q=0.10, or capped coverage, or high single-event share, or a wide interval."
C, Weak
"A not-separable-from-chance read, a skilled read capped by a structural style (concentrated, or holding both outcome legs to resolution), or a skilled read with two or more weaknesses."
D, Insufficient
"Too few resolved positions to read an edge either way; no record evidence to grade."
Grade A specifically requires survival at q = 0.05 with no coverage cap on the record, a diversified book, and an interval no wider than 12 probability points. Each graded wallet also carries one binding caveat naming its largest structural weakness (capped coverage, one event driven, or single-category dependence), or none.
The chance arithmetic on our own cohort
The simplest negative-control logic needs no simulation at all. In the frozen 2026-06-09 scan of our own published top-50 cohort, 35 wallets were testable at a 2.5 percent one-sided threshold, so under a null of zero skill the expected number of uncorrected positives is 35 × 0.025 = 0.875, call it about 0.9. Observed: exactly 1 of 35, which is what noise predicts, and it did not survive the correction at q = 0.10. The cohort's result matches its own chance baseline, and that is exactly what we published at /research/top50-skill-scan. The 500-draw control generalizes the same question to cohorts where the answer is less clean.
The honest-null rule
When the control cannot be computed (no random pool was supplied to a run, or the pool is smaller than the cohort), it is reported as null with the reason stated, never fabricated. A synthesized baseline would quietly poison every number that leans on it, and a reader who cannot see which runs had a real control cannot trust any of them. The same rule governs the rest of the pipeline: an inconclusive result is a publishable result.
Where you see this on the site
Both live in the enterprise cohort audit deliverable, delivered per period with the survivor counts, the negative-control baseline, and the grade column side by side. The per-wallet statistic being controlled and the correction each draw runs through are covered in the realized-edge and false-discovery-rate explainers, and /enterprise describes the engagement itself.
What this does NOT mean
A grade is not a rating of a person and not a forecast; it compresses the strength of the resolved record evidence, nothing more. D is not an insult: it says the record is too thin to grade either way. A null cohort result, where nobody clears and the placebo band explains everything, is a valid, reportable deliverable, not a failed engagement. And the negative control does not certify a method as correct; it only shows whether this result, on this cohort, beats what chance produces on same-sized random lists.
Convexly publishes new methodology research roughly every 6-8 weeks plus the /learn series on a rolling cadence. Get the next paper in your inbox when it ships:
Frequently asked
What is a negative control in a wallet cohort audit?
What do the A/B/C/D grades mean?
Why run placebo cohorts at all?
What is the pool restriction, and why is it disclosed?
What happens when the control cannot be computed?
Related explainers
- /learn/cohort-audit: the deliverable these two pieces anchor
- /learn/luck-share: the published cohort reading that matched its own chance baseline