Research note · Published 2026-08-07 · Pre-registration AsPredicted #295174 (public anonymous PDF)
The guard we built for this failure watched the wrong axis
We pre-registered a forward test of our own market-rating instrument, ran it to its filed analysis point, and it returned a null.
The null is not the finding. The finding is that our registration already contained a branch built for exactly this failure, that branch fired correctly on the axis it watches, and the test still produced a confident-sounding label it had no power to earn. The guard's existence is what made the null sound like a result.
We are publishing the mechanism because the failure mode generalises: a competently designed threshold, guarding the quantity its authors reasoned about, silently blind to the one they did not.
What we filed
AsPredicted #295174, submitted 2026-06-06. The question: do the frozen Market Trust verdict tiers separate Polymarket markets that resolve cleanly from markets that do not?
Two arms, collapsed from four tiers. Higher-trust: use, use_with_caveats. Lower-trust: discount, do_not_cite. The outcome, fixed in advance: BAD if a market was voided, disputed at the oracle, or reached no terminal price; CLEAN otherwise.
The primary test had two conditions, both required: a one-sided 95% Newcombe lower bound on the arm difference clearing zero, and that difference exceeding a resolution-rule-text baseline. One confirmatory analysis, run once, at the first window meeting every filed floor. No optional stopping.
Filed floors: at least 30 resolved forward cards, across at least 3 categories, at least 5 per reported category or a stated caveat, and a variance gate of at least 5 BAD markets.
Four deviation records are on file and are part of the design as it ran: UMA classification (2026-06-07), intake sampling policy (2026-07-03), BAD-unit operationalisation (2026-07-26), writer pagination (2026-07-26). The first matters for reading the outcome definition and is discussed below.
Accrual disclosure, required by our own deviation record. The resolution ledger accrued in two regimes: 13 rows on 2026-07-25, 3,418 on 2026-07-26 after a writer-pagination fix. Most of the analysed corpus arrived in the final days of the window.
What it returned
The filed analysis point is 2026-07-30, the first window meeting every floor. Every floor was met:
| floor | required | observed |
|---|---|---|
| resolved forward cards | >= 30 | 3,876 |
| categories | >= 3 | 65 |
| per-category minimum | >= 5 or a caveat | caveat issued, 34 categories |
| BAD markets (variance gate) | >= 5 | 6 (across 7 BAD card rows) |
The engine returned no_separation, whose pre-registered wording is "Market Trust tiers do not separate on resolution cleanliness."
Read alone, that sentence is a finding about the instrument. It is not one.
The artifact is recomputed daily by cron. We report the 2026-07-30 read because that is the filed analysis point; the 2026-08-03 recomputation returns the same verdict, the same higher arm and the same bound to three decimals. The 2026-07-30 artifact is pinned immutably in the repository so this is checkable.
The number that was in no floor
| arm | markets | clean | clean rate, Wilson 95% |
|---|---|---|---|
| higher-trust | 4 | 4 / 4 | 100% [51.0%, 100%] |
| lower-trust | 1,770 | 1,764 / 1,770 | 99.66% [99.26%, 99.84%] |
Difference +0.34pp. One-sided 95% Newcombe lower bound -0.400.
The higher-trust arm held four markets, and its own clean rate could have been as low as 51% without this data noticing.
Floors are filed in cards; the interval runs on markets, deduplicated because cards on the same market share one resolution outcome. 3,876 cards cover 1,774 distinct markets.
And the arm is plausibly three, not four. The single use-tier market carries a ledger resolution timestamp equal to its scheduled endDate, while the venue records it as actually closing five weeks before our forward window opened. The audit fields added in 2026-06-28 precisely to make a scheduled-close fallback visible rather than silent are NULL on all 4,268 rows, so we cannot currently settle it from our own store. The verdict is unchanged either way, because every attainable outcome for a three-market arm also fails.
Direction carries no information here
The higher arm returned the most extreme result available to it, 4 of 4 clean, and the bound still sat at -0.400.
If the arms were identical at the lower arm's rate, a four-market arm returns 4 of 4 clean 98.65% of the time (Fisher one-sided exact p = 0.9865). Perfection was the modal outcome under a true difference of exactly zero. Any reading of that 100% as encouraging is reading noise.
The test could not have passed
Sweeping every attainable outcome for a four-market arm against the observed lower arm:
| higher arm result | one-sided 95% lower bound |
|---|---|
| 0 / 4 | -0.998 |
| 1 / 4 | -0.939 |
| 2 / 4 | -0.814 |
| 3 / 4 | -0.640 |
| 4 / 4 | -0.400 |
None clears zero. Holding the lower arm at its observed 1,764/1,770 composition, power against every alternative was zero: the confirmatory branch could not fire whatever those four markets did.
For a perfect higher arm to clear zero against a lower arm this clean takes roughly 909 markets, on the same fixed-lower-arm assumption. Seven markets sat in the higher-trust tiers across the entire universe the day before we filed.
The second half of the primary test, beating the rule-text baseline, was satisfied (+0.0034 against -0.0040). At this arm size that means nothing either.
What the null does and does not say
A one-sided lower bound below zero means the contrast is not established. Establishing that the arms are alike is a different test, requiring an equivalence margin stated in advance. We filed none, so equivalence is unavailable in either direction.
Why the filed design could not catch it
The registration anticipated the danger in general. Its sample-size section reasons that bad-resolution events are rare, and it builds a branch called inconclusive_power_limited for exactly the case where a null would be power-starved rather than real.
That branch has two triggers. One is the variance gate on BAD markets. The other fires when an arm is degenerate, defined as holding zero markets.
Six BAD markets cleared the variance gate. Four is not zero. The test passed both guards and produced a confident label.
The floors were global (total resolved, total categories, total BAD) and the failure was per-cell. A global floor cannot see a starved cell.
We want to be precise about the contribution here, because the bare version of this lesson is a template-completion error that any reviewer would name. In a two-proportion comparison, power is governed by the smaller arm; "specify n per condition" is not a discovery. Two things around it are less obvious:
1. No general-purpose preregistration template asks for it. We checked. The OSF Preregistration form asks how many units will be analysed and accepts "an arbitrary number of subjects" as a rationale. AsPredicted asks how many observations will be collected. Most tellingly, the Secondary Data Analysis template (van den Akker et al., Meta-Psychology 5, 2021) is written for the case where the data already exist, and still asks for power against a SESOI rather than for the counts sitting in the dataset.
2. The guard we did build was the sophisticated one, and it still missed. Registered Reports gate on design-answerability at Stage 1 through outcome-neutral tests and positive controls, but that is a reviewer's judgment about whether the measurement works. Neither it nor our own branch asks the arithmetic question.
The check we did not run, and now do
Enumerate every outcome the smallest cell could produce, and ask whether any of them clears the decision threshold. If none does, the study cannot pass on any data, and you know it before collecting a single observation.
This needs no assumed effect size, which is what separates it from a power calculation. It is not a probability that the study will succeed; it is whether success is in the outcome space at all. The nearest existing ideas are cousins, not the same thing:
Design analysis (Gelman and Carlin, Perspectives on Psychological Science 9(6), 2014) computes what a design would produce under an assumed true effect. The prior question is whether the confirmatory branch is reachable regardless.
The minimal statistically detectable effect (Lakens, Sample Size Justification) is the closest relative: the smallest effect that would yield significance at a given n. Lakens asks whether that effect is plausible. Ours is the degenerate case where it is not attainable.
Futility analysis in clinical trials asks mid-trial whether a study can still succeed. This is the pre-data, deterministic limit of the same question: conditional power is zero not because interim data are discouraging, but because the rejection region is empty over the attainable sample space.
Industry A/B testingalready treats "compute the achievable minimum detectable effect, and do not run the test if it exceeds any plausible effect" as a launch gate. That gate is absent from academic preregistration.
The underlying arithmetic is textbook. Sign-test tables print a dash where no critical value exists at a given n and alpha. We found no named methodological principle that turns it into a pre-filing step, which is the gap.
It now runs in our verdict engine on every branch, emitted alongside the verdict and never read by it, because a gate consulting it would be a post-hoc rule our filing does not contain.
Count the unit that is independent, not the rows
A second cardinality trap sits one level below the first, and we walked into it the same week while sizing a follow-up.
Scanning the 1,200 most-recently-closed Polymarket markets for bad resolutions gives 90 events, and every pre-resolution signal looks like a perfect predictor: not one of the 329 markets lacking a published void clause voided. Collapse to the underlying real-world fixture and it is 8 events over 305 fixtures. One abandoned match voids every derivative market it spawned at once; five abandoned tennis matches account for 79 of the 90, and one of them produced 26 by itself.
The row count overstated the independent evidence by roughly 11x. Anyone sizing a test off "90 events" files another design that cannot pass.
So the check has two parts, and the second is the one we keep needing: count the cells, and before that, work out what makes two of them independent.
What we changed
The verdict producer now emits per-arm counts on every branch, unconditionally, so a label cannot travel without its denominator. The committed artifact predates this and does not yet carry it.
The attainability sweep above runs on every branch, diagnostic only.
The filed analysis point is now stamped once and never advanced, and the terminal artifact is pinned immutably. Before this, a daily cron kept recomputing past the filed read with nothing recording which read was filed. We caught this by making the mistake: an earlier revision of this note quoted a later recomputation as the result.
The pre-filing count is written into our analysis framework as a filing-time step. We also tried it as an automated gate and removed it the same day: over a corpus that deliberately preserves superseded drafts unedited, it fired on 12 and then 14 documents, almost all amendments and history. A check that over-fires becomes noise.
Every research draft, pre-registration and methodology claim now gets three independent specialist reviews before it leaves drafting. The first application was to this note, and each lens found something the others did not, including two errors in earlier revisions of this document.
What we are not claiming
We are not claiming Market Trust separates markets on resolution quality. We are not claiming it does not.
And we cannot go back and find out.Under field 7's no-optional-stopping clause, #295174 is spent: it was analysed once, at its filed point, and this question can never be re-asked under it. That is the price, and naming it is what makes the discipline mean anything.
Three further limits a careful reader should have:
The two arms were different market populations. The higher arm was a neg-risk tournament leg, a crypto price threshold, and two markets turning on a defined political term. The lower arm was around 90% sports fixtures and their derivatives. Whatever distinguishes those arms, market archetype is a live candidate, and no amount of extra sample separates it from trust tier.
The breadth floor was met by tag granularity, not breadth. Sixty-five "categories" includes individual players and individual leagues.
One filed BAD condition was structurally unobservable. The oracle-dispute branch could not fire on our data path: the venue field we read reports a terminal state, so a market disputed and then settled reads as resolved. We have never observed a dispute on this path. Under our 2026-06-07 deviation record a disputed-then-upheld resolution is CLEAN in any case, since the market reached the correct terminal price. But "we counted disputes and found none" would be false; the branch could not fire.
Every BAD event we observed was the venue's mid-price void settlement, where all outcomes pay equally. That is not a refund: a holder who paid 0.90 receives 0.50.
Market Trust remains a preview surface. Nothing here promotes it.
Why this note stops where it does
There is a second finding we are deliberately not publishing yet.
Market Trust's top tier requires four of six evidence pillars measured, and 0.138% of scored markets clear it. Almost none of that is a fact about the markets; it is a fact about how much of each market we have measured. The rating is substantially reporting our own coverage while wearing the grammar of a judgement about the venue. That is the inverse of surveillance bias, the known effect where looking harder finds more: here, looking less scores lower. It also has a known fix, which medicine solved with GRADE by reporting the estimate and the certainty as two orthogonal quantities rather than downgrading the estimate when evidence is thin.
We are holding that half because the fix is not shipped. Publishing a live product defect alongside a remedy that renders nowhere would be the same move this note criticises. When the distinction between "we did not assess this" and "we assessed it and it failed" reaches an actual reader, that becomes a different and better artifact than a confession.
Reproducibility
Every number above is a count or an interval over our own stored artifacts. The terminal verdict artifact for the 2026-07-30 analysis point is pinned at an immutable dated path in the repository alongside the filed registration text, the tier definitions, the freeze hash and the four deviation records; the daily latest.json is a rolling file that has since advanced past it. Intervals are computed with the same library the production verdict uses. The pre-registration's public receipt is aspredicted.org/zb2ca7.pdf, listed with every other filing on our pre-registration registry.
Convexly builds independent audit infrastructure for prediction markets. We publish negative results, including our own.