Research note · 2026-08-21

Our pre-registered calibration test returned a null. Here is what we built because of it.

Market Trust v0.2 rated markets on operator-set absolute thresholds. We filed a forward test of those ratings before any outcome was known; it reached its filed analysis point and could not separate the tiers, and we published that null with its full mechanism on 2026-08-07. This page is the revision the null prescribed: the v0.3 lane that replaces absolute thresholds with horizon-matched peer baselines. It also publishes what our own price archive says about Polymarket close prices, because the revision should be judged against measured venue behavior, not assumptions.

What we filed, and what came back

On 2026-06-06 we pre-registered a forward calibration test of Market Trust v1 ratings, the frozen formula behind the v0.2 cards (AsPredicted #295174): would markets rated higher resolve more cleanly than markets rated lower, on resolutions occurring only after the filing? The method, weights, and cutpoints were frozen at filing and are pinned by a content hash; any post-freeze change fails CI.

The test reached its filed analysis point on 2026-07-30 and returned no_separation; the registered null wording is "Market Trust tiers do not separate on resolution cleanliness." That wording must not be read alone. The arms at the filed analysis point were 4 markets (4 of 4 clean) against 1,770 (1,764 of 1,770 clean), and the one-sided 95% lower bound on the arm difference was -0.400: the contrast was never established, which is not the same as the tiers being alike. No attainable outcome of the four-market arm, 0 of 4 through 4 of 4 clean, clears the filed bar at that composition; the published note carries the full sweep. That makes this a structural null: the rating machine never populated its own upper tier, and the result is uninformative either way about the instrument itself, not evidence that the ratings track resolution quality and not evidence that they do not.

We published this result on 2026-08-07, with the full mechanism of why the null is uninformative and the pre-filing check that would have caught it, in The guard we built for this failure watched the wrong axis. The registration's public receipt is at aspredicted.org/zb2ca7.pdf. A pre-registered null is a valid, publishable outcome. Our filed negative-result rule prescribes what follows: publish the null, keep the canary label, revise only under a new frozen version, and never mutate the frozen rows. Our standing rule adds that a revision graduates only under its own future filing. This page is that revision.

Why the machine was degenerate

Of the 18,302 cards the v0.2 machine has produced through 2026-08-21, 98.4% carried the bottom rating and exactly one market ever earned the top one. The cause is structural: v0.2 scored price-support pillars on absolute thresholds, so a 3-cent book was judged by the same spread bar as a 90-cent book, and nearly every market failed. A rating machine that gives almost everything the same grade cannot be calibrated against outcomes, because there is nothing to separate. We published every one of those cards while this was true; each carried its canary label, and this page is the accounting.

v0.3: judged against horizon-matched peers

The v0.3 lane scores a market's book quality as a percentile among peer markets in the same time-to-scheduled-close bucket and the same price-boundary band. Spread is normalized by the cheaper side, so the 3-cent and 90-cent books are finally measured on the same axis, and spread and cheaper-side depth are each scored against their own cell. Every cell must hold at least 50 distinct markets and 200 market-day observations over a trailing 28-day window; a cell below its floors publishes a named shortfall and no number. As of 2026-08-21, 5 of 15 cells clear the floors, and the rest say so.

The lane is live today as a labeled preview block on any market page that has a fresh reading, computed within the last 48 hours; a market with no fresh reading shows nothing rather than a stale number. It is physically separate from the frozen v0.2 lane, it does not feed the Market Trust rating, and it graduates into a rating input only under its own pre-registered filing, after the same kind of forward test that just returned this null.

What close prices actually did

A revision built on scheduled-close anchoring should show its work on real venue behavior, so we measured it: our own orderbook snapshots of resolved Polymarket markets, joined to our own on-chain resolution sweep, anchored to the venue's scheduled close.

The headline is heterogeneity, not a single number. What "price at close" means depends on where the venue schedules close: Esports markets mostly resolve before their scheduled close (67.9% of 393), so the final book has usually watched the match; Sports markets' scheduled close sits at game start (5.9% of 170 resolve early), so the final book is a genuine forecast. Their last-price Brier scores differ, 0.027 [0.016, 0.040] (n=249) against 0.179 [0.151, 0.207] (n=166); the categories also differ in price composition, and pooled numbers blend these regimes.

Day-of prices (the final-24h median, 605 markets) score a Brier of 0.131 [0.117, 0.146] and deviate from the diagonal in the mid range: in the 87 markets where one side's day-median sat at 60 to 70 cents, that side won 80.5% of the time [72.0, 88.5]; the 30-to-40-cent row of the figure is the same markets by complement, not a second finding. That deviation concentrates where the measurement window spans the event itself, so we do not read it as market-wide mispricing. Win rates rise with price under both snapshot rules, with two small local inversions; prices were directionally informative.

Panel A: day-median prices, final 24h before scheduled close

002020404060608080100100209817886151150877881209Price (cents) with contract count per binObserved win rate (%)

605 markets. Brier 0.131, 95% CI [0.117, 0.146]. Mirrored bins are one deviation: ten bins, five free comparisons.

Panel A2: last observed pre-close price (partly post-outcome for in-event categories)

0020204040606080801001004193233468281473329422Price (cents) with contract count per binObserved win rate (%)

612 markets. Brier 0.085, 95% CI [0.071, 0.099]. Mirrored bins are one deviation: ten bins, five free comparisons.

Panel B: close-anchored horizon coverage

Final 24h before close

829

supported

1 day before close

104

below floor, no curve drawn

1 week before close

80

below floor, no curve drawn

1 month before close

29

below floor, no curve drawn

3 months before close

0

below floor, no curve drawn

The publication floor is uniform and operator-set: every interior price bin must hold at least 20 contracts. Only the close bucket clears it at current coverage.

Category split, both snapshot rules

CategoryMarketsBrier, day-median ruleBrier, last-price rule
Esports2490.129 [0.114, 0.148]0.027 [0.016, 0.040]
Sports1660.182 [0.155, 0.208]0.179 [0.151, 0.207]
uncategorized1080.077 [0.044, 0.115]0.069 [0.034, 0.107]
19 further categories below the 50-market floor: gated, no cell published.
Close-anchored read: prices are Convexly's own orderbook snapshots of 605 resolved Polymarket markets with priced coverage in the final 24h before the venue's scheduled close (of 68,026 conditions our collectors have observed since 2026-05-07; the full denominator chain and extraction SQL ship with this figure). Outcomes are Convexly's own on-chain resolution sweep. Each market contributes one price and its complement, so the curve is antisymmetric by construction and intervals (percentile bootstrap, market-clustered, B=2000) should be read on five free bins. The tracked set skews liquid, May-August 2026, mostly esports/sports; pooled numbers blend the three close-semantics regimes shown in the category table. Scheduled close is venue-editable metadata, not a trading halt. Descriptive figure; not a readout of the pre-registered Market Trust calibration test (AsPredicted #295174), whose frozen method, stores, and forward window are untouched here.

Reproduce this

The extraction SQL, the 829-market extraction, the results JSON, and the analysis script (seed 20260821, B=2000; results regenerate byte-identically) ship in the repository, alongside a dated receipt for the card census on this page. The chain's endpoints are stated in the figure caption; the full denominator chain from 68,026 observed conditions down to 605 priced close-window markets ships with the bundle, and the review record for every claim on this page, including the claims we blocked, is committed alongside the data.

What travels with this page

Market Trust remains a canary preview. Ratings are dated, descriptive diagnostics of market quality, never predictions of outcomes and never claims about what a probability should be. The v0.3 preview does not feed any rating. This page summarizes the already-published #295174 verdict and adds nothing to it; the authoritative account is the 2026-08-07 note and the pinned artifact it quotes. Nothing here previews any future filed test.