For external methodology changes filed after 2026-04-25, Convexly attempts to lock hypotheses before analysis and then publishes the receipt status here. External links are shown as verified only when the linked page contains the expected AsPredicted ID, title, and filing date; otherwise they are marked pending, stale, or broken. The verdict + run date + effect-size CI are reported within 24 hours of the test running. The original V1 and V1-M papers were not retroactively pre-registered; they remain frozen-coefficient methodology with ex-ante version-controlled commitment via the SHA-256 audit chain. Failed methodology tests land in the negative-result registry; the audit chain is verifiable in your browser at /research/verify.
Last updated 2026-09-04T12:30:00Z.Receipts checked 2026-08-07T10:48:28Z.15 entries
Receipt health: 7 externally verified/8 pending public URL/0 broken or stale
AsPredicted #287368
Filed 2026-04-25 · Ran 2026-04-27
V1.5 follow-up experiments E2 + E7
Failed (rejected)External receipt verified
E2 per-wallet temporal holdout: ρ = +0.111 [+0.046, +0.175], well below the +0.30 pre-reg threshold. E7 per-quarter IC stability: median ρ = +0.038, only 3 of 5 quarters positive vs ≥5/6 required. Both failed.
Initial in-sample test of skill-weighted aggregation as per-market price prior. 24 aggregator variants tested; all rejected. Cohort substitution amendment filed as #287714.
Cohort substitution from V1 (8,656 wallets) to V1-M (8,778 wallets) to verify the negative result is not cohort-specific. All 24 aggregator variants rejected on V1-M as well; consistent with the original finding.
V2.8.2 wash-filter TOST equivalence test on V1-M Polymarket cohort (Sirolly-adapted)
PassedPublic receipt not verified
Wash-filter robustness check on the V2.8.2 negative result. Brier delta CI [+0.16028, +0.19287] sits inside the pre-registered TOST equivalence range [+0.154, +0.204]. Movement after wash filtering: +0.00243 Brier (1.4% relative). The V2.8.2 finding (skill-weighted aggregation rejected) is robust to wash-trader filtering at composite-z >= 3.0.
CME V0.2 backtest: 90-day walk-forward on Polymarket constraint-projection signals
PendingPublic receipt not verified
90-day walk-forward backtest of the CME V0.2 constraint-projection pipeline pre-registered. Hyperparameters frozen ex-ante (thresholds, sizing, cost model, performance metrics). No hyperparameter tuning based on backtest results allowed by the pre-reg.
Status note 2026-08-28: the 90-day walk-forward run was targeted for 2026-07-29 and has not run; no verdict artifact exists. The date slipped and this entry now says so rather than carrying a past date as a plan. The linked paper carries the same dated slip note (2026-08-13). A new run date will be recorded here when one is set.
AsPredicted #288610
Filed 2026-05-01
V2-Perps Edge Score: skill ranking with CRPS + funding-capture pillars
PendingPublic receipt not verified
Pre-registers the form (4 pillars: CRPS-posture, conviction, discipline, funding-capture) + 7 validation gates for the V2-Perps Edge Score composite. Form locked at a pinned build; coefficients TBD pending Hyperliquid 90-day cohort fit. Composite reduces to V1 / V3b on binary outcomes (Brier-equivalence identity) and extends across crypto perps, equity perps, compute futures, AI benchmark markets, valuation futures, and prediction markets per spec Section 6.
Status note 2026-08-28: the 90-day walk-forward run was targeted for 2026-07-30 and has not run; no verdict artifact exists. The date slipped and this entry records the slip rather than carrying a past date as a plan. A new run date will be recorded here when one is set.
Strictly-prospective realized-vs-control validation of CME signals. H1: realized PnL of the CME-chosen side at the USD 1,000 capacity tier exceeds the mean of K=20 matched-noise controls, one-sided paired permutation at alpha = 0.025, AND the 95% bootstrap CI lower bound for mean paired difference is > 0. Evidence window = 92 calendar days beginning the first full UTC signal-emission day after the filing timestamp; pre-filing/same-day signals excluded. Reports insufficient_sample if fewer than 30 resolved signal/control pairs by the analysis date; no threshold tuning or window extension without a new pre-registration. CME methodology frozen for the window.
Deviation note filed and founder-approved 2026-08-28, while the resolved pair count stood at zero: three input-wiring defects (archive file selection, look-ahead timestamp reference, control join key) are corrected, and the 92-day emission window is declared as 2026-06-01 through 2026-08-31 inclusive, the reading the filed text supports and the shorter of the two candidates. The frozen statistical recipe is unchanged. Full note: the 2026-08-28 deviation note for this filing, held in the private repository and supplied on request.
This verdict is published with its settlement-boundary disclosure and is not quotable as a result; see the disclosure beside the artifacts. Settlement-boundary disclosure
PASSED: persistence_confirmed. Window one, 2026-06-02 to 2026-08-30, closed. First closed-window computation 2026-08-31, recomputed 2026-09-01. Both pre-registered legs pass on the frozen 178-candidate / 3,693-control sets; the field-6 both-in-window rerun agrees, so no partial_persistence downgrade. Terminal readout, dated snapshots, disclosures and the written readout: https://www.convexly.app/research/forward-validation.
Window one CLOSED 2026-08-30 (contiguous with #303724's window two, which opened 2026-08-31). The terminal harness readout was first computed on 2026-08-31, the day after close, and recomputed on 2026-09-01 as late in-window captures landed; per the window-two filing's field-8 commitment it auto-published through the fail-closed window_closed gate on /research/forward-validation as persistence_confirmed, both pre-registered legs passing on the frozen 178-candidate / 3,693-control sets. The artifact was then FROZEN at the 2026-09-01 computation and both computations are served there as dated snapshots with their sha256. The binding caveat travels with the numbers: Kish effective clusters 5.28 against 139 active candidate wallets on the pooled contrast, and the equal-weight contrast is a pre-registered secondary check, a different estimand from the pooled excess. The written narrative readout was published on 2026-09-02 and this entry's verdict was updated in the same commit, so the 24-hour verdict-update rule was missed by two days for this entry; that gap is recorded here rather than papered over. Publication provenance: this project's automated independent-review agent (an in-house AI check with a non-overridable veto, not a third-party or human review) gated the narrative on 2026-09-01 (pass with required edits, applied) and required further edits on 2026-09-02; the founder, the only human in the chain, approved publication on 2026-09-02. Evidence note: the 2026-09-01 terminal-read note for this filing, held in the private repository and supplied on request. SETTLEMENT-BOUNDARY DISCLOSURE (recomputed 2026-09-04): window membership in this test is applied to a market's scheduled end date, not its settlement time, a deviation from the filed word resolve. Recomputed 2026-09-04 on each market's actual settlement time, taken independently from the venue's closedTime and the Polygon settlement block (the two disagree on window membership for 1 of the 211,518 markets both cover and on the settlement day for 16; the counts here are on the Polygon settlement block, which covers all 211,520): 1,549 of 71,529 candidate positions (2.2%) and 23,213 of 443,880 control positions (5.2%) had settled outside the window they were counted in. On the settlement-defined pool the frozen rule returns the same label, persistence_confirmed, with the pooled contrast at +2.96pp (was +2.93pp), leg A p = 0.0029, and the candidate BCa interval +2.08pp to +6.32pp over 127 active wallets (was 139); the Kish effective-cluster caveat was computed on 139 wallets and has not been recomputed on 127. The 10.7 percent figure previously reported here measured a different quantity, positions first archived before their scheduled day, and is superseded by these counts. The opposite direction, positions scheduled outside the window that settled inside it, was never captured by this collector, so it is unmeasured; this verdict is honest to publish but not quotable as a result until the wrongly-excluded direction is measured.
PUBLIC RECEIPT VERIFIED 2026-07-25: the author-created anonymous AsPredicted PDF is live at https://aspredicted.org/xq76gq.pdf and resolves with the filed title, the 2026/06/01 05:46 PT filing timestamp, the 178-candidate / 3,693-control design, and the frozen content hash 0e12fe1ffd, with no author identity exposed (anonymous per AsPredicted blind-review default; URL is permanent and survives any later deanonymization). Filed independent of #294035. The window-one terminal readout is on /research/forward-validation; the in-sample FDR-cleared candidate registry itself is pending a data-release review.
TERMINAL, analysed once at the filed analysis point (2026-07-30, the first window meeting every filed floor: 3,876 resolved forward cards, 65 categories, 6 BAD markets). Verdict `no_separation`, filed wording 'Market Trust tiers do not separate on resolution cleanliness'. THAT LABEL OVERSTATES WHAT THE DESIGN COULD SUPPORT and must not be quoted alone: the higher-trust arm held FOUR markets (4/4 clean, Wilson 95% [51.0%, 100%]) against a lower arm of 1,770 (1,764 clean, [99.26%, 99.84%]), giving a one-sided 95% Newcombe lower bound of -0.400. Holding the lower arm at its observed composition, EVERY attainable outcome from 0/4 to 4/4 returns a negative bound, so the confirmatory branch could not fire on any data that arm could produce. The correct reading is that the contrast is not established, NOT that the arms are alike; no equivalence margin was filed, so equivalence is unavailable in either direction. Under field 7's no-optional-stopping clause this registration is spent and the question cannot be re-asked under it. Four deviation records are on file (UMA classification 2026-06-07, intake sampling 2026-07-03, BAD-unit operationalisation 2026-07-26, writer pagination 2026-07-26).
PUBLIC RECEIPT VERIFIED: the author-created anonymous AsPredicted PDF is live at https://aspredicted.org/zb2ca7.pdf and is the receipt cited in the published research note. CORRECTED 2026-08-10: this field previously read 'public AsPredicted receipt pending. The PDF is private, so this filing has NO receipt a reader can reach and must never be cited externally as pre-registered.' That note went stale when the receipt was made public and was never updated, so this entry contradicted its own receipt_status of verified_external. receipt_status is authoritative and was correct; the note was wrong. The terminal read returned no_separation and is published in full.
AsPredicted #296460
Filed 2026-06-12
Anytime-valid e-process monitors over forward streams (Convexly, 2026)
PendingExternal receipt verified
No result recorded yet. The hypothesis is filed and the test has not been run, so there is nothing to report. The verdict lands here within 24 hours of the run.
CLV-edge forward skill test, wallet side (Convexly, 2026)
Published, not quotableExternal receipt verified
This verdict is published with its settlement-boundary disclosure and is not quotable as a result; see the disclosure beside the artifacts. Settlement-boundary disclosure
CORRECTION 2026-09-04: the pass published on 2026-07-03 (rho +0.159 [+0.063, +0.255], 464 wallets) admitted positions by scheduled end date; under settlement time that artifact did not meet the floors (155 wallets of 300) and was not the confirmatory read the filing defines; its below-floor correlation (+0.145 [-0.023, +0.306]) carries no verdict. PASS at the confirmatory read the filing defines, which under settlement time is the 2026-07-11 archive: Spearman +0.188, 95% CI [+0.071, +0.302], 304 wallets (floor 300, met by four), 20,101 positions; secondary +0.045 [-0.062, +0.153], spans zero. This read is a reconstruction made on 2026-09-04 by re-applying the filed stopping rule to the frozen archive, after the 2026-07-03 pass and the below-floor number at that artifact were both known; the rule admits one answer and both settlement sources return the same archive. Later archives are descriptive only under the frozen rule: sampled weekly through 2026-09-01 the frozen replay returns a pass at eight of nine later archives and not at 2026-07-18 (lower bound -0.003), with rho declining from +0.188 to +0.107 (95% CI [+0.035, +0.179] at 2026-09-01, 741 wallets); the confirmatory read is the highest point of that series. The filing's entry-time criterion is not applied by the frozen replay and cannot be measured from the archive, which records no entry time. REGISTRY CORRECTION 2026-08-13: this entry rendered 'no result recorded' for 41 days after the 2026-07-03 founder-approved publication because the manifest was not updated in the publication pass. The manifest is a second renderer of verdict facts; it must move in the same commit as any verdict publication.
SETTLEMENT-BOUNDARY DISCLOSURE (corrected 2026-09-04): window membership was applied to each market's scheduled end date; the filing said resolve. Recomputed on settlement time from the venue closedTime and the Polygon settlement block (agree on membership for all but 2 of 38,016 markets, on the day for all but 3): 17,561 of 59,945 admitted positions (29.3%) and 10,256 of the 19,740 in the correlation pool (52.0%) had settled before the window opened (wrongly admitted direction); 316 positions on the venue operator (313 on chain; 133 in the correlation pool on both) had been wrongly excluded and are admitted. Two markets without a venue record kept their scheduled date under the venue operator; the chain operator covers all. Artifacts published beside the paper, each naming its inputs by path and sha256. Not quotable as a result until the disclosure sidecar binds the result to this note and to the later-archive series.
AsPredicted #296463
Filed 2026-06-12
Resolved-market calibration bundle, Market Trust side (Convexly, 2026)
PendingExternal receipt verified
No result recorded yet. The hypothesis is filed and the test has not been run, so there is nothing to report. The verdict lands here within 24 hours of the run.
Wallet-skill FDR candidate-set forward-persistence WINDOW TWO (discretionary cohort)
PendingExternal receipt verified
Window two of the strictly-prospective forward-persistence protocol first filed as #294147. Frozen objects are UNCHANGED and pinned by the same content hashes (discretionary snapshot sha256 0e12fe1ffd9e38d364d85ad6530ea5f9fedd42e536a0b0d2b910164f88e99b88; parent extract sha256 cf95c5f1b9a957477f6b3c98bdf235744eff8f34230083a4c8bffd4657ef2ec6): the realized-edge measure mean(won - vwap_prob), the 178-wallet candidate set, the 3,693-wallet control set, and the micro-market exclusion rule. H1 (both legs required, identical in form to window one): (a) RELATIVE, candidate-set pooled forward edge exceeds control-set pooled forward edge, one-sided wallet-label permutation at alpha = 0.025, 10,000 permutations, seed 20260725; (b) ABSOLUTE, candidate-set pooled forward edge has a 95% BCa wallet-cluster lower bound above 0, 2,000 resamples, seed 20260726. Evidence window = the 90 calendar days 2026-08-31 through 2026-11-28, contiguous with window one and sharing no position. Floors unchanged: >=10 forward positions/wallet, >=40 candidate wallets, >=1,000 candidate forward positions, else insufficient_sample. The question this window adds is DURABILITY: whether an edge that persisted across one window persists across a second, and with what attenuation. Only the SIGN and the joint pass rule are registered, never a magnitude; window one's observed values are explicitly NOT registered as an expected effect size. A pre-registered null (no_persistence) or insufficient_sample is a valid, publishable outcome, and field 8 commits to publishing whichever of the four verdicts the frozen rule returns.
Post-close late-capture allowance FIXED 2026-09-02, before any window-two verdict existed and while the window held almost no evidence: the first closed-window computation runs 2026-11-29 (expected_run_at_utc); the computation of 2026-11-30 is the terminal read and freezes the artifact, absorbing two days of late capture, the same allowance window one used. Recorded here so the freeze day is a property of the protocol visible in advance, not a choice made after the fact. SETTLEMENT-BOUNDARY DISCLOSURE (2026-09-03): as in window one, membership is applied to a market's scheduled end date, not its settlement time; the disclosure and the owed settlement-timestamp recomputation apply to this window's verdict when it is read, and until that recomputation lands the verdict will be honest to publish but not quotable as a result.
PUBLIC RECEIPT VERIFIED 2026-07-26 at filing time: the author-created anonymous AsPredicted PDF is live at https://aspredicted.org/228s3x.pdf and resolves with the filed title, the 2026/07/26 03:40 PT filing timestamp, and AsPredicted #303,724. Every paragraph of all eight filed fields was field-diffed against the on-disk filing source and matched verbatim, including both frozen content hashes, both seeds, the four-valued verdict vocabulary, and the standing-cadence commitment. This is WINDOW TWO of the same protocol as #294147: same frozen 178-candidate / 3,693-control sets, same frozen realized-edge measure, same micro-market exclusion, same floors, same joint verdict rule. Its window is 2026-08-31 to 2026-11-28, defined by window one's CLOSE rather than this filing's timestamp, so the two windows are contiguous and share no position. Filed BEFORE window one's terminal read exists; field 8 carries an explicit non-blindness disclosure stating that the interim maturation label was visible at filing and that the residual is the decision to CONTINUE the protocol. External citation of the FILING is receipted; the VERDICT stays founder-gated to the terminal read on/after 2026-11-29.
Filing policy
Filing rule: For post-2026-04-25 methodology changes that affect external claims, Convexly either files a pre-registration before analysis runs or marks the item internal-only / pending-public-url until an external receipt can be verified.
Verdict update rule: When a pre-registered test runs, the verdict + run_at_utc + verdict_summary are updated within 24 hours of the run completing. Verdicts are PASSED, FAILED, or PENDING. Failed pre-registrations are added to the negative-result registry at /research/negative-results.
Supersession rule: When a pre-registration is superseded by an amendment, the original entry is kept (verdict noted as superseded) and the amendment is added as a separate entry. Original entries are never removed.
Audit-chain link: Every entry's audit_chain_anchor field references the SHA-256-hash-chained run identifier in the public audit log at /research/cme/audit_log.jsonl (or paper-specific provenance log). The /research/verify page walks the chain in client-side JavaScript and renders a green stamp if every prev_hash matches its parent's row_hash.