Every pre-registered hypothesis Convexly has tested and rejected, with the AsPredicted ID, the effect-size confidence interval, the specific claim disproved, and the open questions that remain. Indexed in machine-readable form at /research/negative-results.json and verifiable via the audit chain.
Skill-weighted aggregation as a per-market price prior
Rejected
Twenty-four pre-registered skill-weighted aggregator variants (linear, exponential, rank-based, top-k, conviction-weighted, posture-weighted, calibration-weighted, multiplicative, et al.) were tested as forecasts. All twenty-four were worse than the market price. The wash-trader-filtered variant Brier delta CI [+0.16028, +0.19287] sits inside the pre-registered TOST equivalence range [+0.154, +0.204]; movement after wash filtering was +0.00243 Brier (1.4% relative).
Specific claim disproved
Aggregating Polymarket trader forecasts by any skill-weighted shape produces a per-market probability prior that is more accurate than the market price.
Analysis date
2026-04-27 (2026-04-27T23:14:39Z)
Effect-size metric
Per-market Brier delta vs market-implied baseline; positive means the skill-weighted aggregator was worse.
Effect-size CI
+0.179, primary W-EXP beta=4 a=2.0 S-RAW, 95% CI [+0.164, +0.1935]; best of 24 variants still +0.0570, 95% CI [+0.0486, +0.0654].
Legacy effect-size field
+0.179 [+0.16028, +0.19287]
Direction
Aggregator was WORSE than market price by 17.9% Brier on average.
Pre-registered TOST equivalence range
[+0.154, +0.204]
Decision rule
Reject H0 only if delta_market < -0.005 and CI_hi < 0.0. Observed deltas were positive for all 24 variants, so zero variants beat the market-implied baseline.
Receipt status
pending public url
Filed / ran
Filed 2026-04-25, ran 2026-04-26, wash-filter robustness 2026-04-29
What remains open
·Per-wallet ranking (V1, OOF Spearman +0.514, NOT closed by V2.8.2)
·Within-wallet z-score for insider-pattern detection MVP (NOT closed)
·Wash-trading detection (Sirolly methodology, NOT closed)
·Structural / constraint-projection inference (CME V0.1, uses MARKET PRICES not skill-weighted aggregation as forecast input, NOT closed)
Implications for product
Bright line: when evaluating any new commercial proposal, check first 'does it use skill-weighted aggregation as a per-market price prior?' If yes, reject unless there's a substantive reason to believe the prior shape was wrong vs the 24 already-rejected shapes.
Data availability
No public data bundle exists for this negative result. Corrected 2026-08-10: this entry previously carried data_bundle_url = https://www.convexly.app/research/marketalpha-v2/v2-8-2-data-bundle.tar.gz. That file has never existed on this site, so the URL returned 404 for the life of the entry, and the registry page rendered it as a download button unconditionally. V2.8.2 ran on the 8,997-row Polymarket wallet universe, which is held internally and has never been published, so this result is recomputable by Convexly and not by a reader. The prior URL is preserved in this field rather than overwritten, per the addition_rule below.
V1.5 experiment E2 (pre-registered): refit V3b OLS on a 2024-01-01 to 2025-09-30 training window with a 14-day purge, score on a 2025-10-15 to 2026-04-15 held-out window. Frozen V1 coefficients applied to training-window pillars produced Spearman ρ = +0.111 with 95% CI [+0.046, +0.175] and p ≈ 0.001. Positive and the CI excludes zero, but well below the pre-registered ρ ≥ +0.30 pass threshold. The fail is reported per the pre-reg's no-post-hoc-re-analysis rule.
Specific claim disproved
The V1 V3b composite, fit on full-cohort data and applied to training-window pillars, predicts forward held-out PnL at Spearman ρ ≥ +0.30 on the V1-M cohort.
Analysis date
2026-04-27 (2026-04-27T19:17:00Z)
Effect-size metric
Spearman rank correlation between V3b score and held-out signed log PnL.
Effect-size CI
Primary refit V3b: rho = -0.0821, 95% CI [-0.149, -0.006], N = 805 paired wallets. Secondary frozen V1 V3b: rho = +0.1114, 95% CI [+0.046, +0.175].
Legacy effect-size field
Spearman ρ = +0.111, 95% CI [+0.046, +0.175], p ≈ 0.001
Direction
Positive but below the +0.30 pre-registered threshold.
Pass required rho >= +0.30 and CI lower bound > 0. The secondary frozen-coefficient result was positive but below the +0.30 threshold; the primary refit result was negative.
Receipt status
verified external 287368
Filed / ran
Filed 2026-04-25, ran 2026-04-27
What remains open
·V3b as a CROSS-SECTIONAL ranker (the per-period ranking + S4 partial Spearman +0.494 controlling for log capital remain open)
·V3b as a per-wallet TEMPORAL forecast: closed by E2; do not market it as such
Implications for product
V3b is positioned as a cross-sectional ranker, not a per-wallet temporal predictor and not a forecast-aggregation weight. External framing should never claim 'V3b predicts your forward PnL across windows' without citing this E2 rejection.
V1 V3b per-quarter IC stability across the pre-registered window
Rejected
V1.5 experiment E7 (pre-registered): bucket V1-M positions by quarter, recompute V3b with quarter-local standardization, compute Spearman of frozen V3b vs signed log PnL in each quarter. Pre-registered window 2024Q1 to 2025Q2 (six quarters); 5 of 6 quarters present (2024Q1 had 29 wallets, below the ≥5-position cohort filter). Median per-quarter Spearman = +0.038, range across quarters [-0.164, +0.155]; 3 of 5 quarters positive. Pre-registered pass criterion required median ρ ≥ +0.30 AND ≥5/6 positive quarters. Both legs failed.
Specific claim disproved
The frozen V1 V3b composite produces a stable per-quarter Spearman with realized log PnL across the 2024Q1 to 2025Q2 window.
Analysis date
2026-04-27 (2026-04-27T19:17:00Z)
Effect-size metric
Median quarterly Spearman rank correlation of frozen V3b vs signed log PnL.
Effect-size CI
Median per-quarter rho = +0.038 across 5 present quarters; per-quarter range [-0.164, +0.155]. 2025Q1 rho = -0.164, 95% CI [-0.221, -0.102].
Legacy effect-size field
Median per-quarter Spearman = +0.038, range [-0.164, +0.155] across 5 quarters
Direction
Per-quarter IC is unstable; one quarter (2025Q1) is meaningfully negative with CI [-0.221, -0.102].
Pre-registered TOST equivalence range
Pre-registered: median ρ ≥ +0.30 AND ≥5/6 positive quarters. Observed: median +0.038, 3 of 5 positive. Both legs fail.
Decision rule
Pass required median rho >= +0.30 and at least 5 of 6 quarters positive. Observed median was +0.038 and 3 of 5 present quarters were positive.
Receipt status
verified external 287368
Filed / ran
Filed 2026-04-25, ran 2026-04-27
What remains open
·Exploratory extension to 2025Q3 to 2026Q2 (NOT in pre-reg) shows 2026Q1 ρ = +0.281, CI [+0.249, +0.314]; whether IC stability returns in more recent windows is open and pre-registered for replication as #11 in the 12-month plan
·Cross-quarter autocorrelation of per-quarter ρ values is +0.366 lag-1, so when the V3b signal is on it tends to persist; the failed E7 test was on the level of the IC, not its persistence
Implications for product
V3b is not marketed as a stable per-quarter IC predictor. A genuine forward out-of-sample publication (#11 in the 12-month plan, ships 2026-10-01) is the next opportunity to revisit the per-quarter stability claim with a fresh pre-reg. Note: the existing rolling-Spearman series is an in-sample contemporaneous diagnostic with target leakage (score and PnL share positions; conviction is PnL-derived), not a forward test.
Market Trust v1 tiers and resolution cleanliness (forward calibration)
Null: contrast not established
Pre-registered forward test of whether higher Market Trust tiers resolve more cleanly. Analysed once at the filed analysis point (2026-07-30, the first window meeting every filed floor: 3,876 resolved forward cards, 65 categories, 6 BAD markets). The higher-trust arm held FOUR markets (4/4 clean, Wilson 95% [51.0%, 100%]) against a lower arm of 1,770 (1,764 clean, [99.26%, 99.84%]); one-sided 95% Newcombe lower bound -0.400, Fisher one-sided p = 0.9865. Holding the lower arm at its observed composition, every attainable higher-arm outcome from 0/4 to 4/4 returns a negative bound, so the confirmatory branch could not fire on any data that arm could produce. The filed label no_separation overstates what the design could support: the correct reading is that the contrast is NOT ESTABLISHED, not that the tiers are alike. No equivalence margin was filed, so equivalence is unavailable in either direction.
Specific claim disproved
None disproved. The filed hypothesis (higher Market Trust tiers resolve more cleanly) was neither confirmed nor refuted: attainable power against every alternative was zero. Recorded under the registry's failed-or-null rule as a null.
Analysis date
2026-07-30 (2026-07-30)
Effect-size metric
Clean-resolution share, higher-trust arm minus lower arm
Effect-size CI
One-sided 95% Newcombe lower bound -0.400 (4/4 vs 1,764/1,770); two-sided interval not filed
Direction
Not interpretable as a direction: the higher arm's four markets cannot move the bound positive under any outcome.
Decision rule
Analyse once at the first window meeting every filed floor (3,876 resolved forward cards, 65 categories, 6 BAD-unit markets); no optional stopping; the registration is spent after the single read.
Receipt status
verified external
Filed / ran
Filed 2026-06-06, ran 2026-07-30
What remains open
·Whether Market Trust tiers separate on resolution cleanliness. Roughly 909 higher-tier resolved markets would be needed at the observed base rates before the filed contrast could fire, and under the filing's no-optional-stopping clause the question cannot be re-asked under this registration.
Implications for product
Market Trust remains canary_preview. Tier language may not claim separation on resolution quality in either direction; the attainability context travels with any quotation of no_separation.
Addition rule: An entry is added the moment a pre-registered hypothesis is rejected by the test it specified. Entries are never removed; corrections to the analysis are added as additional fields, not by overwriting prior content.
Verification: Each entry exposes analysis_date, effect_size_metric, effect_size_ci, pre_registered_decision_rule, and AsPredicted IDs. Public receipt health is governed by /research/preregistrations.json; pending-public-url entries are not treated as externally verified until their public AsPredicted page resolves to the expected ID, title, and filing date.
Credibility-loaded-term audit: This file passes the credibility-loaded-term audit per the pre-publication claim audit: every effect-size estimate has a CI in the same line; 'rejected' refers to a formal pre-registered test, not informal opinion; 'robust' and 'near zero' do not appear without their CI / TOST counterpart; only pre-registered tests are listed (V1.5 S6 was an exploratory follow-up, not a pre-registered hypothesis, and is therefore not a registry entry).