AI tournament · phase 1

Three AIs forecast the market every morning. We score them on calibration.

Haiku 4.5, Sonnet 5 and Opus 5 receive the same data package on 10 S&P 500 names every day at 12:30 UTC and return, before the open, the probability that each one closes up and how much it will move. At the close we score them with Brier and CRPS and publish the result whatever it says.

Phase 1: registration. 55 resolved predictions per model out of the 200 pre-registered before anything can be claimed. This is the state of the counter, not a result.

What is measured

Calibration, not hit rate: Brier and its decomposition (resolution) for direction; a three-quantile CRPS and the coverage of the 80% band for the size of the move.

What we expect

The pre-registered expectation is that no AI discriminates one-day direction: in this universe it is measured as null. Where value may exist is in how much a name moves, not which way.

What it is not

Not a recommendation or a signal. No model sees headlines or searches the web; only prices, volatility, catalyst news counts, the main macro calendar (Fed, CPI, jobs) and, in variant B, quarterly-results context.

Backtest results (simulation), not real trading. Not investment advice. Past performance does not guarantee future results.

Reading

Sentences generated by fixed rules from the numbers, using the pre-registered thresholds. Nobody writes them by hand.

Updated 2026-09-11 11:26 UTC

Paper P&L

Fixed, pre-registered rule: each name is held long or short in proportion to how far the probability sits from 0.50, equal capital per name, one-session holding, close to close, 10 bp round-trip cost. The basket is the same list of names bought equal-weight. 🏆 = best net return among models with the same sessions; below 250 sessions the leader is noise.

ModelVar.SessionsNet returnbp/daySharpeMax drawdownBasketvs basketp
🏆Sonnet 5A50.04 %0.8-0.0 %-2.73 %2.77 %0.29
Sonnet 5B30.02 %0.5-0.0 %-2.98 %2.99 %
Opus 5A50.02 %0.3-0.0 %-2.73 %2.74 %0.29
Opus 5B30.01 %0.40.0 %-2.98 %2.99 %
EnsembleB30.01 %0.2-0.0 %-2.98 %2.98 %
Haiku 4.5X10.00 %0.00.0 %0.00 %-0.00 %
Opus 5X10.00 %0.00.0 %0.00 %-0.00 %
EnsembleX10.00 %0.00.0 %0.00 %-0.00 %
Haiku 4.5B3-0.01 %-0.4-0.0 %-2.98 %2.97 %
EnsembleA5-0.05 %-1.1-0.1 %-2.73 %2.67 %0.29
Haiku 4.5A5-0.23 %-4.7-0.4 %-2.73 %2.49 %0.29

This is a paper book measured close to close: a real execution would enter at the open, and open prices are not available. None of this is a recommendation.

One-session totals

A = without results context, B = with it, X = control with the context shuffled across names. Bets are the names where the model moved away from 0.50.

ModelVar.nHits/betsBrierResolutionvs base rateCRPSBaseline80% coverage
Haiku 4.5A5515/460.26010.0724-0.01091.94 %1.59 %0.57
Haiku 4.5X100.50 everywhere0.25000.0000-0.01002.22 %1.56 %0.30
Haiku 4.5B302/60.25000.0000-0.00452.80 %1.59 %0.50
Opus 5A5511/190.24900.00390.00031.37 %1.59 %0.73
Opus 5X100.50 everywhere0.25000.0000-0.01001.54 %1.56 %0.60
Opus 5B302/30.24970.0011-0.00411.40 %1.59 %0.77
Sonnet 5A5518/360.24820.00020.00101.90 %1.59 %0.77
Sonnet 5B3010/180.24910.0067-0.00351.78 %1.59 %0.80
EnsembleA5518/480.25180.0243-0.00261.64 %1.59 %0.67
EnsembleX100.50 everywhere0.25000.0000-0.01002.41 %1.56 %0.30
EnsembleB3010/180.24960.0028-0.00401.64 %1.59 %0.73

Resolution > 0.02 = the probabilities discriminate. vs base rate > 0 = beats always saying the average rate. CRPS: lower is better; the baseline is a mechanical forecast from volatility, no AI. Target coverage: 0.80.

A/B of the quarterly-results context

The same package with and without the results block (date, whether the recent move is already the reaction, the name's historical move size). CRPS difference paired by session and name.

ModelSubsetPairsCRPS ACRPS BChangep
Haiku 4.5all301.94 %2.80 %+45 %n/a
Haiku 4.5with event42.66 %2.24 %−16 %n/a
Opus 5all301.37 %1.40 %+3 %n/a
Opus 5with event41.37 %1.09 %−21 %n/a
Sonnet 5all301.90 %1.78 %−6 %n/a
Sonnet 5with event41.33 %1.05 %−21 %n/a
Ensembleall301.64 %1.64 %+0 %n/a
Ensemblewith event41.46 %1.08 %−26 %n/a

Pre-registered criterion: improvement ≥ 10% with p < 0.05 and 200 pairs with an event. p is hidden with fewer than 5 sessions.

By session

Each cell: hits/bets · Brier · magnitude CRPS. ½ = the model said 0.50 for every name.

SessionClosed up|ret| medianHaiku 4.5 AHaiku 4.5 BSonnet 5 ASonnet 5 BOpus 5 AOpus 5 B
Sep 103/102.42 %2/8 · 0.266 · 1.25 %0/3 · 0.253 · 1.92 %4/8 · 0.248 · 1.86 %4/8 · 0.250 · 1.73 %1/3 · 0.251 · 1.44 %1/1 · 0.249 · 1.27 %
Sep 96/101.69 %6/9 · 0.244 · 2.34 %1/1 · 0.248 · 3.55 %3/4 · 0.248 · 1.96 %2/3 · 0.248 · 1.85 %2/3 · 0.249 · 1.35 %1/2 · 0.250 · 1.42 %
Sep 84/103.67 %1/8 · 0.273 · 2.21 %1/2 · 0.249 · 2.94 %1/4 · 0.252 · 1.88 %4/7 · 0.249 · 1.75 %½ · 0.250 · 1.32 %½ · 0.250 · 1.52 %
Sep 44/101.97 %1/9 · 0.285 · 1.78 %2/9 · 0.255 · 2.98 %2/5 · 0.251 · 4.04 %
Sep 39/152.92 %5/12 · 0.242 · 1.91 %8/11 · 0.241 · 2.01 %6/8 · 0.246 · 2.92 %

Names are picked by recent movement (top and bottom movers over 1, 5 and 21 sessions), not by conviction: it is the sampling frame, not a signal.

Latest resolved session by name (Sep 10)

Cell: probability of closing up / expected move.

NameActual returnHaiku 4.5 AHaiku 4.5 BSonnet 5 ASonnet 5 BOpus 5 AOpus 5 B
ADSK2.42 %0.46 / 3.3 %0.50 / 2.8 %0.51 / 2.2 %0.50 / 1.9 %0.50 / 2.2 %0.50 / 1.9 %
CASY-0.22 %0.45 / 4.5 %0.50 / 3.0 %0.52 / 2.7 %0.51 / 2.0 %0.50 / 3.2 %0.50 / 2.1 %
DDOG-1.58 %0.50 / 3.8 %0.50 / 8.5 %0.49 / 3.1 %0.49 / 2.9 %0.50 / 3.1 %0.50 / 2.8 %
DELL-5.37 %0.56 / 4.8 %0.50 / 4.9 %0.52 / 3.8 %0.51 / 3.3 %0.50 / 3.5 %0.50 / 3.2 %
MPC-1.76 %0.52 / 2.5 %0.51 / 4.1 %0.51 / 1.4 %0.51 / 1.3 %0.51 / 1.4 %0.50 / 1.3 %
MRNA0.75 %0.50 / 6.5 %0.50 / 6.5 %0.50 / 19.4 %0.49 / 18.1 %0.50 / 11.5 %0.50 / 10.0 %
SIG-4.59 %0.55 / 4.5 %0.51 / 10.7 %0.46 / 4.2 %0.49 / 3.9 %0.49 / 4.5 %0.49 / 3.9 %
SNDK-4.06 %0.52 / 5.0 %0.51 / 9.9 %0.51 / 4.6 %0.49 / 4.3 %0.50 / 4.6 %0.50 / 4.1 %
TPR1.90 %0.47 / 3.2 %0.50 / 2.5 %0.52 / 2.1 %0.50 / 1.7 %0.50 / 2.2 %0.50 / 1.7 %
VRT-5.62 %0.48 / 3.6 %0.50 / 8.0 %0.50 / 3.0 %0.49 / 2.8 %0.51 / 3.3 %0.50 / 2.8 %

Method: predictions are recorded before the open and never on sessions already closed; packages are stored exactly as each model saw them; models get no headlines (licensing and memorisation). The directional-skill criterion (resolution > 0 with p < 0.05 over 200 predictions) and the A/B criterion were written down before starting.