AI tournament · phase 1
Three AIs forecast the market every morning. We score them on calibration.
Haiku 4.5, Sonnet 5 and Opus 5 receive the same data package on 10 S&P 500 names every day at 12:30 UTC and return, before the open, the probability that each one closes up and how much it will move. At the close we score them with Brier and CRPS and publish the result whatever it says.
What is measured
Calibration, not hit rate: Brier and its decomposition (resolution) for direction; a three-quantile CRPS and the coverage of the 80% band for the size of the move.
What we expect
The pre-registered expectation is that no AI discriminates one-day direction: in this universe it is measured as null. Where value may exist is in how much a name moves, not which way.
What it is not
Not a recommendation or a signal. No model sees headlines or searches the web; only prices, volatility, catalyst news counts, the main macro calendar (Fed, CPI, jobs) and, in variant B, quarterly-results context.
Reading
Sentences generated by fixed rules from the numbers, using the pre-registered thresholds. Nobody writes them by hand.
- Phase 1: 55 resolved predictions per model out of 200. Nothing below is a result.
- Direction: no model beats the base rate; always saying the average rate remains the best forecast, as expected.
- Magnitude Haiku 4.5 A: behind the mechanical baseline (1.94 % vs 1.59 %); plain volatility predicts size better.
- Magnitude Haiku 4.5 B: behind the mechanical baseline (2.80 % vs 1.59 %); plain volatility predicts size better.
- Magnitude Opus 5 A: ahead of the mechanical baseline (1.37 % vs 1.59 %), not yet significant.
- Magnitude Opus 5 B: ahead of the mechanical baseline (1.40 % vs 1.59 %), not yet significant.
- Magnitude Sonnet 5 A: behind the mechanical baseline (1.90 % vs 1.59 %); plain volatility predicts size better.
- Magnitude Sonnet 5 B: behind the mechanical baseline (1.78 % vs 1.59 %); plain volatility predicts size better.
- Magnitude Ensemble A: on par with the mechanical baseline (1.64 % vs 1.59 %).
- Magnitude Ensemble B: on par with the mechanical baseline (1.64 % vs 1.59 %).
- Paper P&L: provisional leader Sonnet 5 A at 0.04 % over 5 sessions vs -2.73 % for the basket of the same names; 4 of 4 models above the basket. Below 250 sessions the leader is noise and changes week to week.
- A/B Haiku 4.5: 30 pairs (4 with an event, out of 200); with results context the CRPS worsens by 45%.
- A/B Opus 5: 30 pairs (4 with an event, out of 200); with results context the CRPS worsens by 3%.
- A/B Sonnet 5: 30 pairs (4 with an event, out of 200); with results context the CRPS improves by 6%.
- A/B Ensemble: 30 pairs (4 with an event, out of 200); with results context the CRPS worsens by 0%.
Updated 2026-09-11 11:26 UTC
Paper P&L
Fixed, pre-registered rule: each name is held long or short in proportion to how far the probability sits from 0.50, equal capital per name, one-session holding, close to close, 10 bp round-trip cost. The basket is the same list of names bought equal-weight. 🏆 = best net return among models with the same sessions; below 250 sessions the leader is noise.
| Model | Var. | Sessions | Net return | bp/day | Sharpe | Max drawdown | Basket | vs basket | p | |
|---|---|---|---|---|---|---|---|---|---|---|
| 🏆 | Sonnet 5 | A | 5 | 0.04 % | 0.8 | – | -0.0 % | -2.73 % | 2.77 % | 0.29 |
| Sonnet 5 | B | 3 | 0.02 % | 0.5 | – | -0.0 % | -2.98 % | 2.99 % | – | |
| Opus 5 | A | 5 | 0.02 % | 0.3 | – | -0.0 % | -2.73 % | 2.74 % | 0.29 | |
| Opus 5 | B | 3 | 0.01 % | 0.4 | – | 0.0 % | -2.98 % | 2.99 % | – | |
| Ensemble | B | 3 | 0.01 % | 0.2 | – | -0.0 % | -2.98 % | 2.98 % | – | |
| Haiku 4.5 | X | 1 | 0.00 % | 0.0 | – | 0.0 % | 0.00 % | -0.00 % | – | |
| Opus 5 | X | 1 | 0.00 % | 0.0 | – | 0.0 % | 0.00 % | -0.00 % | – | |
| Ensemble | X | 1 | 0.00 % | 0.0 | – | 0.0 % | 0.00 % | -0.00 % | – | |
| Haiku 4.5 | B | 3 | -0.01 % | -0.4 | – | -0.0 % | -2.98 % | 2.97 % | – | |
| Ensemble | A | 5 | -0.05 % | -1.1 | – | -0.1 % | -2.73 % | 2.67 % | 0.29 | |
| Haiku 4.5 | A | 5 | -0.23 % | -4.7 | – | -0.4 % | -2.73 % | 2.49 % | 0.29 |
This is a paper book measured close to close: a real execution would enter at the open, and open prices are not available. None of this is a recommendation.
One-session totals
A = without results context, B = with it, X = control with the context shuffled across names. Bets are the names where the model moved away from 0.50.
| Model | Var. | n | Hits/bets | Brier | Resolution | vs base rate | CRPS | Baseline | 80% coverage |
|---|---|---|---|---|---|---|---|---|---|
| Haiku 4.5 | A | 55 | 15/46 | 0.2601 | 0.0724 | -0.0109 | 1.94 % | 1.59 % | 0.57 |
| Haiku 4.5 | X | 10 | 0.50 everywhere | 0.2500 | 0.0000 | -0.0100 | 2.22 % | 1.56 % | 0.30 |
| Haiku 4.5 | B | 30 | 2/6 | 0.2500 | 0.0000 | -0.0045 | 2.80 % | 1.59 % | 0.50 |
| Opus 5 | A | 55 | 11/19 | 0.2490 | 0.0039 | 0.0003 | 1.37 % | 1.59 % | 0.73 |
| Opus 5 | X | 10 | 0.50 everywhere | 0.2500 | 0.0000 | -0.0100 | 1.54 % | 1.56 % | 0.60 |
| Opus 5 | B | 30 | 2/3 | 0.2497 | 0.0011 | -0.0041 | 1.40 % | 1.59 % | 0.77 |
| Sonnet 5 | A | 55 | 18/36 | 0.2482 | 0.0002 | 0.0010 | 1.90 % | 1.59 % | 0.77 |
| Sonnet 5 | B | 30 | 10/18 | 0.2491 | 0.0067 | -0.0035 | 1.78 % | 1.59 % | 0.80 |
| Ensemble | A | 55 | 18/48 | 0.2518 | 0.0243 | -0.0026 | 1.64 % | 1.59 % | 0.67 |
| Ensemble | X | 10 | 0.50 everywhere | 0.2500 | 0.0000 | -0.0100 | 2.41 % | 1.56 % | 0.30 |
| Ensemble | B | 30 | 10/18 | 0.2496 | 0.0028 | -0.0040 | 1.64 % | 1.59 % | 0.73 |
Resolution > 0.02 = the probabilities discriminate. vs base rate > 0 = beats always saying the average rate. CRPS: lower is better; the baseline is a mechanical forecast from volatility, no AI. Target coverage: 0.80.
A/B of the quarterly-results context
The same package with and without the results block (date, whether the recent move is already the reaction, the name's historical move size). CRPS difference paired by session and name.
| Model | Subset | Pairs | CRPS A | CRPS B | Change | p |
|---|---|---|---|---|---|---|
| Haiku 4.5 | all | 30 | 1.94 % | 2.80 % | +45 % | n/a |
| Haiku 4.5 | with event | 4 | 2.66 % | 2.24 % | −16 % | n/a |
| Opus 5 | all | 30 | 1.37 % | 1.40 % | +3 % | n/a |
| Opus 5 | with event | 4 | 1.37 % | 1.09 % | −21 % | n/a |
| Sonnet 5 | all | 30 | 1.90 % | 1.78 % | −6 % | n/a |
| Sonnet 5 | with event | 4 | 1.33 % | 1.05 % | −21 % | n/a |
| Ensemble | all | 30 | 1.64 % | 1.64 % | +0 % | n/a |
| Ensemble | with event | 4 | 1.46 % | 1.08 % | −26 % | n/a |
Pre-registered criterion: improvement ≥ 10% with p < 0.05 and 200 pairs with an event. p is hidden with fewer than 5 sessions.
By session
Each cell: hits/bets · Brier · magnitude CRPS. ½ = the model said 0.50 for every name.
| Session | Closed up | |ret| median | Haiku 4.5 A | Haiku 4.5 B | Sonnet 5 A | Sonnet 5 B | Opus 5 A | Opus 5 B |
|---|---|---|---|---|---|---|---|---|
| Sep 10 | 3/10 | 2.42 % | 2/8 · 0.266 · 1.25 % | 0/3 · 0.253 · 1.92 % | 4/8 · 0.248 · 1.86 % | 4/8 · 0.250 · 1.73 % | 1/3 · 0.251 · 1.44 % | 1/1 · 0.249 · 1.27 % |
| Sep 9 | 6/10 | 1.69 % | 6/9 · 0.244 · 2.34 % | 1/1 · 0.248 · 3.55 % | 3/4 · 0.248 · 1.96 % | 2/3 · 0.248 · 1.85 % | 2/3 · 0.249 · 1.35 % | 1/2 · 0.250 · 1.42 % |
| Sep 8 | 4/10 | 3.67 % | 1/8 · 0.273 · 2.21 % | 1/2 · 0.249 · 2.94 % | 1/4 · 0.252 · 1.88 % | 4/7 · 0.249 · 1.75 % | ½ · 0.250 · 1.32 % | ½ · 0.250 · 1.52 % |
| Sep 4 | 4/10 | 1.97 % | 1/9 · 0.285 · 1.78 % | – | 2/9 · 0.255 · 2.98 % | – | 2/5 · 0.251 · 4.04 % | – |
| Sep 3 | 9/15 | 2.92 % | 5/12 · 0.242 · 1.91 % | – | 8/11 · 0.241 · 2.01 % | – | 6/8 · 0.246 · 2.92 % | – |
Names are picked by recent movement (top and bottom movers over 1, 5 and 21 sessions), not by conviction: it is the sampling frame, not a signal.
Latest resolved session by name (Sep 10)
Cell: probability of closing up / expected move.
| Name | Actual return | Haiku 4.5 A | Haiku 4.5 B | Sonnet 5 A | Sonnet 5 B | Opus 5 A | Opus 5 B |
|---|---|---|---|---|---|---|---|
| ADSK | 2.42 % | 0.46 / 3.3 % | 0.50 / 2.8 % | 0.51 / 2.2 % | 0.50 / 1.9 % | 0.50 / 2.2 % | 0.50 / 1.9 % |
| CASY | -0.22 % | 0.45 / 4.5 % | 0.50 / 3.0 % | 0.52 / 2.7 % | 0.51 / 2.0 % | 0.50 / 3.2 % | 0.50 / 2.1 % |
| DDOG | -1.58 % | 0.50 / 3.8 % | 0.50 / 8.5 % | 0.49 / 3.1 % | 0.49 / 2.9 % | 0.50 / 3.1 % | 0.50 / 2.8 % |
| DELL | -5.37 % | 0.56 / 4.8 % | 0.50 / 4.9 % | 0.52 / 3.8 % | 0.51 / 3.3 % | 0.50 / 3.5 % | 0.50 / 3.2 % |
| MPC | -1.76 % | 0.52 / 2.5 % | 0.51 / 4.1 % | 0.51 / 1.4 % | 0.51 / 1.3 % | 0.51 / 1.4 % | 0.50 / 1.3 % |
| MRNA | 0.75 % | 0.50 / 6.5 % | 0.50 / 6.5 % | 0.50 / 19.4 % | 0.49 / 18.1 % | 0.50 / 11.5 % | 0.50 / 10.0 % |
| SIG | -4.59 % | 0.55 / 4.5 % | 0.51 / 10.7 % | 0.46 / 4.2 % | 0.49 / 3.9 % | 0.49 / 4.5 % | 0.49 / 3.9 % |
| SNDK | -4.06 % | 0.52 / 5.0 % | 0.51 / 9.9 % | 0.51 / 4.6 % | 0.49 / 4.3 % | 0.50 / 4.6 % | 0.50 / 4.1 % |
| TPR | 1.90 % | 0.47 / 3.2 % | 0.50 / 2.5 % | 0.52 / 2.1 % | 0.50 / 1.7 % | 0.50 / 2.2 % | 0.50 / 1.7 % |
| VRT | -5.62 % | 0.48 / 3.6 % | 0.50 / 8.0 % | 0.50 / 3.0 % | 0.49 / 2.8 % | 0.51 / 3.3 % | 0.50 / 2.8 % |
Method: predictions are recorded before the open and never on sessions already closed; packages are stored exactly as each model saw them; models get no headlines (licensing and memorisation). The directional-skill criterion (resolution > 0 with p < 0.05 over 200 predictions) and the A/B criterion were written down before starting.