AI tournament, five weeks in: none beats plain volatility, and an error of ours skewed the A/B
Since 3 September, five language models have been forecasting the stock market in public before it opens. Haiku 4.5, Sonnet 5 and Opus 5, from Anthropic, plus DeepSeek V4.1 Flash and Qwen 3.8 Flash, which joined on 25 September. Every morning at 12:30 UTC, an hour before the New York open, they receive the same data package on ten S&P 500 stocks, with no web search and no headlines, and return the probability that each one goes up that day and how much it will move. After the close we score them with Brier and CRPS and publish the scoreboard whatever it says. We have about 24 scored sessions: between 235 and 245 predictions per model for the three Anthropic models, and 90 and 80 for the other two.
Direction first. No model is distinguishable from simply always saying the average rate of up days. That is no surprise: we wrote it down before starting. What is worth adding is what it does not mean. With about 240 predictions we can only detect a correlation between probability and outcome of 0.18 or more (80% power, p < 0.05). The small effects that get published for this kind of signal, 0.02 to 0.05, need between 2,500 and 20,000 predictions: between one and eight years at the current pace. That nobody discriminates today does not show it cannot be done; it shows we cannot yet see it. Also, Opus answers exactly 0.50 on four names out of five, which is a polite way of not betting.
Where there was more room was in how much each stock moves, so we compare every model with a three-line mechanical forecast: the stock's recent volatility times three quantiles measured before most of the sample was played. In CRPS, lower is better: Haiku 1.62% against 1.19% for the mechanical forecast; Sonnet 1.32 against 1.19; Opus 1.21 against 1.17; DeepSeek 1.34 against 1.16; Qwen 1.29 against 0.90. Opus ties. The other four fall behind. With so few sessions we would not stake much on the statistical significance, but the sign is the same in all five.
There is a simple explanation and we measured it. What each model adds on top of volatility is noise: the correlation between the move it announces (relative to the stock's volatility) and the move that happens runs from −0.03 to +0.09 depending on the model, with intervals that include zero in every case. And the package contains nothing the mechanical forecast does not have: prices, volatility, a count of news with a catalyst. A model that knows nothing the baseline does not know can only match it or spoil it.
So far, as expected. What was not expected sat in the second experiment. Alongside the main test runs a pre-registered A/B: does it help to give the model the quarterly-results context of each stock, that is, whether it reports today, whether it has already reacted, and how much it usually moves on results? Each model receives the same package with and without that block. For weeks the public scoreboard said the block made the prediction worse by between 17% and 43%, depending on the model.
A result saying that more information makes the forecast worse deserves suspicion before it deserves a headline. We recomputed the magnitude figures we publish from the raw rows (they matched) and then looked at where the announcement dates came from. When the earnings calendar had no announcement near a company, the system took the first news item classified as “earnings” as the announcement date. We checked those dates against the real record of earnings reports: of 35 dates taken from news, none matched (within three days) a real report. Of 28 calendar dates, all 28 did.
We looked at which news items had manufactured those dates. An article about the earnings of a delivery platform's couriers. A software product with “eps” in its name. A carmaker's quarterly sales. The transcript, published weeks later, of the previous quarter's call. A note about another company's results that was dragging the first one along. Our catalyst classifier works only with the headline and looks for loose words —“earnings”, “eps”, “Q3”— that appear in far more news than earnings reports. With that false date, the package told the model “today is the reaction day, expect an earnings-sized move”. The model complied, inflated its forecast, and the stock did not move.
Split by where the date came from, the numbers tell a different story. With a calendar date (74 pairs) the block improves CRPS by 2.2%, not significant either way. With a news date (144 pairs) it worsens it by 121%. 74% of the rows in the subset we had pre-registered as “with event” were of the second kind. The headline we were publishing measured an error of ours, not the usefulness of the results context.
Our own pre-registration foresaw it: if the rate of mislabelled states went over 10%, the A/B would be frozen until the date source was fixed. It was 74%. So we did what it says: the code no longer lets a news item set the date of an announcement, the “with event” half of the sample now prioritises calendar announcements, and the count of pairs with an event goes back to zero from the 9 October session. The text of the re-registration was committed to the repository before the first affected run. What was measured before is still published on the tournament page, folded away and labelled as an earlier epoch that does not count towards the criterion. With real announcements, outside earnings season, one to four rows per session qualify; reaching the 200 pairs per model that the criterion requires will take until roughly February 2027.
What we do not conclude. Not that language models are useless for forecasting markets: the package they get gives them no information the mechanical forecast lacks, so there is not much to expect. The test that would make the most sense is to give them new, public information, such as the text of the earnings reports companies file with the SEC, and measure them against the same baseline. If we do that, it will be pre-registered before it starts. And the same classifier feeds the event map and the weekly bulletin, so we will review how much of their measured effect comes from calendar announcements and how much from noise.
None of this is investment advice or a signal: it is a calibration experiment on paper. The numbers, the predictions by session and the full A/B, with its epochs, are at soberquant.com/ai-tournament and in the tournament's public API.
Calendrier hebdomadaire des événements
Cette page est le bulletin : elle est régénérée chaque dimanche avec la semaine à venir. Nous ne l'envoyons pas par e-mail et ne demandons pas votre adresse. Si un jour nous l'envoyons, vous pourrez le demander ici.