← Voltar ao blog

48 days of live forward testing: the full scoreboard, including the strategy that is down 10.6%

Publicado a 2026-08-05

On June 18th we set eleven strategies marking net asset value forward, in simulation (paper), each starting with $100,000 and with no ability to touch the past. We promised to publish the whole scoreboard after a few weeks, win or lose. It has been 48 calendar days — 32 marked sessions, June 18th to August 4th — and this is all of it. One clarification about the benchmark, because it decides how every number below reads: the cap-weighted S&P 500 (SPY) returned 3.29% over the window, with a 3.4% maximum drawdown. But our equity baskets split the money equally across some 25 names, and the fair comparison for an equal-weight basket is not the cap-weighted index but the equal-weighted one: the equal-weight S&P 500 (RSP) returned 5.30%, and the full equal-weight universe 6.94%. We give all three and compare against the right one, not the flattering one.

The good. Our anchor strategy — the S&P 500 Core, the only one with prior survivorship-free validation and academic backing — is up 7.34%. Against the cap-weighted index that is a 4.05-point edge; against the equal-weighted index, its honest comparison, 2.04 points; and against the equal-weight universe, four tenths of a point. We publish all three because keeping only the first is exactly the trick we call out in everyone else. The combined book of weakly correlated premia (inverse-vol weights) is up 3.69% with a 0.8% maximum drawdown: below the index on return, but with a fraction of its shaking. And both Tactical Defensive books did exactly their job: +2.44% and +2.12%, behind the market in a rising leg, with maximum drawdowns of 1.3% and 1.1% against the index's 3.4%. When the market rises, the defensive lags; that is the premium it pays for its insurance, not a bug.

The bad, and the most interesting part. The Core + Trend variant — the same engine with a momentum filter on top, by far the BEST looking one in backtest — is down 10.64%, with a 26.1% maximum drawdown. Thirteen and a half points below the market and eighteen below its plain sibling. It rotated into semiconductors right before the sector selloff of July 2nd. That is this cut's honest headline: the prettier backtest lost to the simpler one in live conditions, and lost badly. This is precisely what a forward A/B is built for.

What we already knew was broken and still is. Our intraday news rig has accumulated 4,906 trades at an average of -4.53 basis points net per trade, a 24.2% hit rate and a p-value of 1.000: no signal. Restricted to real broker fills it gets worse, -12.07 basis points. And among the candidates under observation, the one we labelled in writing back in June as a probable recency mirage is down 6.91%. When the uncomfortable prediction comes true, that gets published too.

Is this what we expected? In character, yes, almost point by point: the anchor leads, the defensives protect and lag in an up leg, the aggressive one swings, the mirage deflates, timing signals stay null. In magnitude, two warnings. The first is counterintuitive: the anchor is doing TOO well. Its expected edge was around 7 points a year on average, and it has booked 2 over its fair benchmark in seven weeks; that is not skill showing up, it is noise in our favour, and anyone annualising it is fooling themselves. The second cuts the other way and is just as uncomfortable: its maximum drawdown has been only 2.7%, against roughly 9% that the backtest showed for a full year. Seven weeks without a scare do not prove the strategy is safe; they prove the sample is too short to have seen the risk yet. A comfortable ride is, if anything, one more reason to distrust the number.

A transparency note on the numbers themselves. On July 4th we found a bug of our own: in the July 1st monthly rebalance, two bots valued at zero the positions still held but dropped from the new candidate list, destroying net asset value and mis-sizing the new book. We fixed it, repaired the databases from backup using that day's real closes, and republished. The post-fix result was less flattering to us than the bug itself: part of what looked like a Core collapse was the bug, but the Core + Trend loss was and remains real. And today, while preparing this cut, a second fault surfaced — smaller, but the same species: four of the bots stamped the NAV date using the server's local time instead of UTC, so their whole series ran one day ahead and even placed marks on Saturdays. The amounts were right; the labels were not. We rebuilt the series from the execution logs and they now line up. We are telling you because anyone who only publishes when the numbers come out clean is not publishing a track record, they are running an ad.

And now what this scoreboard does NOT mean. None of it is statistically significant, and it will not be for months. Our own engine demands date-clustered bootstrap, deflated Sharpe and probability of backtest overfitting before issuing a verdict; 32 sessions do not produce a verdict, they produce an anecdote. The Sharpe ratios these series currently spit out — above 4 on the anchor, above 3 on the defensives — are decorative: annualising seven good weeks is the oldest trick in the industry. We publish this because we committed to publishing the whole book, not because the book says anything yet. We will be back with the next cut, same table, whichever way it goes.

Everything above is simulated (paper): none of these strategies trades real money, none of it is investment advice, and past results — let alone seven weeks of them — do not anticipate future ones.