← Back to blog

We spent a year measuring with an instrument six times too blunt

Published on 2026-08-14

For a year we tested whether a signal worked the way everyone does it: build a basket of the 25 best-scoring names, see how it does against the market, decide. We did that about twenty-five times, across different signals, and almost always got the same answer: nothing. We published those as verdicts. Two weeks ago we found out the problem might not be the signals, but the ruler we were judging them with.

A 25-name basket over a 500-name universe throws away 95% of the information. It only looks at the very top and ignores everything else. The alternative is to measure the rank correlation between the signal and subsequent returns across the WHOLE universe every month — the plain old rank IC — which uses five hundred names instead of twenty-five. We compared both instruments on exactly the same data and the gap is uncomfortable: the smallest effect the basket can tell apart from noise is **44.7 basis points per month**; with the full cross-section it drops to **about 7 equivalent**. Six times more sensitive, with no new data.

That turns every previous null into an open question: was there no effect, or did we fail to see it? It is the genuinely uncomfortable question, because the comfortable answer — 'there was nothing there' — is precisely the one that lets us sleep.

So we measured everything again. Ten signals, original definitions untouched, expected sign fixed from the literature BEFORE looking at any result, over roughly 495 names a month and up to ten years of history. Four fundamental anomalies from our July campaign (net share issuance, asset growth, accruals, fundamental momentum) and six classic families with decades of literature behind them: value two ways, quality two ways, 12-month momentum and low volatility.

**None of the ten survives.** And the numbers are something. Twelve-month momentum, the most cited factor in the history of quantitative finance, prints an IC of **+0.0018**. Low volatility prints **+0.0001**. These are not 'positive but underpowered': they are nailed to zero, in the S&P 500, over a decade. Three of the four fundamental anomalies come out with the **opposite sign** to what the literature predicts.

The one that taught us most is accruals. In July, measured with baskets, it was 'the only one with a pulse': a p-value of 0.049, brushing significance, enough to put it on a watchlist and revisit with more data. With the full cross-section it lands at **p=0.66, sign flipped**. That pulse was not a weak effect: it was an artefact of the tail of the top 25 names. The basket was not merely less sensitive — it was showing us things that were not there.

While we were at it we tested the last candidate we had left: the hypothesis that companies which CHANGE the language of their annual filings underperform afterwards (Cohen, Malloy and Nguyen, Journal of Finance 2020). We downloaded 22,200 SEC documents and computed 19,647 year-over-year similarities. Result: **null**, with a caveat that matters more than the null. Our estimator's confidence interval is [−20.7, +41.4] basis points per month: that range contains zero and it also contains the paper's full effect. Our test could not tell them apart. The honest statement is not 'it does not work', it is **'we are not able to know'**.

And out of that comes the number that has made us think the most. Given how variable the monthly excess of a twenty-five-name basket is, detecting a real 20-basis-point-per-month effect with statistical confidence would take **about fifty years of data**. Not fifty years of one specific anomaly: fifty years of any anomaly that size. Most of what the industry sells as 'factors' is, at this scale, indistinguishable from noise across an entire professional lifetime. That is not pessimism, it is arithmetic.

Along the way we caught one more mirage, the seventh of the project. Filing similarity appeared to predict future **volatility** — not returns, volatility — with a p-value below 0.0001. It would have been useful: volatility is genuinely predictable, and it drives position sizing. We added a control for the previous sixty days of realised volatility and the effect collapsed to **p=0.46**. Companies whose filings do not change are simply the ones that were already calm. Without that control we would now own a shiny new indicator that is really a moving average with extra steps.

We also tested whether it pays to exit each position when it is ready rather than on a fixed schedule. We measured the ceiling by cheating — picking the best exit with perfect foresight — and it came out seventeen times better. Looked enormous. We ran the identical trick on a **randomly chosen** basket and it did even better: +863 basis points against +798 for the one with signal. The entire 'ceiling' was the maximum of five correlated noises. Had we trained a model on it, it would have produced a beautiful backtest for exactly that reason.

So what is left? The verdict does not change: we still find no simple alpha that survives retail costs in large caps. But the **kind** of verdict changes. It used to be 'we looked hard and found nothing', which sounds like fatigue. Now it is 'we measured it with the right instrument, six times more sensitive, and it is still not there' — a different and considerably stronger claim.

And one rule we now apply to ourselves: **the basket is for trading, the IC is for deciding**. Building a concentrated portfolio makes sense once you know a signal is worth something; using it as a measuring instrument cost us a factor of six in sensitivity on every test we ran for a year. It is the kind of mistake that does not hurt while you make it, because it produces nulls — and a null looks an awful lot like a job well done.

We are publishing this because we committed to publishing the whole book, and a methodological mistake of our own is part of the book. Every number above comes from reproducible scripts; the detail, with confidence intervals and deflation for the number of tests, is in the project documentation. Our strategies remain in simulation (paper): none of this trades real money, and none of it is investment advice.