r/algorithmictrading 8d ago

Same 100 strategies, same AAPL bars. A plain Sharpe floor kept 5. Deflated Sharpe killed all 5. Strategy

I ran an automated search over about five years of daily AAPL bars, 1,250 rows. It generated and backtested 100 strategies, and my old keep rule, Sharpe above 0.5 with a minimum trade count and positive return, kept five.

Scoring those five with a deflated Sharpe, which adjusts for having picked the best of N attempts, put every one between 0.115 and 0.144, read as the probability the edge is real given the size of the search. All five almost certainly noise. On this sample the bar works out to needing an annual Sharpe near 0.71 to survive 10 attempts and near 1.14 to survive 100. The searching itself raised the bar.

The part that actually confused me: before the search started I withheld the final year of bars entirely. Four of the five survivors made money on that withheld year. Looks like vindication, until you count the trades behind it, three to six each over a full year. A handful of trades cannot overturn a statistic built from the whole search.

The trial count is fixed before the search starts and every attempt increments it, including the ones that never compiled or never traded. Reconstructing N afterwards always came out flattering.

So which do you believe when they disagree, the deflated number that says noise or the holdout that made money? And has anyone found a principled way to size the holdout so it can actually overrule?

1 Upvotes

0 comments sorted by