Kieran Duff
Home Letters Handbook Order splitting Strategy funnel About Subscribe
Frameworks · Note 024 · 24 Jun 2026

The Lucky Sharpe

If you pick the best Sharpe from a big search, most of what makes that number look good is selection bias. You can even put a rough figure on how much.

The short version
The Lucky Sharpe cover image

Here is something almost nobody who builds strategies wants to hear. If you generate 200 variants of a strategy and purely keep the one with the best Sharpe, a large chunk of what makes that Sharpe look good is luck. You can even put a rough figure on how much.

I went through this properly when I started mining strategies, and it changed how I read every backtest I produce now.

Where the luck comes from

Every variant you test is a draw from a distribution of possible results. Give a strategy zero real edge and run it on real price data, and it will still post some return, because markets have structure and your rules will catch some of that. Run a hundred of those zero-edge variants and a few will look genuinely good.

Often, they're not good because they found anything, but more so because someone always wins the raffle.

The more variants you test, the better your luckiest result looks. Keep only the winner and you have recorded a number that was selected for being lucky. You did it the moment you sorted the results column and took the top row. Statisticians have a name for it: multiple testing.

How much does it actually inflate?

A useful rule of thumb: the expected best of N independent backtests grows with the square root of the log of N. Test 10 variants and your winner looks meaningfully better than its true worth. Test 1,000, which is nothing for a strategy miner, and the gap is wide enough that a completely worthless strategy can post a backtest Sharpe comfortably above 1.0.

Sit with that. A strategy with no edge whatsoever can hand you a backtest you would happily fund, purely because you tested enough rubbish alongside it. That is the trap that eats builders who find genetic optimisers and strategy factories before they understand what those tools do to the statistics.

What I actually do about it

I count the trials honestly. Every variant, every instrument I tried and binned, every "let me test one more idea" across the whole project. When I mine with StrategyQuantX, a single run can throw out thousands of candidates, so the few that make it into my test book have survived brutal testing.

I keep data the optimisation never touches. An out-of-sample block, and an older block from before the in-sample window, the "old OOS" that catches anything curve-fit to recent conditions. Most of the lucky ones fall over right here, which is the actual point of this test.

I expect the live Sharpe to come in under the backtest, and I size the book as if it will. When live underperforms, that is just the selection bias unwinding in real time.

I stop trusting the single best row and judge the whole optimisation surface instead. The row you are most tempted to keep is the one most likely to be lucky, so the real question is whether strong in-sample settings stay strong out of sample across the entire grid, not only at your chosen peak. This is exactly what Opt My Strategy by Martyn Tinsley is built for. It takes the full optimisation and scores the whole parameter space on whether in-sample performance actually maps to out-of-sample performance, with his Walk Forward Correlation method doing the heavy lifting. Across the grid, a high correlation means the edge is structural and turns up in places you never fitted it; a low one means your winner was a lucky spike with nothing underneath. It steers you toward the stable region of the surface, where the edge holds without being fitted to it. That is the whole game.

Martyn and I got deep into this on a livestream I did with him for Darwinex Zero, worth a watch if you want to see it run on a real optimisation rather than just described.

The mindset that matters

Stop treating the best result of a big search as a discovery. Treat it as a suspect. It walked out of a process built to produce flattering outliers, so the burden sits with the strategy to prove it is more than the raffle winner.

The harder you searched, the more sceptical you should be of what you found. The winner of a big search is guilty until proven innocent. Treat it that way and most of your overfitting problem solves itself before a strategy ever touches live capital.

Sources

David H. Bailey & Marcos López de Prado, The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality (Journal of Portfolio Management, 2014).

Martyn Tinsley, Walk Forward Correlation: A Diagnostic for Over-Fitting and Structural Edge in Trading Strategy Optimisation.

Kieran Duff runs XAQP, a systematic strategy live since April 2025 with around $3.7M in capital through Darwinex as of June 2026. He writes about how a systematic book is actually managed.

Disclosure. Personal commentary, not financial advice. Capital at risk. I am an employee of Darwinex; content touching Darwinex products may represent a conflict of interest, disclosed per MAR Article 20.

XAQP figures are point-in-time as of June 2026 and will change.

The Letter

Get the next letter in your inbox.

Please consider subscribing to receive more episodes of The Letter.

Subscribe now