The first backtest returned +220% on AAPL. I didn't believe it, because it came from the same window I'd spent the afternoon tuning against. So I wrote down a protocol, ran it once, and published what came back — including the part that made my own platform look bad.
I'd built a backtesting engine inside AlphaOS, an agentic multi-market trading platform. Real OHLCV, costs charged on both sides at 5bps commission plus 5bps slippage, daily mark-to-market equity, and a buy-and-hold benchmark computed over the identical window.
Before trusting a single number from it, I reimplemented one strategy independently in Python and compared. Total return 27.70% versus 27.70%, end equity $12,769.95 versus $12,770, 3 trades, 66.7% win rate, −11.1% max drawdown. Both implementations agreed to the cent. The engine was sound.
Then I started improving the strategies, and this is where it gets interesting — because the engine being correct is not the same as the results being meaningful.
The baseline golden-cross rule on AAPL returned +27.70% over 7.2 years. Buy-and-hold returned +531.45%. That's a bad strategy, clearly.
So I tried variants. An ATR trailing stop instead of a fixed one. A slower exit on the 200-day rather than the 50-day. Scaled position sizing — 100% / 50% / 0% instead of all-or-nothing. That last one took AAPL from +27.70% to +220.18%, with a lower drawdown than buy-and-hold and a Sharpe of 0.88 against the baseline's 0.53.
An eight-fold improvement. It felt like progress.
I had tested four variants on the same data and kept the two that scored best. That is not research. That is selection on the evaluation set — and every choice I made while looking at a result fitted the strategy a little more tightly to that specific window.
The engine was measuring correctly. It was measuring the wrong thing: how well a rule I had chosen because it fit this data, fits this data.
Split history into two windows. Do all selection and fitting on the first. Then run the chosen configuration exactly once on the second, and accept whatever comes out.
The discipline is in the word once. If you look at the out-of-sample result and then adjust anything, the second window has become a tuning window too, and you've quietly spent your only honest test.
So I wrote the protocol down before running it:
Rank five portfolio variants on 2006–2015 by Sharpe ratio. Take the highest. Run it once on 2016–2026. Report that number regardless of what it says.
Ten US equities, equal-weighted, each sleeve gated by its own trend signal, costs applied to every rebalance. The tuning window contains the global financial crisis; the test window contains a long bull run plus the 2020 and 2022 drawdowns. Both contain more than one regime, which matters — a tuning window with only one regime selects a strategy fitted to that regime.
| Variant | Sharpe | Buy & hold | Edge |
|---|---|---|---|
| base — binary in/out | 0.57 | 0.60 | −0.03 |
| ATR trailing stop | 0.52 | 0.60 | −0.08 |
| slow exit — SMA200 | 0.87 | 0.60 | +0.27 |
| scaled 100/50/0 | 0.56 | 0.60 | −0.04 |
| scaled 100/75/50/0 | 0.55 | 0.60 | −0.05 |
10 US equities · tuning window 2007-07 → 2015-12 · costs 5bps/side
Clear result. The slow-exit variant beat buy-and-hold by +0.27 Sharpe and was the only one above the benchmark at all. If I'd stopped here — as most published backtests do — I'd have written a confident post about how exiting on the 200-day rather than the 50-day avoids whipsaws and adds real risk-adjusted return.
| Return | Max drawdown | Sharpe | |
|---|---|---|---|
| slow exit — the chosen variant | 121.2% | 11.7% | 1.07 |
| Buy & hold | 3,587.6% | 50.0% | 1.17 |
Same 10 equities · test window 2015-10 → 2026-08 · never used for selection
In-sample edge +0.27. Out-of-sample edge −0.10. The advantage didn't shrink — it inverted. The strategy that most clearly beat the benchmark on the tuning window lost to it on data it hadn't seen.
One failed walk-forward could be an unlucky split. So I did the whole thing again on a different asset class, different decade, different volatility regime, no shared data: seven crypto pairs, tuned on 2018–2022 — a full cycle including two crashes — and tested on 2023–2026.
In-sample winner: the scaled variant, Sharpe 1.19 against buy-and-hold's 1.10, edge +0.09.
Out-of-sample: Sharpe 0.51 against 0.71. Edge −0.21. Worse than equities. Zero of five crypto variants held a positive out-of-sample edge — not one.
| Universe | Tuned | Tested | In-sample edge | Out-of-sample | Drawdown saved |
|---|---|---|---|---|---|
| 10 US equities | 2007–2015 | 2015–2026 | +0.27 | −0.10 | 38.3pp |
| 7 crypto pairs | 2018–2022 | 2022–2026 | +0.09 | −0.21 | 16.7pp |
Two independent asset classes, same protocol, same outcome
Look at what in-sample selection rejected on the equities run. The two scaled variants scored −0.04 and −0.05 in-sample — below the benchmark, so I discarded them.
Out of sample they scored +0.04 and +0.06: the only two variants of the five with a genuine positive edge.
It didn't merely fail to identify the best strategy. It actively selected against it — picking the one variant whose apparent edge was an artifact, and rejecting both that generalised.
Across all ten variant-runs, the correlation between in-sample edge and out-of-sample edge was +0.26. Close enough to noise that in-sample ranking told me almost nothing about which strategy would work next.
One thing generalised, without a single exception, in both asset classes and every variant: drawdown reduction.
On equities, out-of-sample drawdowns ran 4.6% to 23.7% against buy-and-hold's 50.0%. On crypto, 16.0% to 36.4% against 53.1%. The chosen variant cut equity drawdown from 50.0% to 11.7% — a 38.3 percentage point reduction that held on data it had never seen.
These rules are a risk-management tool, not an alpha source. They reliably cut drawdown roughly in half. They do not reliably beat buy-and-hold on risk-adjusted return, and every backtest of mine that suggested otherwise was measuring its own tuning window.
That's a narrower claim than the one I started with. It's also the only one that survived a test I couldn't tune.
The engine runs client-side. You can run these backtests yourself in the browser — pick a strategy, a symbol and a period, or sweep all 15 symbols at once. The walk-forward result is printed at the top of that page, above every single-window number it produces.
Source is on GitHub.
Because it's the same failure mode. A demand forecast tuned until it fits last year's history looks excellent and tells you nothing about next quarter. The difference is that a bad forecast takes a quarter to embarrass you, while a bad backtest lies immediately and convincingly.
Markets are the harsher teacher. That's why I built here.