agenttrading

How Long Should You Backtest a Trading Strategy? Years and Trades

July 19, 2026 · Agenttrading · Last updated July 2026

  1. 1
  2. 2
  3. 3
  4. 4
  5. 5
01 THESIS · AS A TESTABLE RULE

02 EVIDENCE · FUNDAMENTALS

03 BACKTEST · GROWTH OF $10,000
Strategy Buy & hold

04 RISK · IN PLAIN ENGLISH

05 VERDICT · HISTORICAL, NOT PREDICTIVE

Past performance does not guarantee future results. Educational analysis only, not financial advice.

Backtest a trading strategy over at least 20 years of data and at least 100 to 200 trades, whichever demands more. The calendar rule makes sure the strategy meets several different market regimes, including 2008 and 2020; the trade-count rule makes sure the result is statistically meaningful rather than a handful of lucky outcomes. A strategy that clears one test but fails the other has not been tested properly.

Two different questions, two different rules

"How long" hides two questions that people constantly conflate. One is about calendar coverage: has this rule seen enough different market conditions to be believed? The other is about sample size: are there enough trades for the average to mean anything? A slow trend-following rule might produce 25 trades in 20 years, which passes the calendar test and fails the sample test. A day-trading rule might generate 4,000 trades in 18 months, which passes the sample test and has never seen a bear market.

TestMinimum barWhat it protects against
Calendar coverage20+ years of adjusted daily dataA strategy that only ever worked in one regime, usually a long bull market
Trade count100 to 200 closed trades minimumAn average built from noise, where two or three trades produce the entire result
Regime spreadAt least two full bear marketsRules that quietly assume dips always recover quickly
Out-of-sample window20% to 30% of the data held backOverfitting: parameters tuned until the past looks perfect

Past performance does not guarantee future results. For educational and informational purposes only. Not financial advice. Consult a licensed advisor.

Why 20 years is the calendar floor

Twenty years of daily data reaching back from today covers the dot-com unwind's aftermath, the 2008 financial crisis, the 2020 crash and recovery, the 2022 rate shock, and long stretches of grinding bull market in between. That range matters because most strategies are regime-dependent without their author realizing it. A buy-the-dip rule tested only on 2010 to 2021 looks superb, because in that window every dip recovered. Run the same rule through 2000 to 2002 or through 2008 and it can spend years underwater.

Ten years is the most common shortcut and the most misleading one, since a decade starting after 2010 contains no serious extended bear market at all. If the data simply does not exist, an ETF that launched in 2015 for instance, the honest response is to say the window is short and treat the result as provisional, not to quietly present it as validated.

How many trades is a valid backtest?

Aim for at least 100 closed trades, and prefer 200 or more, because expectancy and win rate are averages and averages from small samples are unstable. With 30 trades, a couple of outsized winners can create an apparent edge that disappears entirely on the next 30. There is no magic threshold at which a result becomes true, but below roughly 100 trades you are usually measuring luck, and the more parameters your rule has, the more trades you need to justify them.

The removal test

A fast, brutal check: recalculate the result with your single best trade deleted. If the strategy flips from profitable to unprofitable, one trade was carrying the whole thing and the sample is too thin or too outlier-dependent to trust. Do the same with the worst trade. A robust strategy's conclusion survives both edits; a fragile one does not.

How far back should you backtest?

As far back as reliable adjusted data exists for the instrument, with 20 years as the target and the full available history preferred for index and ETF strategies. The important word is adjusted: prices must account for splits and reinvested dividends, or your results will understate total return badly, especially for dividend-paying and bond funds over long windows. Historical stock data covers what that adjustment actually does to a result.

One caveat runs the other way. Market structure has changed, and data from the 1980s reflects wider spreads, higher commissions, and different liquidity than today. Very long histories are useful for regime coverage but should not be treated as a perfect simulation of trading now, particularly for anything short-term or intraday.

Does longer always mean better?

No, and this is where the rule has a genuine edge case. Extending the window helps only while the added data resembles a market you could actually trade. Beyond that, more history can dilute a real structural change: a rule based on market microstructure that only exists post-decimalization gains nothing from 1990s data. The right instinct is to cover as many regimes as possible while asking, honestly, whether the earliest years describe the same game.

Split the window rather than just lengthening it

The higher-value move is usually not more years but better use of the years you have. Hold back the most recent 20% to 30% of data, fit the rule on the earlier portion only, then run it once on the untouched section. If performance collapses out of sample, the rule was fitted to noise. Doing this repeatedly across rolling windows is walk-forward analysis, which is the closest thing backtesting has to an honesty check.

How long should you backtest a trading system before trading it live?

After the historical test passes, most disciplined traders add a forward period of one to three months in paper or at minimal size before committing real capital. Historical data cannot capture your actual fills, your latency, or whether you will follow the rule after five losing trades in a row. That live-but-small phase tests the trader, not the strategy, and it fails more often than people expect.

A practical checklist

  1. Set the window to 20+ years of split- and dividend-adjusted daily data, or the full history of the instrument if it is younger, and say so if it is younger.
  2. Count the trades. Under 100, treat every conclusion as provisional regardless of how good the curve looks.
  3. Confirm at least two bear markets are inside the window. If not, the strategy has never been stress tested.
  4. Charge realistic costs on every trade. A high-frequency rule with no commission or slippage assumption is not a backtest, it is a chart.
  5. Hold back an out-of-sample slice and only look at it once, at the end. Looking early and re-tuning defeats the purpose entirely.
  6. Run the removal test on the best and worst trades and see whether the verdict survives.

Running that checklist by hand takes an evening per idea, which is why most ideas never get tested at all. Typing the rule in plain English into backtesting software runs it across 20+ years of adjusted daily data with costs charged per trade, prints every assumption on the result, and stamps an honest verdict against buy-and-hold, including UNDERPERFORMED when that is what the record says. For allocation and rebalancing rules the same applies on ETF backtesting and portfolio backtesting. The method itself is laid out in how to backtest a trading strategy, and the errors that most often invalidate a long backtest are collected in common backtesting mistakes.

Put it on the bench

Ideas are cheap. Verdicts take a bench.

Agenttrading restates your idea as a testable rule, backtests it on 20+ years of adjusted daily data, and explains the risks in plain English. Honest verdicts, even when the idea loses.

Past performance does not guarantee future results. For educational and informational purposes only. Not financial advice. Consult a licensed advisor.