How the strategies are tested

Markets
Method
Most backtests flatter the strategy. These are the eight steps every strategy on this site goes through before it is published, and the traps each step is there to catch
Author

Louis Foy

Published

October 2, 2026

A backtest asks a simple question: if this rule had been followed in the past, what would have happened? It is easy to get a flattering answer. Prices are known with hindsight, costs are easy to forget, and given enough attempts almost any rule can be tuned until it looks brilliant on the data it was tuned on. The steps below are the ones every strategy in this section goes through. Each exists because skipping it produces a number that will not survive contact with the market.

2. Use the whole market, including the failures

Price data cover every American common stock, about 12,600 of them, including more than 6,600 that have since been delisted through bankruptcy, takeover or decline. Testing only on companies that exist today, known as survivorship bias, quietly removes the losers and inflates returns. The data are daily bars from Massive (formerly Polygon.io); insider trades come straight from the SEC’s quarterly Form 4 datasets.

3. Only use what was known at the time

This is where most backtests go wrong, usually without anyone noticing. Look-ahead bias means letting information from the future leak into a past decision. The rules used here:

  • Act on the next open. A signal based on a day’s close, or on a filing published that evening, can only be traded the next morning. Every strategy enters at the next day’s open.
  • Undo split adjustments before filtering on price. Historical prices are adjusted for later splits, so a $0.40 penny stock that later did a 1-for-1,000 reverse split appears as a $400 share. A “$5 minimum price” filter then lets in exactly the stocks it was meant to exclude. On one strategy this error made results four times worse.
  • Point-in-time company size. Market value is computed from shares outstanding reported before the trade date, times the price on that date, never from today’s figures.
  • Check the data actually exist. Some series start later than they appear to. Exchange-rate history on the data plan used here begins in late 2024, for example; one early test silently dropped every trade before that date until the gap was found.

4. Charge realistic costs

Each trade pays Interactive Brokers’ commission on both the purchase and the sale (the tiered schedule, with its 35-cent minimum, unless stated otherwise), plus a modelled bid-ask spread of 3 to 15 basis points a side, wider for smaller and less liquid shares. Results are rerun with the spread doubled to see whether the edge survives.

Trade size matters more than most backtests admit. A fixed minimum commission is trivial on a £5,000 position but costs about half a percent of a £100 round trip, enough to wipe out a strategy whose average gross gain is under 1%. Every strategy page states the stake it was tested at.

5. Compare with simply buying the market

A strategy that made 10% in a year when the S&P 500 made 25% has not worked. Every trade’s return is set against the S&P 500 (SPY) over exactly the same days. A rule that only makes money in a rising market, and less than the market, is a leveraged index fund with extra steps.

6. Keep data back that the rules never see

Before any testing starts, the most recent stretch of data is set aside as a holdout. Rules are developed on the earlier years, with results checked separately before and after a split date within them. When the rules are final they are frozen and run once on the holdout. If the holdout result is poor, the strategy is rejected, not re-tuned, because re-tuning on the holdout turns it into one more piece of in-sample data.

7. Try to break it

A good result invites suspicion. Each candidate is put through the same checks:

  • Neighbouring settings. If a 3-day hold works but 2 and 4 days do not, the result is probably luck. Real effects change gradually as parameters move.
  • Remove the best trades. Dropping the top 1% of winners shows whether the profit comes from a steady edge or from two lucky lottery tickets.
  • Year by year. An edge that appears in only one year is a feature of that year, not of the rule.
  • Statistical significance. The t-statistic on average returns should be comfortably above 2, and the threshold rises with the number of variants tried. Testing 30 ideas and reporting the best one guarantees a false positive.
  • Count the attempts. Every page says how many variants were tested on the same data, so readers can judge the risk of a fluke for themselves.

8. Paper trade before real money

Strategies that pass are run live without money. An automated scanner reads new data each evening, sends a phone alert with the trades it would make, and records simulated fills at the real opening and closing prices. The paper results are then compared with the backtest. Only when the two agree, over enough trades to mean something, does a strategy graduate to a real account.

What gets published

Both. Strategies that pass are published with their rules, results and weaknesses. Strategies that fail are published too, because a clear negative result, such as the finding that buying gap-ups on heavy volume has lost money since 2021, is as useful as a positive one and far less likely to be a fluke. Nothing here is investment advice, and a positive backtest is a reason to keep testing, not a reason to trade.

Check The trap it catches
Delisted stocks included Survivorship bias
Enter at the next open Trading on information not yet available
Unadjusted prices for filters Split-adjustment look-ahead
Commission + spread, doubled spread test Edges that exist only before costs
Return versus the S&P 500 Mistaking a bull market for skill
Holdout run once Overfitting to the test data
Neighbours, top 1% removed, by year Luck, outliers and one-off regimes
Paper trading Backtest assumptions that fail in practice

Not investment advice. Backtests are hypothetical and past performance is not a reliable guide to future results. See the disclaimer.