Degrees of freedom and the size of the search
Every adjustable element in a system is a degree of freedom, and they multiply rather than add. Suppose a model has six parameters and you try ten plausible settings for each. That is ten to the power of six, one million combinations examined against the same history. Somewhere in a million combinations there is one that looks superb on any dataset, including a dataset of pure noise, because with enough attempts you are no longer finding structure, you are selecting a coincidence. The parameters do not have to look like parameters either. A session filter, a minimum range requirement, a choice of which year to exclude, the decision to use the body rather than the wick, and the length of an average are all dials. Count them honestly. Most people counting their own system find three or four times as many as they expected, and that count is what determines how much of the result can be noise.
The arithmetic of testing many variants
Put a number on the selection effect. Imagine you test twenty variants of an idea on the same data, and imagine none of them has any real edge at all. Give each variant a one in twenty chance of clearing a conventional significance bar by luck alone. The chance that a specific variant fails to clear it is 0.95, so the chance that all twenty fail is 0.95 raised to the twentieth power, which is 0.36. The chance that at least one clears the bar is therefore about 0.64. Test one hundred variants and the same calculation gives 0.95 to the hundredth power, which is about 0.006, so the chance of at least one apparent success rises to roughly 0.99. That is a near certainty, from an idea containing nothing. One blunt correction is to divide the significance level by the number of variants tried, so twenty variants mean each candidate is measured against 0.05 divided by 20, which is 0.0025.
A spike is not an edge, a plateau might be
There is a cheap diagnostic for a fitted parameter. Plot the backtest result against the parameter value across its whole plausible range instead of reading only the best setting. If the chart shows a lonely tall spike surrounded by poor results, the system works at that one number and nowhere near it. Markets do not contain a special truth about a fourteen period lookback that vanishes at thirteen and at fifteen. A spike is a coincidence wearing a label. If instead there is a broad region where results are decent and the peak sits inside that region, the parameter is probably describing something real, and your chosen value is a sensible point in a tolerant zone. The same logic applies to every dial at once, which is harder to visualise and exactly the same test. Prefer the centre of a plateau to the top of a peak, and accept the slightly lower backtest figure that comes with it.
The complexity curve and where it turns
Add rules to a system and the measured result on the fitting data improves without limit. There is always another condition that excludes a past loss. Measured on fresh data the same additions help for a while and then actively hurt, because each new condition is tuned to detail that will not repeat. The result is a curve that rises, turns and falls, and the turning point is the amount of complexity the data can support. It sits much earlier than intuition suggests. The uncomfortable consequence is that the version with the best backtest is almost never the version with the best future, so choosing by backtest score selects systems from the wrong side of the turn. A simpler system with a worse fitted number and a plausible reason for working is the better bet, which is also the argument against treating geometric level systems as signals when no mechanism explains why the level should matter.
The quiet forms that do not look like optimisation
Not all fitting involves a parameter sweep. Look ahead errors creep in whenever a calculation uses information the moment did not have, for example testing against a daily close while assuming an entry earlier that day, or using a revised data series. Instrument selection matters, because a rule chosen since it worked on the chart you happen to watch is fitted to that chart. Exclusion is powerful and feels responsible, since removing an unusual period looks like cleaning the data while it is really removing evidence. Peeking is the worst because it is invisible: you look at the out of sample result, change one thing, look again, and the holdout has quietly become training data. Each of these improves the number without improving the system, and none of them involves the word optimise. Treat any result that depends on a cleaning decision as unproven until it survives without the cleaning.
A budget of trades per parameter
A rough budget keeps the ambition honest. Count the trades in the test record, count the dials you turned, and divide. Two hundred trades and six dials gives about thirty three trades per dial, which is a thin allowance for anything subtle. The point of the budget is not the precise figure, since no formula settles it, but the direction it pushes you. Either collect more trades or use fewer dials, and collecting trades takes real calendar time. This is also why a high proportion of winners is such a seductive distraction during development. It is easy to manufacture by tightening conditions until only the comfortable past trades survive, and the result is a thin sample that cannot support any conclusion at all. A score is not a probability, and a count of surviving historical trades is not evidence that the conditions mattered.
What actually reduces the problem
Four habits do most of the work. Decide first, then test. Write the rule and the reason before you run anything, so the test is a test rather than a search. Keep a holdout you touch once. Carve off a chunk of history, lock it away, and spend it a single time at the end. If you look twice it is no longer a holdout. Prefer rules with a reason. A mechanism that explains why the behaviour should exist gives you something to check when results fade, while a pattern with no mechanism leaves you guessing. Report the search. Say how many variants you tried, not merely which one won, because that count is what anyone needs in order to judge the result. None of these habits improves a backtest. All of them make the backtest mean something closer to what you want it to mean.
The limit of all of this
No procedure proves a system is not overfitted. The best you can do is make overfitting less likely and make its signature easier to spot, which is a weaker claim than most development write ups imply. Parameter sensitivity, honest holdouts and declared search counts all reduce the risk and none of them eliminate it, because the history is finite and you have already seen it. There is a second limitation pointing the other way. Sometimes a system degrades out of sample because conditions changed rather than because it was fitted, and the two look identical from the inside. Separating them needs a reason for the rule and a view on whether that reason still holds. Keep a checklist of the ways a backtest lies, then take small live positions on the live chart, because forward results are the only measurement nobody has fitted.
FAQ
How do I tell overfitting from a market that changed?
You often cannot from the data alone, which is why a stated reason for the rule matters so much. If the mechanism is still plausible and the result merely faded, a change in conditions is credible. If there was never a mechanism, fitting is the simpler explanation and the one to assume by default.
Is a system with fewer parameters always safer?
Safer against fitting, yes, because there is less to tune. It can still be wrong for other reasons, including being too crude to describe anything useful. The goal is not the smallest possible rule set but the smallest one your trade count can actually support, which is usually smaller than planned.
Does the correction for multiple testing apply to discretionary ideas?
In spirit, yes. Every chart you scrolled through looking for a pattern was a test, and you ran far more of them than you counted. The arithmetic cannot be applied cleanly because the number of implicit comparisons is unknown, but the direction of the bias is identical and usually larger.
Can I use the holdout again after I change the system?
Not as a holdout. Once a result from that data has influenced a decision, the data is part of the fitting process whatever you choose to call it. Either accept that you now have no untouched sample, or set aside a new period and wait for it to accumulate, which is slow and honest.
Why does a plateau beat a peak if the peak scores higher?
Because the peak exists at one setting only and the future will not land exactly there. A plateau means nearby settings also work, so small errors in your choice cost little. You trade a slightly lower historical number for a result that does not depend on having guessed a precise value.
ⓘ See these ideas on real price: open the free XAUUSD live chart.