The null hypothesis for a trading system
A significance test starts by writing down the boring explanation and then asking how easily it accounts for what you saw. For a trading system the boring explanation is that the rules have no edge, so the average result per trade is zero before costs and the outcomes you recorded are just the scatter you would get from entering at arbitrary moments. That statement is the null hypothesis, and the whole procedure is an attempt to embarrass it. You never prove it false. You measure how surprising your results would be if it were true, and if they are surprising enough you decide the null is a poor description. Setting the null up carefully is most of the work. A null of zero edge after costs is a different and harder test than a null of zero edge before costs, and a null built from random entries with your actual holding period is harder still. Say which one you are using, because the three produce different verdicts on identical data.
Computing a t statistic from your trades
The standard test for an average divides the average by its standard error. Define the inputs: an average result of 0.15R per trade, a standard deviation of individual results of 1.1R, and 120 trades. The standard error is 1.1 divided by the square root of 120, which is 1.1 divided by 10.954, giving 0.1004. The t statistic is 0.15 divided by 0.1004, which is 1.494. The conventional two sided threshold at a five per cent level, with this many observations, sits near 1.98. So 1.49 does not clear it, and the honest summary is that the data is consistent with no edge. Now change only the trade count. With 400 trades the standard error becomes 1.1 divided by 20, which is 0.055, and the t statistic becomes 0.15 divided by 0.055, which is 2.727. That clears the threshold comfortably. The edge was identical in both cases. The sample did the work, which is the single most important thing to understand about these tests.
What a p-value is not
A p-value is the probability of seeing results at least as extreme as yours if the null were true. It is not the probability that the null is true, and it is not the probability that your edge is real. Those are different quantities and converting between them needs information the test does not contain, chiefly how plausible the idea was before you tested it. That prior plausibility is exactly where trading ideas are weakest, because most patterns people test were found by looking at the same data. A p-value also says nothing about the size of an effect. With enough trades a trivially small edge produces a tiny p-value, and with few trades a large edge produces an unimpressive one. Finally, a value just above the threshold is not evidence of nothing, and a value just below is not proof of something. The threshold is a convention for making decisions, not a boundary in nature, and treating it as one causes most of the confusion that follows.
The permutation test, which fits trading better
There is a test that suits trading more naturally than the textbook formula, because it builds its own null from your own data. The idea is to break whatever link you believe exists while keeping everything else intact. One version shuffles the signal labels across the history, so the same number of entries occurs with the same holding period but at arbitrary moments, then records the result. Repeat a few thousand times and you have a distribution of outcomes from a method with your trade count, your holding period and your cost structure, but no edge. Then see where your real result falls in that distribution. If one random arrangement in two hundred beat your real one, the p-value is about 0.005. This approach handles the awkward features of trade data, including odd distributions and overlapping positions, without assuming normality. It also forces you to state exactly what you are claiming, since the way you shuffle defines the null. The same method is the sensible way to test a seasonal claim, which is why most seasonal patterns in gold deserve a shuffle test before anyone acts on them.
Looking repeatedly and stopping when it looks good
There is a specific way to manufacture significance that feels like diligence. You test an idea, get an unimpressive result, collect another month of data, test again, and keep going until the figure clears the threshold, then stop and write it up. This is called optional stopping and it inflates the false positive rate badly, because you have given yourself many chances to cross the line and only recorded the crossing. The same thing happens when you test the idea on one timeframe, then another, then on a different session, and report whichever version worked. Each look is a test, whether or not you counted it. Two defences exist and both are uncomfortable. Fix the sample size in advance and analyse once. Or declare every look, and apply a correction that accounts for how many you took. This is why a careful note about treating divergence readings with caution matters more than a clever test, since the signal was usually chosen after the fact.
Costs turn a significant result insignificant
Tests on gross results answer a question nobody trades. Keep the 400 trade example with its average of 0.15R, a standard deviation of 1.1R and a standard error of 0.055, which gave a t statistic of 2.727. Now subtract a cost of 0.08R per trade, a figure you would take from your own fills rather than assume. The net average becomes 0.07R. The standard error is unchanged at 0.055, so the t statistic becomes 0.07 divided by 0.055, which is 1.273. That sits well below the 1.98 threshold. The same 400 trades support a confident claim about a gross edge and no claim at all about a net one. This is not a quibble about decimals. It is the usual fate of thin edges, especially those relying on frequent entries, and it is the reason any significance claim should be computed after costs or not at all. Testing gross and then trading net is one of the most expensive habits in system development.
Significance is not size and not stability
A test that clears the threshold tells you the result is hard to explain by scatter alone, under the null you specified, assuming you did not search for it in the same data. Three large qualifications. It says nothing about size, so a significant edge can be too small to be worth the attention after costs and effort. It says nothing about stability, since the test treats the period as one sample and cannot see that the behaviour faded halfway through. And it says nothing about the future, because every observation came from conditions that have ended. Significance is a filter that removes weak claims rather than a licence to act, and the direction of travel is always from a stated reason, to a pre registered test, to small forward positions, to a slowly accumulating record on the live chart. It is also why no test can support an answer to the question of which way gold goes today, which is a single unrepeatable event rather than a sample.
FAQ
What p-value should I require for a trading idea?
There is no correct level, and the conventional five per cent is a borrowed convention rather than a principle. What matters more is the number of ideas you tested, since the threshold must tighten as the search widens. Many people are better served by asking whether the effect is large and explainable than by chasing a specific cut off.
Is a t test valid when trade results are not normally distributed?
Reasonably robust for the average once the sample is moderately large, because the sampling distribution of an average behaves better than the raw data. It copes badly with extreme outliers, which trade results often contain. A permutation test avoids the assumption entirely and is usually the better choice for this kind of data.
Can I test significance on a backtest I developed on the same data?
You can run the arithmetic and the output will not mean what it appears to mean. The test assumes the hypothesis was fixed before the data was seen, and in development it never is. Treat such a figure as a descriptive summary, and reserve any significance claim for a holdout or forward period.
How do I account for having tested many variants?
State the number of variants tried, then tighten the threshold accordingly, for example by dividing the level by that number. The correction is crude and conservative. The more useful discipline is to reduce the number of variants in the first place by choosing ideas with a stated reason rather than sweeping through combinations.
Does an insignificant result mean the idea is worthless?
No. It means this sample cannot distinguish the idea from no edge, which is a statement about the evidence and not about the idea. A genuine but modest edge regularly fails to reach significance in a few hundred trades. The correct response is a smaller position and more data, not abandonment or false confidence.
ⓘ See these ideas on real price: open the free XAUUSD live chart.