The standard error, in one line
Take the results of your trades in R and compute two things: the average, and the standard deviation of those individual results. The average is your estimate of the edge. The standard deviation describes how much the individual trades scatter around it. The standard error of the average is the standard deviation divided by the square root of the number of trades, and it tells you how far your estimate is likely to sit from the truth. Roughly two standard errors either side of the average gives a range that usually contains the real value, which is the band you should quote whenever you quote an edge. The scatter in the denominator is the part people overlook. Trade results in R are widely spread by construction, since a system producing 3R winners and 1R losers has a large spread even when it works. Large scatter plus a small count equals a wide band, and a wide band means the number you are proud of has not been measured yet.
What thirty trades is actually worth
Define the inputs so the arithmetic is checkable. Say the average result is 0.2R per trade and the standard deviation of individual results is 1.2R. At 30 trades the standard error is 1.2 divided by the square root of 30, which is 1.2 divided by 5.477, giving 0.219. Two standard errors is 0.438, so the plausible range runs from about minus 0.24R to about plus 0.64R. That band contains zero, a weak edge, a strong edge and a losing system, all at once. At 100 trades the standard error is 1.2 divided by 10, which is 0.12, so the band is 0.2 plus or minus 0.24, running from minus 0.04 to plus 0.44. Still touching zero. At 400 trades the standard error is 1.2 divided by 20, which is 0.06, and the band runs from 0.08 to 0.32, which finally excludes zero. The edge was identical in all three cases. Only the evidence changed.
Solving for the number you need
Turn the formula round and ask how many trades are required before the band stops containing zero. You want two standard errors to be smaller than the average, so two multiplied by 1.2, divided by the square root of n, must be less than 0.2. That rearranges to the square root of n being greater than 12, which means n greater than 144. Check the boundary: at exactly 144 trades the standard error is 1.2 divided by 12, which is 0.1, and two standard errors is 0.2, precisely equal to the average. So 144 trades is the point where the evidence first separates the result from nothing at all, under these defined inputs. Two consequences follow. A larger spread of individual results pushes that figure up quickly, since it enters the calculation directly. And a smaller edge pushes it up fast as well, because the required count scales with the square of the ratio between spread and edge. Thin edges need very large samples.
The same question asked about a win rate
Proportions follow the same logic with a different formula. The standard error of an observed win rate is the square root of the rate multiplied by one minus the rate, divided by the number of trades. Suppose you observe 60 per cent winners in 50 trades. That is the square root of 0.6 multiplied by 0.4, which is 0.24, divided by 50, giving 0.0048, and the square root of that is 0.0693. Two standard errors is 0.1386, so the plausible range for the true rate runs from about 46 per cent to about 74 per cent. You cannot even tell whether this system wins more often than it loses. Push the same observed rate out to 500 trades and the standard error becomes the square root of 0.24 divided by 500, which is 0.0219, so the band narrows to roughly 56 per cent to 64 per cent. That is a usable estimate. Fifty trades was not, and no amount of confidence in the method changes the width of the band.
How long that takes in calendar time
Convert the counts into weeks and the problem becomes concrete. At three trades a week, 144 trades takes 48 weeks. At one trade a week, which is a realistic rate for a selective method, the same count takes nearly three years. This is the uncomfortable trade off hiding inside the argument for taking only the best setup each week. Selectivity usually improves the quality of each trade and it slows the accumulation of evidence to a crawl, so a highly selective trader may never collect a statistically convincing sample of any single setup type. The honest response is not to trade more in order to gather data, which trades a known cost for an uncertain benefit. It is to accept that your conclusions will remain provisional, to lean on reasoning about mechanism rather than on counts, and to size positions as though the edge might be smaller than measured.
Why the effective sample is smaller than the count
Every formula above assumes the trades are independent and drawn from the same process. Three things break that in practice and all of them shrink the effective sample below the raw count. First, dependence. Trades taken in the same conditions within a few days share whatever that condition was, so five trades in one trending week carry less information than five trades spread across five unrelated weeks. Second, changing rules. If you adjusted the entry criteria at trade sixty, the first sixty trades belong to a different system and cannot be pooled with the rest, however tempting it is. Third, fat tails. The standard deviation itself is estimated from the same small sample, and a few outliers make that estimate unreliable, so the band you compute is itself uncertain. A good test is to count only the trades taken under the exact current rules, in genuinely distinct conditions. That number is usually a fraction of what the journal contains.
What to do with a small sample instead of believing it
Small samples are the normal condition, so the skill is operating sensibly inside one rather than waiting for certainty. Report a range rather than a point whenever you describe your results, because a single figure invites a confidence the data does not support. Size positions for the lower end of that range, not the middle. Lean on a stated mechanism for why the setup should work, since a reason can be reasoned about while a count of twenty cannot. Pool related setups carefully to gain sample, and note that pooling is a decision you should make before looking at the results. Then keep collecting under fixed rules, which means writing the rules into a trading plan and leaving them alone, and recording every trade as it happens on the live chart rather than reconstructing it later. The sample grows slowly and that is the only honest route. Any shortcut just relabels a small sample as a large one.
FAQ
Where does the figure of thirty trades come from?
It comes from a rough convention in introductory statistics about when a sampling distribution becomes roughly normal, not from any claim that thirty observations measure an effect precisely. It has been borrowed and misapplied. For trade results, where the spread between individual outcomes is wide, thirty is nowhere near enough to separate a modest edge from nothing.
Does a longer backtest solve the problem?
It helps with the count and not with the independence, and it introduces a different risk. More history means older conditions, and a rule measured across a very long period is averaged over markets that behaved differently. You gain sample and lose relevance, so the extra data is useful without being a straightforward improvement.
Can I count every trade across different setups together?
Only if you intend to trade them as one system and you decided to pool them before seeing the results. Pooling after the fact, especially dropping the setups that did poorly, is selection rather than measurement. If you genuinely run one process with variations, pool it. If you run several methods, each needs its own sample.
My last twenty trades were all profitable. Does that mean anything?
It means you had a good run, which systems with no edge also produce regularly. Twenty consecutive winners is rare for any realistic win rate, so it more likely indicates that a condition suited your method recently, or that the targets were very close relative to the stops. It is not evidence about the long run.
Is forward testing better than a larger backtest?
Different rather than better, and it is the only data that includes your real fills and your real behaviour. It accumulates far more slowly, so most people need both: history for a rough view of whether the idea has any basis, and a forward record for the figure they actually trust.
ⓘ See these ideas on real price: open the free XAUUSD live chart.