Home / Blog / The equity curve you got was one sample from a distribution
TRADING GUIDE

The equity curve you got was one sample from a distribution

A backtest hands you one line. It feels like the answer, because it is the only one you have seen. Shuffle the order of the same trades and the line changes shape while the end point stays put. Monte Carlo testing is the habit of generating many of those lines on purpose, so you can see how much of your result was sequence and how much was edge.

📅 October 8, 2026⏱ 8 min readBy XAUUSDLiveChart Research Desk
Track gold in real time on the live chartOpen Live Chart →
THE EQUITY CURVE YOU GOT WAS ONE S
XAU/USD…
01

What the method actually does

Monte Carlo testing means rebuilding the same system many times with the randomness put back in. You start from a list of outcomes, one row per trade, each expressed in R so position size drops out of the maths. Then you construct new histories from that list. A reshuffle uses every trade exactly once and changes only the order. A bootstrap draws trades at random with replacement, so some appear twice and others not at all, which lets the composition of the sample vary as well as the sequence. A parametric run ignores the list and draws from inputs you state yourself, such as a win rate and a payoff multiple. Each variant answers a different question. Reshuffling asks what the path could have looked like. Bootstrapping asks what a different draw from the same behaviour could have looked like. A parametric run asks what a system with these stated properties does over a long horizon. None of the three is a forecast. All three describe variability inside a model you chose.

02

Reordering keeps the total and ruins the comfort

Reorder a set of trades and the total is untouched, because addition does not care about sequence. The drawdown cares a great deal. A cluster of losses that lands early arrives while the account is small, so the same string of losses removes a different fraction of equity depending on when it turns up. Move the cluster to the middle and the curve sags after a comfortable start. Move it to the front and the account spends a long stretch underwater before anything works. Both histories hold identical trades and identical profit. Only one of them would have been easy to sit through. A single backtest cannot show you this, because it shows one order out of an enormous number. Ten trades can be arranged in 3,628,800 different orders. One hundred trades can be arranged in a number with more than one hundred and fifty digits. Your backtest picked one of them and then invited you to draw conclusions from it.

Same 100 trades, five orderings, one shared ending0R+20Rtrade 1trade 100deepest drawdown of the fivethe one your backtest showed
03

A streak calculation you can do on paper

Streak length is where a little arithmetic pays. Define a system that loses on six trades out of every ten, so the loss probability is 0.6. The chance that ten specific consecutive trades are all losses is 0.6 raised to the power of ten, which works out at 0.0060, a touch over half of one per cent. Most people stop there and feel safe. Now count the opportunities instead. A run of one hundred trades contains ninety one overlapping windows of ten. Multiply 91 by 0.0060 and the expected number of all losing windows is about 0.55. Stretch the run to two hundred trades and there are one hundred and ninety one windows, so the expected count rises to about 1.15. In that system a ten trade losing run is not a freak. It is roughly what you should plan to meet once. Monte Carlo performs this counting across every streak length at once, which is why its output is a distribution rather than one reassuring figure.

04

Reading the output, which is a tail and not an average

The useful output is never the average ending. Averages are what expectancy already told you. What you want from the run is the shape of the bad end: the distribution of worst drawdown, the longest losing run, the longest time spent below a previous high, and the spread of final equity. Read those at the unpleasant end rather than the middle. If a tenth of the simulated histories show a drawdown you could not hold, then a tenth of the futures that behave like your backtest end with you abandoning the plan at the worst possible moment. That is a sizing problem and not an entry problem, which is why this work belongs next to your risk management rules. It also tends to expose optimistic assumptions about risk to reward, because a long tail of small losses punishes any system that needs a rare large winner to pay for everything else.

05

Reshuffling, bootstrapping and parametric runs

Choosing between the three variants matters more than the number of runs. Reshuffling is the honest minimum. It asks only about order, keeps the exact set of outcomes, and cannot invent a loss larger than any you recorded. That last point is also its weakness, since a system that has never met its worst day will be simulated as though the worst day cannot happen. Bootstrapping loosens the grip by allowing composition to vary, so a single run can contain four copies of your largest loss. It still cannot exceed the largest loss in the sample. Parametric runs escape the sample entirely and let you ask direct questions, for instance what happens if the average loss is a fifth larger than recorded because of slippage. The cost is that you are now testing your assumptions rather than your trades. Run all three and compare them. Where they disagree, the disagreement is the information.

Reshuffle keeps the trades, bootstrap changes the mixtrade list in Rone row per tradereshuffleorder onlybootstrapdraw with replacementparametricinputs you statemany curvesone run eachread the tailworst drawdownlongest losing runtime under waterwhat you take from the tail is position size, not a new entry rule
06

Where Monte Carlo quietly flatters a system

Every version assumes trades are independent and drawn from the same behaviour throughout. Trading breaks both assumptions. Outcomes cluster, because a market condition that suits a system produces several winners in a row and then stops producing any. Positions in one instrument are correlated by definition, so two open trades can really be one bet. Costs are usually held at a constant in the model while in practice they widen exactly when the market moves fastest. The deepest problem sits upstream of all this. The trade list fed into the simulation came from a backtest, and if that backtest was fitted to its own history then the simulation will produce a beautifully detailed picture of a fantasy. Precision about the variability of a wrong number is still wrong. Work through a backtest honesty checklist before you run a single simulation, because this method inherits every flaw in its input and adds a layer of credibility on top.

07

A routine that uses it before the money is on

A workable routine is short. Write down the trade list and the assumptions in one place, including cost per trade and the largest loss you accept as possible. Run a reshuffle, a bootstrap and a parametric version rather than picking one. Record the drawdown you would meet in the worst tenth of histories, then set size so that figure is tolerable rather than merely unlikely. Re-run the whole thing whenever the sample grows by a meaningful amount, and log forward trades as they happen on the live chart so the sample you feed back in is a real record and not a remembered one. Expect the distribution to widen when live trades enter, not narrow. That widening is the method working properly. If the simulated tail and the live experience keep disagreeing in the same direction, your model of the system is wrong, and no number of extra runs will repair it.

Q

FAQ

How many simulations are enough?

A few thousand runs is plenty for reading a tail, and going to a hundred thousand changes little beyond the third decimal place. Precision in the simulation count is cheap and almost never the binding constraint. The binding constraint is the quality and length of the trade list you are resampling, which no amount of extra runs improves.

Does reshuffling change the final profit?

No. A reshuffle uses every trade once, so the total is identical in every run. Only the path differs, which is exactly the point. The method isolates sequence risk by holding the outcomes fixed, so any variation you see in drawdown comes purely from the order in which the same results arrived.

Can Monte Carlo tell me whether my edge is real?

It cannot. It assumes the edge in your trade list is genuine and then explores the paths such an edge could produce. Testing whether the edge exists at all is a separate job for out of sample work and significance testing. Running a simulation on a fitted backtest simply dresses up the original error.

Why does a bootstrap sometimes look worse than a reshuffle?

Because it can draw your largest loss more than once and can leave your largest winner out entirely. That produces histories a reshuffle is unable to build. If the bootstrap tail is far worse than the reshuffle tail, your result depends heavily on a small number of individual trades, which is worth knowing early.

What should I actually change after running one?

Usually position size, occasionally the maximum number of open trades, and sometimes the decision to trade the system at all. The simulation does not improve entries or exits. It tells you what the system can put you through, so the response is a limit you can live with rather than a new rule.

ⓘ See these ideas on real price: open the free XAUUSD live chart.

More from the blog

View all posts →
Join GroupChat