Learn / Risk & Position Sizing

How many trades before you can judge a strategy

A run of twenty trades is consistent with almost any underlying edge, including none. The number of trades needed before results mean anything depends on your strategy’s shape — and high-R, low-win-rate approaches need substantially more than people expect.

Why small samples say nothing

At a 45% win rate, over twenty trades you might win six or win thirteen. Both are entirely ordinary. Six wins looks broken; thirteen looks excellent. The underlying edge was identical.

The general point: the variance of a proportion over a small sample is large enough to swamp the difference between a good strategy and a bad one. So the observed win rate over twenty trades is mostly noise about the true one.

You can feel this without statistics. Losing streaks of six to nine trades are ordinary at typical win rates. If a normal streak can consume a third of your sample, the sample cannot be telling you much.

High-R strategies need more, not less

This is the part that catches people, and it is the opposite of the intuition that bigger wins mean faster confirmation.

A strategy winning 50% at 1R has its return spread evenly across trades. Most trades contribute similarly, so the average stabilises relatively quickly.

A strategy winning 25% at 3R concentrates its return in a minority of trades. Now the sample’s result depends heavily on how many of the rare large winners happened to land in it. Miss them and a genuinely positive strategy looks negative; catch two early and a mediocre one looks excellent.

ShapeReturn concentrationSample needed
60% win, 1RSpread across most tradesSmaller
45% win, 2RModerately concentratedLarger
25% win, 4RIn a quarter of tradesMuch larger
10% win, 10RIn one trade in tenVery large

The rule of thumb that follows: the more your edge depends on rare events, the longer you cannot know whether you have one. Which is uncomfortable, because high-R approaches are attractive for exactly the reason that makes them slow to validate.

What to look at while you wait

You are not helpless with a small sample. You are just looking at the wrong thing if you are looking at profit.

Process compliance. Did you take the trades your rules specified, at the sizes they specified, with stops where they belonged? This is measurable immediately and it is the precondition for the results meaning anything. Results from trades you did not follow your rules on do not evaluate your rules.

Realised versus planned R. The gap is fees, slippage and early exits, and it converges much faster than expectancy because every trade contributes. If your planned 2R is realising as 1.3R, that is actionable on a small sample.

Whether the distribution looks like you expected. If you designed for 3R winners and nothing has exceeded 1.5R, that is information about the setup even before the sample is large.

Execution quality. Slippage on entries versus exits, and on stop-outs specifically — see slippage and market impact. This stabilises quickly.

All of these are measurable well before expectancy is, and all of them are things you can act on.

The stability test

Rather than looking for a threshold count, watch whether your estimate moves.

Compute expectancy after each new trade and look at the series. Early on it lurches with every result. As the sample grows the lurches shrink. When a new trade no longer moves the estimate much, you have something.

If it is still swinging materially on each new trade, the sample is too small — whatever the count is. This has the advantage of adapting automatically to your strategy’s shape, which a fixed number cannot.

What small samples are used for instead

Abandoning strategies too early. A positive-expectancy approach inside an ordinary losing streak looks broken, and the natural response is to change something. You have then replaced a strategy you had not evaluated with one you cannot evaluate either, and reset the sample to zero.

Over-confidence from a good start. Symmetrical and more expensive, because it usually comes with increased size.

Optimisation. Tuning on a sample too small to distinguish signal from noise fits the noise. In an AI-assisted setup this is easier to do accidentally than with parameters, because trying several prompts and keeping the one that worked does not feel like optimisation — see why AI backtests do not survive live trading.

The uncomfortable implication

If you take a few trades a week at a high-R approach, the sample size needed implies months to years before the results mean much.

That is genuinely inconvenient and it does not go away by being ignored. Two honest responses: trade a shape whose edge is spread across more trades, or accept that your evaluation horizon is long and size accordingly — because risk of ruin is about surviving long enough for the edge to express itself, and a long evaluation horizon means a long survival requirement.

FAQ

How many trades do I need to evaluate a strategy?

More than twenty, and how many more depends on the shape. A strategy whose return is spread evenly across trades stabilises faster than one where a minority of large winners carries everything. The practical test is whether your expectancy estimate still moves materially when a new trade arrives — if it does, the sample is too small regardless of the count.

Why do high-R strategies need larger samples?

Because their return is concentrated in fewer trades. If a quarter of your trades produce all the profit, then whether a sample contains the expected number of those trades dominates the result. A genuinely good strategy can look negative over a stretch that simply missed them.

My strategy is losing after 30 trades — should I change it?

Thirty trades cannot distinguish a losing strategy from a winning one inside an ordinary drawdown. What is worth checking at that point is process: did you take the specified trades at the specified sizes with stops where they belonged, and is realised R tracking planned R? Those are answerable now; profitability is not.

What can I measure before the sample is large enough?

Process compliance, the gap between planned and realised R, execution quality including slippage on stop-outs, and whether your outcome distribution resembles what you designed for. All of these converge faster than expectancy and all are actionable.