Subtle & customizable audible cues to indicate when charts turnover, plus world clocks for any market — by & for professional traders
Open the appGuidance last reviewed 2026-08-19
Every statistic in a trading journal is an estimate made from a sample. With enough trades those estimates settle down. With few, they produce numbers that look precise and mean almost nothing.
This is the least interesting page in any analytics tool and the one that governs all the others. A profit factor of 2.4 over eighteen trades is not a finding. It is a coincidence with a decimal point.
| Question | Rough answer |
|---|---|
| Why small samples lie | Trading results scatter far more widely than intuition expects |
| Win rate | ~100 trades for Β±10 points, ~400 for Β±5 |
| Expectancy | Several hundred, more if results are dispersed |
| Maximum drawdown | Never settles β it grows with the record |
| What 20 trades CAN tell you | Whether you followed your own plan |
Take a method that genuinely wins 55% of the time. Over 20 trades, the number of wins you actually see will land somewhere around 11 β but "around" is doing a great deal of work. Runs of 7 wins from 20 and runs of 15 from 20 both happen regularly, and they read as a broken method and a brilliant one.
The intuition that fails is this: we expect randomness to look scattered, so a run of eight losses feels like evidence of something. In a process with a 45% loss rate, a run of eight losses shows up roughly every couple of hundred trades. It is not evidence. It is Tuesday.
β The consequence people act on: small samples do not just produce uncertain answers, they produce confidently wrong ones. Twenty trades will hand you a profit factor, an expectancy and a win rate to two decimal places, and none of them are measurements.
There is no single threshold, because it depends on how widely your results scatter. A method producing every result between β1R and +1.2R settles quickly. One where most trades are small and a rare +12R carries the record may never settle at all.
Rough working figures:
| Statistic | Starts to be meaningful | Reasonably settled |
|---|---|---|
| Win rate | ~100 trades | ~400 trades |
| Payoff ratio | ~100 trades | ~300 trades |
| Expectancy | ~200 trades | 500+ trades |
| Profit factor | ~200 trades | 500+ trades |
| Standard deviation of R | ~100 trades | ~300 trades |
| Maximum drawdown | β | β οΈ Never. It grows with record length by definition |
β οΈ These are conventions, not findings. They come from the arithmetic of sampling error and from what practitioners tend to converge on, and the right number for your method depends on your own dispersion. Treat them as an order of magnitude.
A rule of thumb worth more than the table: the precision of an average improves with the square root of the sample. To halve your uncertainty you need four times the trades. Going from 25 to 100 trades is a real improvement; going from 100 to 120 is not.
Testing several ideas and keeping the best one.
If you evaluate eight setups over thirty trades each, the best-looking one is very likely to be the luckiest rather than the best. You did not find an edge; you ran a competition that variance won.
β Be most suspicious of your best-looking statistic. The smaller the sample, the more likely your standout number is the one luck inflated most β and it is precisely the number you will want to trade larger.
The same applies to optimising parameters. A moving-average length that worked beautifully on your history has been selected for working on that history, which is not the same as working.
You do not need the formulas to use the idea. What matters is that every statistic has a range around it, and the range is wider than people assume.
A 60% win rate over 50 trades is consistent with a true rate anywhere from roughly 46% to 73%. That is not a technicality β the bottom of that range and the top are different businesses. The same 60% over 500 trades narrows to roughly 56β64%, which is an actual measurement.
β The habit worth building: when you see a statistic, ask what it would be if you had been a bit unluckier. If the answer still works, you have something. If a slightly worse run makes it unprofitable, you have a sample.
Take your R-multiples, shuffle the order, and redraw the equity curve. Do it fifty times.
You will get fifty different curves from exactly the same trades β with different maximum drawdowns, different longest losing streaks, and sometimes a very different-looking record. That spread is what variance alone can do to a fixed set of results.
Two things fall out of it, and both are useful:
Your actual drawdown was a draw from that distribution, not a property of the method. Some of the shuffles will be considerably worse than what you lived through, and those orderings were just as likely.
If a meaningful share of the shuffles end unprofitable, the record depends on sequence rather than edge.
β οΈ Reshuffling assumes trades are independent of each other. If yours cluster β several positions on one theme, or revenge trades after a loss β they are not, and the exercise understates the real spread.
This is the part usually missed, and it is the reason to keep records from day one even though the statistics are worthless at that stage.
A small sample cannot measure your edge. It can measure your Behavior.
Things that are perfectly visible after twenty trades:
β Process adherence needs no sample size, because you are not estimating anything β you are counting. Twenty trades is enough to tell you whether you are following your own rules, and that is usually the thing costing money in the first year regardless of edge.
| Term | Meaning |
|---|---|
| Sample | The trades you have, from which everything is estimated |
| Variance / dispersion | How widely individual results scatter |
| Sampling error | The gap between your sample's figure and the true one |
| Confidence interval | The plausible range around an estimate |
| Selection effect | Picking the best of several tests, and being fooled by it |
| Reshuffling | Randomizing trade order to see what sequence alone can do |
| Independence | Whether one trade's result is unrelated to the next |
Nothing here is a metric. It is the caveat attached to all of them β see Trading Performance Metrics Explained for the numbers themselves, and Maximum Drawdown for the one statistic that never settles.
The K-ratio and SQN both fold sample size into the calculation, which is a large part of why they exist.
The trade counts given are conventions and orders of magnitude rather than findings; the right number for any method depends on how widely its results scatter.
Chart Ding is a market clock with customizable alarms for traders — candle-close alerts, world market sessions, and holiday warnings. Open the app.
Nothing on this website is trading or investment advice. Chart Ding is a market clock and alarm tool — it does not recommend trades, evaluate strategies, or take account of anyone's circumstances. Nothing here should be construed as a recommendation or relied on as the basis for a trading decision. Consult a licensed professional before placing real money at risk. Trading involves risk of loss.