Skip to content
Let's All Get Right

Building and Testing a Trading System

Testing Your System: Success, Backtesting, Ratios, and the Journal

How to test a set of trading rules without fooling yourself, why win rate is the wrong scoreboard, and the one number that answers whether a system makes money.


You can't test a system against a goal you never wrote down

You have a trading plan now — written rules for what you buy, when you exit, and how much you risk. The obvious next question is whether it's any good. That question has no answer until you say what "good" would look like, because a test needs something to measure against.

So start there. Testing isn't a thing you do once, at the end. It starts the moment the plan is an idea, and it keeps running for as long as you use it.

"Success" means three different things, and people mix them up

The first is you. Did you follow your own rules? Did you take the entry your plan told you to take, and the exit it told you to take, on the day you didn't feel like it? That's a success you control completely, and it's the only one you control completely.

The second is a trade. And here's the part that reads wrong the first time: a trade can be a success and lose money. If you followed the rules, sized the position the way your plan said, and the trade went against you anyway — the system worked. It just didn't pay this time. No set of rules wins every trade. A plan that expects to is broken before it starts.

The third is the plan itself. This is the one that has to earn its place. If you follow your rules faithfully for months and the money keeps going the wrong way, then you're a disciplined trader running a plan that doesn't work. Discipline isn't the problem there. The plan is.

Success without a benchmark isn't measured at all

Video coming soon

The benchmark as a high-jump bar — what it means to clear it, and what it means to keep missing.

This lesson explains the idea in full without it.

Say your system made 8% last year. Good? You genuinely cannot tell. If the broad market made 4%, you did something. If it made 20%, you spent a year of your attention to fall twelve percentage points behind what you'd have gotten by buying an index fund and going outside.

That's why you need a benchmark — the index you measure your performance against. Think of it as a high-jump bar. Clearing it means your rules added something. Missing it repeatedly means they didn't, and no amount of effort you put in changes that reading.

The bar has to match what you're jumping over. A benchmark should look like the kind of thing your plan trades. If your plan buys large U.S. companies, an index of large U.S. companies — like the S&P 500 — is the honest comparison. If it trades small companies, a small-company index like the Russell 2000 is the fair one. Measuring a small-company system against a large-company index tells you about the difference between those two markets, not about your rules.

Figure

A high-jump bar with a runner's arc drawn over it. Above the bar, labeled 'your system's return'; the bar itself labeled 'benchmark return.' A second, lower panel shows the same arc passing under the bar, labeled 're-examine the plan.'

Two honest complications.

Return isn't the only axis. If your system slightly trails its benchmark but rides through the year with much smaller swings, that may be a trade you'd take on purpose. Less volatility — how much a value jumps around — for slightly less return is a real choice, not a failure. What you can't do is compare returns and pretend the risk column doesn't exist.

And if it trails on both axes, say so out loud. A system that returns less than its benchmark and swings harder than its benchmark has been beaten by a low-fee index fund that required nothing of you. That's not a verdict on you. It's information, and it cost you a year to get, so don't throw it away by refusing to look at it.

Reward and risk: a number you set before you enter

The other way people define success is a reward/risk ratio — your average winner divided by your average loser. Somebody who plans trades at 3:1 is saying they aim to make about three dollars for every dollar they put at risk.

The important word is before. Reward/risk is a planning number. You know your entry and you know where your exit fires, so you know what one share can cost you; you know your target, so you know what one share could make. You compare those two numbers while you can still decline the trade. Computed after the fact, it's a report card. Computed before, it's a filter, and the filter is the whole value.

Backtesting: testing your rules against a past you have already read

Now the test itself. Backtesting means running your rules against historical price data — you walk forward through old charts, take every trade your rules would have taken, and record what each one did. No real money moves. You get a data set at the end.

You don't need anything fancy for this. Old charts and something to write in will do it. The point isn't the tooling; it's a hundred honest rows.

Video coming soon

Why a system has to be tested in a rising market, a falling market, and a sideways one before you know anything about it.

This lesson explains the idea in full without it.

Test it in every kind of weather. This is the part most people skip. A system tested only across a rising market has been tested for one weather condition. Of course it made money — nearly everything made money. You've learned that your rules work when they don't need to. Run the same rules through a falling market and a sideways market before you believe anything about them. What you're after isn't a passing grade; it's a map of when this system works and when it should stay in the drawer. That map is more useful than the profit number.

Figure

Three small price charts side by side, labeled 'rising,' 'falling,' and 'sideways.' Beneath each, a blank scorecard for the same set of rules — same system, three different verdicts.

Three ways a backtest lies to you

You already know what happened. This is the big one. You are testing rules against a history you've read, on charts whose right-hand side you can see. Every judgment call — was that a real breakout, would I have held through that dip — gets made by someone who knows the answer. Hindsight bias is the habit of finding an outcome obvious once you know it. It contaminates every rule you write while looking at the chart that suggested it. The only defense is rules specific enough that a stranger could apply them to a chart you've never seen and take the same trades you'd take. If your rule needs you to be there to interpret it, you haven't tested a rule. You've tested your memory.

Tune enough knobs and you fit the noise. Change the moving average from 50 days to 47 because it scores better. Add a filter that skips Tuesdays because Tuesdays were bad. Keep going and eventually you have a beautiful curve — a set of rules shaped precisely to one stretch of history, matching its accidents rather than anything that repeats. That's overfitting, and it's seductive because it looks exactly like progress. The tell is fragility: nudge a parameter slightly and the results collapse, or the system that shone on one decade is a disaster on the next. Real edges are usually blunt. If your system only works at exactly 47 days, it doesn't work.

You are not the person in the backtest. Clicking through old charts, you take every trade calmly and hold every drawdown without flinching. On the day, with rent money in the position, you are a different animal. You'll hesitate on entries your test took instantly. You'll bail on a loser your test held to the stop. And the mechanics differ too — a backtest fills you at whatever price the chart shows, while a live order fills where it fills. Slippage — the gap between the price you expected and the price you got — is a real cost that a backtest quietly hands you for free. Every one of these makes the test look better than the life.

Two habits push back. Test enough trades that a lucky run can't carry the average — a handful of trades tells you nothing, and something on the order of a hundred starts to mean something. And take every trade your rules generate across the whole window, not the pretty ones. Skipping a period because "conditions were weird" is how you build a system that only works when nothing unusual happens, which is never.

Then, before real money, run the rules forward on charts you haven't read yet — either by waiting and recording what your rules would do, or by paper trading, placing the trades at real prices with money that isn't real. Some brokers let you do this. It removes the money but not the mechanics, and rules that fall apart the moment they meet a future you can't see fall apart there, cheaply.

Win rate is not the scoreboard

Here's where most people's intuition is wrong, and it's worth slowing down for.

Ask someone whether a trading system is good and they'll ask how often it wins. Win rate — the share of your trades that make money — feels like the answer. It isn't, and it isn't even close, because it only counts how often and never how much.

A trade that makes $400 and a trade that makes $4 both go in the win column. A trade that loses $30 and a trade that loses $3,000 both go in the loss column. The tally treats them as equals. Your account does not.

So a system can win 40% of the time and make money. And a system can win 70% of the time and lose money. Both of those are ordinary, and neither is a trick.

Expectancy: the number that actually answers the question

Expectancy puts size and frequency in the same place. It's the average profit or loss per trade once you weight the average win and the average loss by how often each happens:

Expectancy = (win rate × average win) − (loss rate × average loss)

Work it on something concrete. Everything below is invented to make the arithmetic visible — round numbers, fictional tickers, no claim that any real system behaves this way.

Say you backtested System A and it produced ten trades. (A real backtest needs far more than ten; ten is here so you can add the column up yourself.)

TradeTickerResultAmount
1XYZWin+$420
2ABCLoss−$100
3DEFLoss−$95
4GHIWin+$260
5JKLLoss−$110
6MNOLoss−$100
7PQRWin+$180
8STULoss−$90
9VWXWin+$340
10FAUNLoss−$105

Four winners, six losers. So the win rate is 40% and the loss rate is 60% — this system is wrong more often than it's right, and by the usual instinct you'd throw it away.

Add it up instead. The winners total $1,200, so the average win is $300. The losers total $600, so the average loss is $100. That's a reward/risk ratio of 3.0 — winners run three times the size of losers.

Expectancy = (0.40 × $300) − (0.60 × $100) = $120 − $60 = +$60 per trade

Sixty dollars a trade, on average, from a system that loses six times out of ten. Since each loss risks about $100, that's roughly $0.60 of profit per dollar risked. The four winners more than pay for the six losers, because each winner is worth three losers.

Now the mirror. System B wins seven times out of ten — a win rate almost anyone would take.

System ASystem B
Win rate40%70%
Average win$300$50
Average loss$100$200
Reward/risk ratio3.00.25
Expectancy+$60−$25

System B expectancy = (0.70 × $50) − (0.30 × $200) = $35 − $60 = −$25 per trade

System B is right 70% of the time and bleeds twenty-five dollars a trade. It feels wonderful to run — seven wins out of every ten, a steady drip of small victories — and the three losses quietly take back more than the seven wins brought in. This is the shape of a system where somebody moves their stop "just this once" to avoid booking a loss. The win rate stays high. The account drains.

That's the whole lesson in one table. Win rate and reward/risk are two halves of one number, and only the whole number tells you anything. A system with a lot of winners can survive a modest reward/risk — singles and the occasional double. A system with more losers than winners has to hit hard when it hits. Neither is better. Both work. What doesn't work is judging either one by half of itself.

The journal is the honest mirror

A trading journal is a record of your trades and your reasoning — entry, exit, stop, size, result, and, most importantly, why. It can be a notebook or a spreadsheet. The format is not the point.

Here's the point. Memory edits itself. Ask anyone how their trading is going and you'll get the winners in high resolution — the read that worked, the discipline they showed. The losers arrive pre-explained: bad luck, weird news, a fluke. Nobody does this on purpose. That's what makes it dangerous. Confirmation bias — noticing the evidence that says you're right — doesn't feel like bias from the inside. It feels like remembering.

The journal exists because you cannot trust your own account of why you did something. Written at the time, before the outcome exists, your reasoning can't be revised by the result. That's the whole pitch. It's the only record of your thinking that the outcome didn't get to edit.

Which means the reason column is the one that matters and the one people skip. "Bought XYZ at $40" is bookkeeping — your broker has that. "Bought XYZ at $40 because it broke resistance on heavy volume; stop at $37; getting impatient after three flat weeks" is evidence. Six months from now, that last clause is the most valuable thing on the page.

DateTickerInOutStopSizeResultWhy I took itHow I felt
ABC$28.00$26.60$26.50100 sh−$140Breakout above three-month resistance, volume above averageFine — planned trade, stop did its job
DEF$51.00$57.20$48.9060 sh+$372Flag after a strong run; entry on the breakNearly sold early on the first red day

What to ask the pages

Look for streaks. Three or more wins or losses in a row is worth pulling the charts back up. For each one, ask three questions: what was the stock doing, what was its sector doing, and what was the whole market doing? Answer that across a dozen streaks and patterns surface — maybe your rules fall apart whenever the broad market is falling, which is a genuine discovery about when to sit out.

Run the ratios. Win rate, reward/risk, expectancy, on your actual trades rather than your backtest. This is where you find out whether the system you're running is the system you tested.

Ask whether you followed the rules. Honestly. A losing month where you followed every rule and a losing month where you didn't are different problems with opposite fixes, and only the journal can tell them apart.

Check your exits. Stops set too tight get you knocked out of trades that would have worked, turning winners into small losses and quietly wrecking your average win. Stops set too loose let one trade undo six. The journal is where you see which is happening, because it holds both the level you set and what price did afterward.

Where this leaves you

That's the loop, and it's the last thing this course has to give you: define what success would mean, test the rules against history without letting yourself cheat, measure with expectancy rather than a win count, write down what you actually did, and adjust deliberately when the evidence is real rather than when the week was bad.

Be clear about what you have now. You've learned a method and a vocabulary. That's not the same as being good at this, and no course makes anyone good at this — the gap between understanding how a system is tested and running one that works is filled with time, records, and losses you paid for. Plenty of capable people who do all of this correctly still don't beat a low-fee index fund they could have bought and ignored. That's not a discouragement. It's the comparison that keeps you honest, and you now know exactly how to run it.

What this course can honestly claim to have done is take something away: the option of not knowing. You can't tell yourself anymore that a run of winners means the system works, or that a losing trade means you did something wrong, or that you'll remember why you bought it. The tools for finding out are simple, they're all in this lesson, and nothing keeps you from using them.

Key takeaways

  • Define success before you test anything. Following your rules, any single trade, and the plan itself are three different kinds of success — and a trade can lose money while the system works exactly as designed.
  • A return means nothing without a benchmark. Making 8% is good against a market that made 4% and a wasted year against one that made 20%. If your system trails its benchmark and swings harder, an index fund beat you doing nothing.
  • Backtesting is where people fool themselves. You already know what happened, so hindsight contaminates every rule; tune enough settings and you fit one history's noise instead of anything repeatable; and a system tested only in a rising market has been tested in one kind of weather. Past results guarantee nothing.
  • Win rate is not the scoreboard. Expectancy = (win rate × average win) − (loss rate × average loss). A system can win 40% and make money, or win 70% and lose it, because size counts as much as frequency. Positive expectancy describes a long run of trades, never your next one.
  • Keep a journal because your memory doesn't. It edits your winners up and explains your losers away. Reasoning written down before the outcome exists is the only record the outcome can't rewrite.

Check your understanding

Question 1 of 5

Two backtested systems, same number of trades. System A wins 35% of the time; its average win is $400 and its average loss is $100. System B wins 75% of the time; its average win is $60 and its average loss is $250. Which one made money?