How big can you trade a small edge?
A rule that wins a little on average is only useful if you trade it at the right size. Too small and it never gets anywhere. Too big and one bad run ends the account. I took two small edges and asked how the size of each trade decides what they can do, and whether they could pass a funded-account challenge.
- What I did
- I took every trade of two simple rules from my relative value study (71 on the crack spread, 107 on the Treasury butterfly) and measured each one in units of the risk taken. Then I replayed that history thousands of times in random order, risking anything from 0.25% to 25% of the account per trade, under the rules of a typical funded-account challenge.
- Why it matters to a trader
- Prop firms (which fund traders) and multi-manager funds give you capital and a drawdown limit. Inside that limit, the size of your trades decides whether a small edge is tradeable or not.
- What I found
- Each rule wins about three trades in four, but its average loss is bigger than its average win. At 1% risk per trade, no simulated account passes a challenge within a year: each rule trades only 3 to 4 times a year. The chance of passing rises with risk and levels off near half. But the chance of passing and still holding the account a year later peaks at 6% per trade (26% for the crack, 21% for the butterfly) and falls to 5% at 20%. Trading the two rules together lifts the peak to 37%.
- What I could not show
- That real trades would look like these. The trades are few (71 and 107) and the resampling cannot invent a shock worse than the worst one in the sample (-1.84 R). The two rules were flat on most days, so their near-zero correlation says little about a real crisis. And funded-account rules differ from firm to firm: mine are typical, not anyone’s contract.
The idea in plain words
Sizing a trade means choosing how much of the account you lose if you are wrong. In the rules used here, every trade has a stop: a price at which you give up. The distance from the entry to the stop is one R. A trade that makes half of that distance earned +0.5 R. A trade stopped out lost about -1 R. If you risk 2% of the account per trade, then -1 R costs 2% and +0.5 R earns 1%.
The economic story is about arithmetic, not about markets. Losses hurt more than gains help: after losing 10% you need to make 11% to get back. A rule with a small average profit and a few big losses, as these two are, leaves little room. Size it too big and a run of losers takes the account down before the small edge has time to show. Size it too small and it earns too little, too slowly, to matter.
How you trade it. Risk a fixed share of the current balance on each trade. Choose that share from three things: how fast you need to grow, how deep a fall you can accept, and how sure you are of the edge. A common answer is to take a fraction of the Kelly fraction: the risk per trade that makes an account grow fastest over the long run on the trades you have seen.
How it can lose. The Kelly fraction is estimated from a few trades, so it is uncertain. Stops can be gapped through, so a loss can be bigger than 1 R. Costs eat a small edge. And a funded account has its own rules, which end the game early: a drawdown is a fall from a peak, and a challenge asks you to make a profit target before the balance falls by a maximum drawdown or by a daily loss limit.
- Bootstrap
- To see what could have happened, I redraw the real history at random, with replacement, many times. A path is one such redraw. I do it three ways: one trade at a time (which ignores runs of bad trades), in blocks of one quarter, and in whole calendar years (which keeps a bad year together).
- Percentile
- The 5th percentile is the outcome that 5% of paths are worse than. It is a "bad luck" case.
- Time under water
- Days when the balance is below a peak it reached earlier.
Try different sizes
Loading the data
Does the account pass the challenge?
What the account looks like
Does the way history is drawn change the answer?
The same setting, three ways to resample the real trades. Drawing trades one by one ignores runs of bad trades. Blocks and whole years keep them.
| Method | Median worst fall | Worst fall, 95th percentile | Fall beyond the limit | Years under water, median | Longest spell under water, years | Pass | Pass and keep the account |
|---|
Kelly: the bet that grows the account fastest
Replay the real past
What the results say
The trades, in R
Turned into R, the two rules look alike. Most trades win, and the losers are bigger than the winners. That is the usual shape of a mean-reversion trade: many small gains, a few large losses. The second column of the table is why size matters so much.
| Crack | Butterfly | |
|---|---|---|
| Period | 7 July 2008 to 22 September 2026 | 2 January 1991 to 10 September 2026 |
| Trades | 71 | 107 |
| Trades a year | 3.9 | 3.0 |
| Winning trades | 76% | 71% |
| Average win, in R | 0.62 | 0.55 |
| Average loss, in R | -1.20 | -0.99 |
| Average trade (expectancy), in R | 0.19 | 0.11 |
| Worst trade, in R | -1.84 | -1.53 |
| Trades that lost more than 1 R | 21% | 23% |
| Skew (below 0: more big losses than big wins) | -1.10 | -0.97 |
| Kelly fraction | 21.9% | 17.2% |
| Kelly fraction, 90% range | 2% to 39% | 0% to 33% |
| Resamples with no edge at all | 3.3% | 6.6% |
The trades are few: 71 and 107. They are spread over 19 and 36 years, so a typical year has 3 to 4. That is the main limit on this whole page. On average each rule takes a trade every 3 to 4 months.
Growth against risk
The Kelly fraction is 22% for the crack and 17% for the butterfly. That sounds enormous, and it is. At 20% risk per trade, close to the crack’s Kelly fraction, the median worst fall over five years is about 51% of the account, and the balance after five years in an unlucky case (5th percentile) is 0.46 times the start. Growth per trade is zero again at about 1.8 times Kelly, and negative beyond. Half of Kelly is still 11% and 9%.
Three things push the sensible size well below Kelly. First, Kelly is estimated from few trades. Resampling whole years of trades, 90% of the estimates for the crack lie between 2% and 39%, and 7% of the butterfly resamples show no edge at all. Second, the estimate does not hold up out of sample. If the Kelly fraction is estimated only from earlier trades and then used on the next one, the butterfly ends at 0.48 times the start, with a worst fall of 73%. Half of that estimate ends at 0.88; a fixed 5% ends at 1.36. Third, a funded account stops you at a drawdown of 5% to 10%, not at the 50% that full Kelly can reach.
| Kelly estimated only from earlier trades | Crack: final balance | Crack: worst fall | Butterfly: final balance | Butterfly: worst fall |
|---|---|---|---|---|
| Full Kelly | 1.94 | 91% | 0.48 | 73% |
| Half of Kelly | 2.52 | 55% | 0.88 | 45% |
| A quarter of Kelly | 1.77 | 30% | 0.99 | 24% |
| Fixed 1% risk | 1.11 | 4% | 1.07 | 4% |
| Fixed 5% risk | 1.65 | 18% | 1.36 | 18% |
| Kelly from all trades (uses the future) | 4.27 | 70% | 1.59 | 57% |
Each row starts with a balance of 1 after the first 20 trades and risks the stated amount on each later trade, using only what was known before it. The last row cheats on purpose, to show how much the future flatters the result.
The funded-account challenge
The table uses typical rules: reach +10% before the balance falls 10% below its start or drops 5% in one day, within a year. If you pass, you keep going for another year with the same limits. Firms differ, so treat these numbers as typical of what challenges use, not as any firm’s terms. The rule values are assumptions (Idea); the simulation of them is Tested.
| Risk per trade | Crack: pass | Crack: pass and keep | Butterfly: pass | Butterfly: pass and keep | Both: pass | Both: pass and keep |
|---|---|---|---|---|---|---|
| 1% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% |
| 2% | 0.8% | 0.8% | 0.8% | 0.8% | 0.1% | 0.1% |
| 3% | 7.6% | 7.2% | 5.8% | 5.7% | 2.7% | 2.7% |
| 5% | 31.6% | 23.8% | 18.6% | 16.8% | 15.2% | 15.0% |
| 6% | 38.8% | 25.9% | 25.4% | 21.2% | 22.6% | 21.7% |
| 8% | 47.5% | 22.5% | 32.8% | 19.4% | 36.2% | 32.3% |
| 10% | 47.3% | 15.1% | 41.0% | 15.5% | 47.3% | 37.3% |
| 15% | 56.4% | 9.6% | 42.1% | 6.0% | 50.5% | 19.4% |
| 20% | 51.7% | 5.1% | 36.8% | 2.0% | 50.8% | 9.7% |
"Pass and keep" is the share of accounts that pass and then go a full year without breaking a rule. Bold: the best row of each column that the grid contains (the grid has 18 steps, so the true peak lies somewhere near). For "both", each rule risks half of the stated amount. 5,000 paths per cell, quarter blocks. Repeating one cell with five different seeds gave a pass rate between 30.8% and 34.9% for 5% risk on the crack, so a percentage here is good to about 2 points.
Three things stand out. First, small sizes do nothing. On average the crack earns 0.19 R a trade. At 1% risk that is 0.19% of the account per trade and about 0.7% a year, so a 10% target takes about 14 years. In a one-year window nobody passes. Second, the pass rate does not peak, it levels off. It climbs with risk to about 44% to 56% and stays there, because at large sizes the first one or two trades decide everything: it is a coin flip. Third, keeping the account does peak. The chance of passing and holding on for a year is highest near 6% risk per trade and collapses beyond 10%. At 20%, 90% of the crack accounts that pass fail within the next year.
Time is the hidden limit. The rules trade rarely, so the time window matters more than anything else. With a time limit of three years the best chance of passing and keeping the account is 56% for the crack (at 4% risk), 45% for the butterfly and 66% for both, against 26%, 21% and 37% with one year. With a limit of three months it is at most 6%.
| Stress case: best chance of passing and keeping the account | Crack | Butterfly | Both |
|---|---|---|---|
| Typical: target 10%, drawdown limit 10% | 26% at 6% | 21% at 6% | 37% at 10% |
| Target 8%, drawdown limit 10% | 33% at 6% | 28% at 6% | 44% at 10% |
| Target 8%, drawdown limit 5% | 24% at 6% | 19% at 6% | 26% at 8% |
| Target 10%, drawdown limit 5% | 19% at 6% | 14% at 6% | 21% at 8% |
| Drawdown measured from the highest balance | 22% at 5% | 17% at 6% | 27% at 8% |
| No daily loss limit | 35% at 10% | 28% at 10% | 40% at 12.5% |
| Time limit 3 months | 5% at 12.5% | 4% at 8% | 6% at 10% |
| Time limit 3 years | 56% at 4% | 45% at 6% | 66% at 7% |
Each cell is the best of the 18 risk levels tried, with 3,000 paths per level. All other rules are as in the typical case.
Costs and slippage
The trades already include a flat cost. To see how fragile the edge is, I took a further amount off every trade. It does not take much.
| Taken off every trade | Crack: average trade | Crack: Kelly | Crack: pass and keep at 5% | Butterfly: average trade | Butterfly: Kelly | Butterfly: pass and keep at 5% |
|---|---|---|---|---|---|---|
| Nothing | 0.19 R | 22% | 25% | 0.11 R | 17% | 14% |
| 0.05 R | 0.14 R | 17% | 20% | 0.06 R | 10% | 11% |
| 0.1 R | 0.09 R | 11% | 17% | 0.01 R | 1% | 8% |
| 0.2 R | -0.01 R | 0% | 10% | -0.09 R | 0% | 3% |
Two rules instead of one
The daily profit of the two rules has a correlation of 0.004 over their 5,039 common days (95% range -0.02 to 0.03). That looks like a free lunch. It is not quite one, for a reason worth understanding: both rules are flat on most days, and a day of zeros pulls a correlation towards zero whatever happens when both trade. On the 608 days when both held a trade, the correlation of their daily profit was -0.03. On weekly totals it is -0.09 (range -0.15 to -0.03), and on monthly totals -0.11.
Did it rise in a crisis? In the three stress years of the sample the weekly correlation was -0.16 in 2008, -0.08 in 2020 and -0.24 in 2022. It did not rise. But the two rules held a trade on the same day on only 27, 74 and 41 days of those years, and the sample only starts in 2006. Three episodes cannot rule out a jump in correlation in a crisis that looks different from them.
| Correlation of the two rules’ profit | Value | 95% range | Sample |
|---|---|---|---|
| Daily, all common days | 0.004 | -0.02 to 0.03 | 5,039 days |
| Daily, days when both hold a trade | -0.03 | n/a | 608 days |
| Weekly totals | -0.09 | -0.15 to -0.03 | 1,058 weeks |
| Monthly totals | -0.11 | n/a | 244 months |
| Weekly, 2008, 2020 and 2022 together | -0.17 | n/a | 3 years |
| Weekly, all other years together | -0.06 | n/a | 18 years |
| Months when both lost, against if independent | 10 of 219 | n/a | 12.6 expected |
The combined book (the two rules traded together). "Both" means each rule risks half the stated amount on every trade. Over the same window (7 July 2008 to 22 September 2026) the Sharpe ratio is 0.41 for the crack, 0.31 for the butterfly and 0.52 for both. The Sharpe ratio is the average daily profit divided by how much it jumps around, scaled to a year: above 1 is good, 0.3 to 0.5 is small. These Sharpe ratios use the version with the stop, measured in R and over this window, so they differ slightly from the figures on the crack and butterfly pages. Measured on raw profit instead of R, scaled to equal volatility, the book comes out at 0.44, so the two routes roughly agree. But the 95% ranges (resampling whole years) are wide: -0.04 to 0.93 for the crack, -0.14 to 0.72 for the butterfly, and 0.10 to 0.93 for both. A Sharpe of 0.3 on this many trades is inside the noise, and that is why I do not trust it as a number. What is more solid is that the book swings less: its daily volatility is 0.70 of the average of the two rules, close to the 0.71 (one over the square root of two) that two unrelated rules would give.
| Same window, 5% risk | Median worst fall, 5 years | Fall beyond 10% | Pass | Pass and keep | Real history: final balance | Real history: worst fall |
|---|---|---|---|---|---|---|
| Crack alone | 14.6% | 82% | 31.6% | 23.8% | 1.82 | 19.6% |
| Butterfly alone | 12.7% | 73% | 22.7% | 21.7% | 1.46 | 16.2% |
| Both, half the risk each | 8.3% | 30% | 15.2% | 15.0% | 1.67 | 10.7% |
| Both, equal volatility | 11.6% | 68% | 30.5% | 28.2% | 2.05 | 15.0% |
"Equal volatility" risks 0.71 of the stated amount on each rule, so that the book’s daily volatility equals the average of the two rules alone. Quarter blocks, 5,000 paths. "Real history" runs the actual days once, not a simulation.
At the same stated risk, the book has a shallower worst fall (8.3% against 14.6% and 12.7%) and a much lower chance of a fall beyond 10% (30% against 82% and 73%). That is the diversification. It also earns less, so the chance of passing is lower at the same stated risk. At equal volatility the book passes about as often as the better single rule and holds on more often (28% against 24% and 22% at 5%).
Does the way history is drawn matter?
I expected the naive draw, one trade at a time, to understate the risk, because real losses come in runs. In these trades they barely do: after a loss, the next crack trade lost 31% of the time (after a win: 22%), and for the butterfly 27% against 29%. A runs test finds nothing (z-scores -0.9 and 0.0). So the three methods give similar answers. The clearest difference is time under water, which is longer when whole years are kept together.
| 3% risk, 5 years | Method | Fall beyond 10% | Median worst fall | Years under water | Longest spell, years |
|---|---|---|---|---|---|
| Crack | Independent trades | 39% | 8.9% | 3.6 | 2.0 |
| Quarter blocks | 37% | 8.9% | 3.8 | 1.9 | |
| Whole years | 44% | 9.9% | 3.7 | 2.0 | |
| Butterfly | Independent trades | 31% | 8.1% | 3.9 | 2.3 |
| Quarter blocks | 32% | 8.4% | 4.2 | 2.4 | |
| Whole years | 31% | 8.1% | 4.3 | 2.5 | |
| Both | Independent days | 7% | 5.6% | 4.6 | 1.9 |
| Quarter blocks | 3% | 5.0% | 4.3 | 1.7 | |
| Whole years | 6% | 5.8% | 4.3 | 1.8 |
Stress replays
The demo above replays the real worst calendar year (the lowest return at the risk you choose, which can differ from the worst year in profit per barrel on the crack page) and the real worst 10 trades in a row at whatever risk you choose. For the crack the worst year was 2011: at 3% risk it would have lost 7.4%, with a fall of 9.5% from its peak. At 5% risk the same year loses 12.2%. For the butterfly it was 1998 (9.3% lost at 3%). The real history also shows how long a drawdown can last. At 1% risk the crack account spent up to 6.0 years below an earlier high, and the butterfly account 10.4 years.
In which situations could it make sense?
Each point is marked Tested when I measured it in this data, or Idea when it is how the market works in theory and I did not test it. The samples are small, so read the numbers lightly. "Make sense" here means: when could you size a small edge up, and when should you stay small.
When it could make sense
Economic context
- Idea A calm macro backdrop: the central bank on hold, inflation steady, growth ordinary. Spreads wander and come back, the small wins keep coming, and falls stay shallow, so a larger size is easier to hold.
- Idea Plenty of liquidity. Costs stay low and stops are filled close to their level, so a lost trade stays near -1 R.
- Tested Two rules from different markets, with different macro drivers (Treasury yields and oil products). Their weekly profit had a correlation of -0.09, and in 2008, 2020 and 2022 it was -0.16, -0.08 and -0.24. Together they lifted the best chance of passing and keeping the account from 26% and 21% to 37%.
Technical context
- Tested When the risk per trade is modest. With a daily loss limit of 5%, no crack account failed on that limit at 4% risk or below. At 8% it was 21% of them.
- Tested Losses did not come in runs. After a loss, the next crack trade lost 31% of the time and after a win 22%. For the butterfly the numbers were 27% and 29%. So one-at-a-time sizing does not need a rule for streaks.
- Tested A long time window. Rules that trade rarely need time. With a three-year limit instead of one year, the best chance of passing and keeping the account rose from 26% to 56% for the crack and from 37% to 66% for both rules.
When it can hurt
Economic context
- Idea A central bank shock, such as the 2022 rate hikes, that reprices a whole curve at once. The butterfly’s stop can be gapped through when every yield moves together.
- Idea An oil shock, such as March 2020, when demand collapsed. Refining margins jump and stay away from their average, so the crack rule keeps losing on the same side.
- Tested In 2008, 2020 and 2022 the crack rule averaged -0.08 R over 11 trades, against 0.24 R in other years. The butterfly did the opposite: 0.41 R over 17 trades against 0.05. With so few trades I cannot say which is typical.
- Tested The edge can fade. The crack rule averaged 0.26 R in its first 35 trades and 0.11 R in the last 36, and its Kelly fraction went from 32% to 13%. A size chosen on the first half would have been too big for the second. (The butterfly went the other way: 0.06 to 0.15.)
- Idea A firm changes its rules, for example how the daily loss is measured (on balance or on equity with open profit) or by adding a trailing drawdown. A size tuned to one set of rules can fail under another.
Technical context
- Tested Gaps. 21% of crack trades and 23% of butterfly trades lost more than 1 R, the loss the stop was meant to cap. The worst lost 1.84 R and 1.53 R. A stop is a plan, not a guarantee.
- Tested Costs. Taking 0.1 R off every trade cuts the Kelly fraction from 22% to 11% for the crack and from 17% to 1% for the butterfly. Taking 0.2 R off removes the edge.
- Tested Time under water. Even at 1% risk, the real crack account spent up to 6.0 years below an earlier high and the butterfly 10.4 years. A funded account rarely gives you that long.
- Idea A Sharpe of 0.3 is hard to trade for a funded account. To earn 10% a year at a Sharpe of 0.3, you need volatility of about 32% a year, and one bad year at one standard deviation then costs more than the whole drawdown limit.
- Idea Volatility clustering. Volatility comes in bursts, so a stop set at two standard deviations of the last year can be too tight just when a shock starts. That is one reason firms cut risk after a drawdown. In these trades a loss did not make the next loss more likely, so the reason to cut risk is survival, not a change in the edge.
- Idea Leverage and margin. At 5% risk per trade the position is large next to the account, and exchange margin can be called before the stop is reached. I did not model margin.
- Idea Event risk: the weekly US oil inventory report, an OPEC decision, a CPI print or a Fed meeting can move a spread by several stops between two closes. Daily closes hide this.
Where are we today? This page has no live reading. Position sizing is a rule for a chosen edge, not a signal: it does not say what to trade today. The trades and profits it uses are fixed files that change only when the two relative value rules are re-run, so they are not refreshed every day. The last trades in them closed on 22 September 2026 (crack) and 10 September 2026 (butterfly), and the last daily profit is from 22 September 2026. What each rule holds now is on the trade tickets page. The demo shows a warning when its data is more than 45 days old.
How much should you trust it?
The settings were not picked for being the best. The risk levels are a fixed grid, the challenge rules are round numbers, and I tried 44 settings in all (variants of the trades, of the rules and of the time limit), shown above or below. I did not pick a winner.
Take away the best. Removing one year or one trade moves the numbers by about a third at most, and the edge stays positive.
| Crack: trades removed | Trades | Average trade, R | Kelly | Pass | Pass and keep |
|---|---|---|---|---|---|
| all trades | 71 | 0.19 | 22% | 31% | 25% |
| without the best year (2019) | 63 | 0.13 | 15% | 25% | 18% |
| without the worst year (2022) | 68 | 0.22 | 26% | 34% | 29% |
| without the best trade | 70 | 0.17 | 20% | 29% | 23% |
| without the worst trade | 70 | 0.22 | 26% | 33% | 27% |
| without 2008, 2020 and 2022 | 60 | 0.24 | 30% | 30% | 28% |
| Butterfly: trades removed | Trades | Average trade, R | Kelly | Pass | Pass and keep |
|---|---|---|---|---|---|
| all trades | 107 | 0.11 | 17% | 17% | 14% |
| without the best year (2008) | 99 | 0.08 | 13% | 12% | 10% |
| without the worst year (2016) | 104 | 0.13 | 20% | 17% | 15% |
| without the best trade | 106 | 0.10 | 16% | 15% | 13% |
| without the worst trade | 106 | 0.12 | 20% | 17% | 15% |
| without 2008, 2020 and 2022 | 90 | 0.05 | 8% | 7% | 6% |
Pass rates are at 5% risk, with 3,000 paths, drawing trades one at a time.
Out of sample. Split in halves, the crack rule averaged 0.26 R in the first half and 0.11 R in the second. The butterfly averaged 0.06 and 0.15. The Kelly fraction estimated on the earlier trades only ranged from 13% to 38% (crack) and from 2% to 34% (butterfly) as trades were added. That is how unsteady the number is.
No look-ahead. R is profit divided by the stop distance, which is set on the day the trade opens. The tests rewrite the end of the data and check that nothing earlier moves.
What I did not model. Margin and financing. The bid-offer spread beyond the flat cost already in the trades. Intraday moves (a real daily loss limit counts open profit and loss during the day; I use closing balances). Trades bigger than the worst in the sample: resampling cannot invent a shock that never happened. A change in how often the rules trade. Sizing two rules that overlap in time with each other in mind. And the real rules of any firm.
What this is not
This is research. Nothing here has been traded with this sizing, and nothing here is advice. The results describe what would have happened to two rules in the past, replayed in random order. That is not a promise about the future, and it is not a plan for a real account.
Source and tests at
github.com/Nicolas8330/position-sizing.
Every figure on this page comes from one run of
examples/make_results.py on 30 September 2026. The last trade used is from
22 September 2026 and the last daily profit from 22 September 2026. The trades come from the
relative value page; see also the
crack spread and butterfly pages.