Recovery factor, Sortino and time underwater: what the durability numbers mean
Every finished walk-forward, robustness and portfolio run carries a panel called Living with it. The stat cards above it tell you what the strategy totalled. These tell you what holding it would have been like, which is a different question and often a more important one.
Net profit and win rate say nothing about whether you'd still have been trading the system by the end.
Time underwater
Two figures: the longest unbroken stretch below a previous equity high, and the share of the whole period spent below one.
Depth is only half of a drawdown. A 15% fall recovered in three weeks and one that grinds on for two years are completely different experiences, and the second is what makes people abandon a system that was still working. Max drawdown alone can't tell them apart.
A portfolio we tested spent 97.6% of its life below a previous high, including one unbroken stretch of 491 trading days — close to two years. Its headline figures looked unremarkable rather than alarming: $44,006 profit, profit factor 1.02. Nobody reading those two numbers would have guessed at the two years.
Flat days at a peak don't count as underwater, so a strategy that simply trades rarely won't score badly for inactivity. This is genuine time below a prior high.
Recovery factor
Net profit ÷ worst drawdown. What did you earn for the worst moment you had to sit through?
| Value | Reading |
|---|---|
| Below 1 | You made less than your worst drawdown |
| 1 | You made roughly what you risked at the worst point |
| 2–3 | Where most real portfolios of retail strategies land |
| Above 3 | Genuinely strong |
| Negative | It lost money; the ratio isn't telling you anything else |
The portfolio above scores 0.29 — risking $3.45 to make $1.
Two ways this number flatters a strategy.
It rewards long tests. Profit accumulates across the whole period, while the worst drawdown is a single event. Ten years of data grows the numerator and often leaves the denominator alone, so a ten-year recovery factor and a two-year one aren't comparable. Always read it next to the test length.
It uses the worst drawdown that happened to occur, and the future one is usually deeper. This is the bigger problem, and the Robustness stage exists to measure it. One strategy we tested scored 1.02 on its backtest drawdown of $8,407. Reshuffling the same trades 2,000 ways:
| Drawdown used | Recovery factor |
|---|---|
| The backtest's $8,407 | 1.02 |
| Median resample, $11,056 | 0.78 |
| Bad-luck 95th percentile, $21,401 | 0.40 |
Same strategy, same profit. The backtest simply got a lucky ordering.
Why not Calmar? Calmar is the better-known cousin — CAGR ÷ max drawdown — but it needs a stated account size to turn profit into a percentage, and a futures P&L stream doesn't have one. You can trade the same strategy in a $25,000 or a $250,000 account. Recovery factor is capital-free, so it works on any run. Set your capital in a portfolio standard and the drawdown-percentage rule gives you the Calmar-style view alongside it.
Sortino
Return per unit of downside volatility, annualised from daily P&L.
Sharpe treats a huge winning day as risk, which is not how anyone experiences it. Sortino counts only the downside, so a strategy with occasional large gains isn't penalised for having them.
| Sortino | Reading |
|---|---|
| Below 1 | Poor — return is small next to downside risk |
| 1–2 | Acceptable |
| 2–3 | Very good |
| Above 3 | Excellent |
The portfolio above scores 0.19. In practical terms that means its yearly profit is about a fifth the size of its bad-day variability: roughly $20,000 a year of profit against an annualised downside deviation near $105,000. For reference, simply holding a broad equity index has historically produced somewhere around 0.5–1.0 on this measure.
The consequence worth understanding. There's a rule of thumb that confirming an edge is real rather than luck takes about (2 ÷ Sharpe)² years of live data. That portfolio's daily Sharpe is 0.11, which works out at over three centuries. The honest reading isn't "this is a weak strategy" — it's you could trade this for a lifetime and never accumulate enough evidence to know whether it works. Every good year and every bad year would be indistinguishable from noise.
Comparing Sortino against Sharpe is informative in itself. Much higher Sortino means the volatility is mostly upside, which is the good kind. When both sit far below 1, the shape hardly matters — there isn't enough return to justify either.
Two Sharpes, and why
You'll see Sharpe in the stat cards and Sharpe (daily) in the durability panel. They often disagree, and the daily one is the one to read.
The original divides average profit per trade by the standard deviation per trade, then annualises by √252 — which is only correct if the strategy happens to average one trade a day. A scalper's is inflated; a swing trader's is understated. It's kept because older runs are stored with it, and changing what a saved number means is worse than leaving it.
Sharpe (daily) and Sortino are both built from a daily equity curve with flat days included. That makes them comparable across holding periods: a system trading five times a day and one trading five times a month can sit on the same page and be judged on the same scale.
Using them
These are inputs to a decision, not a verdict. Rough guidance:
- Recovery factor below 1 — you're risking more than you're making. Rarely worth deploying whatever the profit figure says.
- Longest time underwater — compare it honestly against how long you'd actually keep funding something that isn't working. Most people overestimate this.
- Sortino below 1 — expect not to be able to tell whether it's working for a very long time.
- All three together, against limits you set — that's a portfolio standard, which turns these readings into a pass/warn/fail you define.
Filtering the Optimize leaderboard on them (1.0.57)
Two of these can now keep settings off the Optimize leaderboard before you ever look at them. Both are in the Optimize box on the strategy page:
- Max drawdown, % of net profit (default 50). A setting whose worst drawdown was larger than this share of its net profit is not ranked, and neither is one that lost money. It is the recovery factor turned around: 50 % means a recovery factor of at least 2 — the setting made at least twice its worst drawdown. Profit keeps adding up over a longer test while the worst drawdown usually does not, so the same strategy scores about half as much over two years as over one; pick the share with the length of your test in mind. 0 turns it off.
- Longest time underwater, trading days (off by default). A setting that once went longer than this many trading days without a new equity high is not ranked. About 21 trading days make a month.
The leaderboard shows all three numbers for every row — drawdown as a % of net, the share of days under water, and the longest stretch under water — so you can see how close the survivors came to the line.
Why the share of days under water is shown but not used as a filter: nearly every profitable setting spends most of its days below its last high — 66 to 98 % across a range of real optimizations — because a strategy is only out of the water on the day it sets a new high. And since flat days at a peak don't count, a strategy that trades rarely scores well on it simply by trading less. The longest stretch is the one that tells you what holding the strategy would have felt like.
The Max drawdown $ limit stays beside them: a prop-firm evaluation fails at a fixed dollar drawdown however large the profit, so the two limits answer different questions and can be used together.
These limits apply to Optimize (and to each timeframe of a comparison), not to walk-forward learning windows or audits, where a share of four months' profit would mean something stricter than the same share of two years'.
And remember every one of them describes the past, on the specific ordering of history that happened. Cost sensitivity and the Robustness stage exist to ask how much of it survives contact with a different one.