Case study · RSI-2 reversion · BTC + SOL · 2022-08-11 to 2026-07-05

A strategy that lost $72,983. We tried to save it.

An RSI-2 mean-reversion system: 11,888 real fills across BTC and SOL over 3.9 years, winning 54.6% of its trades and still finishing −$72,983. We ran it through the same audit every NoxarQuant journal gets. The audit cut the losses hard. It also answered a question most tools never ask: was this worth trading at all?

Take every signal
−$19,393
3,567 trades in the held-back window
Obey the frozen verdicts
−$4,965
same window, same fills, NDC sat out
Losses the audit removed
$14,428
74% less bleed — and still red. That is the finding.
84%

The verdicts predicted forward, not backward.

16 of the 19 conditions flagged NDC kept losing across 3,567 trades that did not exist when the verdict was written. That is the claim everything below rests on: not that the audit can describe a past, but that it survives contact with data it has never seen.

Hard wall
Verdicts frozen on trades to 2025-05-16. Everything tested after is a market the classifier never saw.
Frozen, not refitted
Fixed on the first 8,321 trades, then never recomputed.
Lookup only
Each of 3,567 later trades inherits its group’s old verdict. No peeking forward.
Real fills
Executed trades with fees. No synthetic entries, no survivorship pick.
01
Splitting the book by time, not by luck
The first 8,321 trades are the learning half; the last 3,567 are held back as a live test. On the learning half alone, the audit grouped every trade by the conditions it was taken in — strategy | session | volatility | RSI | trend — and issued a verdict per group. Those verdicts were then locked.
0
CORE
proven edge
1
PROMISING
building a case
83
FORMING
no verdict yet
19
NDC
no demonstrable consistency

Read the first cell again. Across 8,321 trades and nearly three years, not one set of conditions cleared the bar for a proven edge at any point, on any slice of the data. The proven-edge column is empty and stays empty.

02
What the verdicts were worth in money
Three ways to trade the held-back window: take everything, take everything except the NDC conditions, or take only the NDC conditions. Identical trades and fills throughout. The only variable is whether you listened to a label written years earlier.
Cumulative P&L on held-back trades — 2025-05-16 to 2026-07-05
every point is a real trade, in sequence
0k-5k-10k-15k-20k
Skip NDC  −$4,965Everything  −$19,393NDC only  −$14,428
03
Every way you could have traded it
The obvious follow-up: if the audit sorts conditions by quality, what if you had traded only the good ones? Here is every filter applied to the same held-back window. The answer to the CORE question is the one that matters, and it is not a number.
Strategy in the test windowtradesnet$/tradewin%
Take every signalno filter3,567−$19,393-5.4452.4%
Skip NDCobey the frozen verdicts1,691−$4,965-2.9454.9%
CORE onlytrade only proven-edge conditions0no trades
PROMISING onlyconditions still building a case3+$29+9.5933.3%
FORMING onlyno verdict yet1,687−$4,990-2.9655.0%
NDC onlywhat obeying the audit removes1,876−$14,428-7.6950.1%

The CORE row is empty because no CORE condition exists. Nothing in 11,888 trades ever cleared the bar for a proven edge, so that row has no trades to show. It is not bad luck in one window; it is the shape of the whole book. PROMISING fired 3 trades in thirteen months — too few to conclude anything from, and we are not going to pretend otherwise. Every filter that produced a meaningful number produced a negative one.

04
Every flagged condition, checked one by one
The 84% headline could still hide a single lucky group, so here is the full working: every NDC condition with enough fresh trades, its average before the wall and after it. Sixteen kept bleeding. Three did not, and they are shown in place rather than dropped.
84%
16 of 19 conditions flagged NDC stayed net-negative on trades that did not exist when the verdict was written.
Condition flagged NDClearn $/ttest $/tn
ASIA|HIGH|30_70|FLAT-12.94-18.38218 → 97held
NEW_YORK_PM|HIGH|30_70|FLAT-19.52-18.07187 → 88held
NEW_YORK_AM|MEDIUM|30_70|FLAT-25.91-16.72266 → 100held
LONDON|HIGH|30_70|FLAT-22.41-15.0799 → 38held
NEW_YORK_PM|LOW|30_70|FLAT-16.41-14.50148 → 48held
NEW_YORK_AM|MEDIUM|30_70|DOWN-19.28-11.5397 → 63held
AFTER_HOURS|MEDIUM|30_70|FLAT-10.84-11.16408 → 190held
AFTER_HOURS|HIGH|30_70|FLAT-12.45-10.38177 → 70held
ASIA|LOW|30_70|UP-5.02-9.61153 → 38held
ASIA|LOW|30_70|FLAT-4.79-8.47353 → 143held
ASIA|MEDIUM|30_70|FLAT-12.21-7.12519 → 217held
NEW_YORK_PM|MEDIUM|30_70|DOWN-7.79-5.26141 → 74held
LONDON|LOW|30_70|FLAT-5.17-4.81494 → 190held
AFTER_HOURS|LOW|30_70|DOWN-6.30-3.50116 → 56held
LONDON|MEDIUM|30_70|FLAT-6.14-2.66486 → 215held
NEW_YORK_PM|MEDIUM|30_70|FLAT-11.50-0.50304 → 132held
ASIA|LOW|30_70|DOWN-21.26+3.65131 → 55flipped
NEW_YORK_PM|MEDIUM|30_70|UP-12.92+7.52167 → 49flipped
NEW_YORK_PM|LOW|30_70|UP-17.35+11.1450 → 13flipped
05
The honest shape of it
Every group with enough data, plotted by its learning-half average against its test-half average. Red is NDC. The mass sits low and left, and the top-right quadrant — conditions that made money before and kept making it — is effectively empty. Three flagged groups did drift back above water, and they are shown here rather than dropped.
Every condition group: past vs future, $/trade
dot size = trades in the test window · shaded band = stayed negative
future avg $/trade ↑ past avg $/trade →in -1.9 / out 7.96 (n96)in -5.02 / out -9.61 (n38)in -4.1 / out -13.19 (n78)in -2.8 / out -2.38 (n72)in -5.17 / out -4.81 (n190)in -2.1 / out -0.72 (n105)in 4.25 / out -43.84 (n18)in -16.41 / out -14.5 (n48)in -12.21 / out -7.12 (n217)in -6.14 / out -2.66 (n215)in -25.91 / out -16.72 (n100)in -4.53 / out -7.43 (n24)in 0.68 / out -1.02 (n44)in -11.5 / out -0.5 (n132)in -10.04 / out 4.24 (n87)in 1.7 / out -3.57 (n98)in -1.6 / out -11.19 (n95)in -22.41 / out -15.07 (n38)in -7.79 / out -5.26 (n74)in -6.3 / out -3.5 (n56)in -2.23 / out -6.36 (n143)in -12.94 / out -18.38 (n97)in -12.45 / out -10.38 (n70)in 5.94 / out 10.53 (n33)in -12.92 / out 7.52 (n49)in -10.84 / out -11.16 (n190)in -19.28 / out -11.53 (n63)in -6.62 / out -7.25 (n43)in -4.63 / out 7.22 (n57)in -4.79 / out -8.47 (n143)in -17.35 / out 11.14 (n13)in -0.48 / out -7.6 (n56)in -6.47 / out 1.22 (n16)in 0.99 / out -8.29 (n104)in -19.52 / out -18.07 (n88)in -1.18 / out -20.15 (n45)in 4.39 / out 15.9 (n31)in 3.32 / out -10.85 (n50)in 6.83 / out -10.15 (n67)in -21.26 / out 3.65 (n55)in -5.71 / out -13.61 (n38)in -9.19 / out 1.7 (n27)in 5.97 / out -30.4 (n31)in -0.21 / out 19.13 (n39)in 3.99 / out 2.46 (n56)
flagged NDCeverything else
06
The autopsy: where the money actually went
One more question the raw fills can answer: was the strategy even wrong? Reconstruct every trade's gross price move — exit minus entry, times size — and compare it with what actually landed after costs.
Gross price P&L
+$22,121
the strategy made money at the price level — 65.6% of trades moved in its favour
Total costs
−$95,104
modelled at $8.00 per trade flat — 0.04% per side on this book’s fixed $10k notional
Net
−$72,983
what the account actually saw

1,299 trades — 10.9% of the book — called the direction correctly and still lost money. Average loss on those: $3.58, fee-sized, not thesis-sized. Longs lost $36,213 and shorts lost $36,771 — near-identical, because costs do not care which way you traded. And the equity curve declines almost perfectly linearly across four years: no blow-up, no regime break, just a constant per-trade drag.

Keep both win rates in view: 65.6% gross, 54.6% net. The eleven-point gap between them is the cost drag expressed as a rate — the 1,299 right-but-lost trades, seen from the other side.

So the diagnosis is sharper than “no edge.” This strategy has a real gross edge of about +$1.86 per trade — and pays a modelled $8.00 per trade to express it. The edge is not absent; it is 4.3× too small to pay its own costs. And the conclusion is robust to the cost assumption: even at half these costs the edge is still 2.2× too small. Break-even sits at $1.86 per trade — this book turns profitable at 77% lower execution costs, or 4.3× more gross edge, and no amount of condition-filtering changes that arithmetic.

Where this evidence runs out

The 84% above is a real out-of-sample result, and it is also a small one. Three things a careful reader should hold against it, stated here rather than left to be found.

Small n, and not independent

19 conditions is not many, and they are not 19 separate experiments. They partition one strategy on two instruments over one period, so they share the same underlying market and the same execution. Treat this as one strong observation, not 19 corroborating ones.

The misses share a shape

Of the three that reverted, two are NEW_YORK_PM in an UP trend. Of the three NDC conditions in an UP trend, two flipped — while 12 of the 16 that held sit in a FLAT trend. The failures are not evenly scattered; NDC looks weakest in trending conditions.

One dimension did nothing

All 19 flagged conditions fall in the same RSI band, 30_70. On this strategy the RSI dimension separated nothing, so four of the five conditions carried the result. A different strategy would likely lean on them differently.

And the honest gap: this shows the audit cutting losses on a strategy that was already doomed. It does not yet show the reverse — that it confirms a real edge on a system that works. Until a profitable book is put through the same wall, a fair reader can ask whether NDC is detecting a broken strategy or simply flagging volatile conditions. That case study is the next one to run, and it should be published whichever way it lands.

Why the answer was to stop

We could have kept going. Drop two more conditions. Tighten a threshold. Add a filter. Keep trimming until the curve finally clears zero and the whole thing looks like an edge.

That move has a name: overfitting. Every filter added to make a past look profitable is fitted to noise — the specific, never-repeating accidents of that exact history. It reads as a strategy. It dies the moment it meets a market it has not memorised. And it is dangerous precisely because it does not feel like a mistake. It feels like work.

The audit already told us what tinkering would have hidden. Cutting NDC removed $14,428 of real bleed, 74% of the damage, and the 84% persistence rate says those verdicts were predictive rather than descriptive. But zero conditions ever earned a proven-edge verdict, the gross edge is 4.3× smaller than the per-trade cost, and the best honest version still finishes red. There was no edge underneath. There was a losing system in which some conditions lost faster than others.

The most valuable thing an audit can tell you is when to stop. This one said: do not chase it.
Audit your own strategy
upload your fills · freeze the verdicts · find out if it is worth chasing
Method and limits. 11,888 executed trades, split 70/30 by trade count, not by date, so both halves carry comparable weight. Verdicts were computed on the learning half only and never recomputed; test-window trades inherit their group’s verdict by lookup. Persistence is measured on NDC groups with at least 10 test-window trades. Every trade carries a fully resolved condition set: no trade in this study sits in a bucket with a missing dimension. Costs are a model, not reconstructed exchange fills: $8.00 per trade flat, equivalent to 0.04% per side at this book’s fixed $10,000 notional — taker-fee territory; the 4.3× conclusion survives halving it. Verdicts were computed under the classification gates current as of 2026-07-30 (SQN and sample floors); a scoring-display update shipped the same day changed how small samples are scored, not how conditions are classified, so every number on this page reproduces in the product today. This is one strategy on two instruments and is not evidence about any other. Past results never guarantee future ones.