I found a very clean pattern on five-minute boards. The backtest was steadily positive and it held up for a long time. It was entirely manufactured by the labels — the ruler I judged wins and losses with was not the ruler the market settles on.
The rule was simple, and on five-minute boards the backtest showed +2 to +6 percentage points of excess. Change the entry basis and it was still there. Change the subset and it was still there. Split the sample and it was still there.
Any finding of the "it survives every cut" variety should raise one suspicion first: that it is not in the data at all, but in the measuring instrument.
To judge a five-minute board you need to know whether that window went up or down. And "up or down" depends entirely on which price you compare:
| The research ruler | The settlement ruler | |
|---|---|---|
| Price used | the exchange spot close | the oracle's final-60-second time-weighted average |
| Compared to | both against the window open | |
| Most of the time | they agree | |
| Near the boundary | they disagree | |
Judging every five-minute window with both rulers: 207 of 2,051 windows disagree — an unconditional mislabel rate of 10.1%.
Those 207 are not scattered randomly. They cluster in one family of windows: open and close sitting almost on top of each other, where any small late move rewrites the outcome. I call them knife-edge windows.
207 / 2,051 is over every window — an unconditional rate. The conditional disagreement rate within the knife-edge family is much higher, but I never computed it, so it is not quoted here.
Why state this explicitlybecause I got this denominator wrong once, describing the whole-sample rate as the knife-edge rate. Written here so it does not happen a third time.
For errors to cancel, they must be independent of your selection rule. Here the two are tightly coupled.
The windows that selection rule picks out are not evenly distributed across the sample — they concentrate heavily in one family: windows where open and close sit almost together, where a small move in the last minute rewrites the result.
Which is exactly the knife-edge family, and exactly where the two rulers most often part. That is not coincidence, but why the two coincide is internal to the strategy and not opened here — the argument does not need it.
① the selection concentrates in knife-edge windows → ② knife-edge windows are where the rulers most often disagree → ③ and in that family the disagreement is not half-and-half — it leans one way → ④ so the rule systematically stands on the side the bad label favours.
The error is rectified by the selection rule. Noise that should have been symmetric is filtered into a one-way bias, and it condenses into a stable-looking +2 to +6pp that is not there.
Splitting the 207 mislabelled windows by direction: 125 favour me, 82 favour the reverse, a net lean of 43 windows.
43 ÷ 2,051 = 2.1% of windows systematically judged backwards. And one window judged backwards is a ±100 percentage point swing in P&L (a win becomes a loss or the reverse). Multiply:
2.1% × 100pp ≈ 2.1pp of inflation
And the backtest's observed excess was +2 to +6pp. The lower bound matches exactly — the size of the "finding" equals the size of the instrument bias. That is effectively the verdict.
In other words: how often the rulers disagree is not the point. Whether the disagreement has a direction is. 10% of symmetric disagreement is harmless; 2% of directional disagreement is enough to fabricate a finding.
A comparison between two rulers cannot say which is right. So I took the disputed windows to a third party that cannot be argued with: the settlement written on-chain.
| Batch | Result |
|---|---|
| All 207 disputed windows on the short tenor | 207 : 0 |
| A second batch of 31 disputed windows from the other tenors' audit | 31 : 0 |
238 cases, not one exception. The chain sides with the settlement ruler every time. So this is not "two defensible views" — one of them is simply wrong, and it was mine.
The immediate question after finding an instrument fault is how far the contamination reaches. So I audited all four tenors at identical granularity:
| Tenor | Disagreement | Directional? | Verdict |
|---|---|---|---|
| 5 minutes | 10.1% | yes, 125:82 | finding void |
| 15 minutes | 3.1% | symmetric | unaffected |
| 1 hour | 0% | — | unaffected |
| 4 hours | 0.4% | — | unaffected |
The mechanism explains the gradient: the final 60 seconds is 20% of a five-minute window and 0.4% of a four-hour one. The shorter the tenor, the more likely "an average of the last stretch" and "the close" disagree.
Only one died; the rest did not need redoing. Bounding the contamination precisely is as important as finding the fault — otherwise you either discard sound work or keep unsound work.
The principle adopted, in one line: the ruler was wrong — repair it, then re-audit every line it ever judged with the correct one.
So the repair was not just "stop using the bad label". It was: build a settlement-truth label table, collect it on an ongoing basis, backfill the history — 7,330 windows — and re-score.
Rescoring the full history against settlement truth collapsed the family's excess to approximately zero, and several model-free baselines changed sign outright:
| Old ruler | Correct ruler | |
|---|---|---|
| A model-free baseline | +2.72pp | −1.23pp |
| Another baseline | −3.82pp | +0.13pp |
A baseline that contains no model at all should reflect only the market. If changing the basis flips it from clearly positive to clearly negative, the problem cannot be in the strategy — it is in the measurement.
This is the most underrated use of null models: they are simultaneously a control for the strategy and a calibration check on the instrument. A strategy moving is attributable to the strategy; a null model moving can only be the ruler.
Then a full retrain against the corrected target — same features, same protocol, only the label changed — returned four negative cells out of four, with contemporaneous nulls near zero. Both halves negative. Verdict: dead on the corrected ruler.
That combination is what makes the kill clean: it is not that a broken instrument produced ugly readings — the strategy is dead on the repaired instrument too.
| Action | Why |
|---|---|
| Verify the ruler before the strategy | settle one market by hand and check your label agrees |
| Score against settlement, not a proxy | the market pays on settlement; anything else is a proxy |
| Distrust "it survives every cut" | robustness across cuts is also what an instrument bias looks like |
| Watch the nulls for sign changes | a model-free baseline moving indicts the measurement |
The first is nearly free and I did not do it: take one settled market, work out its outcome by hand from the stated terms, and check your pipeline agrees. One market, ten minutes.
A pattern that survives every cut is not necessarily robust — it may simply mean you are measuring the instrument rather than the market. Verify the ruler before you trust the reading, and remember: disagreement rate is not the danger, direction is.