LZLZL/Prediction markets/Research method
FREEBUILD IT F · MethodFlagship

Your backtest won,
but the ruler was wrong

2026-08-21 · a real instrument failure, from finding to artefact to re-audit

I found a very clean pattern on five-minute boards. The backtest was steadily positive and it held up for a long time. It was entirely manufactured by the labels — the ruler I judged wins and losses with was not the ruler the market settles on.

1The symptom: a finding that was too good

The rule was simple, and on five-minute boards the backtest showed +2 to +6 percentage points of excess. Change the entry basis and it was still there. Change the subset and it was still there. Split the sample and it was still there.

Any finding of the "it survives every cut" variety should raise one suspicion first: that it is not in the data at all, but in the measuring instrument.

2Two rulers, measuring different things

To judge a five-minute board you need to know whether that window went up or down. And "up or down" depends entirely on which price you compare:

The research rulerThe settlement ruler
Price usedthe exchange spot closethe oracle's final-60-second time-weighted average
Compared toboth against the window open
Most of the timethey agree
Near the boundarythey disagree

Judging every five-minute window with both rulers: 207 of 2,051 windows disagree — an unconditional mislabel rate of 10.1%.

Those 207 are not scattered randomly. They cluster in one family of windows: open and close sitting almost on top of each other, where any small late move rewrites the outcome. I call them knife-edge windows.

⚠ On the denominator

207 / 2,051 is over every window — an unconditional rate. The conditional disagreement rate within the knife-edge family is much higher, but I never computed it, so it is not quoted here.

Why state this explicitlybecause I got this denominator wrong once, describing the whole-sample rate as the knife-edge rate. Written here so it does not happen a third time.

3Why the errors do not cancel out

For errors to cancel, they must be independent of your selection rule. Here the two are tightly coupled.

The windows that selection rule picks out are not evenly distributed across the sample — they concentrate heavily in one family: windows where open and close sit almost together, where a small move in the last minute rewrites the result.

Which is exactly the knife-edge family, and exactly where the two rulers most often part. That is not coincidence, but why the two coincide is internal to the strategy and not opened here — the argument does not need it.

So a loop forms

① the selection concentrates in knife-edge windows → ② knife-edge windows are where the rulers most often disagree → ③ and in that family the disagreement is not half-and-half — it leans one way → ④ so the rule systematically stands on the side the bad label favours.

The error is rectified by the selection rule. Noise that should have been symmetric is filtered into a one-way bias, and it condenses into a stable-looking +2 to +6pp that is not there.

How large is the lean? Computable directly

Splitting the 207 mislabelled windows by direction: 125 favour me, 82 favour the reverse, a net lean of 43 windows.

43 ÷ 2,051 = 2.1% of windows systematically judged backwards. And one window judged backwards is a ±100 percentage point swing in P&L (a win becomes a loss or the reverse). Multiply:

2.1% × 100pp ≈ 2.1pp of inflation

And the backtest's observed excess was +2 to +6pp. The lower bound matches exactly — the size of the "finding" equals the size of the instrument bias. That is effectively the verdict.

In other words: how often the rulers disagree is not the point. Whether the disagreement has a direction is. 10% of symmetric disagreement is harmless; 2% of directional disagreement is enough to fabricate a finding.

4Ask the chain: 207 to 0

A comparison between two rulers cannot say which is right. So I took the disputed windows to a third party that cannot be argued with: the settlement written on-chain.

BatchResult
All 207 disputed windows on the short tenor207 : 0
A second batch of 31 disputed windows from the other tenors' audit31 : 0

238 cases, not one exception. The chain sides with the settlement ruler every time. So this is not "two defensible views" — one of them is simply wrong, and it was mine.

5What about the other tenors

The immediate question after finding an instrument fault is how far the contamination reaches. So I audited all four tenors at identical granularity:

TenorDisagreementDirectional?Verdict
5 minutes10.1%yes, 125:82finding void
15 minutes3.1%symmetricunaffected
1 hour0%unaffected
4 hours0.4%unaffected

The mechanism explains the gradient: the final 60 seconds is 20% of a five-minute window and 0.4% of a four-hour one. The shorter the tenor, the more likely "an average of the last stretch" and "the close" disagree.

Only one died; the rest did not need redoing. Bounding the contamination precisely is as important as finding the fault — otherwise you either discard sound work or keep unsound work.

6The ruling: fix the instrument, then re-audit everything it judged

The principle adopted, in one line: the ruler was wrong — repair it, then re-audit every line it ever judged with the correct one.

So the repair was not just "stop using the bad label". It was: build a settlement-truth label table, collect it on an ongoing basis, backfill the history — 7,330 windows — and re-score.

7After the repair

Rescoring the full history against settlement truth collapsed the family's excess to approximately zero, and several model-free baselines changed sign outright:

Old rulerCorrect ruler
A model-free baseline+2.72pp−1.23pp
Another baseline−3.82pp+0.13pp
⚠ A baseline changing sign means the ruler is broken, not the strategy

A baseline that contains no model at all should reflect only the market. If changing the basis flips it from clearly positive to clearly negative, the problem cannot be in the strategy — it is in the measurement.

This is the most underrated use of null models: they are simultaneously a control for the strategy and a calibration check on the instrument. A strategy moving is attributable to the strategy; a null model moving can only be the ruler.

Then a full retrain against the corrected target — same features, same protocol, only the label changed — returned four negative cells out of four, with contemporaneous nulls near zero. Both halves negative. Verdict: dead on the corrected ruler.

That combination is what makes the kill clean: it is not that a broken instrument produced ugly readings — the strategy is dead on the repaired instrument too.

8Four things that prevent this

ActionWhy
Verify the ruler before the strategysettle one market by hand and check your label agrees
Score against settlement, not a proxythe market pays on settlement; anything else is a proxy
Distrust "it survives every cut"robustness across cuts is also what an instrument bias looks like
Watch the nulls for sign changesa model-free baseline moving indicts the measurement

The first is nearly free and I did not do it: take one settled market, work out its outcome by hand from the stated terms, and check your pipeline agrees. One market, ten minutes.

9One line to keep

A pattern that survives every cut is not necessarily robust — it may simply mean you are measuring the instrument rather than the market. Verify the ruler before you trust the reading, and remember: disagreement rate is not the danger, direction is.

EvidenceCheck it yourself

Two rulers research used exchange spot close vs open; the market settles on the oracle's final-60-second time-weighted average vs open.
The 10.1% all 2,051 windows judged under both; 207 disagreed. An unconditional whole-sample rate — not the knife-edge subset's conditional rate, which was never computed.
Direction 125 favouring me against 82 the other way; net 43 windows, 2.1% of the sample, which at ±100pp per window is about 2.1pp of inflation — matching the lower end of the observed finding.
On-chain review 207:0 on the disputed windows and 31:0 on a second batch. 238 cases, no exceptions.
Contamination bounded four tenors audited at identical granularity: 5m 10.1% directional (void) · 15m 3.1% symmetric · 1h 0% · 4h 0.4%, tracking the final-segment share of the window.
The repair a settlement-label table, ongoing collection, and 7,330 windows backfilled.
Re-audit same features and protocol, target swapped: four cells of four negative, nulls near zero, both halves negative.
Checked 2026-08.

NextWhere to go

F · METHOD
Null models: the baseline that changed sign
E · UP/DOWN
The real-money line judged on this ruler
A · MECHANICS
Resolution sources differ by category
F · METHOD
Replays must state their bias
This is an educational and research record. It is not investment advice, promises no returns, and offers no personalised trading recommendations. Rules and API behaviour are per the official documentation; this page states when it was checked and both can change without notice. Prediction markets are restricted or unavailable in some jurisdictions — confirm your own before taking part.