LZLZL/Prediction markets/Research method
FREEBUILD IT F · MethodFlagship

Gates and the single read

2026-08-21 · why "just one more look" destroys an experiment

Once the criteria are written, two things decide whether they matter: how the bar was set, and how many times you let yourself look. The second is the counter-intuitive one — repeated checking manufactures significance by itself, even if you never change a number.

1Four gates in parallel, none redundant

GateBlocks
① Magnitude: gross excess ≥ cost line"there is an edge but it does not pay costs"
② Significance: z ≥ 2.0"big enough, but it is luck"
③ Beat the nulls: all three"it is actually the market, not the strategy"
④ Stability: halves agree in sign"one stretch is carrying everything"

Each blocks a different false positive and none substitutes for another. My own final read is the cleanest illustration: ③ and ④ PASS, ① and ② FAIL.

⚠ Reading only ③ and ④ gives the opposite conclusion

"It beat three null models and both halves agree" sounds like something real. But the magnitude was +1.56pp against a 2.25pp cost line.

Beating random betting is not beating costs. A strategy at +1.56pp of excess with 2.25pp of cost is not "slightly profitable" — it is losing with certainty on every single trade.

"It beat the null" is the sentence people use to convince themselves. It establishes only that the strategy is not pure noise; it says nothing about covering the fee.

2The single read: why once

The data is sitting there. Why not look more often?

because looking is a test look twenty times and you have run twenty tests — each with its own chance of passing by luck
PracticeWhat it actually is
Accumulate to n, read onceone test
Check daily, stop when goodmany tests, reporting only the best
Not good? "give it more time"the same, and it can never fail

That last row is the crux: an experiment that permits "give it more time" has no failure state. It has only "succeeded" and "not yet" — which is not an experiment, it is waiting.

3You may look; you may not act

AllowedForbidden
watch the current numberchange criteria because of it
watch trends, check data qualitystop early or extend because of it
stop for instrument failurestop because the number is ugly

The third row's distinction matters: stopping because the instrument is broken and stopping because the result is bad are different acts. The first is mandatory — the data is dirty and continuing is pointless. The second is cheating. The test: "if the number were good, would I still stop for this reason?"

4When early reading is legitimate

One case: when the remaining sample cannot mathematically change the outcome.

My final read worked exactly that way. Trigger was n=2,500; the actual read was n=2,468 (98.7%). The justification recorded in the verdict:

to lift excess from +1.56 to +2.25, the remaining 32 windows would each need about +55pp which is at the theoretical maximum for a single window — impossible

An earlier kill used the same arithmetic: at n≈460 with excess −1.8pp against a +4pp gate, the remaining ~40 windows would each need +71pp, above the single-window ceiling.

This has a strict condition attached

It may only be used to kill early, never to declare success early.

Because "can no longer pass" is a strictly computable fact, while "is now certain to pass" never is — later samples can always deteriorate.

The asymmetry is deliberate: killing early saves continued burn, declaring early saves testing you should still be doing. One is a gain; the other is a risk.

5Setting the sample size

Threshold and sample size are set together, never apart:

z = excess ÷ standard error, and standard error ∝ 1/√n larger n makes the same excess more significant

In practice you invert it: the size of edge you want to detect determines the sample you need. Mine estimated a standard error of about 1.0pp at n≈2,500, so z ≥ 2.0 corresponds roughly to "excess of at least 2pp". That lines up with the 2.25pp threshold computed from costs — the two gates cross-check each other rather than being independently invented.

A practical corollaryif the sample you need far exceeds what you can ever collect, the line should be abandoned before it starts — you will never obtain data capable of a conclusion. That judgement is free, and almost nobody makes it.

6One line to keep

Gates must be computed, run in parallel, and frozen in advance; the data is read once. An experiment that permits "give it more time" has no failure state — and something that cannot fail cannot demonstrate success either.

EvidenceCheck it yourself

Four gates in parallel any one failing is a failure; my final read passed ③④ and failed ①②.
Final read n=2,468 (98.7% of the 2,500 trigger); excess +1.56pp against a 2.25pp gate; z=1.62 against 2.0.
Early-read arithmetic the remaining 32 windows would each need about +55pp — at the single-window theoretical ceiling, so impossible.
An earlier instance at n≈460, excess −1.8pp against a +4pp gate; remaining ~40 windows would need +71pp each, above the ceiling.
Gates cross-checking standard error about 1.0pp at n≈2,500 makes z ≥ 2.0 correspond to roughly 2pp, consistent with the 2.25pp computed from costs.
Checked 2026-08.

NextWhere to go

F · METHOD
Preregistration: freeze the bar first
F · METHOD
Kill rules and pre-claimed failure modes
F · METHOD
Null models
E · UP/DOWN
The line these gates killed
This is an educational and research record. It is not investment advice, promises no returns, and offers no personalised trading recommendations. Rules and API behaviour are per the official documentation; this page states when it was checked and both can change without notice. Prediction markets are restricted or unavailable in some jurisdictions — confirm your own before taking part.