LZLZL/Prediction markets/Research method
FREEBUILD IT F · MethodFlagship

Preregistration
freeze the bar before you look

2026-08-21 · including one of mine that admits it is not clean

Preregistration does one thing: write down what counts as success before you see the result. It sounds like paperwork. It is the only tool that stops one particular class of error — the class whose defining feature is that you will never notice you committed it, because every step looked reasonable.

1What it prevents

Not lying. Deceiving yourself honestly. The usual sequence:

StepWhat you thoughtWhat happened
1"let's just run it and see"no success criterion exists
2"that number is poor, maybe the sample is small"keep running, wait for a better one
3"let me try slicing it differently"a dozen subsets appear
4"this subset looks great"you picked the largest number in the noise
5"found an edge"you found the selection process

Every step is intuitive; together they manufacture a false finding. And you will genuinely believe it, because you did not lie at any point.

A live record of my own

In one analysis I sliced twelve lines × two halves = 24 subsets and reported the good ones. I wrote a warning into the document at the time: "apart from a single z=3.28, every z falls between 1.2 and 1.9 — this is hypothesis generation, not a result."

Writing that took ten seconds, and it told everyone afterwards — me included — that those numbers could not be traded on. That is the minimum viable version of preregistration: record that you selected.

2What has to be frozen

ItemWhy in advance
Which line is judgedstops swapping horses: running five and reporting the best
The thresholdstops "just short, but the direction is right"
The sample size to read atstops stopping at the prettiest moment
When you readstops repeated peeking until a good number appears
What counts as failurefailure must be as explicit as success
The date you wrote itmakes "before or after" a verifiable fact

The one that killed my line had four thresholds: gross excess ≥ a computed cost line, significance z ≥ 2.0, beat three contemporaneous null models, and both halves of the sample agreeing in sign — all frozen seven days before the read.

3Thresholds are computed, not chosen

"Gross excess ≥ 4%" is a number pulled from the air. A correct threshold is derived from costs:

threshold = fee paid + measured slippage − rebates confirmed to arrive every term needs a measured source; none may be estimated

And write down how the threshold may move, in advance. Mine said: "if measured slippage > 0, raise accordingly." Measured slippage came in at +0.8pp, so the threshold automatically rose from 1.45pp to 2.25pp — a clause firing, not a criterion being revised after the fact.

⚠ And I still made an error here, worth writing out

That 1.45 was "1.75 of fee minus 0.30 of rebate stack". The 0.30pp never arrived — not one cent.

Add it back and the true threshold was 1.75 + 0.8 = 2.55pp. In other words, I spent money I had not received while computing my own bar.

Rulea cost line may contain only items confirmed to occur. This error's direction is always to lower the bar, which is why it will never be caught by its own consequences.

4★ A preregistration that admits it is dirty

The first paragraph of mine reads, verbatim:

"23% of the sample already seen — disclosed: the single judged line was designated at n=569, partly influenced by the then-leading value; everything below is frozen a priori."

Plainly: the line under judgement was chosen after peeking at almost a quarter of the sample, and partly because it happened to be ahead.

That is a real contamination. It sits in the first paragraph rather than hidden. The reasoning:

Result
No preregistrationthe contamination exists but is invisible; nobody can characterise it later
A preregistration pretending to be cleanthe contamination is concealed — worse than none
One that admits itthe contamination becomes a quantity you can argue about

A preregistration that admits its contamination still beats not having one — it at least lets a reader ask "how big is this, and could it overturn the conclusion?" And that question cannot be asked at all unless it is written down.

5The day of the read

Preregistration earns its keep at exactly one moment: when the data arrives and you dislike it.

My final read: threshold 2.25pp, read +1.56pp; threshold z ≥ 2.0, read z = 1.62 — two fails. But the other two criteria passed (beat three nulls, halves agreed in sign).

Reading only the passes, it is easy to conclude "there is an edge, just not a big one". Criteria 1 and 2 exist to block that conclusion.

Three escape routes deliberately not taken

Do not raise the z threshold ("2 is 2") ② Do not reopen a segment  ③ Do not delete a single row of data.

The bar was set seven days earlier, so when the number came it was no longer mine to move. A line that needs its bar moved to survive was never alive — moving it only defers the discovery, and it spends real money in the meantime.

6The minimum viable version

No tooling required. Before your next attempt, write four lines in a notes file and save it:

#Write
1which one I am judging (one line, not a family)
2how much sample before I read, and I read once
3what number counts as success (computed from costs)
4today's date

Four lines, two minutes. It will not stop every error, but it stops the most expensive class: "in hindsight I think this should count as a success."

7One line to keep

Preregistration is not there to prove anything to anyone else. It is there so that the future you cannot lie to the present you. Its only moment of use is when the data lands, you do not like it, and the bar is no longer yours to move.

EvidenceCheck it yourself

The four criteria gross excess ≥ computed cost line · z ≥ 2.0 · beat three contemporaneous nulls · halves agree in sign. Frozen seven days before the read.
Threshold derivation fee 1.75pp + measured slippage 0.80pp (n=662) − rebate stack 0.30pp = 2.25pp; the 0.30pp measured zero, so the true bar was 2.55pp.
Contamination disclosure the file's opening paragraph states "23% of the sample already seen" and "partly influenced by the then-leading value".
Final read n=2,468 (98.7% of the 2,500 trigger); +1.56pp / z1.62 → criteria 1 and 2 FAIL; 3 and 4 PASS.
Escape routes not taken z threshold unchanged, no segment reopened, no data deleted.
The 24-subset warning one analysis sliced 12 lines × 2 halves and recorded at the time that this was hypothesis generation, not a result.
Checked 2026-08. Full ledger in the autopsy.

NextWhere to go

F · METHOD
Gates and the single read
F · METHOD
Kill rules and pre-claimed failure modes
E · UP/DOWN
The line this preregistration killed
F · METHOD
Your backtest won, but the ruler was wrong
This is an educational and research record. It is not investment advice, promises no returns, and offers no personalised trading recommendations. Rules and API behaviour are per the official documentation; this page states when it was checked and both can change without notice. Prediction markets are restricted or unavailable in some jurisdictions — confirm your own before taking part.