OrderFlowAi
CORTEXDOMHeatmapReversalIcebergVolumeDeltaFlow PulseNT8SymphonieResearch
Cortex AI StrategyAutonomous TradingWyckoff AnalysisChart AnalysisTrading JournalLearn Order FlowDetect Iceberg OrdersUse Volume ProfileAI Software Comparison
My accountStart trialDE
OrderFlowAi / Research
RESEARCH · LAB NOTEBOOK

We publish what we measure. Even when it goes against us.

Every hypothesis is pre-registered before the data exists — question, sample and decision rule are fixed before anything is computed. This notebook shows, week by week, what we tested, what the data said and what we changed or discarded as a result. Negative findings get equal space.

43pre-registrations
54findings we retracted ourselves
12notebook entries
23 Sept 2026last updated
OrderFlowAi Researchupdated weekly
01 · HOW TO READ THE NOTEBOOK

Three steps, always in this order.

A finding only counts here if the rule was fixed before we looked at the data. If you compute first and pick the metric afterwards, market data will almost always show you something — and almost never something that holds.

01 · PRE-REGISTERED

Question, sample, metric and threshold are written down with a date. Nothing about them changes afterwards.

02 · WHAT THE DATA SAID

The result with sample size and baseline — also when it refutes the hypothesis.

03 · CONSEQUENCE

What we built, rebuilt or discarded as a result. A finding without a consequence is just a number.

HOLDSDOES NOT HOLDOPENREBUILTCORRECTED
02 · ENTRIES

The notebook — latest week first.

Entries are added weekly and never polished after the fact. If a finding later turns out to be wrong, the entry stays and the correction appears below under “Corrected”.

21 Sept 2026 – 23 Sept 2026OPEN

The new signal chain: CORTEX alone decides — and may end the trade

PRE-REGISTERED

Fixed on Sept 22, before the first data existed: more than 6.6 entries per day, a CORTEX-exit share between 10 and 30%, verdict only after 5 trading days and 30 trades. Old and new chain are never pooled.

WHAT THE DATA SAID

No verdict yet. Measured before the rebuild: 6 of the model's 49 inputs are practically silent, and no single feature separates up from down (AUC 0.47–0.51). Three throttles made the signal chain 3.3 s slow.

CONSEQUENCE

Throttles removed (≈ 0.2 s by calculation), training moved to its own process, fresh learning state for everyone with release 0.5.30. Not built, because it measured too weak beforehand: a bracket width taken from the bracket head (AUC 0.54).

Pre-registration: s469-kette-umbau · 22 Sept 2026
14 Sept 2026 – 20 Sept 2026DOES NOT HOLD

The weekend study's last candidate falls — despite a 65.6% hit rate

PRE-REGISTERED

In a daily forward test since Aug 29, frozen: “flush absorption” (price drops ≥ 12 ticks in 10 s under strong buying → long entry). Decision at n ≥ 50 and Wilson lower bound ≥ 55% — with positive expectancy.

WHAT THE DATA SAID

At n = 90 the candidate hits 65.6% (lower bound 55.3%) — and loses $5.17 per trade. On Sept 14 (n = 52) the threshold was briefly met; the rule has no stopping at the first crossing.

CONSEQUENCE

Discarded by rule since Sept 18, as were the other two candidates. Lesson for every future pre-registration: a hit rate without expectancy is not an edge.

14 Sept 2026 – 20 Sept 2026DOES NOT HOLD

The “fireworks” in the HUD: spectacular, but not predictive

PRE-REGISTERED

Four displays in the app each got a baseline that simply bets against the trend or with the last move. They only hold if the Wilson lower bound beats that baseline — one verdict after 30 resolved events.

WHAT THE DATA SAID

Reversal signal 57.0% (lower bound 46.8%) vs 48.2%. Takeover at the inside 58.5% (n = 41). A price where pressure was absorbed holds 48.6% vs 46.5%. Strongest HUD events 45.7% vs 50.1%. None holds.

CONSEQUENCE

All stay displays — with honest legends (“pressure absorbed here — no proven barrier”). Found and fixed on the way: the wall tracker counted cancellations as absorption.

Pre-registration: s449-cortex-rev · 14 Sept 2026 | s449g-absorptions-konto · 14 Sept 2026 | s449i-hud-ereignisse · 14 Sept 2026
14 Sept 2026 – 20 Sept 2026DOES NOT HOLD

Highly significant, but too small: repeating prints as an iceberg hint

PRE-REGISTERED

Do ≥ 5 equal-size prints at one price within 60 s indicate real reloading in the order book? Required in advance: at least 10 percentage points difference to the control and p < 0.05, once, on the first 5 days.

WHAT THE DATA SAID

20,170 candidates vs 18,649 controls: 11.4% vs 8.5% — p ≈ 10⁻²², but only 2.9 percentage points.

CONSEQUENCE

No display in the app, no parameter search on the same days. The example shows what a pre-set effect size is for: significant does not mean relevant.

Pre-registration: s447d-clip-wiederkehr · 12 Sept 2026
7 Sept 2026 – 13 Sept 2026HOLDS

The exchange as referee: our delta was 30% wrong

PRE-REGISTERED

For the first time with real CME order data as ground truth: the rule holds if it assigns ≥ 90% of volume correctly on ≥ 6 of 7 days. Switch only with a confirmation day after deployment (≥ 95%, per-minute r ≥ 0.95).

WHAT THE DATA SAID

Quote rule 97.5% of volume correct, the tick rule used until then only 69.9% — with the wrong daily sign on 4 of 7 days. With a locked quote every neighbour rule is below chance. Confirmation day Sept 11: 97.0%, per-minute r = 0.985.

CONSEQUENCE

Delta switched in indicator, app and AI; locked prints count as “side not determinable” instead of guessed. Old and new days are never pooled — the learning state was restarted.

Pre-registration: s441-je-trade-join · 10 Sept 2026 | s441-gesperrt-naechst · 10 Sept 2026 | s443-quote-regel-umsetzung · 10 Sept 2026
7 Sept 2026 – 13 Sept 2026DOES NOT HOLD

Absorption and exhaustion: largely an artefact of classification

PRE-REGISTERED

“Just a label” if the tick version and exchange truth agree on the sign ≥ 90% of the time; “noise” below 75%. Second question: does the true exhaustion predict the price target better (AUC ≥ 0.53)?

WHAT THE DATA SAID

Signs agree only 74.1% of the time; the tick version reports exhaustion in 41–49% of seconds, the truth in 34%. And the true exhaustion does not hold either: AUC 0.510 vs 0.510.

CONSEQUENCE

The correct rule makes the measure true, but not useful. For the model restart that means: it costs no signal, because there was none.

Pre-registration: s441-buch-wahrheit · 10 Sept 2026 | s441-exh-wahrheit-ziel · 10 Sept 2026
7 Sept 2026 – 13 Sept 2026DOES NOT HOLD

The replay machine: would the strategy have won the other way round?

PRE-REGISTERED

105 replayable trades, tick by tick on our own tape, including costs. A variant only holds with a positive net result and a Wilson lower bound above its own break-even line. One look per variant.

WHAT THE DATA SAID

Original −1.76 ticks per contract. Reversed direction −1.73. No time stop, T1 = stop, no break-even lock: −1.36 to −1.72. Larger scale: −2.01 to −0.89. Brackets that grow with speed: −1.86 / −1.87. No variant holds.

CONSEQUENCE

The entries themselves carry no edge — no management rescues them. Strategy not reversed, brackets not enlarged. The shortcut “flip the sign” (+0.95 ticks) was retracted before any code changed.

Pre-registration: s437-inversion · 9 Sept 2026 | s439-bracket-struktur · 9 Sept 2026 | s440-skala · 10 Sept 2026 | s440-vollgas-bracket · 10 Sept 2026
31 Aug 2026 – 6 Sept 2026DOES NOT HOLD

Order flow describes the move — but does not predict the next one

PRE-REGISTERED

Fixed since July 29: the sign of the order-flow imbalance must hit the next mid-price move with a skill ≥ 1.25 against the chance line. If this stage falls, the whole cascade above it falls.

WHAT THE DATA SAID

19 days, 77,972 windows: hit rate 49.2–50.2%, skill 0.995 on all four horizons. The simultaneous move, however, is explained almost completely by the same order flow (median R² 0.893).

CONSEQUENCE

On Sept 1 stages 5 to 8 were shut down — exactly as fixed in advance. A planned order-flow entry gate was never built.

Pre-registration: s394 · 29 Jul 2026
31 Aug 2026 – 6 Sept 2026HOLDS

A well-known iceberg trader's theses against 33 days of data

PRE-REGISTERED

Exploratory, but with a fixed hurdle: candidate only above the 95th percentile of a corrected control, not worse on ≥ 60% of days and an effect of ≥ 2 percentage points.

WHAT THE DATA SAID

Sweep zones pull price back (on 25 of 33 days), sweeps through a 60-minute extreme reverse more often (55.4% vs 51.8%). Falling short: iceberg zones as anchors and the entry rule “ATR + 15%”.

CONSEQUENCE

Two candidates for a later pre-registration on new days — no code in strategy or CORTEX. Experience supplies hypotheses, not the yardstick.

Pre-registration: s429-ereigniszone · 31 Aug 2026
31 Aug 2026 – 6 Sept 2026DOES NOT HOLD

“Open road after full throttle”: the market picture does not survive the data

PRE-REGISTERED

Seven states of an intuitive market picture (full throttle, refuel stop, jam, wall) against time-matched control moments: holds above the control's 95th percentile and on ≥ 60% of days.

WHAT THE DATA SAID

436,088 seconds, 19 days: after a refuel stop the move continues less often than chance (46.5% vs 58.8%, better on 0 of 19 days). Large walls, on the other hand, measurably hold (break-through 42.9% vs 49.5%).

CONSEQUENCE

No phase module built. The opposite signature became a new, separate pre-registration on fresh days — instead of turning thresholds until the picture fits.

Pre-registration: s421-autobahn · 31 Aug 2026 | s421-autobahn-rev2 · 31 Aug 2026
24 Aug 2026 – 30 Aug 2026DOES NOT HOLD

The only positive learnability trace does not replicate

PRE-REGISTERED

A sequence model (GRU) had shown a weak trace in August (AUC 0.548 on 2 test days). Replicated only at p ≤ 0.01 on days that analysis never saw — with a working positive control.

WHAT THE DATA SAID

6 fresh days: AUC 0.5045, p = 0.366. Positive control sharp (0.796), placebo silent — an edge from AUC ≈ 0.605 would have been found. New feature sets failed as well (also in the repeat on Sept 11).

CONSEQUENCE

The August trace counts as small-sample chance. Four model classes are empty on this question; no model restart for new features.

Pre-registration: s415-sequenz-oos · 26 Aug 2026 | featrev-s415 · 26 Aug 2026
17 Aug 2026 – 23 Aug 2026DOES NOT HOLD

Filtering out small prints does not improve the delta

PRE-REGISTERED

Paired test against the raw delta: confirmed only at ≥ 3 percentage points difference and p ≤ 0.005. The exploration day on which the idea arose was excluded.

WHAT THE DATA SAID

Delta from clearly classifiable volume only: −0.55 percentage points. From prints of 5 contracts and more only: +0.63 percentage points (p = 0.40). Divergences under high rule agreement do not hold either (p = 0.19).

CONSEQUENCE

The raw delta stays — a month later the exchange truth showed that the problem lay elsewhere: in the rule itself (week 37).

Pre-registration: s407h-identifizierbarkeit · 18 Aug 2026 | s407h-divergenz · 18 Aug 2026
03 · CORRECTED

Findings we retracted ourselves.

Every error is recorded with a date and the correct statement. The register is the reason the remaining numbers can be trusted.

22 Sept 2026
The AI signal was said to be 16 times older than our own cost line allows.

The cost line measures the age of the order book, not of the signal — the comparison does not hold. What stands: the signal chain was shortened from 3.3 s to about 0.2 s.

10 Sept 2026
The strategy was said to trade against the direction — reversed it would be profitable.

In a tick-accurate replay both directions lose (−1.76 and −1.73 ticks per contract). The mirrored trade has its own stop.

10 Sept 2026
The tick rule was said to be the better estimate of the aggressor side.

Against real exchange data the quote rule reaches r = 0.983 per minute, the tick rule r = 0.677. The earlier number had compared the rule with its own fields.

8 Sept 2026
A sign error was said to explain why the strategy's direction was anti-predictive.

There was no flipped sign, but disagreement between classification rules on stale quotes — a measurement problem, not a bug.

2 Sept 2026
The AI's bracket head was said to separate strongly, at twice the base rate.

Against the correct trivial line (0.695) its accuracy of 0.631 was below it; it had been compared with a positive rate.

1 Sept 2026
Order-flow direction was said to predict the next move (skill 1.5).

Over 19 days the pre-registered skill was 0.985–1.004 — no predictive value. A different quantity had been computed.

15 Aug 2026
The regime router was said to be merely mis-tuned.

Its label is independent of the true regime (kappa −0.04 to −0.07) — a different threshold changes nothing.

29 Jul 2026
A size evidence of 75.4% was said to prove the aggressor side.

A placebo hit 75.2% — the information content was 0.2 percentage points. Since then every rate needs a null hypothesis.

28 Jul 2026
The rising learning curve was said to show that the model learns.

The rise followed the falling share of flat market phases; against the day-dependent chance line 4 of 5 sessions were negative.

15 Jul 2026
The entry edge was said to depend on the regime and only vanish in the mixed pool.

Across 26,940 samples no market regime showed an entry edge (p = 0.29 / 0.30 / 0.80).

04 · THE REPORT

The Search for the Edge — the full research report.

The report brings it all together: how a statistical edge would have to be proven, which traps we fell into, what is established today and what is still missing. It does not claim that an edge has been found.

Key statements

  1. We have not found an edge — and we say so. The confirmatory main result is negative; every number in the report is measured, every refuted one is marked as refuted.
  2. The rule comes before the calculation. 43 pre-registrations fix question, sample and threshold before any result exists.
  3. We retracted 54 of our own findings — including some that had already been reported as a success. A selection is listed below under “Corrected”.
  4. The measurement base is checked against the exchange: since September our delta matches the true aggressor side 97% of the time — before, it was 70%.
  5. The goal remains a verifiable analysis AI — defined not by features but by the fact that every statement can be recomputed.
Read online43 pages · 23 Sept 2026PDF · print editionPDF · 23 Sept 2026PDF · presentationPDF · 23 Sept 2026Deutsche Ausgabe43 Seiten
05 · CHANGE LOG

When this page changed.

  • 23 Sept 2026Page launched: 12 entries (August–September), 10 corrections from the retraction register, research report edition of Sept 23 (43 pages) as web and PDF version.

OrderFlowAi is analysis software. Nothing on this page is investment advice or a promise of future results. All measurements come from our own recordings (ES futures, Level 2) and a simulation account.

The tools we measure with are the same ones you see in OrderFlowAi.

START 14-DAY FREE TRIAL
0
Zum Inhalt springen
OrderFlow Ai
Home
OrderFlow Ai
Home
Home