Every experiment, and how each one is scored

Retail trading runs on claims that are almost never checked: strategies with impressive win rates, screenshots of the good days, backtests fitted to the past. So we take a claim, write the rule down with a date before we look, grade it forward on public data, and publish the verdict — including when the answer is embarrassing.

Rule frozen before the fact  •  Graded forward, winners and losers  •  Recomputable from public data  •  Published either way

What a verdict is allowed to say — the five outcomes

“Failed” is three different findings wearing one word, and the difference decides what you should do next. A claim that was tested properly and lost is not the same as one the data was never able to answer. These are the only five things we let ourselves conclude, and every one of them gets published.

HELD UPRegistered in advance, graded forward, and the effect survived out-of-sample. Nothing has reached this yet.
DIDN’T HOLD UPRegistered in advance, graded forward, and the effect went to zero or the other way. The interval excludes zero on the wrong side. This is a real refutation, not an absence of evidence.
INCONCLUSIVERan to its verdict date, and the interval still includes zero. The claim is neither confirmed nor refuted — and saying so is not the same as saying it failed.
UNDERPOWEREDToo few observations to detect the effect the registration said we were looking for. The honest answer is that this test could not have found the answer either way.
RUNNINGRegistered, window open, no verdict yet. The result is not known to us either.

One rule holds all five together: the label is chosen by the interval we published in advance, not by how the number reads afterwards. Worked example, including the arithmetic: Gate 1. The CORRECTION tag further down is not a sixth verdict — it marks a field note about a mistake in our own apparatus, which is a different kind of finding from a verdict on a claim.

EXPERIMENT 01DIDN’T HOLD UP

Our own low-float momentum screen

We started with ourselves. We built a screener for low-float stocks gapping up on heavy volume, pre-registered the rule we believed in, and ran it forward for seven weeks. It lost — significantly, and in the opposite direction to the one we predicted. Here is the whole thing, including the number that flipped sign as the sample grew.

registered 2026-06-23 · verdict 2026-07-29 · n=309 · Δ −3.0pp vs baseline · 95% CI [−4.4, −1.5]
EXPERIMENT 02RUNNING

The 2-period RSI "high win rate" mean-reversion trade

One of the most widely published retail strategies of the last twenty years: buy a large-cap when the 2-day RSI collapses while it's still above its 200-day average. It is usually sold on its win rate. We are testing whether a high win rate survives contact with expectancy — on liquid names, where trading costs can't be blamed for the answer.

registered 2026-07-29 · first verdict expected ~2026-09 · universe frozen · control group logged daily
EXPERIMENT 03RUNNING

The MACD bullish crossover

Arguably the most widely taught indicator signal in retail trading — it ships on every charting platform and appears in essentially every beginner course. Which makes it the least likely thing in the world to still contain an edge, and exactly the sort of universally-believed claim nobody bothers to check. Registered prior: no edge, roughly 1 in 6 that it clears.

registered 2026-07-31 · same 40-name universe · day-matched control · live results update automatically
EXPERIMENT 04RUNNING

The turn-of-month effect

A 40-year-old claim from the academic calendar-anomaly literature, still sold today in seasonality newsletters: returns cluster in the last few and first few trading days of the month. It was the strongest survivor of our 25-test calendar sweep — in which, corrected for multiple testing, nothing survived. So it gets the only test that counts: forward, window frozen in advance, on QQQ.

registered 2026-08-06 · QQQ decides, SPY replication · time-matched control · first verdict ~2027-09
EXPERIMENT 05RUNNING

Overnight vs intraday — the market that only pays while it's closed

Over 27 years, essentially all of QQQ's return came between the close and the next open; the hours the market was actually open made nothing. Firms sold this as "night effect" ETFs, and the ETFs died. We registered it forward as an attribution claim — where returns happen, explicitly not a tradeable edge — with each day's intraday leg as its own control.

registered 2026-08-06 · paired per-session control · attribution only, binding scope · first read ~2027-01

Field notes — the mistakes, written up

Errors we found in our own record. Both of these made our numbers look better than they were, and both are written up in full.

FIELD NOTEMETHOD

The number nobody reports: how many strategies did you test before this one?

Mine 200 zero-edge strategies over 30 trades and the median best "win rate" is 73%. A hit rate you are shown cannot be read without two numbers you are almost never shown — how many candidates were searched, and how many trades each one got. We ran the arithmetic, checked our own first instinct (it was wrong), and surveyed four backtesting tools for either number.

58-test sweep, 0 survive · 25-test sweep, 0 survive · four products surveyed, none report the count
FIELD NOTECORRECTION

We logged 128 picks after the opening bell

The whole record rests on picks being written down before the market opens. On seven occasions our scheduler drifted past the bell and logged anyway, and we didn't notice for six weeks — an outside reader found it in our public CSV. What broke, what it did to the numbers, why we won't claim it flattered us, and the gate that makes it impossible now.

found 2026-08-04 by an outside review · 128 picks / 7 cohorts excluded · headline −2.68% → −3.40%
FIELD NOTE

The bug that makes your backtest look good — and why you can't see it

Historical prices are silently rewritten every time a company splits. Mix a price you stored with one you download later and you invent returns that never happened — and in cheap stocks the error is always in your favour. We found it in our own code, where it had been inflating our numbers for weeks. How it works, why it survives review, and the six checks that catch it.

found 2026-07-29 while auditing Experiment 01 · cost us a +8.0% baseline that was really −2.9%

All field notes →

How an experiment works here

  1. The claim gets written down first. Rule, universe, entry, exit, cost assumption, and the exact bar it has to clear — all frozen with a date, in a public file, before any of the data it will be judged on exists.
  2. It runs forward only. No backtest ever counts as evidence. Only outcomes that happen after the registration date are scored.
  3. It is graded against a control. "It made money" is not a result. The question is always whether it beat the honest alternative — doing the simple thing instead.
  4. The statistics are done properly. Confidence intervals, a correction for the fact that repeated bets on the same names are not independent, and an explicit count of how many results you'd expect to look good by chance alone.
  5. The verdict is published either way. That is the whole point. A method we hoped would work and didn't is a more useful result than one more success story.

What should we test next?

  1. If there's a strategy, indicator or claim you keep seeing sold — and you've never seen anyone check it honestly — tell us and it goes in the public queue. We pick from it.
  2. Two rules: it has to be specific enough to write down as a rule before the fact, and we test techniques, not people.

We are not selling signals, and nothing here is advice. The output of this project is verdicts on public claims — including our own, which is where we started and where the record is worst.

Get every verdict by email

A verdict lands every few weeks. We'll email you each one, free, the day it publishes — and nothing else. No signals, no picks, no offers. Verdicts stay free and public for everyone either way; this just means you don't have to remember to check back.

One confirmation email first — you're not on the list until you click it. Privacy

Rules are frozen in HYPOTHESES.md; every verdict is recorded in the audit log; the raw data is downloadable from the track record. Method and standards: the method · principles.
Educational and informational only — not investment advice, and not a broker-dealer.
Evidence Validation · Score & metrics · Trust & Data
Tools Calculator · Vetting Guide · 7-Check Scorecard
Legal Disclaimer · Privacy · Terms