The number nobody reports: how many strategies did you test before this one?
Search through enough variations of pure noise and one of them will look like an edge. That is not a warning about dishonest people; it is arithmetic, and it applies to us. Here is the arithmetic, the two numbers that make a mined result readable, and what we found when we went looking for a backtesting tool that reports either of them.
Start with ours
In August we ran a sweep of calendar effects on QQQ: 25 variations of “does the market behave differently at this point in the week or month.” One of them looked good.
Corrected for the fact that we had run 25 tests, none of them survived. Not the good-looking one, not any of them.
We had already done the same thing at larger scale. A 58-test sweep across selection filters and entry timing — every plausible way of choosing which names the screen keeps, and when to buy them. Zero survived the same correction. Across the project to that point: 33 registered hypotheses, none passed.
None of those sweeps produced a strategy. What they produced was a habit: before believing the best result in a batch, count the batch.
Why the count is the whole story
Run one test on a strategy with no edge, at the usual 5% threshold, and you have a 1-in-20 chance of a false positive. Run twenty, and the chance at least one comes back “significant” is 64%. Run fifty-eight — our sweep — and it is 95%. At two hundred variations, finding a winner is not evidence of anything. It is arithmetic. You were always going to find one.
| Variations tested | Chance at least one looks significant | Threshold you would actually need |
|---|---|---|
| 5 | 23% | 1 in 100 |
| 20 | 64% | 1 in 400 |
| 25 | 72% | 1 in 500 |
| 58 | 95% | 1 in 1,160 |
| 200 | >99.99% | 1 in 4,000 |
Assumes the tests are independent. Real strategy variants are correlated — twenty tweaks of one moving average are closer to one test than to twenty — which makes the true figures lower. Lower, not absent: correlation shrinks the penalty, it never removes it, and nobody who skips the correction entirely is doing so because they measured the correlation.
What a mined win rate looks like when there is nothing there
Take a coin. It has no edge — it is a coin. Call each flip a trade, and let it trade thirty times.
Now mine it. Generate 200 “strategies” that are just 200 independent runs of that same coin, and keep the best one. The median best win rate is 73%.
Not the luckiest case — the median. Half the time it is higher. With 20 variations instead of 200 the median best is still 67%. Give each strategy 100 trades instead of 30 and 200 variations still yields 63%.
That is the range in which mined strategies are marketed. A 68% win rate on a few dozen trades is not a suspicious result. It is the expected result of searching hard enough through pure noise.
The number that cannot be read on its own
Here is where this gets more interesting than a simple debunk, and where we had to correct our own first instinct.
One strategy-generation product publishes a customer’s account of using it: he screens for a profit factor above 1.6, a win rate above 65%, and a return-to-drawdown ratio above 3, and reports that about one strategy in every million iterations clears that bar. He offers it as evidence the product works.
We got this wrong first
Our first reaction was that one-in-a-million is exactly what noise produces. We checked, and that is wrong. Under a zero-edge model — each trade a coin flip of equal size — here is how often pure noise clears that same three-part filter:
| Trades per strategy | Noise passes | Expected hits per million searched |
|---|---|---|
| 20 | 8.6% | 85,600 |
| 30 | 4.3% | 43,400 |
| 50 | 1.6% | 16,100 |
| 100 | 0.19% | 1,900 |
| 200 | 0.003% | 30 |
At 30 trades, noise clears that filter forty-three thousand times per million. One hit per million would mean the search was finding less than chance. At 200 trades, noise clears it about 30 times per million — so one-in-a-million is genuinely selective.
The same hit rate is either meaningless or impressive depending on a number that was not reported. Without the trade count, it cannot be read at all — in either direction.
Across a plausible range of sample sizes the noise rate moves by more than three orders of magnitude. That is the real finding, and it is a more useful one than “mining is bad.” Mining is a legitimate way to generate hypotheses. What makes a mined result readable is publishing the two numbers that let someone else do this arithmetic: how many candidates you searched, and how many observations each one got. Almost nobody publishes either.
Where the category actually stands
We looked at four consumer-facing backtesting and strategy-generation products. We could not find one that reports, alongside a winner, how many candidates that winner beat — or that discounts the result for the size of the search.
What they do offer is real and worth crediting. One splits data into training and test halves, validates forward on rolling out-of-sample segments, runs a Monte Carlo stress test on the trade sequence, grades the result with a four-way verdict that includes “looked good in-sample only,” and makes you tick a box acknowledging a strategy is fragile before it will save it. Another adds a noise test that rebuilds the price series a thousand times over. A third advertises Monte Carlo, walk-forward matrices, parameter permutation and automatic overfitting tests. This is serious engineering against overfitting one strategy.
None of it addresses the search that selected that strategy. A train/test split scores the finalist; it does not penalise the tournament that produced the finalist. And the tools that search hardest disclose the search least — one advertises combining and verifying millions of entry and exit conditions on its front page, and reports no count anywhere in its output.
The fourth is the telling one. It publishes a genuinely good educational page on multiple-testing correction — and states plainly that the difficulty is bookkeeping, because correction requires knowing how many tests were run. It is right. Its own tooling does not keep that book either.
Scope of that claim, stated honestly
We read one of these four products’ shipped application code in full. For the other three we read public documentation, marketing and support material, not their in-app output. A count could exist in a screen we have not seen. If you work on one of these and we have missed it, tell us and we will correct this note in public, which is what we do with our own errors.
The check, in ten seconds
Ask: “How many variations did you test before you showed me this one — and how many trades did each one get?”
- Two numbers, even large ones, mean they were counting, and you can do the arithmetic above.
- “We tested thoroughly” is not a number.
- “Just this one,” from a product with an optimizer, a miner, or a parameter sweep, is not true.
- “Why does that matter?” is the answer. Move on.
And the follow-up that costs nothing if they were honest: what would you have shown me if this one had not worked? A pre-registered process has an answer — the failures, published. A mined one does not, because the failures were never written down.
This is check #4 on the seven checks any service should pass — rules frozen in advance, or fitted to the past. If you want to put a service through all seven, the scorecard takes two minutes and nothing you enter leaves your browser.
What we do about it
Every hypothesis is registered with a date and an evaluation window before we look. When we sweep, we publish the sweep — all 58, all 25, not the survivor. When a sweep produces nothing, that is a result and it goes up.
And the case we are least comfortable with, which is why it is here: the turn-of-month effect was the strongest survivor of that 25-test calendar sweep, in which, corrected for multiple testing, nothing survived. We did not publish it as a finding. We registered it forward as Experiment 04, on a frozen window, with the verdict date set in advance — because the only way to redeem a mined result is to stop mining and let it face data that did not exist when you found it.
That verdict is due in 2027. We do not know what it says. That is the point.
Reproduce every number above
Seeded, so the figures are stable. Runs in under a minute in stock Python.
import random
from math import comb
# 1. chance at least one of N independent tests clears a 5% threshold under the null
for N in (5, 20, 25, 58, 200):
print(N, f"{(1 - 0.95**N)*100:.2f}%", "threshold 1 in", round(1/(0.05/N)))
# 2. median best win rate from mining N zero-edge strategies of n trades each
def best_of(N, n, trials=20000, seed=7):
rng = random.Random(seed)
b = sorted(max(sum(rng.random() < 0.5 for _ in range(n)) for _ in range(N))
for _ in range(trials))
return b[len(b)//2] / n
for N, n in ((20, 30), (200, 30), (200, 100)):
print(f"best of {N} over {n} trades -> {best_of(N, n)*100:.1f}%")
# 3. how often zero-edge noise clears win>=65% AND PF>=1.6 AND net/maxDD>=3
def pass_rate(n, trials=200000, seed=11):
rng = random.Random(seed); hits = 0
for _ in range(trials):
w = eq = peak = mdd = 0
for _ in range(n):
r = 1 if rng.random() < 0.5 else -1
w += r > 0; eq += r
peak = max(peak, eq); mdd = max(mdd, peak - eq)
l = n - w
if w/n < 0.65: continue
if l and w/l < 1.6: continue
if mdd == 0 or eq/mdd >= 3: hits += 1
return hits/trials
for n in (20, 30, 50, 100, 200):
p = pass_rate(n)
print(f"n={n:<4} noise passes {p*100:7.3f}% -> {p*1e6:>9,.0f} per million searched")
Educational and informational only — not investment advice, and not a broker-dealer. ThePickLog is operated by AMD Ventures, LLC (Florida). The products described above are not named on purpose: the point is the category, not any vendor, and every figure attributed to a product is paraphrased from that product’s own public material. The simulation is a deliberately simple zero-edge model with equal-sized trades; real strategies have asymmetric payoffs, which changes the numbers but not the direction of the argument.
If you are checking someone else’s track record for this, start with the 7 checks every service should pass.
Evidence Validation · Score & metrics · Trust & Data
Tools Calculator · Vetting Guide · 7-Check Scorecard
Legal Disclaimer · Privacy · Terms