EXPERIMENT 04 · REGISTERED 2026-08-06 · RUNNING

Does the market pay you at the turn of the month?

One of the oldest claims in the calendar-effect literature: stock returns cluster in the last few and first few trading days of each month. It has been in academic journals since the 1980s and is still sold today in seasonality newsletters and "smart timing" content. We froze the window definition before the fact and are grading it forward on QQQ.

STATUS
Registered and running. No verdict yet.

Every trading session is classified in advance by the exchange calendar — turn-of-month or not — and graded automatically at each close. A verdict needs at least 30 post-registration turn-of-month sessions and at least 12 complete turn-of-month cycles; the verdict is computed on a single pre-declared date, 1 September 2027, and then locked. Until then the report shows the running numbers and no verdict at all — see the amendment note below. Running numbers live in the calendar experiments report.

Where this came from

We ran an exploratory sweep of 25 declared calendar patterns over 27 years of QQQ — day-of-week effects, month effects, options-expiration weeks, pre-holiday days, and more. With an honest correction for having run 25 tests at once, none of them survived. That is the normal result and we're saying so up front.

But exploration is for generating candidates, not believing them — and this was the strongest one. The turn-of-month gap (about +9.6 vs +3.5 basis points per day) pointed the same direction in all three decades of the backtest and in the S&P 500 as well. Roughly 58% of QQQ's summed daily return arrived in a third of its days. A forward test is how a candidate earns or loses belief, so here it is.

The rule, frozen

The claim
QQQ's mean close-to-close return over the last 4 plus first 3 trading days of each month exceeds its mean over all other sessions.
Asset
QQQ decides the verdict. SPY is graded in parallel purely as a replication read — reported, never a pass criterion.
Window definition
Derived from the exchange calendar itself, so there is no signal to capture and no logging race — whether a future session is turn-of-month is already fixed. The month currently in progress only grades its first-3 sessions, because its "last 4" cannot be known yet and our outcome files are append-only.
Control
Time-matched: turn-of-month sessions are scored against the average of all other sessions in the same forward window. (Our usual day-matched control compares a pick against its universe; for a timing claim on a single asset, the honest control is other days, not other tickers.)
Costs
The claim is graded at the day level, gross. Expressed as a strategy it trades ~12 round trips a year, so at ETF-scale costs the answer cannot be decided by fees either way — one reason this candidate was chosen.
Window
Only sessions strictly after 2026-08-06 count. The 27-year backtest is context, never evidence.

What it has to clear

1. At least 30 post-registration turn-of-month sessions graded, and at least 12 complete turn-of-month cycles.
2. Turn-of-month mean and median both above the other-days mean and median. Both, not either.
3. A 95% confidence interval on the difference that excludes zero, computed treating each turn-of-month cycle as the unit of evidence — consecutive sessions are not independent, and the last four sessions of one month plus the first three of the next are a single unbroken run of market time.
4. The direction holding across at least three consecutive weekly snapshots.
5. All of the above assessed once, on the pre-declared verdict date of 1 September 2027 — then written down and never recomputed.

Win rate is reported but is explicitly not a pass criterion.

What we expect to happen

Forty years of publication is forty years of arbitrage opportunity. Our registered expectation is that, more likely than not, the forward excess is indistinguishable from zero — we put the chance it clears the full bar at roughly one in three. That is the highest prior we have ever registered, and deliberately so: this is the strongest survivor of a 25-test sweep, and registering it forward instead of believing the backtest is the entire point of this site.

A pattern that looked good for 27 years and stops working the day you test it forward is the single most common story in retail trading. We are checking, in public, whether this is one of those.

Amended one day after registration — before any data existed

What changed, and why we are telling you

The day after registering this experiment we ran an adversarial review of our own grading code and found six defects. Four were plain bugs and we fixed them. Two changed what this experiment has to clear, so they are amended on the record.

The confidence interval was too narrow. We computed the control average once and subtracted it as a fixed number, so none of the control group's own uncertainty made it into the interval. Simulated against pure noise, our “95% interval” was wrong about 12% of the time instead of 2.5%.

And the bar was being re-tested at every look. The report recomputed a verdict on every run, so across a year of looks a claim with no real effect had roughly a one-in-five chance of printing a pass at least once. Both experiments now require a minimum number of independent cycles, not just sessions — and, because a floor only delays the first look rather than limiting how many you take, each one now has a single pre-declared verdict date. The verdict is computed once, on that date, written to a file, and never recomputed; re-running our own code on later data cannot turn a null into a pass. That is why the verdict moved from December 2026 to 1 September 2027.

The residual, stated plainly: even after the fix, at the 12-cycle floor the interval is wrong about 4.2% of the time against a nominal 2.5%. That is a property of this kind of statistic at any sample size reachable in a sane window, so rather than pretend otherwise: read a bare “clears the bar” here as about a 1-in-24 chance of being a false positive, not 1-in-40.

Not one row of data had been graded when this was written, so nothing here was chosen after seeing a result — which is the only reason amending is legitimate rather than fatal. The full diff is in HYPOTHESES.md and in the git history.

How you'll be able to check it

Every session's grade is appended to experiments/EXP0405-CAL-outcomes.csv — public and append-only. The frozen rule lives in HYPOTHESES.md as H-EXP04, dated, with the constants in calendar_eval.py. The running numbers regenerate into the calendar experiments report, and the verdict will be recorded in the audit log and written up on this page. If we change any constant, the test is void and we'll say so.

What this experiment is not

It is not a recommendation to trade the turn of the month or anything else, and nothing here is advice. Whichever way it lands, the verdict describes this window definition, on QQQ, over this period — not calendar effects in general.

Get every verdict by email

A verdict lands every few weeks. We'll email you each one, free, the day it publishes — and nothing else. No signals, no picks, no offers. Verdicts stay free and public for everyone either way; this just means you don't have to remember to check back.

One confirmation email first — you're not on the list until you click it. Privacy

Educational and informational only — not investment advice, not a recommendation, and not a broker-dealer. ThePickLog is operated by AMD Ventures, LLC (Florida).
← Experiment 03: the MACD crossover · Experiment 05: overnight vs intraday →
Disclaimer · Privacy · All experiments