Anyone selling market seasonality should be able to answer three questions: is your grid just data-mining, has your edge already decayed, and does anything you found have a reason to exist? We ran all three against our own published numbers. Two came back in our favour. The first came back against us in 7 of our 19 markets, and every one of them is named below.
The most damaging paper to a product like this one is Sullivan, Timmermann & White, Dangers of data mining: the case of calendar effects in stock returns (Journal of Econometrics, 2001). They built a large universe of calendar rules, applied White's Reality Check, and showed that calendar effects which look decisive alone are far weaker once you account for every rule that could have been tested. That is a precise description of what we do: we search a weekday × hour grid across 19 markets and publish the survivors.
So we built the null. It preserves each bucket's real sample size and each symbol's own overall up-rate, destroys only the weekday-hour structure — the one thing we claim exists — and draws 4,000 times. The second constraint matters more than it looks: we measure z against a flat 50%, so a market that closed up 52.3% of all hours has every bucket shifted upward, and that alone manufactures significant hits with no hour-of-week structure whatsoever. Drawing the null at the symbol's own rate charges us for that.
| Market | Buckets tested | Own up-rate | Found | Noise expects | p |
|---|---|---|---|---|---|
| EUR/USD | 121 | 50.1% | 11 | 0.3 | <0.00025 |
| GBP/USD | 121 | 49.9% | 10 | 0.3 | <0.00025 |
| USD/CHF | 120 | 49.9% | 16 | 0.3 | <0.00025 |
| USD/JPY | 121 | 50.7% | 9 | 0.6 | <0.00025 |
| AUD/USD | 121 | 50.1% | 12 | 0.3 | <0.00025 |
| USD/CAD | 121 | 49.8% | 16 | 0.4 | <0.00025 |
| NZD/USD | 121 | 50.0% | 13 | 0.3 | <0.00025 |
| EUR/JPY | 121 | 50.9% | 7 | 0.9 | <0.00025 |
| GBP/JPY | 120 | 51.1% | 9 | 1.2 | <0.00025 |
| EUR/GBP | 120 | 49.2% | 16 | 0.5 | <0.00025 |
| Gold (XAU/USD) | 120 | 51.0% | 22 | 1.3 | <0.00025 |
| Silver (XAG/USD) | 120 | 49.5% | 35 | 0.5 | <0.00025 |
Publishing only the first table would be the exact behaviour the paper warns about. Here is the rest of it.
| Market | Found | Noise expects | p | |
|---|---|---|---|---|
| BTC/USD | 3 | 0.8 | 0.0432 | marginal |
| ETH/USD | 2 | 0.5 | 0.0818 | fails |
| NAS100 | 3 | 1.8 | 0.2595 | fails |
| SPX500 | 2 | 1.6 | 0.4955 | fails |
| GER40 (DAX) | 1 | 1.0 | 0.6362 | fails |
| US30 (Dow) | 0 | 1.1 | 1.0000 | fails |
| UK100 (FTSE) | not one bucket reaches our 300-sample minimum — largest 193 | cannot be tested | ||
The four testable equity indices are indistinguishable from a noise grid of the same shape. US30 comes back at p = 1.0000 — it found 0 significant buckets where noise alone expects 1.1. Note also that NAS100 and SPX500 both close up 52.3% of all hours, so their null means are the highest in the table: much of the little they do show is upward drift being read as hour-of-week structure. The null caught that, which is why it was built that way.
The same arithmetic runs in the opposite direction, and it is the plainest answer to why our numbers look boring next to a screener's. Our bar is |z| ≥ 3 on at least 300 samples, which fixes the smallest lean each market is capable of flagging at its median bucket:
| Market | Median bucket | Smallest lean we can flag |
|---|---|---|
| Gold (XAU/USD) | n = 1,205 | 54.3% |
| USD/CAD | n = 1,184 | 54.4% |
| EUR/USD | n = 1,180 | 54.4% |
| AUD/USD | n = 1,146 | 54.4% |
| EUR/JPY | n = 1,044 | 54.6% |
| GBP/USD | n = 1,040 | 54.7% |
| Silver (XAG/USD) | n = 1,034 | 54.7% |
| NZD/USD | n = 1,031 | 54.7% |
| USD/JPY | n = 910 | 55.0% |
| GBP/JPY | n = 831 | 55.2% |
| USD/CHF | n = 644 | 55.9% |
| US30 (Dow) | n = 567 | 56.3% |
| GER40 (DAX) | n = 532 | 56.5% |
| EUR/GBP | n = 487 | 56.8% |
| BTC/USD | n = 429 | 57.2% |
| ETH/USD | n = 410 | 57.4% |
| SPX500 | n = 341 | 58.1% |
| NAS100 | n = 340 | 58.1% |
| UK100 (FTSE) | — | 61.3% |
We sell a market whose data cannot clear our own bar. Its largest weekday × hour bucket holds 193 samples against our 300 minimum, across 5 covered years. The dashboard already says so twice — the signals card reports that nothing clears the bar, and the tile reads "Years with data: 5" — but it belongs here too, stated plainly rather than left for a user to discover. This site does not offer 19 equally-supported markets. The FX majors and metals are where this dataset is strong.
The literature is blunt about this. Schwert found the weekend, size and value effects weakening after the papers documenting them were published (NBER w9277); McLean & Pontiff tracked 97 anomalies and found systematic post-publication decay; Urquhart & McGroarty report that calendar anomalies have essentially gone since the 1980s. We publish one pooled number per bucket across 2003–2026, so an effect that died in 2015 would still show — and our sample-size column, the site's main honesty device, would make a dead edge look more trustworthy rather than less.
Split at 2014. First the harsh, selected question: of the buckets the product flags today, how many survive on the recent half alone?
The same run reports that of the buckets significant in the first era, only 18% are still significant in the second. Read cold, that is a decay story. It is not one. Those buckets were selected on first-era data: any bucket near the boundary gets into that set only when first-era noise happens to favour it, so the set is guaranteed to regress whether or not anything changed. That is the winner's curse, not market efficiency. Reporting it as decay would have been a confident, well-presented, wrong finding.
The honest measure has to be unselected: take every bucket, measure how far its up-rate sits from 50%, and compare eras. Counts need one more correction — the second era holds 22% more bars and so detects the same effect more easily, so its z is scaled by 0.905 before the two are compared.
| Market | 2003–2013 | 2014–2026 | Change | Significant, era 1 | Significant, era 2 |
|---|---|---|---|---|---|
| EUR/USD | 2.84pp | 2.84pp | 0% | 8 | 9 |
| GBP/USD | 2.50pp | 2.89pp | 16% | 3 | 8 |
| USD/JPY | 3.02pp | 2.74pp | -9% | 5 | 5 |
| AUD/USD | 3.02pp | 2.74pp | -9% | 14 | 5 |
| USD/CAD | 2.89pp | 2.76pp | -4% | 13 | 8 |
| NZD/USD | 3.11pp | 3.24pp | 4% | 11 | 8 |
| EUR/JPY | 2.38pp | 2.97pp | 24% | 1 | 8 |
| GBP/JPY | 2.60pp | 3.21pp | 23% | 5 | 9 |
| Gold (XAU/USD) | 3.13pp | 3.22pp | 3% | 13 | 14 |
| Silver (XAG/USD) | 5.98pp | 4.00pp | -33% | 41 | 19 |
| All ten | 3.15pp | 3.06pp | -3% | 114 | 93 |
| Excluding silver | 2.83pp | 2.96pp | 4% | 73 | 74 |
73 significant buckets in the first era against 74 in the second, measured without picking winners and with the power advantage removed. That is stability to within one bucket across 23 years.
Silver is the one market of the ten that genuinely decays: 5.98pp to 4.00pp, and 41 significant buckets down to 19. On its own it turns a 4% ex-silver figure into -3% across all ten. Part of that is a data-quality artifact and the numbers say so: silver's mean absolute hourly move was 1.32% in 2003, 1.21% in 2004, 0.97% in 2005, against a median of 0.23% across every year since 2012. A 5-fold difference in a metal whose realised volatility was not 5 times higher then looks like the early feed, not the market. Silver's early-era buckets should not be trusted, and silver is the one symbol where our pooled number overstates what is live today.
We are not measuring the same object. Those papers study day-of-year and day-of-month calendar anomalies in equity indices; we measure hour-of-week structure in FX. A January effect is a calendar quirk with no ongoing mechanism — once published, capital competes it away and it stays away. An hour-of-week effect is anchored to when desks are actually open: session opens, the liquidity cycle, the daily rollover. You cannot arbitrage away the fact that London and New York are both open at 14:00 UTC. That is a hypothesis consistent with our result, not something this test proves.
The best-specified claim in trading folklore is month-end rebalancing: index trackers benchmarked to MSCI and the major bond indices must rebalance their currency hedges when equity values move, and they transact at the WM/Reuters fix on the last trading day of the month. It has what the weaker folklore lacks — a named mechanism, a precise time, and a prediction about which markets should respond.
The control is the same fix hour on every other trading day, not an ordinary hour: the fix is itself elevated, and measuring against ordinary hours would credit month-end with the fix's own effect. Significance is a Welch t on log|move|, because hourly ranges are right-skewed enough that a raw-magnitude test reports whichever group caught the biggest handful of hours.
| Market | Month-end fix vs ordinary fix | t | Direction (z) |
|---|---|---|---|
| USD/CHF | 150% | 4.67 ✓ | 2.50 |
| EUR/USD | 133% | 4.40 ✓ | -2.09 |
| AUD/USD | 132% | 4.14 ✓ | -1.43 |
| GBP/USD | 133% | 3.82 ✓ | -0.26 |
| USD/CAD | 126% | 3.35 ✓ | -0.02 |
| USD/JPY | 129% | 3.03 ✓ | 1.45 |
| EUR/GBP | 138% | 2.92 ✓ | -1.98 |
| EUR/JPY | 117% | 2.06 ✓ | 0.01 |
| NZD/USD | 117% | 1.95 | -0.90 |
| GBP/JPY | 115% | 1.58 | 0.63 |
| GER40 (DAX) | 109% | 0.60 | -0.36 |
| Silver (XAG/USD) | 105% | 0.60 | -0.33 |
| Gold (XAU/USD) | 96% | -0.54 | -1.53 |
| US30 (Dow) | 94% | -0.54 | 0.22 |
Read the bottom rows first. Every one of the ten FX pairs leans positive and 8 clear t ≥ 2. Not one of the non-FX markets shows anything — GER40 (DAX), Silver (XAG/USD), Gold (XAU/USD), US30 (Dow) sit at or below their ordinary fix hour. That is the mechanism validating itself. Month-end rebalancing is currency-hedge rebalancing; it should flow through FX pairs and it should not touch gold or the equity indices themselves. It doesn't. A data-mined artifact would not respect an asset-class boundary. This is the strongest single piece of evidence we hold, and it comes from where the effect is absent.
The data is Dukascopy hourly bars, 2003–2026, aggregated into public cubes you can download from this site (example) — the same files the dashboard reads. The three tests on this page are scripts in the repository: test-snoop.js, test-era.js and test-monthend.js. Every table here is generated from their output, so a data refresh updates the page rather than the page being maintained by hand. Measured 2026-08-13.
The same discipline applied to other people's claims lives at /myths, and per-symbol write-ups at /blog.