Evidence

The three tests that could have killed this product

Anyone selling market seasonality should be able to answer three questions: is your grid just data-mining, has your edge already decayed, and does anything you found have a reason to exist? We ran all three against our own published numbers. Two came back in our favour. The first came back against us in 7 of our 19 markets, and every one of them is named below.

187significant buckets found
14expected from pure noise
13.7×the ratio
7markets that fail our own test

1. Is the grid data-mined?

The most damaging paper to a product like this one is Sullivan, Timmermann & White, Dangers of data mining: the case of calendar effects in stock returns (Journal of Econometrics, 2001). They built a large universe of calendar rules, applied White's Reality Check, and showed that calendar effects which look decisive alone are far weaker once you account for every rule that could have been tested. That is a precise description of what we do: we search a weekday × hour grid across 19 markets and publish the survivors.

So we built the null. It preserves each bucket's real sample size and each symbol's own overall up-rate, destroys only the weekday-hour structure — the one thing we claim exists — and draws 4,000 times. The second constraint matters more than it looks: we measure z against a flat 50%, so a market that closed up 52.3% of all hours has every bucket shifted upward, and that alone manufactures significant hits with no hour-of-week structure whatsoever. Drawing the null at the symbol's own rate charges us for that.

MarketBuckets testedOwn up-rateFoundNoise expectsp
EUR/USD12150.1% 110.3<0.00025
GBP/USD12149.9% 100.3<0.00025
USD/CHF12049.9% 160.3<0.00025
USD/JPY12150.7% 90.6<0.00025
AUD/USD12150.1% 120.3<0.00025
USD/CAD12149.8% 160.4<0.00025
NZD/USD12150.0% 130.3<0.00025
EUR/JPY12150.9% 70.9<0.00025
GBP/JPY12051.1% 91.2<0.00025
EUR/GBP12049.2% 160.5<0.00025
Gold (XAU/USD)12051.0% 221.3<0.00025
Silver (XAG/USD)12049.5% 350.5<0.00025

All 12 FX and metals markets clear the null by a wide margin — and clear it at 0.0028, which is 0.05 divided by the 18 markets tested, so the win is not an artifact of running the test 18 times. A grid that had merely been searched hard would land near the "noise expects" column. This one is nowhere near it.

The 7 markets we fail

Publishing only the first table would be the exact behaviour the paper warns about. Here is the rest of it.

MarketFoundNoise expectsp
BTC/USD30.8 0.0432marginal
ETH/USD20.5 0.0818fails
NAS10031.8 0.2595fails
SPX50021.6 0.4955fails
GER40 (DAX)11.0 0.6362fails
US30 (Dow)01.1 1.0000fails
UK100 (FTSE)not one bucket reaches our 300-sample minimum — largest 193cannot be tested

BTC/USD sits under 0.05 alone but not under the corrected bar, which is what "marginal" means here: interesting, not established. We are not going to call it a result.

The four testable equity indices are indistinguishable from a noise grid of the same shape. US30 comes back at p = 1.0000 — it found 0 significant buckets where noise alone expects 1.1. Note also that NAS100 and SPX500 both close up 52.3% of all hours, so their null means are the highest in the table: much of the little they do show is upward drift being read as hour-of-week structure. The null caught that, which is why it was built that way.

This is a power result, not an absence result, and the distinction is the honest part. An index would need a bucket leaning 58% before our own bar could call it. Real effects of three to five points could sit in the indices untouched and this test would never see them. "We cannot demonstrate hour-of-week structure in equity indices" is the correct claim. "There is none" is not.

Nothing at 51% ever reaches our board

The same arithmetic runs in the opposite direction, and it is the plainest answer to why our numbers look boring next to a screener's. Our bar is |z| ≥ 3 on at least 300 samples, which fixes the smallest lean each market is capable of flagging at its median bucket:

MarketMedian bucketSmallest lean we can flag
Gold (XAU/USD)n = 1,205 54.3%
USD/CADn = 1,184 54.4%
EUR/USDn = 1,180 54.4%
AUD/USDn = 1,146 54.4%
EUR/JPYn = 1,044 54.6%
GBP/USDn = 1,040 54.7%
Silver (XAG/USD)n = 1,034 54.7%
NZD/USDn = 1,031 54.7%
USD/JPYn = 910 55.0%
GBP/JPYn = 831 55.2%
USD/CHFn = 644 55.9%
US30 (Dow)n = 567 56.3%
GER40 (DAX)n = 532 56.5%
EUR/GBPn = 487 56.8%
BTC/USDn = 429 57.2%
ETH/USDn = 410 57.4%
SPX500n = 341 58.1%
NAS100n = 340 58.1%
UK100 (FTSE) 61.3%

A bucket at 51% cannot reach our board however real it is — the sample cannot carry it. That is a limitation, and it is also the reason a number that does reach the board is worth reading.

UK100 (FTSE) cannot produce a single publishable number

We sell a market whose data cannot clear our own bar. Its largest weekday × hour bucket holds 193 samples against our 300 minimum, across 5 covered years. The dashboard already says so twice — the signals card reports that nothing clears the bar, and the tile reads "Years with data: 5" — but it belongs here too, stated plainly rather than left for a user to discover. This site does not offer 19 equally-supported markets. The FX majors and metals are where this dataset is strong.

2. Has the edge already decayed?

The literature is blunt about this. Schwert found the weekend, size and value effects weakening after the papers documenting them were published (NBER w9277); McLean & Pontiff tracked 97 anomalies and found systematic post-publication decay; Urquhart & McGroarty report that calendar anomalies have essentially gone since the 1980s. We publish one pooled number per bucket across 2003–2026, so an effect that died in 2015 would still show — and our sample-size column, the site's main honesty device, would make a dead edge look more trustworthy rather than less.

Split at 2014. First the harsh, selected question: of the buckets the product flags today, how many survive on the recent half alone?

143buckets we flag today
95%still lean the same way
61%still clear |z| ≥ 3 on half the data

The second figure is stronger than it looks: halving a sample cuts z by about √2 even when the effect is perfectly stable, so a bucket at z = 3.5 pooled is expected to land near 2.5 on one half and fail this test while being entirely real.

The trap we nearly published

The same run reports that of the buckets significant in the first era, only 18% are still significant in the second. Read cold, that is a decay story. It is not one. Those buckets were selected on first-era data: any bucket near the boundary gets into that set only when first-era noise happens to favour it, so the set is guaranteed to regress whether or not anything changed. That is the winner's curse, not market efficiency. Reporting it as decay would have been a confident, well-presented, wrong finding.

The honest measure has to be unselected: take every bucket, measure how far its up-rate sits from 50%, and compare eras. Counts need one more correction — the second era holds 22% more bars and so detects the same effect more easily, so its z is scaled by 0.905 before the two are compared.

Market2003–20132014–2026ChangeSignificant, era 1Significant, era 2
EUR/USD2.84pp 2.84pp0% 89
GBP/USD2.50pp 2.89pp16% 38
USD/JPY3.02pp 2.74pp-9% 55
AUD/USD3.02pp 2.74pp-9% 145
USD/CAD2.89pp 2.76pp-4% 138
NZD/USD3.11pp 3.24pp4% 118
EUR/JPY2.38pp 2.97pp24% 18
GBP/JPY2.60pp 3.21pp23% 59
Gold (XAU/USD)3.13pp 3.22pp3% 1314
Silver (XAG/USD)5.98pp 4.00pp-33% 4119
All ten3.15pp3.06pp -3%11493
Excluding silver2.83pp2.96pp 4%7374

73 significant buckets in the first era against 74 in the second, measured without picking winners and with the power advantage removed. That is stability to within one bucket across 23 years.

Silver is our weakest data, and we would rather say it first

Silver is the one market of the ten that genuinely decays: 5.98pp to 4.00pp, and 41 significant buckets down to 19. On its own it turns a 4% ex-silver figure into -3% across all ten. Part of that is a data-quality artifact and the numbers say so: silver's mean absolute hourly move was 1.32% in 2003, 1.21% in 2004, 0.97% in 2005, against a median of 0.23% across every year since 2012. A 5-fold difference in a metal whose realised volatility was not 5 times higher then looks like the early feed, not the market. Silver's early-era buckets should not be trusted, and silver is the one symbol where our pooled number overstates what is live today.

Why we differ from the literature, without claiming to refute it

We are not measuring the same object. Those papers study day-of-year and day-of-month calendar anomalies in equity indices; we measure hour-of-week structure in FX. A January effect is a calendar quirk with no ongoing mechanism — once published, capital competes it away and it stays away. An hour-of-week effect is anchored to when desks are actually open: session opens, the liquidity cycle, the daily rollover. You cannot arbitrage away the fact that London and New York are both open at 14:00 UTC. That is a hypothesis consistent with our result, not something this test proves.

3. Does anything we found have a reason to exist?

The best-specified claim in trading folklore is month-end rebalancing: index trackers benchmarked to MSCI and the major bond indices must rebalance their currency hedges when equity values move, and they transact at the WM/Reuters fix on the last trading day of the month. It has what the weaker folklore lacks — a named mechanism, a precise time, and a prediction about which markets should respond.

The control is the same fix hour on every other trading day, not an ordinary hour: the fix is itself elevated, and measuring against ordinary hours would credit month-end with the fix's own effect. Significance is a Welch t on log|move|, because hourly ranges are right-skewed enough that a raw-magnitude test reports whichever group caught the biggest handful of hours.

MarketMonth-end fix vs ordinary fixtDirection (z)
USD/CHF 150% 4.67 ✓2.50
EUR/USD 133% 4.40 ✓-2.09
AUD/USD 132% 4.14 ✓-1.43
GBP/USD 133% 3.82 ✓-0.26
USD/CAD 126% 3.35 ✓-0.02
USD/JPY 129% 3.03 ✓1.45
EUR/GBP 138% 2.92 ✓-1.98
EUR/JPY 117% 2.06 ✓0.01
NZD/USD 117% 1.95-0.90
GBP/JPY 115% 1.580.63
GER40 (DAX) 109% 0.60-0.36
Silver (XAG/USD) 105% 0.60-0.33
Gold (XAU/USD) 96% -0.54-1.53
US30 (Dow) 94% -0.540.22

Read the bottom rows first. Every one of the ten FX pairs leans positive and 8 clear t ≥ 2. Not one of the non-FX markets shows anything — GER40 (DAX), Silver (XAG/USD), Gold (XAU/USD), US30 (Dow) sit at or below their ordinary fix hour. That is the mechanism validating itself. Month-end rebalancing is currency-hedge rebalancing; it should flow through FX pairs and it should not touch gold or the equity indices themselves. It doesn't. A data-mined artifact would not respect an asset-class boundary. This is the strongest single piece of evidence we hold, and it comes from where the effect is absent.

Busier is not the same as up. On direction, only 2 of 14 markets reach |z| ≥ 2 and they point opposite ways — about what 14 tests produce by chance. Month-end tells you the hour will be busier. It does not tell you which way. The dashboard flags it as activity, and says so.

Two parts of the folklore do not survive. "GBP/USD is the preferred expression" is not supported — GBP/USD and EUR/USD are 133% and 133%, and the largest effect is USD/CHF. And the documented flow concentrates in a five-minute window while we measure a whole hour, so what we see is a diluted shadow of the real thing.

What we will not publish

How to check us

The data is Dukascopy hourly bars, 2003–2026, aggregated into public cubes you can download from this site (example) — the same files the dashboard reads. The three tests on this page are scripts in the repository: test-snoop.js, test-era.js and test-monthend.js. Every table here is generated from their output, so a data refresh updates the page rather than the page being maintained by hand. Measured 2026-08-13.

The same discipline applied to other people's claims lives at /myths, and per-symbol write-ups at /blog.

See the numbers that survived →

Free on the four major pairs, full 23 years, with the sample size and a significance flag on every figure.