Every COT product on the market sells positioning flow as a signal in its own right. We tested that claim on 12 markets and 16,284 market-weeks and it does not survive. What survives is flow plus a trigger — two of the three combinations we tested, and COTVault prints the failure permanently beside the edge so the comparison is never out of sight.
Flow is urgency, not size. It is the weekly change in the large-speculator net position, expressed as a z-score against the previous 104 weekly changes. A reading of −2 means specs sold harder this week than in all but a few percent of the last two years. A very large position that did not move produces a flow of zero.
That distinction matters, and it is why Flow is a different instrument from COT Classic. Classic measures where the position sits. Flow measures how hard it moved to get there. They can disagree completely — a crowded position that stopped growing reads extreme on Classic and silent on Flow.
The industry reading is that heavy spec buying should be followed and heavy spec selling faded, or some mirror of it. Tested across 12 markets, 2000–2026, on a 13-week horizon, measuring excess return against each market's own drift:
| The industry claim | n | excess | hit | markets | verdict |
|---|---|---|---|---|---|
| Heavy SELLING (z ≤ −2), faded long | 485 | +0.36% | 50.9% | 8/12 | fails |
| Heavy BUYING (z ≥ +2), faded short | 525 | +0.03% | 51.0% | 5/12 | fails |
| 4-week flow momentum, followed long | 757 | −0.50% | 46.6% | 5/12 | fails |
| 4-week flow momentum, followed short | 706 | +0.27% | 50.0% | 6/12 | fails |
| 3+ weeks of buying, followed long | 2,638 | −0.05% | 48.5% | 9/12 | fails |
| 3+ weeks of selling, followed short | 2,663 | +0.10% | 50.7% | 6/12 | fails |
None of them cleared the gauntlet. That null is not hidden in this manual — it is printed permanently on the data panel, on every chart, beside whatever figure is live. The comparison is the product.
Three combinations cleared the first gauntlet: positive excess return, a 95% confidence interval excluding zero, at least 8 of 12 markets positive, and still positive after a 2010 walk-forward split. Two of them survive the stricter test below. The third does not, and ships switched off.
A 13-week claim means two signals a week apart share twelve of their thirteen weeks of return — they are very nearly the same observation. And EUR, GBP, AUD, NZD, CAD and CHF are six views of one dollar, so a week where all six fire is close to one observation, not six.
An ordinary bootstrap treats both as independent evidence and reports an interval several times too narrow. Every figure on this page instead resamples contiguous blocks of calendar weeks across all markets at once, which prices both in. It is a stricter test than anything else in this category applies — and it cost one of the three surviving setups. The point estimates did not move; the confidence did.
| Setup | dir | n | excess | hit | markets | 95% CI | IS → OOS |
|---|---|---|---|---|---|---|---|
| Heavy BUYING + stochastic ≥ 80 | SHORT | 121 | +2.04% | 54.5% | 9/12 | 0.12 – 4.11 | +0.37 → +3.15 |
| Heavy SELLING + pivot average turns up | LONG | 63 | +2.05% | 66.7% | 10/12 | 0.47 – 3.55 | +4.70 → +0.52 |
| Heavy SELLING + bullish 14/75/200 fan — not confirmed, ships off | LONG | 243 | +1.06% | 53.9% | 9/12 | −0.43 – 2.58 | +1.91 → +0.11 |
The stochastic setup. It is the only one of the three that did not decay out of sample — it improved, from +0.37 fitted before 2010 to +3.15 on 73 fresh signals after it. Its sample is reasonable at n=121 and it clears zero under the stricter test, at p = 0.036. Note the lower bound sits at 0.12: the edge is real, and it is not large.
COTVault leads with it for that reason, and not because it carries the largest headline number.
The pivot-average setup shows 66.7% and +2.05% — the best-looking pair of figures in the file. It rests on 63 signals, and it gives back almost all of its edge out of sample: +4.70 fitted before 2010 against +0.52 after. It ships, it is switched on, and the panel flags it as thin whenever it is the live setup. Of the two setups that survive, it is the weaker evidence — not the stronger.
The fan setup did not survive at all. Its point estimate is the same +1.06% it always was, on the largest sample of the three, but once overlapping windows are counted properly the interval runs −0.43 to 2.58 and spans zero, at p = 0.17. The walk-forward said the same thing all along: +1.91 fitted before 2010 against +0.11 after. Under an ordinary bootstrap it cleared, and it was originally shipped switched on. It now ships off, and remains as a toggle only so the pattern stays visible.
| Trigger | Definition |
|---|---|
| Pivot average turns up | The 3-bar average of (H+L+C)/3 crosses above its own previous value. A cross, not merely "higher than last week". |
| Stochastic overbought | 9-3-3 stochastic K ≥ 80. |
| Bullish fan | EMA 14 > EMA 75 > EMA 200, stacked. |
All three are computed on weekly data whatever your chart period, so the reading does not change when you switch timeframes. Every setup is evaluated on a report bar, because every row of the study is a weekly report row — evaluating between reports would count the same extreme several times on a daily chart and inflate the live tally.
Each surviving setup has an obvious mirror. All three were tested and none of them work, so all three ship off, behind a single toggle that exists so you can turn them on and watch them fail.
| Mirror | n | excess | markets | verdict |
|---|---|---|---|---|
| Heavy BUYING + pivot average turns down, short | 45 | −0.11% | 9/12 | negative |
| Heavy SELLING + stochastic ≤ 20, long | 76 | −0.65% | 6/12 | negative |
| Heavy BUYING + bearish fan, short | 203 | +0.43% | 6/12 | fails breadth & CI |
When they are on, a mirror fires as a grey cross — deliberately colourless, because slate in this suite means measured and rejected.