COT net position change, week over week

Every COT product on the market sells positioning flow as a signal in its own right. We tested that claim on 12 markets and 16,284 market-weeks and it does not survive. What survives is flow plus a trigger — two of the three combinations we tested, and COTVault prints the failure permanently beside the edge so the comparison is never out of sight.

1 · What flow actually is

Flow is urgency, not size. It is the weekly change in the large-speculator net position, expressed as a z-score against the previous 104 weekly changes. A reading of −2 means specs sold harder this week than in all but a few percent of the last two years. A very large position that did not move produces a flow of zero.

That distinction matters, and it is why Flow is a different instrument from COT Classic. Classic measures where the position sits. Flow measures how hard it moved to get there. They can disagree completely — a crowded position that stopped growing reads extreme on Classic and silent on Flow.

2 · The claim this tool refuses to make

Flow direction on its own is not an edge

The industry reading is that heavy spec buying should be followed and heavy spec selling faded, or some mirror of it. Tested across 12 markets, 2000–2026, on a 13-week horizon, measuring excess return against each market's own drift:

The industry claimnexcesshitmarketsverdict
Heavy SELLING (z ≤ −2), faded long485+0.36%50.9%8/12fails
Heavy BUYING (z ≥ +2), faded short525+0.03%51.0%5/12fails
4-week flow momentum, followed long757−0.50%46.6%5/12fails
4-week flow momentum, followed short706+0.27%50.0%6/12fails
3+ weeks of buying, followed long2,638−0.05%48.5%9/12fails
3+ weeks of selling, followed short2,663+0.10%50.7%6/12fails

None of them cleared the gauntlet. That null is not hidden in this manual — it is printed permanently on the data panel, on every chart, beside whatever figure is live. The comparison is the product.

3 · What does survive — flow plus a trigger

Three combinations cleared the first gauntlet: positive excess return, a 95% confidence interval excluding zero, at least 8 of 12 markets positive, and still positive after a 2010 walk-forward split. Two of them survive the stricter test below. The third does not, and ships switched off.

How these intervals are built, and why it matters

A 13-week claim means two signals a week apart share twelve of their thirteen weeks of return — they are very nearly the same observation. And EUR, GBP, AUD, NZD, CAD and CHF are six views of one dollar, so a week where all six fire is close to one observation, not six.

An ordinary bootstrap treats both as independent evidence and reports an interval several times too narrow. Every figure on this page instead resamples contiguous blocks of calendar weeks across all markets at once, which prices both in. It is a stricter test than anything else in this category applies — and it cost one of the three surviving setups. The point estimates did not move; the confidence did.

Setupdirnexcesshitmarkets95% CIIS → OOS
Heavy BUYING + stochastic ≥ 80SHORT121+2.04%54.5%9/120.12 – 4.11+0.37 → +3.15
Heavy SELLING + pivot average turns upLONG63+2.05%66.7%10/120.47 – 3.55+4.70 → +0.52
Heavy SELLING + bullish 14/75/200 fan — not confirmed, ships offLONG243+1.06%53.9%9/12−0.43 – 2.58+1.91 → +0.11

Which of the two survivors to actually trust

The stochastic setup. It is the only one of the three that did not decay out of sample — it improved, from +0.37 fitted before 2010 to +3.15 on 73 fresh signals after it. Its sample is reasonable at n=121 and it clears zero under the stricter test, at p = 0.036. Note the lower bound sits at 0.12: the edge is real, and it is not large.

COTVault leads with it for that reason, and not because it carries the largest headline number.

Where the pretty number is the weak one

The pivot-average setup shows 66.7% and +2.05% — the best-looking pair of figures in the file. It rests on 63 signals, and it gives back almost all of its edge out of sample: +4.70 fitted before 2010 against +0.52 after. It ships, it is switched on, and the panel flags it as thin whenever it is the live setup. Of the two setups that survive, it is the weaker evidence — not the stronger.

The fan setup did not survive at all. Its point estimate is the same +1.06% it always was, on the largest sample of the three, but once overlapping windows are counted properly the interval runs −0.43 to 2.58 and spans zero, at p = 0.17. The walk-forward said the same thing all along: +1.91 fitted before 2010 against +0.11 after. Under an ordinary bootstrap it cleared, and it was originally shipped switched on. It now ships off, and remains as a toggle only so the pattern stays visible.

4 · The triggers, precisely

TriggerDefinition
Pivot average turns upThe 3-bar average of (H+L+C)/3 crosses above its own previous value. A cross, not merely "higher than last week".
Stochastic overbought9-3-3 stochastic K ≥ 80.
Bullish fanEMA 14 > EMA 75 > EMA 200, stacked.

All three are computed on weekly data whatever your chart period, so the reading does not change when you switch timeframes. Every setup is evaluated on a report bar, because every row of the study is a weekly report row — evaluating between reports would count the same extreme several times on a daily chart and inflate the live tally.

5 · Measured and rejected — the three mirrors

Each surviving setup has an obvious mirror. All three were tested and none of them work, so all three ship off, behind a single toggle that exists so you can turn them on and watch them fail.

Mirrornexcessmarketsverdict
Heavy BUYING + pivot average turns down, short45−0.11%9/12negative
Heavy SELLING + stochastic ≤ 20, long76−0.65%6/12negative
Heavy BUYING + bearish fan, short203+0.43%6/12fails breadth & CI

When they are on, a mirror fires as a grey cross — deliberately colourless, because slate in this suite means measured and rejected.

6 · Data and honesty notes

← COTVault — the full COT dashboard
COTVault research  ·  see it live on the COT Flow board  ·  not investment advice