COTVault Premium — COT Flow

Every COT product on the market sells positioning flow as a signal in its own right. We tested that claim on 12 markets and 16,284 market-weeks and it does not survive. What survives is flow plus a trigger — two of the three combinations we tested, and this indicator prints the failure permanently beside the edge so the comparison is never out of sight.

1 · Quick start

Flow is reported in the futures contract's terms and converted here to chart terms: on USDJPY read off the yen leg — and on USDxxx generally — specs buying the contract is specs selling the chart, and the sign is flipped to match. On a cross, the base currency's leg is read alone.

2 · What flow actually is

Flow is urgency, not size. It is the weekly change in the large-speculator net position, expressed as a z-score against the previous 104 weekly changes. A reading of −2 means specs sold harder this week than in all but a few percent of the last two years. A very large position that did not move produces a flow of zero.

That distinction matters, and it is why Flow is a different instrument from COT Classic. Classic measures where the position sits. Flow measures how hard it moved to get there. They can disagree completely — a crowded position that stopped growing reads extreme on Classic and silent on Flow.

3 · The claim this tool refuses to make

Flow direction on its own is not an edge

The industry reading is that heavy spec buying should be followed and heavy spec selling faded, or some mirror of it. Tested across 12 markets, 2000–2026, on a 13-week horizon, measuring excess return against each market's own drift:

The industry claimnexcesshitmarketsverdict
Heavy SELLING (z ≤ −2), faded long485+0.36%50.9%8/12fails
Heavy BUYING (z ≥ +2), faded short525+0.03%51.0%5/12fails
4-week flow momentum, followed long757−0.50%46.6%5/12fails
4-week flow momentum, followed short706+0.27%50.0%6/12fails
3+ weeks of buying, followed long2,638−0.05%48.5%9/12fails
3+ weeks of selling, followed short2,663+0.10%50.7%6/12fails

None of them cleared the gauntlet. That null is not hidden in this manual — it is printed permanently on the data panel, on every chart, beside whatever figure is live. The comparison is the product.

4 · What does survive — flow plus a trigger

Three combinations cleared the first gauntlet: positive excess return, a 95% confidence interval excluding zero, at least 8 of 12 markets positive, and still positive after a 2010 walk-forward split. Two of them survive the stricter test below. The third does not, and ships switched off.

How these intervals are built, and why it matters

A 13-week claim means two signals a week apart share twelve of their thirteen weeks of return — they are very nearly the same observation. And EUR, GBP, AUD, NZD, CAD and CHF are six views of one dollar, so a week where all six fire is close to one observation, not six.

An ordinary bootstrap treats both as independent evidence and reports an interval several times too narrow. Every figure on this page instead resamples contiguous blocks of calendar weeks across all markets at once, which prices both in. It is a stricter test than anything else in this category applies — and it cost this indicator one of its three setups. The point estimates did not move; the confidence did.

Setupdirnexcesshitmarkets95% CIIS → OOS
Heavy BUYING + stochastic ≥ 80SHORT121+2.04%54.5%9/120.12 – 4.11+0.37 → +3.15
Heavy SELLING + pivot average turns upLONG63+2.05%66.7%10/120.47 – 3.55+4.70 → +0.52
Heavy SELLING + bullish 14/75/200 fan — not confirmed, ships offLONG243+1.06%53.9%9/12−0.43 – 2.58+1.91 → +0.11

Which of the two survivors to actually trust

The stochastic setup. It is the only one of the three that did not decay out of sample — it improved, from +0.37 fitted before 2010 to +3.15 on 73 fresh signals after it. Its sample is reasonable at n=121 and it clears zero under the stricter test, at p = 0.036. Note the lower bound sits at 0.12: the edge is real, and it is not large.

This indicator leads with it for that reason, and not because it carries the largest headline number.

Where the pretty number is the weak one

The pivot-average setup shows 66.7% and +2.05% — the best-looking pair of figures in the file. It rests on 63 signals, and it gives back almost all of its edge out of sample: +4.70 fitted before 2010 against +0.52 after. It ships, it is switched on, and the panel flags it as thin whenever it is the live setup. Of the two setups that survive, it is the weaker evidence — not the stronger.

The fan setup did not survive at all. Its point estimate is the same +1.06% it always was, on the largest sample of the three, but once overlapping windows are counted properly the interval runs −0.43 to 2.58 and spans zero, at p = 0.17. The walk-forward said the same thing all along: +1.91 fitted before 2010 against +0.11 after. Under an ordinary bootstrap it cleared, and it was originally shipped switched on. It now ships off, and remains as a toggle only so the pattern stays visible.

5 · The triggers, precisely

TriggerDefinition
Pivot average turns upThe 3-bar average of (H+L+C)/3 crosses above its own previous value. A cross, not merely "higher than last week".
Stochastic overbought9-3-3 stochastic K ≥ 80.
Bullish fanEMA 14 > EMA 75 > EMA 200, stacked.

All three are computed on weekly data whatever your chart period, so the reading does not change when you switch timeframes. Every setup is evaluated on a report bar, because every row of the study is a weekly report row — evaluating between reports would count the same extreme several times on a daily chart and inflate the live tally.

6 · Measured and rejected — the three mirrors

Each surviving setup has an obvious mirror. All three were tested and none of them work, so all three ship off, behind a single toggle that exists so you can turn them on and watch them fail.

Mirrornexcessmarketsverdict
Heavy BUYING + pivot average turns down, short45−0.11%9/12negative
Heavy SELLING + stochastic ≤ 20, long76−0.65%6/12negative
Heavy BUYING + bearish fan, short203+0.43%6/12fails breadth & CI

When they are on, a mirror fires as a grey cross — deliberately colourless, because slate in this suite means measured and rejected.

7 · Reading the pane

The data panel, row by row

RowWhat it shows
Flow zThis week's flow z-score, a meter of its extremity, and the state in words: HEAVY BUYING, HEAVY SELLING or an ordinary week.
TriggersAll three triggers at once — pivot average, stochastic, fan. Each is coloured only when it coincides with the flow extreme its setup needs.
Weekly changeThe raw change in net contracts behind the z-score, in chart terms.
The claimThe published excess, hit rate, breadth and confidence interval for the live setup — or for the headline setup when nothing is running. The line below it carries n and the walk-forward split.
ClockHow far through its 13 reports a running setup is. The claim is a 13-week claim; this is the window it was measured over, counted in reports rather than chart bars.
on this chartThe live tally for this symbol — see §9, it is raw direction and not excess.
Flow ALONE / buy sideThe permanent null. The sell-side and buy-side figures for flow with no trigger, printed whether or not a setup is live.

Armed is not a setup

When flow passes the threshold and no trigger comes with it, the pane still washes gold — because gold marks a state, and heavy flow is a state. It is not an instruction. That is the exact condition the rest of the industry sells as a signal, and the exact condition that measured +0.36%.

8 · The colour law

ColourMeaning
GOLDA state. Flow is at an extreme. Never a direction.
TEALA measured bullish edge is live.
REDA measured bearish edge is live.
SLATENeutral — or measured and rejected.

9 · The on-chart tally is not the published figure

They measure different things, on purpose

The panel's on this chart row resolves each finished setup on raw direction over 13 reports: did price finish higher for a long, lower for a short. The published figure is excess return against each market's own long-run drift — which a chart cannot compute, because it has no access to the other eleven markets or to the full 26-year sample.

Expect the two to differ, sometimes widely on a single market. The live tally will usually read higher, because raw direction keeps whatever drift the market had over the period while the published figure subtracts it. A chart showing 72% against a published 54.5% is not the indicator beating its own claim — it is a market that trended, measured with the drift left in.

The published column is the claim; the live tally is a sanity check on this one symbol.

10 · Alerts

AlertFires when
SHORT setup firesHeavy buying and the stochastic is overbought, on a report bar. The headline setup.
LONG setup fires (pivot average)Heavy selling and the pivot average turns up. Message carries the n=63 caveat.
LONG pattern fires (fan — not confirmed)Heavy selling with the fan stacked bullish. Fires only if you switch the setup on; the alert message carries the not-confirmed caveat.
Flow extreme with NO triggerThe armed state. Included so you can watch how often it leads nowhere.
New COT report ingestedThe weekly CFTC release has landed.

11 · Panel states you may see

12 · Data and honesty notes

← COTVault — the full COT dashboard
COTVault  ·  COTVAULT.COM