Methodology

How the rotation reading is built

A point-in-time, survivorship-free dataset and a pre-registered, filtered state model — plus an onset flag for a new rotation and a panel of recognized gauges backtested against the rotation institutions actually executed. The same discipline behind asymmetricbeta.com, applied to sector rotation.

The question

Formally this is a state-estimation problem. The latent state is rotation phase; the observables are noisy sector-level returns and breadth. The deliverable is a calibrated posterior probability over rotation phase at daily frequency, using data through today only — with two kinds of noise removed first: market-wide movement, and the distortion from a single dominant constituent (a cap-weighted sector index is hostage to its largest holding).

Data — owned, not borrowed

The project owns its price feed on one consistent, single-source basis rather than reading a shared, multi-vendor table. That discipline is deliberate — mixing vendors with no cross-source resolution lets one bad print corrupt returns invisibly. Each return is computed within a single source; where two sources meet, they are joined by return, not by splicing incompatible levels.

  • Prices — 6.97M daily bars from a delisting-complete vendor, 1990-2026. The companies that left the index are still present, so the dataset is free of survivorship bias — the single most important property for this question, because names that leave the index are systematically the worst performers and cluster in exactly the sectors rotating out.
  • Membership & weights — S&P 500 constituents and float weights as they actually stood each quarter, joined by ISIN, never ticker (tickers silently recycle — the same symbol can point at two different companies over time).
  • Cross-checks — prices reconcile to a second independent feed at 99.85% agreement; cap-weighted sector returns reconcile to independent sector-ETF benchmarks at 0.978 correlation. Disagreements are flagged, never silently overwritten.

Denoising

Before modeling, each sector return is taken net of its own trailing beta to the market, leaving a market-denoised residual. Alongside it we compute breadth (how many members participate) and concentration (how much of the move is one name). A genuine rotation shows broad participation; a fake one driven by a single mega-cap does not — and the breadth/concentration measures are the direct discriminators.

The mechanical ruler

To detect a rotation you must first define one. The definition is pre-registered and frozen before any model is fit: cross-sectional dispersion above a trailing percentile, for a minimum run of days, with sufficient breadth in the leading sectors and low rank-correlation to the prior ordering. This mechanical rule is an external reference the model never sees — so "state 2 is a rotation" becomes checkable, not asserted.

The model

The reading comes from a discrete hidden Markov model (three states) over the sector observables. Two disciplines matter most:

  • No look-ahead. The probability at each date uses data through that date only — a filtered estimate. Smoothed (full-sample) states look dramatically better and are useless in real time; they are never used to score anything.
  • Calibration. The raw model score is mapped to an honest probability by isotonic regression fit on out-of-sample data, so "0.30" means what it says.

The model is unsupervised — it never sees the mechanical ruler. That one of its states independently lines up with the ruler is the evidence that the state is real rather than an artifact of fitting.

Validation

Rotations are rare, so a single train/test split cannot confirm anything. The evidence is built in two stages:

  • Blocked, purged cross-validation across 1990-2026 (the primary evidence) — contiguous blocks with a gap at each boundary so nothing leaks through the trailing windows. Median out-of-sample AUC 0.909.
  • A single sealed-holdout test on 2021+, run exactly once with everything frozen on the earlier data: AUC 0.8687, calibration error 0.0046.

Every result is reported against honest benchmarks — the unconditional base rate (14.6%), the mechanical rule alone, and a simple dispersion threshold — so the latent-state machinery has to earn its complexity. The model learned from 104 rotation episodes, spanning every regime a reader would name.

Is a new rotation starting? — the onset flag

"Are we in a rotation?" and "is a new one just beginning?" are different questions. The state model answers the first. The young-rotation flag answers the second. The frozen definition requires five consecutive qualifying days, so a rule-follower cannot know on day one — on day k<5 the flag forecasts the probability this nascent run reaches five and confirms. The training data was free: every past run that died at one to four days is a labelled negative the pipeline was otherwise discarding. On the sealed 2021+ holdout it scores AUC 0.812 against 0.736 for a run-length-only null — and beating that null is the entire test, because "runs that have lasted longer tend to continue" is just arithmetic.

Why onset belongs one grain down — and its honest limit. Industry leadership reorders before the sector aggregate does; on the sealed holdout the industry signal crossed first in the large majority of episodes. But the aggregate industry probability is descriptive, not predictive — it is elevated roughly half the time and carries essentially no forward lift on its own, so we publish it as texture, never as a starting-gun. The genuinely useful form of "which rotation is starting" is per-industry: naming the specific industries heating up right now. That is the direction we are validating next, and every per-industry signal is held to the same 13F backtest before it earns a place on the page.

Regime metrics — and how we backtested them

The reading above is one model. On the dashboard we also publish a panel of recognized rotation and regime measures, each computed from the same survivorship-free sector data and shown as its percentile versus its own history — so you can watch the gauges you trust rather than a single number:

  • Rotation Intensity — volume-spike AND turbulence — the 13F-backtested winner (ρ +0.51 vs 13F)
  • Turbulence — covariance-adjusted unusual moves (Chow/Kritzman) (ρ +0.43 vs 13F)
  • RRG spread — dispersion of sector relative-strength (de Kempenaer) (ρ +0.41 vs 13F)
  • Volume intensity — abnormal volume — the one input independent of price (ρ +0.35 vs 13F)
  • Dispersion — how far apart sector returns spread (ρ +0.29 vs 13F)
  • Leadership turnover — churn in the top-3 sectors (ρ -0.34 vs 13F)
  • Absorption ratio — PCA coupling — low = rotation-friendly (Kritzman) (ρ +0.08 vs 13F)
  • Avg correlation — realized pairwise — low = decoupled (ρ +0.05 vs 13F)

Why we backtested, and how. A rotation gauge earns its place only if it matches rotation that actually happened, so we built an independent yardstick from ground truth: the sector bets institutions actually made. From ~580 funds' quarterly 13F filings, 2013-2026 we reconstructed the real quarter-by-quarter rotation the crowd executed — the change in each sector's share of managers' consolidated holdings, net of the drift prices alone would cause — and scored each measure by how closely it tracked that real rotation (Spearman ρ) across 51 quarters. The ρ vs 13F shown on each gauge is that score.

To keep the search honest we held out-of-sample across 5 train/test splits (test rho 0.69-0.75), and we ran a shuffle-null — letting the threshold optimizer loose on scrambled targets to measure how much a good-looking score is pure luck. The winner, Rotation Intensity (days carrying both a volume spike and elevated turbulence), cleared a shuffle-null (p=0.014); a 3-signal gate overfit (train up, test down), tracking the real 13F rotation at ρ≈+0.51 — above the model's +0.33.

The model and Rotation Intensity — how they relate. They are not two estimates of the same thing. The model is a calibrated probability of a rotation regime (a state); Rotation Intensity is a percentile of realized rotation magnitude versus history, from volume and turbulence — inputs the model never sees. Because they are independent, agreement is genuine corroboration and divergence is diagnostic, not a bug: a calm model with elevated intensity is often realized churn running ahead of the regime signature, while an elevated model with quiet intensity is a dispersion pattern not yet backed by heavy volume. We deliberately do not collapse them into one number — a naive combination (gating the model on volume) tracked the actual 13F rotation worse than either alone (ρ 0.28 vs 0.51), and the threshold search, given the option, declined to gate on the model. Publishing both, each with its score, is the honest resolution.

Honest limits. The 13F yardstick is itself imperfect — institutions closely resemble the market and disclose ~45 days late — the sample is 51 quarters, and this measures rotation intensity (that it is happening), not which sector wins. Which gauge to trust is a choice; that is why the panel shows them all, each with its score.

What we claim — and don't

Detection is strong; forecasting is deliberately modest. Identifying whether we are in a rotation today is confirmed out-of-sample and well-calibrated. Predicting a rotation before it is visible is much harder — the forward skill is real but small and decays within weeks, and we publish that plainly rather than dress it up. A weak forecast reported honestly is worth more than a strong one that isn't true.

Research and educational content only. Not investment advice, not an offer or solicitation, and not a recommendation to buy or sell any security.