How Much Do Time-Series Anomaly Detection Metrics Actually Disagree?
Null Models and Cluster-Aware Inference for Rank-Flip Statistics
The audit recomputes TSB-AD-M, 25 models on 180 series drawn from 17 source collections, and a six-dataset, seven-detector grid. Every number on this page is regenerated from committed artifacts by one command.
Finding 1Against its chance level, the metrics mostly agree
A pairwise flip rate compares two orderings of the same models. Two metrics that share no information at all would disagree on half the pairs, so the reference point is 0.50, not 0. Permuting one metric across models 200 times gives a null of 0.5002.
Finding 2Most of the disagreement is near-ties
Flips are usually pooled across all pairs, including pairs the metrics can barely tell apart. Conditioning on one metric's gap is not enough: the flips that survive sit where the other metric barely separates the pair. Requiring both metrics to separate a pair by 0.20 keeps 13.3% of pairs, and 2.2% of those flip.
Finding 3Evaluation series are not independent
The short-anomaly ratio, the covariate most often used to explain disagreement, takes one value inside 11 of the 12 collections that hold more than one series. It is a collection label, so a series-level correlation counts the same collection many times. With collections as the unit, the correlation disappears and the interval widens.
Finding 4The same test, applied to our own metric
We had earlier proposed a composite metric that interpolates between the point and segment scores. At both ends of its weight it collapses to one of its inputs, so stratifying results by that weight measures the definition rather than the data. We report it as uninformative by construction. After the same corrections, one covariate still stands with collections as the unit: segment count, ρ = −0.62 (p = 0.008).
ProtocolSix checks that cost no new experiments
- Report the null. State the random-ranking baseline next to any disagreement statistic. For pairwise flip rates it is 0.5.
- Stratify by both margins. Report the flip rate among pairs both metrics separate by a stated threshold, and the share of pairs that keeps.
- Cluster at the source. Resample source collections or datasets, not series, and report the clustered interval.
- Check identifiability first. Before explaining disagreement with a covariate, check that it varies within clusters.
- Leave one collection out. Report how a structural association moves when each collection is removed in turn.
- Do not stratify a composite by its own weight. At the extremes it is one of its inputs.
ReproduceEvery number from committed artifacts
No dataset download is needed. The command below recomputes the audit, regenerates the manuscript's numbers, and fails if anything drifts from what is committed.
git clone https://github.com/mandu5/structure-aware-tsad-evaluation cd structure-aware-tsad-evaluation pip install -e . make verify
Raw access-controlled datasets are not redistributed; see data access.
CiteCitation
@misc{ko2026metricsdisagree,
title = {How Much Do Time-Series Anomaly Detection Metrics Actually Disagree?
Null Models and Cluster-Aware Inference for Rank-Flip Statistics},
author = {Ko, Youngmin},
year = {2026},
note = {Under review at the Journal of Data-centric Machine Learning Research (DMLR).
Code: https://github.com/mandu5/structure-aware-tsad-evaluation}
}
Earlier version
@misc{ko2026pointmetricsmislead,
title={When Point Metrics Mislead: Structure-Aware Evaluation Reveals Conditional Ranking Shifts in Time Series Anomaly Detection},
author={Ko, Youngmin},
year={2026},
note={Manuscript. Submitted to NeurIPS 2026 Evaluations and Datasets Track. Superseded by a reframed manuscript under review at DMLR (Journal of Data-centric Machine Learning Research), submitted September 2026}
}