Evaluation audit · time-series anomaly detection

How Much Do Time-Series Anomaly Detection Metrics Actually Disagree?

Null Models and Cluster-Aware Inference for Rank-Flip Statistics

Youngmin Ko · Under review at the Journal of Data-centric Machine Learning Research (DMLR), submitted September 2026
Point-level and segment-level metrics are said to rank anomaly detectors differently about a third of the time. That number is usually read against zero. Read against its chance level of 0.50, the two metrics agree on 68.6% of model pairs, and where both metrics clearly separate a pair they disagree on about 2%.
0.31 vs 0.50Observed rank-flip rate against a random-ranking null
2.2%Flip rate among pairs both metrics separate by at least 0.20
11 of 12Multi-series collections where the usual explanatory covariate never varies

The audit recomputes TSB-AD-M, 25 models on 180 series drawn from 17 source collections, and a six-dataset, seven-detector grid. Every number on this page is regenerated from committed artifacts by one command.

Finding 1Against its chance level, the metrics mostly agree

A pairwise flip rate compares two orderings of the same models. Two metrics that share no information at all would disagree on half the pairs, so the reference point is 0.50, not 0. Permuting one metric across models 200 times gives a null of 0.5002.

00.10.20.30.40.50.6 observed 0.3145 chance 0.5002 confident 0.022 shaded: 95% interval clustered over source collections
TSB-AD-M, AUC-ROC against Affiliation-F1: 15,151 of 48,180 model pairs flip. The observed rate sits well below chance.

Finding 2Most of the disagreement is near-ties

Flips are usually pooled across all pairs, including pairs the metrics can barely tell apart. Conditioning on one metric's gap is not enough: the flips that survive sit where the other metric barely separates the pair. Requiring both metrics to separate a pair by 0.20 keeps 13.3% of pairs, and 2.2% of those flip.

All pairs AUC-ROC gap < 0.01 AUC-ROC gap ≥ 0.20 Both metrics ≥ 0.20 apart 0.3145 0.4743 0.2015 0.0220 chance 0.50
Flip rate within margin strata on TSB-AD-M. Near-ties flip at close to chance; confidently separated pairs almost never do.

Finding 3Evaluation series are not independent

The short-anomaly ratio, the covariate most often used to explain disagreement, takes one value inside 11 of the 12 collections that hold more than one series. It is a collection label, so a series-level correlation counts the same collection many times. With collections as the unit, the correlation disappears and the interval widens.

Correlation of the short-anomaly ratio with the flip rate by seriesby collectionone collection left out 0.324 0.056 (p = 0.827) 0.090 95% interval for the flip rate 00.10.20.30.4 resampling seriesclustering collectionssix-dataset grid
Clustering over collections is the honest unit here. On the six-dataset grid the clustered interval runs from 0.07 to 0.42, too wide to state a magnitude.

Finding 4The same test, applied to our own metric

We had earlier proposed a composite metric that interpolates between the point and segment scores. At both ends of its weight it collapses to one of its inputs, so stratifying results by that weight measures the definition rather than the data. We report it as uninformative by construction. After the same corrections, one covariate still stands with collections as the unit: segment count, ρ = −0.62 (p = 0.008).

ProtocolSix checks that cost no new experiments

  1. Report the null. State the random-ranking baseline next to any disagreement statistic. For pairwise flip rates it is 0.5.
  2. Stratify by both margins. Report the flip rate among pairs both metrics separate by a stated threshold, and the share of pairs that keeps.
  3. Cluster at the source. Resample source collections or datasets, not series, and report the clustered interval.
  4. Check identifiability first. Before explaining disagreement with a covariate, check that it varies within clusters.
  5. Leave one collection out. Report how a structural association moves when each collection is removed in turn.
  6. Do not stratify a composite by its own weight. At the extremes it is one of its inputs.

ReproduceEvery number from committed artifacts

No dataset download is needed. The command below recomputes the audit, regenerates the manuscript's numbers, and fails if anything drifts from what is committed.

git clone https://github.com/mandu5/structure-aware-tsad-evaluation
cd structure-aware-tsad-evaluation
pip install -e .
make verify

Raw access-controlled datasets are not redistributed; see data access.

CiteCitation

@misc{ko2026metricsdisagree,
  title  = {How Much Do Time-Series Anomaly Detection Metrics Actually Disagree?
            Null Models and Cluster-Aware Inference for Rank-Flip Statistics},
  author = {Ko, Youngmin},
  year   = {2026},
  note   = {Under review at the Journal of Data-centric Machine Learning Research (DMLR).
            Code: https://github.com/mandu5/structure-aware-tsad-evaluation}
}
Earlier version
@misc{ko2026pointmetricsmislead,
  title={When Point Metrics Mislead: Structure-Aware Evaluation Reveals Conditional Ranking Shifts in Time Series Anomaly Detection},
  author={Ko, Youngmin},
  year={2026},
  note={Manuscript. Submitted to NeurIPS 2026 Evaluations and Datasets Track. Superseded by a reframed manuscript under review at DMLR (Journal of Data-centric Machine Learning Research), submitted September 2026}
}