This paper shows that near-perfect aggregate IDS metrics can mask complete failure on a supported attack category. Using 60 reproducible, leakage-controlled experiments, it establishes transparent partitioning and per-category reporting as essential for valid ML-based IDS benchmarking.
