This paper’s principal contribution is to demonstrate, on the DataSense IIoT benchmark, that overall accuracy is an insufficient basis for evaluating intrusion-detection models and to provide a class-wise, error-oriented evaluation framework that surfaces deployment-critical failures which aggregate metrics conceal—most notably the asymmetric binary error profile of the best model (485 missed attacks versus only 144 false alarms) and the directional multiclass confusions (Recon→Benign and DDoS→DoS) together with the systematically weak Bruteforce class. Beyond this, the work delivers a controlled, uniform comparison of nine deep-learning architectures (feed-forward, recurrent, convolutional, hybrid, and residual) under a single leakage-controlled pipeline, validates the observed performance differences statistically through Cochran’s Q and Holm-adjusted pairwise McNemar tests rather than treating them as incidental, and applies SHAP and LIME at the appropriate interpretive scope (global and class-selective attributions alongside local per-prediction explanations), framed explicitly as associational rather than causal. Collectively, these contributions show that class-wise behavior and specific misclassification patterns—not aggregate accuracy alone—should guide the assessment of deep intrusion-detection models for the Industrial Internet of Things.
