The main contribution of this study is not a new architecture but a clearer picture of what existing ones actually achieve on IoMT traffic. By training seven deep-learning models under a single disclosed protocol on CICIoMT2024 and comparing them with McNemar testing under Holm correction, we show that the differences between architectures, while statistically significant across all twelve pairwise comparisons, are often practically small once the size of the test partition is taken into account. More importantly, the results expose a gap that weighted reporting conceals: the best multiclass model reaches 98.01% weighted F1 but only 58.81% macro F1, and fails to detect the rare Recon Ping Sweep class at all. Since much of the surrounding literature reports weighted or overall accuracy alone, this gap suggests that current performance figures for IoMT intrusion detection may be considerably more optimistic than per-class behaviour warrants. LIME explanations of selected predictions, together with an explicit note that the feature-sequence reshaping used by six of the models carries no genuine temporal meaning, are included to keep the interpretation of these results appropriately bounded.
