Ensemble Deep Learning with Test-Time Augmentation for Breast Ultrasound Classification on BreastMNIST

Our main contribution was a methodical testing of three different models of heterogeneous ensemble CNN–Transformer—ConvNeXt-Tiny, EfficientNetV2-S and Swin-Tiny—using test-time augmentation (TTA) with soft voting, all optimized for the BreastMNIST classification task. The ensemble performs with a test accuracy of 92.31%, higher than that of the best individual model, and class-wise error and inter-model agreement analyses are included to give better estimates of prediction reliability. Critically, the paper formalizes this contribution as an empirical ensemble evaluation rather than a new learning algorithm.