Accurate tomato leaf disease recognition can be
overstated when visually repeated images occur across data
partitions or evaluation is restricted to familiar sources. This
study presents a leakage-controlled classification framework that
combines EfficientNetV2-S, the Convolutional Block Attention
Module (CBAM), an inductive Graph Convolutional Network
(GCN), and feature-level fusion. Starting from 31,450 cleaned
images representing 11 tomato leaf classes, a perceptual-hash
audit identified 1,237 repeated-image groups involving 2,539
images. After removing 24 cross-class conflicts and one corrupted
image, 31,425 images were partitioned group-wise into 21,990
training, 4,718 validation, and 4,717 held-out test samples, with
no verified SHA-256, perceptual-hash, or source-group overlap
across the partitions. The attention-enhanced network produced
1,280-dimensional image embeddings, while a cosine-neighbor
graph with k = 5 and an inductive graph network generated 128-
dimensional relational features.
