We design an attention-gated feature-fusion ensemble
that learns, on a per-sample basis, how much to trust
CNN-derived versus transformer-derived features, rather
than relying on static combination weights or late-stage
decision fusion
We design an attention-gated feature-fusion ensemble
that learns, on a per-sample basis, how much to trust
CNN-derived versus transformer-derived features, rather
than relying on static combination weights or late-stage
decision fusion