The paper’s main contribution is a reframing of the aggregation criterion: it replaces the geometric question (“is this update an outlier?”) with a behavioral one (“does this update improve rare-class recall on a balanced probe?”). This dissolves the imbalance–robustness tension, since an honest rare-class client and a poisoner are indistinguishable by position but separable by behavior. RAVEN operationalizes this as probe scoring → tempered-softmax reputation → blended aggregation, paired with a federated-native Balanced-Softmax objective. Empirically it holds minority recall at 0.770 under strong ALIE collusion where Krum collapses to zero, with the margin widening as the attack strengthens.
The second, arguably more durable contribution is negative: the probe-optimal white-box adversary drives every evaluated defense to zero minority recall, and the natural geometric fix is shown to be structurally unavailable, since honest non-IID clients reach 0.99 cosine similarity against colluders’ 1.00. This converts what could have been a method-specific weakness into a boundary result on the entire aggregation-defense class, and points future work toward signals outside the update geometry — secret rotating probes, system-level attestation.
