Detecting Financial Fraud Rings: Graph Motifs, Learned Detection, and a Base-Rate-Normalised Evaluation

1. Benchmark hygiene changes the answer. AMLworld’s final week has laundering at 59% purity against a 0.1% norm, holding 12.7% of all positives. A standard chronological split puts it in test. Removing it moves the graph-over-tabular advantage from 2.37× to 5.43×.

2. Precise graph queries degrade a GNN when injected as features. Detectors at up to 15.2× lift standalone cost 21% average precision and doubled seed variance as node attributes. Mechanism: they fire on 1.2% of accounts, so the columns are ~99% zero and dilute denser signal. This is a boundary condition on Blanuša et al., who report the opposite — their patterns attach to transactions and feed a tree ensemble, which can branch on a sparse indicator for free.

3. Two metadata columns beat 33 graph-derived ones. Ownership linkage — how many accounts share a legal owner — gave +10.7% AP and cut variance fourfold. Dense and semantic outperformed sparse and structural.

4. Base-rate-normalised lift, and the retention split. AP is uninterpretable across subsets of differing prevalence. Applied to leave-one-typology-out, recall retention is 93.5% but lift retention 45.5% — the model still finds an unseen coordination pattern, but ranks it far less confidently. Either figure alone misleads.

5. Empirical threshold selection for motif queries. Sweeping stored values rather than guessing more than doubled lift on three of four detectors — fan-out 6.2× → 15.2×, fan-in 2.7× → 12.9×