Federated aggregation strategies for credit card fraud detection under differential privacy
Öz
Context—Card fraud cost the global financial system over $28 billion in 2019, and losses have risen every year since. Banks hold the transaction data needed for collaborative fraud detection, but privacy regulations such as GDPR and KVKK prevent cross-institutional data sharing. Federated Learning keeps raw data local: each institution trains on its own records and sends only model parameters to a shared coordinator. Differentially-Private Stochastic Gradient Descent (DP-SGD) counters gradient inversion attacks by clipping per-sample gradients and adding calibrated noise before parameters leave the client, yielding a record-level (ε, δ) guarantee whose strength depends on the accumulated privacy budget ε. How the aggregation strategy behaves under this noise has not yet been studied.
Objective—We compare five aggregation strategies—FedAvg, FedProx, cosine-similarity aggregation, FedAvg-DWA, and FedAdam—across two real-world benchmarks, two heterogeneity levels, four DP-SGD noise multipliers, and 10 random seeds per configuration. We ask whether the relative ranking of strategies survives DP noise, and what DP actually costs in deployed model behavior.
Method—510 federated training runs cover five aggregation configurations, two datasets (Kaggle ULB: n = 284807, 0.17% positive rate; IEEE-CIS: n = 590540, 3.50%), two Dirichlet non-IID levels (α ∈ {0.5, 0.1}), and σ ∈ {0.5, 1.0, 1.5, 2.0}. The base learner is a three-layer MLP with LayerNorm. Privacy accounting uses the Opacus RDP accountant at δ = 10⁻⁵. Performance is reported at both the fixed 0.5 threshold and a validation-tuned F₁-maximizing threshold. All pairwise comparisons carry Bonferroni-, Holm-, and Benjamini–Hochberg-corrected p-values, and all 200 realized client-shard partitions are characterized directly.
Results—On Kaggle ULB under IID partitioning, federated F₁@best-t (0.801–0.804) matches the centralized baseline (0.799 ± 0.036). Under mild non-IID without DP (α = 0.5), FedAvg-DWA narrowly leads (< 0.015 on F₁@best-t), though neither test survives correction. In the ULB privacy sweep, 7 of 40 Wilcoxon tests show nominal significance but none survive family-wise correction; performance consistently orders with FedAvg-DWA leading and FedAdam trailing. On IEEE-CIS, this ordering amplifies, and the extreme pair remains Holm-significant under DP (adjusted p = 0.020). While AUC holds near 0.95, calibration drift collapses precision at the fixed 0.5 threshold; a validation-tuned threshold recovers F₁ loss without privacy cost. Both calibration effects replicate on IEEE-CIS (ɛ = 8.10 ± 5.30 at σ = 1.0).
Conclusion—Aggregation-rule differences under DP-SGD are small and dataset-dependent in magnitude, consistent in direction, and on the ULB sweep none survives correction; outcome differences arise mainly from post-training threshold tuning. We also uncover a previously unreported interaction: Opacus's DP step quietly overrides the standard loss-augmented FedProx setup, making FedProx numerically equivalent to FedAvg, verified by bit-identical per-round global weight trajectories across seeds.
Anahtar Kelimeler
Kaynakça
- E. Ileberi, Y. Sun, Z. Wang, “Performance evaluation of machine learning methods for credit card fraud detection using SMOTE and AdaBoost”, IEEE Access, 9, 165286–165294, 2021. https://doi.org/10.1109/ACCESS.2021.3134330.
- A. Ali, S. A. Razak, S. H. Othman, T. A. E. Eisa, A. Al-Dhaqm, M. Nasser, T. Elhassan, H. Elshafie, A. Saif, “Financial fraud detection based on machine learning: A systematic literature review”, Applied Sciences, 12(19), 9637, 2022. https://doi.org/10.3390/app12199637.
- European Parliament and Council of the European Union, “Regulation (EU) 2016/679 (General Data Protection Regulation)”, Official Journal of the European Union, L119, 1–88, 2016. https://gdpr-info.eu/.
- H. B. McMahan, E. Moore, D. Ramage, S. Hampson, B. A. y Arcas, “Communication-efficient learning of deep networks from decentralized data”, Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS), Fort Lauderdale, FL, USA, 20-22 April 2017, 1273–1282. https://doi.org/10.48550/arXiv.1602.05629.
- M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, L. Zhang, “Deep learning with differential privacy”, Proceedings of the 23rd ACM Conference on Computer and Communications Security (CCS), Vienna, Austria, 24-28 October 2016, 308–318. https://doi.org/10.1145/2976749.2978318.
- C. Dwork, A. Roth, “The algorithmic foundations of differential privacy”, Foundations and Trends in Theoretical Computer Science, 9(3–4), 211–487, 2014. https://doi.org/10.1561/0400000042.
- N. R. Shanbhog, K. S. Totad, A. R. Hanchinal, A. P. Bidargaddi, “Fraud detection in financial transactions using deep learning: A comparative study”, Proceedings of the 5th International Conference on Emerging Technologies (INCET), Belgaum, India, 24-26 May 2024. https://doi.org/10.1109/INCET61516.2024.10593486.
- F. Z. El Hlouli, J. Riffi, M. A. Mahraz, A. El Yahyaouy, H. Tairi, “Credit card fraud detection based on multilayer perceptron and extreme learning machine architectures”, Proceedings of the International Conference on Intelligent Systems and Computer Vision (ISCV), Fez, Morocco, 09-11 June 2020, 1–5. https://doi.org/10.1109/1ISCV49265.2020.9204185.
Ayrıntılar
Birincil Dil
İngilizce
Konular
Makine Öğrenme (Diğer), Veri ve Bilgi Gizliliği, Veri Mühendisliği ve Veri Bilimi
Bölüm
Araştırma Makalesi
Erken Görünüm Tarihi
5 Eylül 2026
Yayımlanma Tarihi
-
Gönderilme Tarihi
12 Mayıs 2026
Kabul Tarihi
27 Ağustos 2026
Yayımlandığı Sayı
Yıl 2026 Sayı: Advanced Online Publication