A Fair Adversarial Training Framework for Jointly Improving Fairness and Adversarial Robustness in Artificial Intelligence Systems

K. T. Iorsamba, S. Ahmad, A. O. Isah, M. D. Noel

Abstract


Artificial Intelligence (AI) systems deployed in high-stakes domains are increasingly challenged by algorithmic bias and adversarial vulnerabilities, yet these two problems are commonly studied independently. This study investigates algorithmic bias as a measurable security vulnerability and evaluates mitigation strategies that jointly address fairness and adversarial robustness. An explanatory quantitative experimental design was employed using the FairFace benchmark dataset. Four mitigation approaches were evaluated against a Baseline model: Domain-Adversarial Neural Networks (DANN), Projected Gradient Descent Adversarial Training (PGDAT), Fair Adversarial Training (FAT), and Fair Adversarial Attack Learning (FAAL). Model performance was assessed using the Balanced Subgroup Error Rate (BSER), Adversarial BSER, and a proposed Attack Surface Ratio (ASR) metric. The Baseline model achieved an overall accuracy of 42.19% and exhibited substantial subgroup disparities, with Clean BSER values ranging from 0.3254 for Black Males to 0.8561 for Middle Eastern Females, confirming severe intrinsic algorithmic bias. Under adversarial attack, vulnerable subgroups experienced significantly increased exploitation risk, with ASR values reaching 2.49 for Black Males and 2.37 for White Males, demonstrating that algorithmic bias enlarges the attack surface of AI systems. DANN achieved the strongest fairness improvement, raising overall accuracy to 67.47% and significantly reducing subgroup disparities, but simultaneously increased adversarial vulnerability, with the Black Male subgroup recording the highest ASR value of 5.14. PGDAT produced the strongest robustness performance, improving accuracy to 56.40% and reducing ASR by 39.32% for Black Males and 25.85% for White Males, although fairness disparities persisted. The integrated FAT model produced moderate joint improvements but failed to achieve statistically significant gains over the Baseline, while FAAL exhibited severe optimisation instability, producing extreme ASR values of 12.78 and 16.50 for the Southeast Asian Male and Female subgroups, respectively. A one-way ANOVA revealed significant differences among models for fairness performance (F = 12.0750, p = 0.00000021), whereas differences in adversarial robustness were not statistically significant (F = 2.0239, p = 0.1014). These findings provide empirical evidence that algorithmic bias functions as an exploitable security vulnerability and demonstrate a persistent trade-off between fairness and robustness in AI systems.


Full Text:

PDF

References


Alvarez, J. M., Bringas Colmenarejo, A., Elobaid, A., Fabbrizzi, S., Fahimi, M., Ferrara, A., Ghodsi, S., Mougan, C., Papageorgiou, I., Reyero, P., Russo, M., Scott, K. M., State, L., Zhao, X., & Ruggieri, S. (2024). Policy advice and best practices on bias and fairness in AI. Ethics and Information Technology, 26(31).

Akinyemi, G., Momoh, M. O., Bulama, H. I., & Abdullahi, M. H. (2025). The Role of Artificial Intelligence in Cybersecurity for Financial Fraud Detection and Prevention. ATBU Journal of Science, Technology and Education, 13(2), 148-153.

Athalye, A., Carlini, N., & Wagner, D. (2018). Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. Proceedings of the 35th International Conference on Machine Learning (ICML 2018).

Bagdasaryan, E., Veit, A., Hua, Y., Estrin, D., & Shmatikov, V. (2020). How to backdoor federated learning. Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics, 108, 2938–2948.

Benz, P., Zhang, C., Karjauv, A., & Kweon, I. S. (2021). Robustness may be at odds with fairness: An empirical study on class-wise accuracy. NeurIPS 2020 Preregistration Workshop.

Buolamwini, J., & Gebru, T. (2018). Gender shades: Intersectional accuracy disparities in commercial gender classification. Proceedings of the 1st Conference on Fairness, Accountability, and Transparency, 81, 77–91.

Chai, J., & Wang, X. (2023). To be robust and to be fair: Aligning fairness with robustness. arXiv:2304.00061.

Chai, J., Jang, T., Gao, J., & Wang, X. (2025). On the alignment between fairness and accuracy: From the perspective of adversarial robustness. Proceedings of the 42nd International Conference on Machine Learning (PMLR 267), 7107–7130.

Cohen, J. M., Rosenfeld, E., & Kolter, J. Z. (2019). Certified adversarial robustness via randomized smoothing. Proceedings of the 36th International Conference on Machine Learning, 97, 1310–1320.

Du, M., Tang, R., Fu, W., & Hu, X. (2022). Towards debiasing DNN models from spurious feature influence. Proceedings of the AAAI Conference on Artificial Intelligence, 36(9), 9521–9528.

European Commission. (2024). European Union Artificial Intelligence Act. Official Journal of the European Union.

Kärkkäinen, K., & Joo, J. (2021). FairFace: Face attribute dataset for balanced race, gender, and age for bias measurement and mitigation. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV 2021), 1548–1558.

Langer, M., Baum, K., et al. (2025). Secure human oversight of AI: Exploring the attack surface of human oversight. arXiv:2509.12290.

Ma, X., Wang, Z., & Liu, W. (2022). On the tradeoff between robustness and fairness. Advances in Neural Information Processing Systems, 35.

Madry, A., Makelov, A., Schmidt, L., Tsipras, D., & Vladu, A. (2018). Towards deep learning models resistant to adversarial attacks. Proceedings of the 6th International Conference on Learning Representations (ICLR).

Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2019). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6), 1–35.

Mehrabi, N., Naveed, M., Morstatter, F., & Galstyan, A. (2021). Exacerbating algorithmic bias through fairness attacks. Proceedings of the AAAI Conference on Artificial Intelligence, Workshops.

National Institute of Standards and Technology. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). U.S. Department of Commerce.

Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453.

Serna, I., Morales, A., Fierrez, J., & Obradovich, N. (2022). Sensitive loss: Improving accuracy and fairness of face representations with discrimination-aware deep learning. Artificial Intelligence, 305, 103682.

Tran, C., Zhu, K., Van Hentenryck, P., & Fioretto, F. (2024). On the effects of fairness to adversarial vulnerability. Proceedings of the 33rd International Joint Conference on Artificial Intelligence (IJCAI-24).

Xu, H., Liu, X., Li, Y., Jain, A. K., & Tang, J. (2021). To be robust or to be fair: Towards fairness in adversarial training. Proceedings of the 38th International Conference on Machine Learning, 139, 11492–11501.

Zhang, Y., Liu, Q., & Wang, C. (2024). Towards fairness-aware adversarial learning (FAAL). Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 12345–12354.

Zhi, H., Yu, H., Li, S., Zhao, X., & Wu, Y. (2025). Towards fair class-wise robustness: Class optimal distribution adversarial training (CODAT). Complex & Intelligent Systems, 11, 403.


Refbacks

  • There are currently no refbacks.