Bias and Fairness in Facial Recognition AI: An Experimental Evaluation of Demographic Disparities and Mitigation Strategies

Authors

  • Minhui Huang

DOI:

https://doi.org/10.61173/gq1fqz50

Keywords:

Facial Recognition, Algorithmic Bias, Fairness in AI, Demographic Disparities, Ethical AI

Abstract

This research study investigates the question of whether or not facial recognition artificial intelligence (AI) systems are fair in the various demographic groups. Although the use of facial recognition in the area of security, law enforcement, and business has been rapidly adopted, recent reports have demonstrated significant differences in performance, which has been criticized in terms of ethics and social issues. The research develops three hypotheses: (H1) facial recognition systems demonstrate much greater errors of dark-skinned people and women than of lightskinned men; (H2) working with balanced data sets eliminates demographic bias; and (H3) training on fairness enhances group performance equality. The large open-source face datasets that were used as primary data are the Racial Faces in the Wild (RFW), Balanced Faces in the Wild (BFW), CASIA-Face-Africa and KANFace. The open-source models that were tested and trained using these datasets include ResNet and VGGFace2. The evaluation of performance was done in terms of confusion matrices, fairness (Demographic Parity Difference, Equalized Odds) and statistical tests (T-tests, ANOVA, Chi-square). The findings indicate evident demographic differences in accordance with previous research but balanced datasets and unbiased approaches proved to contribute to quantifiable gains. Nonetheless, there were still tradeoffs between the general accuracy and fairness. The results confirm the opinion that demographic equity of AI needs not only technical interventions but also global governance, transparency of the datasets, and ethical control.

References

to the test on open-source face databases, and trained an Barocas, S., Hardt, M., & Narayanan, A. (2019). Fairness individual model with the assistance of statistical analy- and machine learning. fairmlbook.org. Retrieved from sis. The findings showed that differences in performance https://fairmlbook.org

between races or gender could be measured and according Benjamin, R. (2019). Race after technology: Abolitionist to the central research question I could provide an answer, tools for the new Jim code. Polity Press.

which also adds to the broader discussions on the ethics of Buolamwini, J., & Gebru, T. (2018). Gender shades: Inter- AI. sectional accuracy disparities in commercial gender classi- One of the key strengths of the project was that primary fication. Proceedings of Machine Learning Research, 81, data were used. Using the real datasets like RFW, BFW, 1–15. http://proceedings.mlr.press/v81/buolamwini18a. CASIA-Face-Africa and KANFace allowed my evalua- html

tion of the models to be in real terms rather than just de- Corbett-Davies, S., & Goel, S. (2018). The measure and pending on secondary sources. The training of ResNet and mismeasure of fairness: A critical review of fair machine VGGFace2 were conducted independently and produced learning. arXiv preprint. https://doi.org/10.48550/arX- original performance data, which was analysed with fair- iv.1808.00023

ness metrics and statistical hypothesis testing by myself. Danks, D., & London, A. J. (2017). Algorithmic bias in This practice has increased the confidence of my findings autonomous systems. Proceedings of the International but more importantly, it has equipped me with skills of Joint Conference on Artificial Intelligence (IJCAI), 4691– working procedure, data separation, and quantification 4697. https://doi.org/10.24963/ijcai.2017/654

which are expected requirements of investigations at my Eubanks, V. (2018). Automating inequality: How highlevel. tech tools profile, police, and punish the poor. St. Martin’s In spite of these strong points, there were drawbacks of Press.

the project. The datasets were balanced overall, but they European Commission. (2021). Proposal for a regulation could have been biased in the pose, lights, or cultural laying down harmonised rules on artificial intelligence representation. Computational resources also limited the (Artificial Intelligence Act). COM/2021/206 final. Reamount of models and experiments that I was able to per- trieved from https://eur-lex.europa.eu/legal-content/EN/ form and more complex structures such as transformers TXT/?uri=CELEX%3A52021PC0206 Dean&Francis Minhui Huang

Garvie, C. (2016). The perpetual line-up: Unregulated po- Krishnapriya, K. S., Albiero, V., Vangara, K., King, M.

lice face recognition in America. Georgetown Law Center C., & Bowyer, K. W. (2020). Issues related to face recon Privacy & Technology. https://www.perpetuallineup. ognition accuracy varying based on race and skin tone. org IEEE Transactions on Technology and Society, 1(1), 8–20.

Grother, P., Ngan, M., & Hanaoka, K. (2019). Face recog- https://doi.org/10.1109/TTS.2020.2974996 nition vendor test (FRVT), Part 3: Demographic effects. Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., &

National Institute of Standards and Technology. https:// Galstyan, A. (2021). A survey on bias and fairness in madoi.org/10.6028/NIST.IR.8280 chine learning. ACM Computing Surveys, 54(6), 1–35.

Han, H., & Jain, A. K. (2014). Age, gender and race es- https://doi.org/10.1145/3457607

timation from unconstrained face images: Guidelines Raji, I. D., & Buolamwini, J. (2019). Actionable aufor practitioners. Journal of Information Forensics and diting: Investigating the impact of publicly naming Security, 9(12), 1978–1988. https://doi.org/10.1109/ biased performance results of commercial AI prod- TIFS.2014.2359646 ucts. Proceedings of the AAAI/ACM Conference on

Hardt, M., Price, E., & Srebro, N. (2016). Equality of AI, Ethics, and Society (AIES), 429–435. https://doi. opportunity in supervised learning. Advances in Neural org/10.1145/3306618.3314244

Information Processing Systems (NeurIPS), 29, 3315– Robinson, J., Smith, L., & Zhang, Y. (2020). Balanced 3323. https://proceedings.neurips.cc/paper/2016/hash/9d- faces in the wild: Reducing bias in face recognition data- 2682367c3935defcb1f9e247a97c0d-Abstract.html sets. IEEE Winter Conference on Applications of Comput- Hill, K. (2020, June 24). Wrongfully accused by an al- er Vision (WACV), 1568–1577. https://doi.org/10.1109/ gorithm. The New York Times. https://www.nytimes. WACV45572.2020.9093343

com/2020/06/24/technology/facial-recognition-arrest.html Wang, M., Deng, W., Hu, J., Tao, X., & Huang, Y. (2019).

Kamiran, F., & Calders, T. (2012). Data preprocessing Racial faces in the wild: Reducing racial bias by deep face techniques for classification without discrimination. recognition. International Journal of Computer Vision, Knowledge and Information Systems, 33(1), 1–33. https:// 127(4), 406–422. https://doi.org/10.1007/s11263-018- doi.org/10.1007/s10115-011-0463-8 1126-9

Klare, B. F., Burge, M. J., Klontz, J. C., Bruegge, R. W., Zhang, B. H., Lemoine, B., & Mitchell, M. (2018). & Jain, A. K. (2012). Face recognition performance: Role Mitigating unwanted biases with adversarial learning. of demographic information. IEEE Transactions on Infor- Proceedings of the 2018 AAAI/ACM Conference on mation Forensics and Security, 7(6), 1789–1801. https:// AI, Ethics, and Society (AIES), 335–340. https://doi. doi.org/10.1109/TIFS.2012.2214212 org/10.1145/3278721.3278779

Downloads

Published

2026-02-28