Project Assignment (1500 words)

profileh66umi6q
presentation.pdf

Bias and fairness in Machine Learning

Motivation

Wide Application scenarios of ML systems

Face recognition system

Speech recognition system

Intrusion Detection System

Autonomous Driving

Automatic information management system

Wireless communication

Is there any ethic issue?

Machine learning pipeline

Data Machine learning

algorithms Data-Driven Decision

Making

Dataset bias Algorithm fairness

Questions: ◦ What is the bias for ML datasets and how it affects the decision making process? ◦ What is the fairness for ML algorithms and how it affects the decision making process? ◦ Our contribution: Try to distinguish a biased or unfair issue on real-life dataset and find out corresponding

solutions.

Bias for datasets Definition: When scientific or technological decisions are based on a narrow set of systemic, structural or social concepts and norms, the resulting technology can privilege certain groups and harm others [BiasFairness18].

Classification [BiasClass]: ◦ Sample bias ◦ Exclusion bias ◦ Measurement bias ◦ Recall bias ◦ Observer bias ◦ Racial bias ◦ Association bias

[BiasFairness18] Bias and Fairness in AI/ML models https://fpf.org/wp-content/uploads/2018/11/Presentation-2_DDF-1_Dr-Swati-Gupta.pdf [BiasClass] 7 Types of Data Bias in Machine Learning https://lionbridge.ai/articles/7-types-of-data-bias-in-machine-learning/ [Survey19] Mehrabi, Ninareh, et al. "A survey on bias and fairness in machine learning." arXiv preprint arXiv:1908.09635 (2019).

Example - IMAGENET sample bias [Survey19]:

Fairness for algorithms[Fairness18] Definition[Intro17]: ◦ No Universal definition • Unawareness • Demographic Parity • Equalized Odds • Predictive Rate Parity • Individual Fairness • Counterfactual fairness

Example – COMPAS algorithm[Fairness18]: ◦ A machine learning system used by U.S officials

to do recidivism prediction ◦ Suppose to be a fair algorithm but actually show

bias against minority groups

[Intro17] A Tutorial on Fairness in Machine Learning https://towardsdatascience.com/a-tutorial-on-fairness-in-machine-learning-3ff8ba1040cb [Fairness18] Chouldechova, Alexandra, and Aaron Roth. "The frontiers of fairness in machine learning." arXiv preprint arXiv:1810.08810 (2018).

Related datasets[Survey19] Dataset Name Size Area Reference

UCI Adult dataset 48842 income records Social A. Asuncion and D.J. Newman. 2007. UCI Machine Learning Repository. (2007). http://www.ics.uci.edu/$\sim$mlearn/

German credit dataset 1000 credit records Financial Dheeru Dua and Casey Graff. 2017. UCI Machine Learning Repository. (2017). http://archive.ics.uci.edu/ml

Pilot parliaments benchmark dataset

1270 images Facial images Joy Buolamwini and Timnit Gebru. 2018. Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification. In Proceedings of the 1st Conference on Fairness, Accountability and Transparency (Proceedings of Machine Learning Research), Sorelle A. Friedler and Christo Wilson (Eds.), Vol. 81. PMLR, New York, NY, USA, 77–91. http://proceedings.mlr.press/v81/buolamwini18a.html

WinoBias 3160 sentences Coreference resolution

Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018. Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods. (2018). arXiv:cs.CL/1804.06876

Communities and crime dataset

1994 crime records Social M Redmond. 2011. Communities and crime unnormalized data set. UCI Machine Learning Repository. In website: http://www. ics. uci. edu/mlearn/MLRepository. html (2011).

COMPAS Dataset 18610 crime records Social J Larson, S Mattu, L Kirchner, and J Angwin. 2016. Compas analysis. GitHub, available at: https://github. com/propublica/compas-analysis[Google Scholar] (2016).

Recidivism in juvenile justice dataset

4753 crime records Social Manel Capdevila, Marta Ferrer, and Eulália Luque. 2005. La reincidencia en el delito en la justicia de menores. Centro de estudios jurídicos y formación especializada, Generalitat de Catalunya. Documento no publicado (2005).

Diversity in face dataset 1 million images Social Michele Merler, Nalini Ratha, Rogerio S Feris, and John R Smith. 2019. Diversity in Faces. arXiv preprint arXiv:1901.10436 (2019).

Recent Related works Category Name Citations Reference

Survey A Survey on Bias and Fairness in Machine Learning

258 Mehrabi, Ninareh, et al. "A survey on bias and fairness in machine learning." arXiv preprint arXiv:1908.09635 (2019).

Fairness in machine learning: A survey

10 Caton, Simon, and Christian Haas. "Fairness in Machine Learning: A Survey." arXiv preprint arXiv:2010.04053 (2020).

Bias Ethical Implications of Bias in Machine Learning

38 Yapo, Adrienne, and Joseph Weiss. "Ethical implications of bias in machine learning." Proceedings of the 51st Hawaii International Conference on System Sciences. 2018.

Identifying and Correcting Label Bias in Machine Learning

41 Jiang, Heinrich, and Ofir Nachum. "Identifying and correcting label bias in machine learning." International Conference on Artificial Intelligence and Statistics. PMLR, 2020.

Understanding Bias in Machine Learning

3 Gu, Jindong, and Daniela Oelke. "Understanding bias in machine learning." arXiv preprint arXiv:1909.01866 (2019).

Fairness Fairness in machine learning 117 Barocas, Solon, Moritz Hardt, and Arvind Narayanan. "Fairness in machine learning." Nips tutorial 1 (2017): 2.

The frontiers of fairness in machine learning

133 Chouldechova, Alexandra, and Aaron Roth. "The frontiers of fairness in machine learning." arXiv preprint arXiv:1810.08810 (2018).

Improving fairness in machine learning systems: What do industrial practitioner need?

135 Holstein, Kenneth, et al. "Improving fairness in machine learning systems: What do industry practitioners need?." Proceedings of the 2019 CHI conference on human factors in computing systems. 2019.

Q&A

  • Bias and fairness in Machine Learning
  • Motivation
  • Machine learning pipeline
  • Bias for datasets
  • Fairness for algorithms[Fairness18]
  • Related datasets[Survey19]
  • Recent Related works
  • 幻灯片编号 8