Answer 3 questions after reading each of 3 articles, each article have 3 pages.

Jason07070926
Data_Rights_and_Wrong_Langkj-r-Bain-2018-Significance.pdf

Robert Langkjær-Bain is a freelance journalist and regular contributor to Significance. He was previously editor of Lux magazine and deputy editor of Research magazine.

This year, the ethics of data hit the headlines like never before. Even those already mistrustful of social media companies, and their mass collection of personal data, were shocked by the Facebook–Cambridge Analytica scandal, which saw a third-party app developer gain access to the profile information of millions of people and share it with a company engaged in political campaigning.

But for those concerned with data ethics, Facebook’s response was just as revealing as the fact of the leak itself. What was described by the media as one of the biggest data breaches in history (bit.ly/2ragdUb) was, in Facebook’s view, not a data breach at all. The company argued that users had consented to share their data, and that of their friends, with the app in question – it was what was done with the data next that was the problem (bit.ly/2D6AAvP).

This is a story that says a lot about the state of data ethics in 2018. If we cannot even agree on what constitutes a data breach, how can we decide where to draw the line between acceptable and unacceptable uses of personal information?

Hype versus fear Ethics is the study of right and wrong – of the moral principles that underlie what we do. The field of data ethics has built on the foundations of computer ethics and information ethics to investigate the morals of data collection, processing, analysis and application.

Things were simpler back when data was scarce. But as technology has raced ahead, old principles governing the protection of personal data – such as a commitment to anonymity and informed consent – have begun to feel irrelevant or unworkable. There are genuine worries about how data is used or shared, and how information is harnessed by algorithms and artificial intelligence.

In the media, fear about the threats posed by “big data” is matched only by hype for the promised benefits. If you believe the headlines, data is perennially on the brink of either curing all the world’s ills or unleashing a dystopian nightmare. How do we steer the right course?

High on the list of ethical concerns is privacy. History provides chilling examples of how sensitive information can be used against the people it came from when privacy breaks down (see box, page 15), and data-based discrimination remains a valid concern today.

In a 2015 study, a group of researchers, including Allesandro Acquisti of Carnegie Mellon University’s Heinz College, sent out identical job applications from candidates whose social media profiles revealed them to be either Christian or Muslim. The Muslim candidates received 13% fewer callbacks – suggesting the employers looked up the candidates and judged them on what they found online (bit.ly/2sSaqnB).

There are other examples, including teachers losing their jobs based on crude analyses of annual fluctuations in students’ test scores, and people being denied credit – or offered higher interest rates – for living on the wrong side of the tracks. In many cases, those affected will have no idea that this kind of thing is happening to them. In 2015, researchers at Carnegie Mellon University found Google’s search results pages showed ads for prestigious jobs more often to men than to women (bit.ly/278sbHF).

In her 2016 book, Weapons of Math Destruction, data scientist Cathy O’Neil describes how bias and discrimination can be baked into algorithms by the choices made about what data to use and how to interpret it – especially when proxy measures are substituted for other things that are harder to quantify.1

Take predictive policing as an example. At the risk of oversimplifying, police forces feed data on past crimes into an algorithm to help them decide where crimes are likely to take place and where best to deploy resources. This seems like an efficient and effective use of existing information, but the American Civil Liberties Union has warned that using data on past crimes to predict future crimes cements bias, because police statistics “primarily document law enforcement’s response” to reports of crime – rather than all instances of crime, whether reported or not (bit.ly/2aLpG5L).

As Andrew Ferguson puts it in his 2017 book, The Rise of Big Data Policing, “Data is not blind. Data is us, just reduced to binary code.”2

DATA AND

“Big data” is gradually revealing its potential to both help and harm us. But the rights and wrongs of what to do about this are hard to pin down. Robert Langkjær-Bain looks at the work under way to ensure the data revolution does not leave ethics behind

SIGNIFICANCE12 December 2018

IN DETAIL

© 2018 The Royal Statistical Society

So, if data is inherently biased, how do we use it fairly and ethically?

Where to begin? An obvious starting point for defining ethical use of data is to ask the public. But this can generate more questions than answers. For example, people’s stated opinions and actual behaviour regarding privacy are notoriously inconsistent. Professor Sofia Olhede of University College London’s Department of Statistical Science says: “People aren’t truthful to themselves or others about how they value privacy. They say they value it, but actually they’ll give it up for something quite small. They want to have all the privacy and all the benefits [of giving it up].”

Another of Alessandro Acquisti’s studies looked at the efficacy of privacy warnings, using a survey of students about cheating in exams. Not surprisingly, participants were less forthcoming when they were warned that their answers would be shared with their tutors. But Acquisti found that a delay of just 15 seconds between the warning and the beginning of the survey was enough to make participants behave as if they had not been warned (bit.ly/2aaTLS4). This fickle, forgetful approach to privacy will not have gone unnoticed by companies whose business relies on getting our data.

However, other research suggests people may not be as hopelessly inconsistent as they seem. Professor Helen Nissenbaum of Cornell Tech showed in a 2016 study that surveys could be used to “crowdsource” the prevailing privacy norms among a given population for a given situation, with a strong level of consensus (bit.ly/2aSpMJ2).

Nissenbaum’s work has also called into question the idea that public data should generally be available as openly as possible. A 2017 study found that most people saw no problem with a car dealer asking a potential customer about their marital status, or how much their home is worth, but did not like the idea of the dealer seeking that same information from an official database.3 What matters to people most is not just whether the data is deemed “sensitive” or not, but the

entire context: what the data says, who it is being shared with, in what way, and for what purpose.

Nissenbaum and others advocate a whole new way of looking at privacy. Anonymity and informed consent remain the go-to tools for managing privacy, but are increasingly hard to achieve in the age of big data. In a 2014 article, Nissenbaum and Solon Barocas write that anonymity and consent face “virtually intractable challenges”.4 They argue that these principles were never really up to the task of mediating the nuanced relationship between data subject and data collector, but were simply the best options we had. And then big data came along and rendered them irrelevant anyway. Dr Julia Lane of New York University’s Center for Urban Science and Progress says: “It’s very clear that we can no longer protect data in the way it’s disseminated. We have to be able to think about ethical use.”

Key to using data ethically is the notion of fairness, of treating people justly and without discrimination. It is not difficult to find definitions of fairness; the hard part is choosing which one to use. Arvind Narayanan, a computer science professor at Princeton, squeezed 21 definitions into one short conference tutorial (bit.ly/2SuHyQm), and that is just scratching the surface, he says.

Lack of statistical bias is one measure, but this is only a first step – bias can creep in through the choice of data studied or the way it is interpreted. We can seek “blindness” to attributes such as gender and race, but we know that in real life this does not guarantee a lack of bias either. And in the case of predictions, we can measure positive predictive value or the rate of false positives or false negatives to see if our predictions are hitting the mark.

Some definitions fall into the category of individual fairness, others aim to achieve group fairness – and the two must be balanced. Different definitions of “fair” will suit different purposes at different times.

An example of this is the controversy over an algorithm called COMPAS, which is used in the US justice system to predict whether defendants seeking bail are likely to reoffend, assigning them a risk rating on a 1–10 scale. COMPAS takes data

13December 2018 significancemagazine.com

R IGH TS WRONGS

13December 2018 significancemagazine.com

from a person’s criminal record as well as their own answers to a set of questions about their past, home circumstances, experiences with crime and so on. It does not consider race – but many of the things it does consider vary greatly between ethnic groups. When the investigative journalism organisation ProPublica looked into how COMPAS predictions panned out, it found the algorithm was more likely to wrongly label a black defendant as high risk, and more likely to wrongly label a white defendant as low risk (bit.ly/2yTLLU1).

On the face of it, that seems like an example of racial bias baked into an algorithm. But in its defence, the company behind COMPAS explained that, for any given point on its 1–10 risk scale, the chance of repeat offending was roughly equal for black and white defendants (bit.ly/2Rcm75e). In that sense, it is fair.

Then along came a reanalysis of ProPublica’s data by Stanford University researchers, which divided white and black defendants into two risk categories: low (points 1–4) and medium/high (points 5–10). They found that within each risk category, the proportion of defendants who reoffend is broadly the same regardless of race, but that the overall rate of reoffending for black defendants is higher than for white defendants. The researchers say that these two facts make the disparity highlighted by ProPublica a statistical inevitability. Given that reoffending is higher overall among black defendants, more black defendants than white defendants will be labelled high risk. And that means that, overall, a greater proportion of blacks will be wrongly labelled high risk.5

The challenge in assessing an algorithm like COMPAS is that there is no single perfect way to judge fairness. Measuring positive predictive value tells you how many criminals have been kept off the streets, but not how many defendants were needlessly kept behind bars; to figure that out, you need to count false positives, while counting false negatives tells you how many criminals were let loose. But which is most useful to know? That depends on whether you are a prison superintendent, a neighbourhood cop, or a human rights campaigner. As for the defendant, only one metric matters: their chances of being freed.

A culture of explanation It would be fascinating to know more about the inner workings of algorithms like COMPAS. But the companies behind them tend to guard their secrets closely – which brings us to the next data ethics issue: transparency. In a world where important decisions are made by data processes that few understand, how do we make sure they are fair?

Professor Luciano Floridi of the University of Oxford, who leads the Oxford Internet Institute’s Digital Ethics Lab, says we need to get away from the notion of algorithms and artificial intelligence as “black boxes” that are too complicated or sensitive to be explained. An algorithm is “not magic”, Floridi says, “it’s a piece of engineering”. However, Floridi also warns against pursuing transparency for its own sake. “We need to be clear about what we want it for.”

Investigating the inner workings of algorithms and AI is like trying determine the cause of a traffic jam, he says. Myriad factors are at play, but at some point you have to decide how much detail you want to go into. “Why each individual driver or pedestrian is where they are at that particular moment, I don’t care,” he says. “We can go as far as we want in explaining the behaviour – it’s a matter of whether we want to.”

Besides, it is not always necessary to open up the black box if real-life outcomes make it clear that bias is present. A 2016 paper by Google research scientist Moritz Hardt, together with Eric Price of the University of Texas at Austin and Nati Srebro of the Toyota Technological Institute at Chicago, assessed fairness by looking at whether attributes about data subjects could be identified from results – in other words, whether you could tell with 100% certainty whether a person was, say, white or black, based on a decision made by an algorithm. If so, something is obviously wrong (bit.ly/2Rd9pDy).

Floridi says we need to create “a culture of explanation”, where people are able to ask questions such as “What’s going on?”, “Who is accountable if it’s not working properly?” and “Can you make sure it’s not going to happen again?”. The European Union may have taken a step in that direction with its new General Data Protection Regulation, which includes a so-called “right to explanation” for automated decisions. However, the real value of this right has yet to be tested, and critics say its scope is limited.

“DATA ARE NOT JUST NUMBERS

WITHOUT MEANING OR CONTEXT ”

SIGNIFICANCE14 December 2018

Less talk, more action Are we any closer to putting all this new ethical thinking into practice? Julia Lane of New York University says: “I’m a little bit worried because there has been a lot of talk but not much by way of actual operational suggestions. Everyone agrees that it’s important to be ethical in the way you handle data, but can we move towards a concrete description of what that is and how it’s measured?”

Lane says one promising piece of work (which she contributed to) is a proposed data science oath, modelled on the Hippocratic Oath, and published in a report on data science education from the US National Academies (bit.ly/2Re2fPi). It includes commitments that “I will respect the privacy of my data subjects”, “I will remember that my data are not just numbers without meaning or context, but represent real people and situations and that my work may lead to unintended societal consequences”, and “I will always look for a path to fair treatment and non-discrimination”.

These ideas fed into the Community Principles on Ethical Data Practices developed last year by Bloomberg, BrightHive and Data for Democracy (which can be seen and signed at datapractices.org). Among other things, the principles call for data scientists to consider consent (even if they cannot secure it explicitly), to be alert to bias, to foster diversity, and to protect data against being deanonymised.

This year in the UK, the government announced the formation of a new Centre for Data, Ethics and Innovation, and set out a data ethics framework for the public sector, covering issues such as clarity about public benefits, proportionate use of data, and accountability.

Meanwhile, the newly formed Ada Lovelace Institute, backed by the Nuffield Foundation, is working to research and promote ethical practices in data science and AI, in collaboration with other bodies including the Royal Statistical Society and the Alan Turing Institute.

Olivia Varley-Winter, interim programme manager at the Ada Lovelace Institute, says its aim is to “move principles through to practice”. A big part of this involves building bridges and bringing together thinking from a variety of sectors and fields. “A lot of stakeholders now care about this, and there is a need for conferring to see how different values interact,” she says. “Different fields have different frameworks for understanding issues such as accountability and bias, so we want to convene diverse voices to create a shared understanding of the ethics of data and AI.” Varley- Winter hopes that by equipping industry with knowledge and skills, businesses will feel more confident about engaging with ethics. “Right now I’d say there is a lot of interest but not a lot of confidence,” she says.

Floridi argues for what he describes as “translational ethics”. In other words: “You need to develop the tools and apply them at the same time”. He says: “If you’re on a raft and it’s sinking, you need to fix the raft while at sea. You can’t go to dock to fix it and put it back in the water.” Researchers can get to work on “important, fundamental principles” while also making a “real difference”, he adds.

The road ahead The outlook for the future is far from clear, but Floridi believes we can be “moderately, mildly optimistic”. “To take a long-term perspective, we have moved from a state of denial [where technology providers would say], ‘We’re not doing anything wrong, we’re just providing a service, technology is neutral’, etc. No one in their right mind still says that with a straight face. But moving from denial into acceptance doesn’t mean we’ve safely reached the third step of doing something really profound; for example, changing business models, playing the role of a responsible corporate citizen. Today, awareness is everywhere. Whether it leads to solutions is a long road.”

The key, says Lane, will be measurement – of how people really feel and behave, of what the harms and benefits resulting from data are, and of what works to address them. “Right now I don’t think of [data ethics] as being a scientific field, because it’s just a lot of people talking,” she says. “But we’re statisticians, so let’s measure it.” n

References 1. O’Neil, C. (2016) Weapons of Math Destruction. New York: Crown. 2. Ferguson, A. (2017) The Rise of Big Data Policing. New York: New York University Press.

3. Martin, K. and Nissenbaum, H. (2016) Privacy interests in public records: An empirical investigation. Harvard Journal of Law & Technology,

forthcoming. bit.ly/2SpGQE5

4. Barocas, S. and Nissenbaum, H. (2014) Big data’s end run around anonymity and consent. In J. Lane, V. Stodden, S. Bender and

H. Nissenbaum (eds), Privacy, Big Data and the Public Good (pp. 44–75).

New York: Cambridge University Press.

5. Corbett-Davies, S., Pierson, E., Feller, A. and Goel, S. (2016) A computer program used for bail and sentencing decisions was labeled biased

against blacks. It’s actually not that clear. Washington Post, 17 October.

wapo.st/2JgzbE5

6. Black, E. (2001) IBM and the Holocaust. Largo, MD: Crown Books.

Data privacy: a casualty of war During the Second World War, governments on both sides of the Atlantic used census data to target their own citizens.

The Nazis had no qualms about using census data to identify Jews, Gypsies and other groups considered undesirable. As described in a 2001 book by Edwin Black, this was done with the help of a German subsidiary of US computing giant IBM, which supplied punch-card technology that allowed census data to be recorded and processed.6 Ultimately, the data helped to facilitate the genocide of millions.

Meanwhile, in the United States, the Census Bureau shared data from the 1940 census with other branches of government to help identify Japanese Americans, at a time when many in this community were being sent to internment camps (bit.ly/2Rg36rc). The Bureau provided block-level data in several states, and in at least one case provided individual data – all made possible by special wartime powers granted under a law that has since been repealed.

For decades the Bureau denied releasing any data, but when evidence was uncovered in 2000, it issued a public apology. It stressed that the activities were lawful at the time, and that it was “legally required to assist in the war effort”. However, this chapter of history has helped fuel recent concerns about the 2020 census and the proposed addition of a question on citizenship status (ti.me/2Rg3wFE).

IN DETAIL

15December 2018 significancemagazine.com