Discussion: Ethical Issues in Criminal Justice Research

profileSuccess50
WK6READINGONLY2.pdf

Learnrng q imilar to the research methods we examined in Chapter 10, the dataused for\) the methods described in this chapter often rely on secondary data.Impor- tantly, the methods covered in this chapter can be used both for investigative purposes and for basic research purposes. The rise of social media platforms such as Facebook and Tivitter have provided both researchers and practitioners with an abundance of data from which to draw inferences.

In this chapteq we will first introduce you to social nenvork analysis and demonstrate how it has been applied to research as well as law enforcement investigative practices. We next discuss mapping generally (and crime mapping specificrll, and examine how this technique has also been applied to police practice as well as general research. And finally, we will conclude the chapter with an introduction to the concept of big data, which you will learn is not merely an extremely large and dynamic set of data but also refers to techniques used to extract information from these datasets.

SOCIAL N ETWOR K ANALYSIS

It is virtually impossible today not to be part of social networks. Everyone you interact with, including those in the virtual as well as in the real world ) are part of your social network. We inherently think about the world in terms of these networks, including such networks as familial and friendship nenvorks, other students in your major in college, people who work out at the samo time as you at the Wm, and the many "friends" you llr,ay not actually know in real life on platforms such as Facebook or those who similarly liked some- thing on Twirter. The method of social network analysis (SNA) has increas- ingly been used since the Internet and these social media platforms emerged. There are entire textbooks devoted to SNA, so the goal of this chapter is to introduce you to the basics, along with a few case studies that highlight their applicability in the field. SNA is not one type of method but is an approach to analysis and a set of methodological techniques that help researchers describe and explore relationships that individuals and groups have with each other (Scott 2017).

Social networks are tFpes of relationships that can include many dif- ferent forms, such as face -to-face and online interactions, digital economic transactions, interaction with a criminal justice agenqr, geopolitical relations among nation states, and so on. As you can see, there are numerous tFpes

1,,

2i,

ffi.ffii -fl*H

Understand how social network analysis and crime mapping can be used for intelligence-Ied policing as well as basic research.

Describe the different research questions crime mapping can answer.

, Understand how computer technology has ushered in

t .a.

our ability to analyze big data and the effects this has had on criminal justice- related research.

':!: Be able to see the connection between research and investigative policing.

l.'. '',' ' 1

An ai:prtath io a*al',rsis and a s*t *f rrethr:d*lr:$ir:al i*clinir;ues that help

r*s*arrh*r-q de $rrihe anri *xpltr* r* I ati t ii s lr r ps ih at i:at ir i ri rl i v i rl u a I s a;r d

0rri.Jt]. havr uviih *ar:h r:ther,

: :, I'l$es of relaiionsitips t*;ii tari irtr:lurit iil.fin1i rJiffrre ni l*rrns, n*t:h as fa*e-r*-fam

ilnd *irlinr intrra*ti*ns, riigital e c*n*mir t r ait s ar: i i r: r.t s, i li i * ra cl i * n'r,, i t ii a * ff {ii {t'ii ; uslic* i1#*nl;y, il**pr:l itr*al r'* latir:rns

ai?1finfi naiicrn stat*s, and s* rn.

SOCIAL N ETWORK A NALYSIS, CRIME MAPPING, AND BIG DATA

of networks available for analysis. The most important component of any network is that it is relational. That is an important assumption. So far in this book, we have talked about variables that measure attributes of the units of analysis, such as the behavior of people, the crime rates of cities, and so on. Relational data measure the contacts, connections, attachments, and ties that relate one unit to the next (Scott 2017).Ar a result, these relation al data are not properties of any particular unit (e.g., individual, group, .ity) but are"relational systems" of units that are created by connecting pairs of interacting units. Importandy, then, it is the technique used to describe and examine relational data that is the key to SNA. Many of the traditional methods that we have already discussed in this book can also be used to collect relational data. For example, surveys, interviews, participant observation, and secondary data can all be used to generate relational data for SNA' as we will see in the case studies that follow.

Although literally thousands of articles examined aspects of social structure in the early twentieth century, one of the first graphical applications of social networks was created by Jacob Moreno to examine friendship choices. In his classic book,LVho Shall Suntiae? Moreno (1953) describes his definition of sociometry as being in accordance with its eqrmology from Latin and Greek, "with the emphasis . . . on the second half of the tem,'metrum,'meaning measure, but also on the first half, 'socius' meaning companion" (51). Instead of focusing exclusively on the individual or exclusively on an aggregate entity, Mereno believed that the relationship between individuals within a group must also be examined. Using a sociometric test, which required individuals to choose their associates from a group in which they were a member, attractions ar,d.repubions were determined. For example, if your instructor wanted to understand the social strucflrre of the class you are in right now, she or he might ask each stu- dent to hypothetically choose among the students whom they wanted to have sit next to them (attraction) and whom they would like to have moved to another class (repulsion). Responses received from each individual in the group could then be graphed in asociogram, which is a way of representing social configurations, with individuals (or some other unit) represented by points and their social relationships to one another depicted by lines (Moreno 1953). The formal terminology that describes these graphs as well as the units and relationships therein are called several things, depending on the discipline. The social sciences generally call the basic unis in a graph nodes (sometimes called ac'tors or aertices) and nodes are connected by relations (sometimes called. ties,links, arcs, or edges). As noted above, relationship data such as these can be collected from many places, such as ofEcial records, Facebook friends, and so on

ffang, Keller, and Zheng 2017). SNA usually consists of at least two datasets. The first is called the nodelist, where all

of the units of observation are stored. The second defines the relations between these units. One of the most common t)apes of relations data is called an adjacency matrix (sometimes called nenlork matix), wherein the nodes constitute both the rows and the columns and the cells specif, if and what kind of relationship exists between the nodes at the intersection of each row and column. We are going to stick with the simplest case of a binary network, which only distinguishes whether a relation does or does not exist between a pair of nodes (Yang et al. 2017). An example of a network graph and the nodelist and adjacenry matrix upon which it is based is presented in Exhibit 11.1. fu you can imagine, nodes and relations can be much more complicated than this simple example, and special software is required to mathematically describe the numerous networks that emerge from such data. A discussion of these issues is beyond the scope of this text, but we want to provide you with some excit- ing case studies of how SNA is being used in research related to criminology and criminal justice.

Belational data:

l,fi*,t.gu r*$ ihs *oiltact$,

* # fi il o ct i * rz*, atta*hrfi e nt$,

arsd,li** tha+, r*la.tfr olifi ilnit

tc thc noxl,

Sociogram:

h grast?t rspr*$e*linU th*

s*r ia I {r*ilfi # * ralicns, with

individuai* {*r **fflfr rslltrl unil} re pro$*ntoii by tr:aints nnr| t**i r st:*ial r*latir:nsh i ps

ir *rl* afiilth*r *,r,pi*r,*rlbu ! ! n*s,

Nodes:

The hasic ufiits ie,$,, p**tJl*\

in a sr:ci'a"l n*t,u*rk gravh,

sfiffistiff1** {:all*r} a#*rs *r '{#rtlc#s,

Belations:

Th* c*nn*tti*n* i* a n*tu'l*rk

#r ",r$h, g t){t1*,lifft *,* {:'dll * d i i * *q,

liuks, fr{{;s, frr xdgss,

Nodelist:

Th* data **t {;t}rfi.'dirrirrU thr nr:d*$ iufiits rst rsb**rttxti**1

f*r a stcial n*t"rt'*rk anal *in,

Adjacency matrix:

fi, rlala,s*t ra nta i n i n g

i nf * r ntali rs n ab rs trt tlt * r*latir.rns hrhrs*n th* units cf c rne rv atitsrt,sorilot i rilo$

r,a"il*rj a fiatw#rk ffiatrix,

Binary network:

*isir nU u ishe*,ruh*th*r rr

rrlati*nslrip d*cs ar d*cs n*t exist lietwc#r] rifid{ir,

t:::: iii:::t::.::;:t: ::: ::: ::: ::: ::: ::: ::::: :::,::::::i!:i.i:!:t!:::!:!:i.::i.i:i:::i:;:i:i:i:::i:!:i.i:i.i:i i::.i::!::!: :. :!: :l :

[i *E, ffi.iiiri:.ii.ii.ii.i,.ii.'.i.i.ii.r,.i.,.,,i.,.i.,.,.,r,,.i .i. :ii:i:ii:::::l:i::t:l::ii:ra:a:::;:l;i:::ii:i:::i::::i::::,:i::i::ittiiiii:ii:i::t:ti::::::i:iilitli:lt::a::ilil:ilii;::

.:.::i:.::j.:::.:::::,:::.::i::::.:::.:.:.:::::1:::i::j:::.:::.:::::::.::::::ii:::.:::.:::::::.j::.:::.:::j:::::.::::..:::i.::::

::::i:i : :l:::::j::::::::::t: ::: :: :::ri.:il::l::i::1::

: i :i::i::: :::::i:1:: :::: : : : : ::: : i::: ::: :::.:::.:: .l:: '::i:.:

i:ii:::t:it:iti::tit::i;i::::::j:ti:i!::i!:i::i:ii:i::i::it:i::i::i:ii:i:lli::i:i::i:i,:iii:i:tit|t::iiiji;::i:l:i:i

i::i::::i:i:i::l:r::i::i::::::i:u:j;:i:::::i:l::::i:::;:t::11:jiil:i::itiiij:1il:r::::::t::;:::iitilij:t::tjj:ii:t::;

i::::::::i:t::i:l::::::!:::::::li::l::i::ir::::;i:l:i:::i:i::i:i;it;i:::lt:i:i:ti:it:iiiliiiti:ii:ir:i:tiilii:i:::i:i::i::!lti:t:r:

::.::::: ::j::: ::. . :r:::+1 ::::: : : : ::1: :1 : :::!:::i :i :ll:::::il:i:::i:::ii.i.i:li:;:i:::!:i:i:i:i. :!. ::: : .: .. :::

'-.

ou've learned, tcards,

hapter content,

CHAPTER 11 . SOCIAL NETWORK ANALYSIS, CRIME MAPPING, AND BIG DATA 329

Nodelist A Andrei B Barbara C Chris D Den n is

E E rica F Fannv

G Gal ina

H H ans

I oor J Jennv

Adiacencv Matrix A B C D E F G H I J

A n 1 0 0 1 0 0 1 1 0 B 1 0 0 0 1 0 0 0 0 U

C 0 0 0 0 1 0 0 0 U 0

D 0 il 0 U 0 tl 0 il 0 n

E 1 1 1 0 0 0 0 0 0 0

F 0 0 0 o 0 0 0 0 1 0

G 0 0 0 0 0 0 0 0 0 0

H 1 0 il 0 n U 0 U 0 0

I 1 U IJ 0 0 1 0 U 0 1

J 0 0 0 0 0 0 0 0 1 0

Source: Adapted from Yang, Ke1ler, and Zheng2017 .

CASE STUDY

Networks of Terrorist Cells

On September 1 l, 2001, nineteen members of the Islamic extremist group al-Qaeda hijacked four airplanes and carried out suicide crashes into four places in the United States.The targets included the north and south towers of the World Tiade Center, the Pentagon, and Wash- ington, D.C. (in the last instance, the passengers fought the hijackers, preventing the plane from hitting its target and causing it to crash into an empty field in Pennsylvania instead). Almost 3,000 people were killed, including all of the passengers on the planes along with the 19 hijackers and many hundreds ofpeople in the targeted buildings, which included rescuers. About a month before the 9/ll attack Zac*ias Moussaoui, a French citizen of Moroccan descent, and sometimes referred to as the 20th hijacker,was arrested after he raised suspicion at a flight school in Oklahoma by requesting information on flying a 747. Motssaoui was eventually indicted and found guilty in 2006 of six charges, including conspirary to com- mit acts of terrorism. Information from his indictrnent (United States of Arnerica o. Zacarins

330 SECTION lV . TOPICAL RESEARCH DESIGNS

Exhibit L1.1 Example of a Network Graph With lts Nodelist and Adjacency Matrix

Moussaoui 2001), along with other information uncovered by Tbe Nera York Times and the Washington Post have been used to conduct social network analyses of the terrorist cell that carried ofig/ll. One of the first attempts was made by Krebs (2002), who created one of the first network graphs of the terrorist network. An adapted snippet of his graph is displayed in Exhibit 11.2.

Without going into the advanced statistical analysis performed by Krebs to describe the strength of the relations for the terrorists, many conclusions can be drawn from this graph, including the meticulousness with which the hijackers kept their identities unknown even from each other. Krebs explains, "Many pairs of team members were beyond the horizon of observability. . . . Keeping cell members distant from each other, and from other cells,

li*.ii 'ffi\iin

*ed Belras

'il,nn"J,N*,*)* ,r,,,, \ \ 'fi;Alr,[\

ilitiiiiili nl I I I lglrl I\l.lq,lll lLr! q,l

'ffifur=-fu Mamduh Mahmud Salim

Mamoun Dar kazanli

jed Moqed

lid AI - Mihdhar

Bandar Alhaz mi

Mohammed Abdullah

Faisal Al Salmi

,al- Marabh

Flaed HijaZi

Ahmed Alnami

Mohamed Abdi

! Flight AA#1 1 - Crashed into WTC North ffi Flight AA#77 - Crashed into Pentagon

ffi Flight UA#93, - Crashed into field in Pennsylvania Hlrii Flight UA#1 75, - Crashed into WTC Southl+;4i+X v $l$+i Associated with Hijackers

Source:Adapted from Krebs 2002, Figure 4, p. 50.

-Bin al-shibh

Stam Suqami Y/ Ahtfud.Al

Alhaz mi

CHAPTER 11 . SOCIAL NETWORK ANALYSIS, CRIME MAPPING, AND BIG DATA 331

Exhibi t 77.2 Partial Network Graph of the gt77 Terrorist Attackers

r El Mot assadeq

minimizes damage to the network if a cell member is captured or otherwise compromised" (2002,46).Kreb's graph also confirms the fact that MohamedAtta was the likely leader of the cell, as he has the most relations with other nodes.

Krebs analysis shows the benefits of SNA when putting together a case for prosecution. It also highlights the inherent difficulty of using SNA for preventing or uncovering secret illegal networls. Krebs concludes,

The best solution for network disruption may be to discover possible suspects and then, via snowball sampling, map their ego networks-see whom else they lead to, and where they overlap. To find these suspects it appears that the best method is for diverse intelligence agencies to aggregate their information-their individual pieces to the puzzle-into a larger emergent map. (51)

While using SNA for investigative purposes has its challenges, our next case study demon- strates the merits of doing so.

CASE STUDY

Finding a Serial Killer The Green River serial killer (GRK) killed his first victim, 16-year-old Wendy CofEeld, in Kings County, Washington. Her body was found inJuly of 1982 in the Green River, which became the name given to the then-unidentified killer. After 48 other murders, Gary Leon Ridgway was finally charged as the serial killer in September 2001, despite the fact that he was on the list of suspecrc much earlier. Media interest of the cases generated thousands of leads, and these leads compounded with every new victim, which resulted in a huge amount of data to examine. However, large amounts of data are not the only hindrance to solving a case. Like all of us, police detectives can have cognitive biases (see Chapter 1 for a list of common biases), and when combined with an overabundance of information coming from the public, both reliable and unreliable, investigations can go awry.

In a recent papeg Bichler, Lim, and Larin Q0l7) have demonstrated how SNA can be used ' to aid in connecting the pieces ofa growing body ofevidence. Bichler and her colleagues col-

lected data about the Green River murders from multiple sources, including newspaper reports, bools, and court transcrips of Ridgway's trial. They then performed a SNA analysis of the evi- dence over time to determine if SNA could have prevented investigators from keeping another man on top of their suspect list instead of Gary Ridgway. They state, "We argue that by identify- ing which actors shift in structural position during the investigation, it may be possible to reduce the damaging effects of tunnel vision, emphasis on specific evidence, and intuition" (141). The goal of their analysis was to find the connections between the victims and the suspects, along with the places they frequented. Because these murders Iikely involved strangers, Bichler et al. (2017) explain that other sources of information must be added to find clusters. They state,

People sharing social space will emerge when we link wimesses, friends, and associ- ates to the places they frequent, but finding individuals positioned between different

::m'm'ff ;ffi ffi:'-,"i'ff :x[;i;,H:*tH"'suspectinacrimeseries

The categories of the data Bichler and her colleagues examined included victims, suspects, investigatory involvement (e.g., witnesses, body finders), family and associates, body disposal sites, last seen locations, and other investigative material. Because there were different levels of

SECTION lV . TOPICAL RESEARCH DESIGNS332

geographical aggregation (e.9., a red light district, a specific hotel), all data points were placed within a census track Without going into the statistical deails of the study, they were particu- larly interested in "middleman" rtrho connected others not direcdy linked to each other. This is statistically measured with a betweenness centrality score. The researchers also created multiple graphs using data over time, beginningwith graphs using data thatwas available from the beginning of the investigation and creating more graphs until the data were exhausted (about 30 months into the investigation). Exhibit 1 1.3A displays a graph using data from the first six months of the investigation while Exhibit 11.3B depicts the investigation at 18 months, after which there was not much new data to incorporate. Line thiclness indicates the number of shared places between pairs, diamonds depict suspects, gray circles indicate victims, and other circles represent other case nodes (e.g., witnesses, body disposal sites).

Source: Adapted from Bichler, Lim, and LarinQAfl),Table 3,p.1,46.

Betweenness centrality sc0re:

A rtatisti * that rn*ax*r** lltr: *xle*ttr y;hi*h nurk* *'*nrz**tl* r,sinr,r rgsd** t*al arr n*t rlir*r,,,,,,1 link*rlt* ***lt *lh*r i'tt s,**i*l **ttit*rk analysis,

(A)

CHAPTER 11 . SOCIAL NETWORK ANALYSIS, CRIME MAPPING,,AND BIG DATA 333

Exhibit 11.3 Network Graph of Nodes in the Green River Killer Investigation at 6 and 18 Months

CASE STUDY

Using Google Earth to Track Sexual Offending Recidivism While the GIS software that was utilized by Fitterer et al. (2015) has many research advan- ages for displaying the spatial distributions of crime, other researchers have begun to take advantage of other mapping tools, including Google Earth. One such endeavor was con- ducted by Duwe, Donnay, and Jbwksbury (2008), who sought to determine the effects of Minnesota's residency restriction statute on the recidivism behavior of registered sex offend- ers. Many states have passed legislation that restricts where sex offenders are allowed to live. These policies are primarily intended to protect children from child molesters by deterring direct contact with schools, day care centers, parks, and so on. Most of these satutes are applied to all sex offenders, regardless oftheir'offending history or perceived risk ofreoffense.

The impact of such laws on sexual recidivism, however, remains unclear. Duwe et al. (2008) attempted to fill this gap in our knowledge.They examined224 sexoffenders who had been reincarcerated for a new sex offense between 1990 and 2005 and asked several research questions, including "Where did offenders initially establish contact with their victims, and where did they commit the offense?" and "What were the physical distances between an offender's residence and both the offense and first contact locations?" (488). The research- ers used Google Earth to calculate the distance between an offender's place of residence, the place where first contact with the victim occurred, and the location of offense.

Duwe and his colleagues (2008) investigated four criteria to classifr a reoffense as pre- ventable: (1) the means by which offenders established contact with their victims, (2) the

cHApTER 11 . SOCTAL NETWORK ANALYStS, CRTME MAPPING, AND BtG DATA 339

Big data:

h,i#{'i lar7* datasel

t*,rJ,,r:* ni* i ns th r: r: ran d*

oi cases), a***ssilslrt i';t c {}ffi{} ute r - r * ?Ld,?Ll:l * f * r n, th at" is tls*rl t* re,t*al patt*rir$, tr*nris, and ass**iatirrrs

l;e1tru**n ur.uriabl*s ruilh **w

c * m fi u te l" +, **hn*l * g';,

distance between an offender's residence and where first contact was established (i.e., 1,000 feet, 2,500 feet, or I mile), (3) the type of location where contact was esablished (e.g., was it a place where children congregated?), and (4) whether the victim was under the age of 18. To be classified as preventable tfuough housing restrictions, an offense had to meet ceftain criteria. For example, the offender would have had to establish direct contact with a juvenile victim within one mile of his residence at a place where children congreg?.te (".g., pr.lc school).

Resuls indicated that the majority of offenders, as in all cases of sexual violence, viaimized someone theyalreadytnew. Oily35% ofthe sexoffenderrecidivists esablishednewdirectcontact with a victim, but these victims were more likely to be aduls than children, and the contact usually occurred more than a mile away from the offender's residence. Of the few offenders who direcdy esablished new conact with a juvenile victim within dose proximity of their residence, none did so near a school, a parlg a playground, or other locations included in residential restriction laws.

The authors concluded that residenry restriction laws were not that effective in prevent- ing sexual recidivism among child molesters. Duwe et al. (2008) stated,

Why does residential proximity appear to matter so litde with regard to sexual reof- fending? Much of ithas to do with the pattems of sexual offendingin general. . . . Sex offenders are much more likely to victimize someone they know. For example, one of the most common victim--offender relationships in this study for those who victim- ized children] was that of a male offender developing a romantic relationship with a woman who had children. . . . Theyused their relationships with these women to gain access to their victims. . . . It was also common for offenders to gain access to victims through babysitting for an acquainance or co-worker. (500)

Clearly, the power of mapping technologies has changed not only the way law enforcement ofiicials are preventing crime but also the way in which researchers are examining the factors related to crime and crime control.

BIG DATA

When do secondary data become what is now referred to as big data? Big data is a somewhat vague term that has been used to describe large and rapidly changing datasets and the analytic techniques used to extract information from them. It generally refers to data involving an entirely different order of magnitude than what we are used to thinking about as large data- sets. For our purposes, big data is simply defined as a very large dataset (e.g., contains thou- sands ofcases), accessible in a computer-readable form, that is used to reveal patterns, trends, and associations among variables. The technological advancements in computing power over the past decade have made analyses of these huge datasets more available to everyone, includ- ing government, corporate, and research entities alike. Importandy, many researchers now contend that "big daa holds great promise for improving the efEciency and effectiveness of law enforcement and security intelligence agencies" (Chan and Moses 2017,299).

Here are some examples of what now qualifies as big data (Mayer-Schcinberger and Cukier 2013): Facebookusers upload more than 10 million photos everyhour and leave a comment or click on a "like" button almost three billion times per day; YouTirbe users upload more than an hour of video every second; Twitter users are sending more than 400 million nveets per day. If all this and other forms of stored information in the world were printed in books, one estimate in 2 0 1 3 was that these books would cover tfie face of the eath 52 layers thick That's "big."

All this information would be of no more importance than the number of grains of sand on the beach except that these numbers describe information produced by people, available to social scientists, and manageable with today's computers. Already, big data anallses are being used to predict the spread of flu, the behavior of consumers, and the prevalence of crime.

340 sEcnoN rv . ToprcAL RESEARCH DESTGNS

Here's a quick demonstration: We talked about school shootings in Chapter 1, which are a form of mass murder. We think of mass murder as a relatively recent phenomenon, brit you may be surprised to learn that it has been written about for decades. One way to examine inquiries into mass murder is to see how frequendy the term mass murderhas appeared in all the boolis ever written in the world. It is now possible with the click of a mouse to answer that question, although with two key limitations: We can only examine bools written in English (and in a few other languages) and, as of 2014, we are limited to "only" one quarter of all books ever published-a mere 30 million books (Aiden and Michel 2013).

To check this out, go to the Google Ngrams site ftttps://books.google.com/ngrams), type in mass murder and serial murder, and check the case-insensitizte box (and change the end- ing year to 2015). Exhibit I 1.7 shows the resulting screen (if you dont obtain a graph, try using a different browser). Note that the height of a graph line represents the percentage that the term represents of all words in books published in each year, so a rising line means greater relative interest in the term, not simply more boo[s being published. You can see that mass ruarder emerges in the early 20th century, while serial ruurder did not begin to appear until much later, in the 1980s. It's hard to stop checking other ideas by adding in other terms, searching in other languages, or shifting to another topic entirely. Our next case study illumi- nates how law enforcement agencies are also harnessing big daa to make predictions.

Ngrams:

F:'*q ue*cy g raphs, pr*clu*erJ

hy fi**gie '.s dataha**, rf all

"1,;r:rris *rint*t|in mrre lhatt

rtrr* l.ltirrJ *t th* 'tu*rld'n h**ks {}v*t ttmfr (with r*vcrafi* still expanding),

0.0000600%

0.0000550%

0.0000500%

0.0000450%

0.0000400%

0.0000350%

0.0000300%

0.0000250%

0.0000200%

0.0000150%

0.0000100%

0.0000050%

0.0000000%

, -'t ,: 'i'

..l'ii

'' :rti '

t::

il

1 990

mass murder (All)

serial murder (All)

1900 1910 1920 1930 1940 1950 1960

Sour ce : Goo gle B o oks Ngram Viewer, http: / h ooks. go o gle. com/n grams.

CASE STUDY

1970 1 980

Predicting Where Crime Will Occur

If you have seen tlle film Minoity Report, yotrhave gotten a far-fetched glimpse of a world where people are arrested for criminal acts that they are predicted to do, not that they have actually done. FOX also had a television series called Minority Report that was based on the same premise. While crime predictions in tlrese shows are based on clairvoyants (people who can see into the future) and not real data, law enforcement agencies are beginning to use big data to predict both future behavior in individuals and, as we saw with crime mapping, areas where crime is likely to occur in the future.

fu we highlighted earlier, crime mapping allows law enforcement agencies to estimate where hot spots of crime are occurring-where they have been most likely to occur in the past. Caplan and Kennedy (2015) from the Rutgers School of CriminalJustice have pioneered

CHAPTER 11 o SOCIAL NETWORK ANALYSIS, CRIME MAPPING, AND BIG DATA 341

Exhibit 11,.7 Ngram of, Mass Murder and Serial Murder

Risk-terrain modeling (BTM):

lvkcl*linu that uses iJata frrnt

se'Je ral $ourcfis to prcrjici

the prohahility i:f rrirn* *c*u rrin7 in the futur#, usir:# the unrierlying ta*t*rs

al tlt* *n tir *rtn *rti that ;ir* a$$o{:iat*d with iile;,*l b*hauiar,

a new way to forecast crime using big data called risk-temain modeling (RTM). Using daa from several sources, this modeling predicts the probability of crime occurring in the future using the underlying factors of the environment that are associated with illegal behavior. The imporant difference between this and regular crime mapping is that it takes into account features of the area that enable criminal behavior.

The process weights these factors (which are the independentvariables) and places them into a final model that produces a map of places where criminal behavior is most likely to occur. In this way, the predicted probability of future crime is the dependent variable. This modeling is essentially special-risk analysis in a more sophisticated form than the early maps ofthe Chicago school presented earlier. Kennedy and his colleagues (2012) explain:

Operationalizing the spatial influence of a crime factor tells a story, so to spefi about how that feature of the landscape affects behaviors and attracts or enables crime occurrence at places nearby to and far away from the feature iself. When certain motivated offenders interact with suitable targets, the risk of crime and victimization conceivably increases. But, when motivated offenders interact with suitable targets at cerain places, the risk of criminal victimization is even higher. Similarly, when certain motivated offenders interact with suitable targets at places that are not conducive to crime, the risk ofvictimization is lowered. (24)

Usingdau frommanysources, RTM statisticallycomputes the probabilityofparticularkinds of criminal behavior occurring in a place. For example, Exhibit 11.8 displays a risk-terrain map thatwas produced for Irvington, NewJersey. From the map, you can see that several variables were included in the model predicting the potential for shootinp to occur, including the presence of gangs and drup, along with other infrastructure information such as the location of bars and liquor stores.Whywere these some ofthe factors used? Because prwious research and police data indicated that shootings were more likely to occur where gangs, drup, and these businesses were present. This does not mean that a shooting will occur in the high-risk areas, it only means that it is more Iikely to occur in these areas compared to other areas. RTM is considered a use of big data because it examines multiple dausets that share geographic location as a common denominator.

CASE STUDY

Predicting Recidivism With Big Data fu you learned in Chapter 2, the Minneapolis DomesticViolence Experiment and the National Institute ofJustice's Spousal Abuse Replication Projecg which were experiments to determine the efEcacy of different approaches to reducing recidivism for intimate partner violence (PV), changed the way IPV was handled by law enforcement agencies across the United States and across the globe. No longer were parties simply separated at the scene, but mandatory arrest policies were implemented in many jurisdictions across the country, which "swamped the sys- tem with domestic violence cases" fl&illiams and Houghton 200+, +38).In fact, some states now see thousands of perpetrators arrested annually for assauls against their intimate partners. Many jurisdictions are attempting to more objectively determine whether these perpetrators present a risk of future violence should they be released on parole or probation.

Williams (2012) developed one instrument to determine this risk, which is called the Reaised Domestic Wolence Screming Insnwmmt (DVSI-R).To determine the effectiveness of the DVSI-R in predicting recidivism, Stansfield and Williams (2014) used a huge dataset that would be deemed big dau, since it contains information on29,317 perpetrators arrested on familyviolence charges in 2010 in Connecticut and is continuouslyupdated for recent arrests

342 SECTION IV . TOPICAL RESEARCH DESIGNS

and convictions. To measure the risk of recidivism for new family violence offenses OIFVO), the DVSI-R includes eleven items: Seven measure the behavioral history of the perpetratoq including such things as prior nonfamily assaults, arrests, or criminal convictions; prior fam- ily violence assaults, threats, or arrests; prior violations ofprotection orders; the frequency of family violence in the previous six months; and the escalation of family violence in the past six months. The other four items include substance abuse, weapons or objects used as v/eapons, children present during violent incidents, and ernployment status. The range of the DVSI-R is from 0 to 28, with 28 representing the highest risk score.

To examine how the DVSI-R predicted future arrests, Stansfield and Williams (2014) used two measures of recidivism during an lS-month follow-up: rearrests for NFVOs and rearrests for violations of protective or restraining orders only. Results indicated that of the over 29,000 cases, nearly one in four Q3o/") perpetators were rearrested, w:rrh l+% of those rearrested for a violation of a protective order. Perpetrators who had higher DVSI-R risk scores were more likely to be rearrested compared to those with lower risk scores. As you can see, using big data to improve decision making by criminal justice professionals is not a thing of the future, it is happening now. The availability of big data and advanced computer technologies for its analysis mean that researchers can apply standard research methods in exciting new ways, and this trend will only continue to grow.

Source: Obtained from personal correspondence with Leslie Kennedy and Joel Caplan.

CHAPTER 11 . SOCIAL NETWORK ANALYSIS, CRIME MAPPING, AND BIG DATA 343

Exhibit 11.8 Risk-Terrain Map

V'r:::::;;:i;::,a;,1;:ii:::i :::irr''i:;::.

g ,iw...,,,.

'.6 j:ii:f,.jri;:'i:j;iil:iiili:l: n:' r;:i':i .1

':':::::::::::!:':::r:l::il:::::: ..:::::::.:t:..,.{ ,',i::::l:::i:::i:i::::i i*'e

ETHICAL ISSUESWHEN USING BIG DATA

Subject confidentiality is a key concern when original records are analyzed with either secondary data or big data. Whenever possible, all information that could identi$, individuals should be removed from the records to be analped so that no link is possible to the identities of living subjects or the living descendans of subjects (lluston and Naylor 1996,1698). When you use data that have already been archived, you need to find out what procedures were used to preserve subject confidentiality. The work required to ensure subject confidentiality probably will have been done for you by the data archivist. For example, the Inter-University Consortium for Political and Social Research (ICPSR) examines all data deposited in the archive carefully for the possibility of disclosure risk All data that might be used to identiS, respondents are altered to ensure confidentiality, including removal of information such as birth dates or service dates, specific incomes, or place of residence that could be used to identifiz subjects indirecdy (see http://www.icpsr.umich.edu/icpsrweb/content/ICPSWaccess/ restricted/index.htrnl). If all information that could be used in anyway to identiS, respondents cannot be removed from a data set without diminishing data set quality (e.g., by prevent- ing links to other essential data records), ICPSR restricts access to the data and requires that investigators agree to conditions of use that preserve subject confidentiality. Those who violate confidentiality may be subject to a scientific misconduct investigation by their home institution at the request of ICPSR (ohnson and Bullock 2009, 218). The IJK Data Archive provides more information about confidentiality and other human subjects protection issues at its website (https://www.ukdaaservice.ac.ul/manage-data/legal-ethical).

It is not up to you to decide whether there are any issues ofconcern regarding human subjects when you acquire a dataset for secondary analysis from a responsible source. The Institutional Review Board (RB) for the protection of human subjects at your college or uni- versity or other institution has the responsibility to decide whether they need to review and approve proposals for secon dary data analysis. Data quality is always a concern with secondary data, even when the data are collected by an official govefirment agency, and even when the data are "big." Researchers who rely on secondary data inevitably make rade-offs between their ability to use a particular dataset and the specific hypotheses they can test. Ifa concept that is critical to a hypothesis was not measured adequately in a secondary data source, the study might have to be abandoned until a more adequate source of data can be found. Alter- natively, hlpotheses or even the research question iself may be modified to match the ana- lytic possibilities presented by the available data (Riedel 2000, 53). For instance, digital data may be unrepresentative of the general population due to socioeconomic differences berween those who use smartphones or connect to the Internet in other ways and those who live offline, as well as because of the difference between datasets to which we are allowed access and those tlat are controlled by private companies (I-ewis 2015). Social behavior online may also not reflect behavior in the everyday world.

Political concerns intersectwith ethical practice in secondary data analyses. How arerace and etbnicity coded in the U.S. Census? You learned in Chapter 4 that changing conceptu- alizations of race have affected what questions are asked in the census to measure race. This daa collection process reflects, in part, the influence ofpolitical interest groups, and it means that analysts using the census data must understand why the proportion of individuals choos- ing otber as their race and the proportion in a muhiracial category has changed. The same types of issues influence census and other government statistics collected in other countries.

Big data also creates some new concerns about research ethics. When enormous amounts of data are available for analysis, the procedures for making data anonyrnous no longer ensure that it stays that way. For example, in 2006, AOL released (for research purposes) 20 million search queries from 657,000 users after all personal information had been erased and only a unique numeric identifier remained to link searches. Ilowever, staff of The Nru York Times conducted analyses of sets of search queries and were able to quickly identi{, a specific

TOPICAL RESEARCH DESIGNS344 SECTION IV

\

individual user by name and location, based on the searches. The collection of big data also makes surveillance and prediction of behavior on a large scale possible. Crime control efforts and screening for terrorists now often involve developing predictions from patterns identi- fied in big data. Without strict rules and close monitoring, potential invasions of privary and unwarranted suspicions are enormous (Mayer-Schtinberger and Cukier 2013).

CONCLUSION

In our data-driven world, data generally----and the methods examined in this chapter specifically- are increasingly being used in research and in intelligenceJed policing. For example, using SNA and crime mapping are now common techniques used in large police agencies to not only respond to crime but to also reduce, disrupt, and prevent it. As one police investigator explained,

Any data that we can collate online, whether it be that online evidence that may indi cate the commission of offense or assist in making a nernrs, a link to that offense, such as photographs, emails, text messages. . . . We use any data that we can get our hands on laudrlly, certainly, to assist in our investigations. (Chan and Moses 2017 ,305)

The use of such methodological techniques is requiring police academies to incorporate daa analysis components into their training. So even ifyou do not plan to become a researcher yourself, you will likely be required to make sense of data in virtually any career you pursue. Hopefully, you should now have the knowledge and skills required to find and use secondary data and to review analyses of big data to answer applied criminal justice and criminological research questions.

D,, :,Rov:i ew :kG[ : te,FIn:s, w! t h,e F I as.hca rds,

11ffi .ed-efi ffi

'1ffi

fi ffi ..,..1'..........4.4.0.

1....11B.effi E.U#fi .ei$$i..lQent *ffi .',' .:. r:: t.....71 1l'l ., , ,: , ,, ,'' SCOfe, J,)J: ::,:::,:,,,,,, :,:,:,, ::

,,Bi$.d.hta,, ,, 340i,

,,. , , , .,Binarv n,e ork '," Ji,y'';),

1ii;1i.; ;Cffiffie.i.1i;ffiaflpin 1.,l.. ....l3l3..4':.l l. I

iGddgrdfi Hic...iufd#ffifi ti6t...'Syi teml

,.(GIS\..3.34;,, ffi#lligeffEd=tr d..$.dliAi

$..,i...i.i'. 338:

Ngrarns, 341 ,:: '

No,deIis:t,,,:329.:..,,

Relationaldata 3'2.9 : , : ' Relationa'329 , ' ,.

: ,

Sstt:'ri,in ffideHil

...... mfl....l..l,.',.3..42 Social nenvork analySis '328 ', ' ,

Socioerarfi 329 : '

SNA,uses relation,al data to:examine the patterns in i s,ocibl fbladonshipC, that, individuals and gioups have'

Cnm.e. rn:tpping| :fof res,earch, purposes is generally

used, to,identifir .the,sp,atial

distribution of Crime, along

y.ith the.1oc1,al in{idators, (su9,! ,Tr p,?n.rry,and soiial',,

disoigani2atil6fi),i.... e[.].i. e1...$t rlg..;ilistfibfited.11acrd$S .

::..i. . -..: I :

aieas, (e,g., neighborhobds, ce'nsui *acti),, : , ' ,

Uaing h"g. dataiets, often ,ieimed. big dtta,,,to,: :, ,,,,, , determine ffifi $1;.6,fi1 ;,';. f ;5...;1{fl..$6[iiiil[;.i.; ih. n.dffiQfia,,

of,advanced computer technology,lBi$,data and ' ,, 1 advanced comprrtir technology hrr.,helped, to, , ,,, advance crime mapping, te,chniques in, cre,ative,,,ways, i*.etU.a!n$...te*...fi1.udicti#u'..ffid.Aa,ti g'..l.lechniquei,,suCh

aSRTM. i: i, , ., i,,- ,', ,':' : , ': ,: ;::

.3.il.5,

,, ,S' rch,the sitb for,,informa,tion on crime mappinb, .

, and you will find rhat the site Containi,,a mnlaiftde

, ,,,:,:,,:,: ,,, :..:,

:,,,,:,,: i.'l...l..i..ffi6...u...s.llcdnsus..nlrueut..

$md.. a$6.....(htrFu/l'

;,,,ceflsus:$oV) contains extenslVe. reports.,of censu,s data,

, , :i- Iudin:$,population datal economic indicators,,and

,, other infbrmation, :acquired through the IJ. S,,,Census

:

,:;**erous' subject5,,.and topi4t, that can: b,e'used,to

, , *tke comparisons ,among different states or cities. , Fihd the,QuickFa,cts option ,and choose your own :

,.:,:,.,srrt.. xo#,pl.t : tne.,counr1r in,:which ybi" ,tive and inpy ', down several statis,tics of interest. Rbpeat this ,Process ,

,,for, other,:counties, in your, ste,te,. Use,the data you, have , ,Collected,to compaie,your county with other counties

.ni,$.....il]ii]$ieHfi:::-..'.lil.ea;.that,is,;s:.ffiu......te]c6*:d$]]ii:ll.;;.........

.o,f..,phone...'..c ll$;..Eflftef

,..$obts, or,picture$,...teken,,.bs...,...'........,.

indM =#els::,i i::itheii.i..iid,ailt

li#es1...ffiat..li ite oas..i$ uld. be imposed re$arding access,:to,,and use',of such data

once they have becom. ,ggre$tua into massive i

datasets?, Is removing explicit,identifieis sufficiert,

H 61ee .affi.{.ltrheil.'... .d.esl,...a.cc.e$S......t6....bi$..: *tu...+iolate.....ri$hts.

toorivacv,? , , r : ,' : I ,J ,-,-, :'

trn J anu a*, 2 0 12, Face b o ok conducte d an' experirnent in,*hich,emotional cues were,,manipulated for 6,$Q;003

t::

5:lt

1... .u$ f$;....S0Ue...S i, e .s,ti:$ fies tnd photos on Facebookb..,.... homepage containing man, p:r*. words, while others sa#.' e$bdne;. pleeStfit.,#ords..ffieimessa$e$,,,$dilr.,,,...,. subsequendy bl these users were,a'litde more likely to re ec[,&e.emo.tional, el,bf e #Or s $ey...had been chosen randornlrr to see. When this experiment was

:

Scieires (Iftr*ir, G"iitory, and Hr"io.k20i4;, so*e people were outraged. What do you think of the ethicsr).

:/;..a,

Dataset Description

201,2states d'ata,sav This stat€,-level dataset compriles o{fiCial statistics'from varrious offidial sources, such as the census, health department records, and police departments, lt includes basic demographic data, crime rates, and incidence rates for various illnesses and inf ant mortality for entire states.

DescriptionVariable Name

regio,h9

p.e.rr,f' A m., $ O Ve qt.y The percentage of f am ilies below the poverty line in a state

murderRt The murder rate reported per 100,000 for a state

cHAPTER 11 . SOCTAL NETWORK ANALYSIS, CRTME MAPPtNG, AND BtG DATA 347