Data Analysis Project

profilegeoffreyzhan
TheDataGameControversiesinSocialScienceStatistics4th1.pdf

2

The Data Game

3

The Data Game

Controversies in Social Science Statistics

Fourth Edition

Mark H. Maier Glendale Community College

and

Jennifer Imazeki

4

San Diego State University

5

First published 2013 by M.E. Sharpe

Published 2015 by Routledge 2 Park Square, Milton Park, Abingdon, Oxon OX14 4RN

711 Third Avenue, New York, NY 10017, USA

Routledge is an imprint of the Taylor & Francis Group, an informa business

Copyright © 2013 Taylor & Francis. All rights reserved.

No part of this book may be reprinted or reproduced or utilised in any form or by any electronic, mechanical, or other means, now known or hereafter invented,

including photocopying and recording, or in any information storage or retrieval system, without permission in writing from the publishers.

Notices No responsibility is assumed by the publisher for any injury and/or damage to

persons or property as a matter of products liability, negligence or otherwi se, or from any use of operation of any methods, products, instructions or ide as

contained in the material herein.

Practitioners and researchers must always rely on their own experience and knowledge in evaluating and using any information, methods, compounds, or

experiments described herein. In using such information or methods they should be mindful of their own safety and the safety of others, including parties for whom

they have a professional responsibility.

Product or corporate names may be trademarks or registered trademarks, and are used only for identification and explanation without intent to infringe.

Library of Congress Cataloging-in-Publication Data

Maier, Mark, 1950– The data game : controversies in social science statistics / by Mark H. Maier and

Jennifer Imazeki.—4th ed. p. cm.

Includes bibliographical references and index. ISBN 978-0-7656-2979-1 (hardcover : alk. paper)—ISBN 978-0-7656-2980-7 (pbk. : alk. paper) 1. Social sciences—Statistical methods. 2. Social problems— Statistics. I. Imazeki, Jennifer. II. Title.

HA29.M236 2012

6

300.1′5195—dc23 2012013223

ISBN 13: 9780765629807 (pbk) ISBN 13: 9780765629791 (hbk)

7

1.

2.

Contents □□□□

List of Figures, Tables, and Boxes

Preface to the Fourth Edition

Acknowledgments

Introduction

The Purpose of This Book

How to Use This Book

Demography

Data Sources

U.S. Census

American Community Survey

Vital Statistics

Controversies

Population Counts

Privacy in the Census: Double-Edged Sword?

Will There Be a Population Boom?

Undocumented Immigrants

Race and Ethnicity

How Big Is the Gay Population?

Households and Families

Summary

Case Study Questions

8

3.

4.

Housing

Data Sources

U.S. Census

American Housing Survey

Other Industry Data

Price Data

Controversies

Housing Crisis

Homeownership

Racial Discrimination by Banks

Geographic Units

Segregation

Is Your City the Best Place to Live?

Summary

Case Study Questions

Health

Data Sources

U.S. Health Surveys

Other Government Surveys

Private Surveys

Worldwide Data

Controversies

Infant Mortality

Abortion

Are We Living Longer?

How to Measure Longevity

Cancer

9

5.

6.

HIV/AIDS

Are Americans Getting Fatter?

Chicago Heat Wave: What Caused the Tragedy?

What’s Unsafe on the Road? Speed, Texting, Teens, Motorcycles, or Alcohol?

Drug Use

Benefit-Cost Analysis

Summary

Case Study Questions

Education

Data Sources

National Center for Education Statistics

U.S. Census Bureau

Other Surveys

State Data

Controversies

Poor Data

Educational Attainment

Testing

Charter Schools: Are They More Effective Than Regular Public Schools?

Teacher Compensation

Summary

Case Study Questions

Crime

Data Sources

Uniform Crime Reports

10

7.

National Crime Victimization Survey

National Incident-Based Reporting System

Controversies

UCR, NCVS, or NIBRS?

Crime Is Down—And We Don’t Know Why

Are There More Female Criminals?

Human Trafficking: How Often Does It Occur?

Where Is Crime the Worst?

Rape

Does Poverty Cause Crime?

Why Is the Black Crime Rate So High?

Does Prison Pay?

Hate Crimes

Does Capital Punishment Deter Murder?

More Guns/More Crime or Less Crime?

What About White-Collar Crime?

Summary

Case Study Questions

The National Economy

Data Sources

U.S. Commerce Department

U.S. Labor Department

U.S. Federal Reserve Board

U.S. Small Business Administration

International Statistics

Controversies

Which GDP?

11

8.

9.

Adjusting GDP Growth for Inflation

Problems with GDP

Underground Economy

Intercountry Comparisons

Has China Caught Up with the United States?

Measuring Productivity

The Savings Rate

Consumer Confidence

International Statistics

Summary

Case Study Questions

Wealth, Income, and Poverty

Data Sources

Wealth

Survey of Income and Program Participation

Indirect Estimates from Tax Records

Direct Counts

Income

Controversies

Wealth

Income

Are We Better Off?

Poverty

Summary

Case Study Questions

Labor Statistics

Data Sources

12

10.

U.S. Bureau of Labor Statistics

U.S. Census Bureau

Controversies

Unemployment

The Minimum Wage and Jobs

Unions

Is the Workplace Safe?

Squirrel Cage or Easier Times?

International Labor Statistics

Summary

Case Study Questions

Business Statistics

Data Sources

Public Corporations

Privately Held Corporations

Small Businesses

Aggregate Statistics

Controversies

Who Is the Biggest of Them All?

Are the Big Too Big?

Is Small Beautiful?

Were the Bailouts a Success?

How Much Profit?

How Now Dow?

Picking Stock Winners

Summary

Case Study Questions

13

11.

12.

Government

Data Sources

U.S. Budget

Military Spending

Money

Prices

Voting Data

Controversies

How Much for the Military?

How Much for Welfare?

How Big Is the Deficit?

A Debt Monster?

Taxes

Measuring Money

Inflation

Currency Rates

Fewer Voters?

Summary

Case Study Questions

Public Opinion Polling

Data Sources

Private Polling Organizations

Media Polls

Research Centers

Controversies

Sampling

Survey Design

14

13.

Predicting Elections

Polling Standards

Summary

Case Study Questions

Conclusions

Poor/Missing Data

Improved Data

Conflicting Definitions

Comparing Apples and Oranges

Comparing Apples and Orangutans

The Mathematics of Social Science Statistics

The Reporting of Controversies

Biased Analysis: Slanting the Numbers

What’s a Researcher to Do?

Index

About the Authors

15

3.1

4.1

5.1

5.2

7.1

7.2

8.1

11.1

2.1

3.1

3.2

4.1

5.1

7.1

8.1

9.1

10.1

List of Figures, Tables, and Boxes

□□□□

Figures

Case-Schiller and FHFA Indices, 2000–2011

U.S. Infant Mortality, 1970–2010

College Enrollment by Gender, 1970–2009

SAT Scores, 1972–2009

U.S. Productivity 1970–2010

U.S. Savings Rate 1980–2010

Are We Better Off?

U.S. Federal Taxes 2009

Tables

Types of Households

If Mortgage Tax Deduction Is Eliminated

The First Shall Be Last

Infant Mortality by Country

U.S. Performance on International Tests

GDP Revisions (Real Annual GDP Growth for January–March 2010)

Percent in Poverty Under Different Poverty Measures

Unemployment (November 2010)

U.S. Corporate Size, 2011

16

2.1

2.2

2.3

2.4

2.5

2.6

3.1

3.2

4.1

4.2

4.3

4.4

4.5

4.6

5.1

5.2

5.3

5.4

6.1

6.2

6.3

6.4

6.5

6.6

7.1

7.2

Boxes

Undercount in History

Interstate Migration

Black “Insanity”: An Argument for Slavery

Second Largest “Ethnic” Group: “No Response”

If It’s Tuesday, I’m Swedish

Head of Household

American Community Survey Versus American Housing Survey

Housing Price Indices and Housing Quality

Are Americans “Food Insecure”?

The Oldest Old

Likelihood of Breast Cancer

How Much Did You Eat?

Did Depenalization Increase Drug Use?

Cost of Tamper-Proof Closures

Schools Are Only as Bad as We Think They Are

Are Women Crowding Men Out of College?

Are Boys Better at Math and Science?

Simpson’s Paradox

Is New York City Telling the Truth?

Misleading Boundaries

Do Sex Offenders Repeat Their Crime?

Wrongful Incarceration

Are You More Likely to Be Murdered by an Acquaintance Than a Stranger?

Who Pays for Research on Guns?

What’s the Cost of a Holiday?

Comparing Apples and Oranges

17

7.3

8.1

8.2

8.3

8.4

9.1

9.2

9.3

9.4

9.5

10.1

10.2

10.3

10.4

11.1

11.2

11.3

11.4

11.5

12.1

12.2

12.3

12.4

Errors in Trade Data

Mean, Median, and Mode

Measuring Wealth

I’ve Got a Secret

Who Is the Wealthiest of Them All?

How to Survey the Unemployed

Is Unemployment Structural or Cyclical?

Too Few U.S. Engineers?

How Many Jobs in a Lifetime?

Are Men Overworked?

Which Is Larger: GM or Switzerland?

Top of the World (2011)

North American Industrial Classification System

Do 90 Percent of New Businesses Fail?

How Many Iraqi Civilian Deaths?

Do the Rich Pay 90 Percent in Taxes?

Taxation at the Margin: Working More Does Earn You More

Annual Inflation Rate

Effect of Rounding

Bias from Self-Reporting

Randomizing Response Techniques

Black Candidates/White Voters

What Do Women Want?

18

Preface to the Fourth Edition □□□□

This book, The Data Game, fills the need for a companion text to courses on the collection and use of social science data. It is suitable for undergraduate courses in statistics and research methods, as well as for first courses in statistics offered in many social science graduate programs.

An analysis of statistical source material is readily justified and often called for by social science practitioners. As illustrated throughout this book, many public policy debates arise because of different interpretations of the underlying data. For topics as diverse as the size of the middle class and the crime rate, ambiguities in the data produce statistics that appear to support opposite positions. Cases such as those included here bring statistics to the real world, demonstrating to students the critical role that data play in a wide variety of social issues.

The subject areas in this book include demography, housing, health, education, crime, the national economy, wealth and poverty, labor, business, government, and public opinion polling. For the fourth edition of The Data Game, Jennifer Imazeki serves as a new coauthor with Mark Maier, bringing together economists with several decades of teaching experience and expertise across the subjects covered in the book. Maier and Imazeki have updated and revised each chapter, adding new data sources and controversies to better illustrate key concepts.

Students in all disciplines will benefit from this interdisciplinary coverage because much research involves social statistics that transcend narrow subject areas, and because almost every example teaches a lesson that has relevance for all the social sciences. Several common questions recur across the disciplinary spectrum represented in this book: How do the popular media misinterpret social statistics? Why are some social statistics continually cited, even though they are widely known to be misleading? Why are some data collected in abundance, while many critically needed data are missing? Why are the categories into which data are organized so critical for statistical analysis? In addition, fundamental

19

statistical techniques are reviewed, including the difference between surveys and complete counts, the interpretation of index numbers, the use of means and medians, and the use of absolute and relative measures. These issues are probed further in case study questions at the end of each chapter, and they are summarized in the final chapter.

In brief, this book is an invitation to social science research for students, whether as future practitioners in social science or as enlightened citizens. The Data Game demonstrates the excitement, the frustrations, and always the importance of social statistics as an instrument for understanding and changing the world in which we live.

20

Acknowledgments □□□□

We owe profound thanks to the specialists who offered guidance to the four editions of The Data Game. These generous individuals include Randy Albelda, James Campen, Robert Carson, Todd Easton, Matthew Edel, Lou Ferlerger, Mona Field, David M. Gordon, Theodore Joyce, Peter L. Maier, Scott R. Maier, Richard McGahey, William D. Mosher, Michelle Naples, Robert Pollin, John Queen, Michael Schiller, David Stern, Michael Swerdlow, Robert Untermann, Thomas E. Weisskopf, and Steven White. Jennifer would also like to acknowledge suggestions and research assistance from several students in her class, The Collection and Use of Data in Economics, particularly Rahmann Brown, Rolan Desmarais, Noe Garcia, Nabyl Leon, Brian Moore, and Jaime Roberts. At M.E. Sharpe we thank George Lobell and Stacey Victor. Jennifer would particularly like to thank Joey for the many, many evenings of discussion and helping me keep things in perspective. Mark sends greetings to Sam and Julia, once students, and now professionals in their own fields, with whom I now can share ideas. And, thank you to Anne, my life partner, for support, wisdom, and just being there.

21

The Data Game

22

1

Introduction □□□□

The Purpose of This Book

Interpreting social statistics can be frustrating. It seems as if there are numbers to prove anything—even entirely opposite points of view. For example, there are statistics to “prove” that the average U.S. family is becoming richer and that it is becoming poorer; that U.S. students are learning more and that they are learning less; that Americans vote less often or just as much as in the past; and that illegal drug use is rising and that it is decreasing.

A quite natural inclination is to reject all statistical results. After all, why trust any number if equally convincing numbers prove precisely the opposite conclusion? This cynical view is said to have been summed up by the nineteenth-century British politician and writer Benjamin Disraeli, who —according to American humorist Mark Twain—listed, in descending order of credibility, “lies, damn lies, and statistics.” Indeed, examples abound in which politicians, journalists, and policymakers fit statistics to their preconceived ideas. This book provides hints to alert readers to ways in which statistics can be misused.

Statistics are more than just sophisticated lies. In most cases the source of contradictory numbers is sincere disagreement between experts. If we can find out why the experts reach different conclusions, we will understand much more about the problem being analyzed. Consider the example of life expectancy, seemingly an easily measurable statistic and indeed one for which nearly complete data is available. However, U.S. life expectancy, about 78 years in 2009, turns out to be controversial when used to evaluate the nation’s health care system. Does the United States rank near the bottom of all developed countries? Or, if one omits early

23

deaths primarily from homicides and accidents, does the United States have the longest life expectancy? Moreover, there are unexplained anomalies in the statistics, such as high longevity for Mexican immigrants, perhaps because some return to Mexico to die, or because the unhealthy are less likely to leave the land of their birth in the first place, or perhaps simply because this group is less likely to smoke than others.

Other statistical controversies presented in this book teach a similar lesson. For example, experts disagree about whether the death penalty deters murder, whether mortgage lending is affected by racial discrimination, and whether taxes are becoming more unfair. No book of any reasonable length could answer these or any of the many other policy questions raised in the following chapters. Rather, the intent here is to show why well-respected researchers are able to reach such contradictory conclusions. In some cases such understanding will help us decide which side is correct; often it is less important to decide which side is correct than to uncover the complex measurement problems that underlie a given issue.

Another purpose of this book is to help researchers, both students and more experienced practitioners, use social statistics correctly. Consider the following hapless case:

A social science researcher wanted to study the effect of military spending on jobs. Do communities with large military contractors benefit from increased employment, as advocates of military spending argue, or does military spending create relatively fewer jobs than other kinds of government spending, as critics of military spending have charged? To answer this question, the researcher obtained records of military contracts from the U.S. Department of Defense (DoD); the contracts were arranged by the city in which the contractor was located. To measure the number of jobs generated in a community by the presence of the military, the researcher consulted publications of the U.S. Bureau of Labor Statistics listing employment by location. Armed with a microcomputer statistical package and all the latest knowledge about statistical probability, the researcher was ready to punch in the numbers and find the answer to his question.

Then the project stalled. Just about every piece of data showed signs of being problematic. The DoD numbers were unusable because they listed contracts by the year in which they were awarded, which was not necessarily the year in which they were spent. To make matters worse, the location where the contract was awarded was not necessarily the location

24

where people were hired; in fact, many contracts were subcontracted to other companies of unknown locale. There were problems with the employment data as well: When one employer dominated the industry, the data were not available on the grounds that the information would betray that company’s trade secrets. Finally, data from the U.S. Defense and Labor departments were incompatible because of different definitions of location—the “cities” and “metropolitan areas” under consideration in each survey were not necessarily the same.

One of us was the ill-informed researcher in this case. (Okay, actually it was Mark.) But we are not the first researchers to see good ideas flounder because of unusable data. It is a recurring complaint in the social sciences that researchers, from the student in training to the advanced scholar, do not know enough about the data they use. By examining the pitfalls encountered by previous researchers, this book will help today’s users of social statistics be more aware of which data sources are available and what limitations they possess. If—prior to undertaking the failed research on military spending—Mark had been aware of the Census Bureau confidentiality problems that frequently arise when researching corporations (see Chapter 10), he would not have expected to find employment data for large firms that dominate a single city’s industry. Similar examples will serve as cautionary tales for other researchers.

In summary, The Data Game is written for two groups: general readers and researchers. First, it will help everyone who is confused by statistics that seem to prove everything and anything. By sorting out the reasons behind seemingly contradictory statistics, we can better understand the issues under debate. Second, this book will assist researchers in assessing the problems of their underlying data. Without such knowledge, many social science projects will fail, as in the case of the military spending research; even worse, other projects will proceed without sufficient caution as to the data’s limitations.

How to Use This Book

Each of the chapters in this book is devoted to a single subject: demography; housing; health; education; crime; the national economy; wealth, income, and poverty; labor statistics; business statistics; government; and public opinion polling. Although students of a particular field will find the chapter in that area most useful, the book is intended to be read as a whole. Social scientists work within their own narrow

25

specialty at considerable cost. Most projects use data from outside a narrow discipline, and the data may have limitations that are unknown to the researcher. For example, almost every area in social science measures variables on a per-person basis, a calculation that presupposes accurate population data, which is not necessarily a warranted assumption (as is discussed in Chapter 2 on demography). Similarly, geographic units such as Metropolitan Statistical Areas (Chapter 3) and corrections of price data for inflation (Chapter 11) are common throughout social science research. Thus it is useful for researchers to consult chapters beyond their specialty.

Each chapter in The Data Game opens with a brief overview of the “Data Sources” for a particular area of social science. These sections will acquaint readers with the names of the most important government and private-sector data sources, as well as the statistics they publish, and in many cases offer an illustrative “data sample.” The organizations, major publications, and websites for these data sources are listed in a table at the opening of each chapter.

Following the “Data Sources” section are “Controversies,” a series of debates about the use of statistics in each area. No attempt has been made to cover every debate in each field; instead, controversies were selected based on their relevancy to recent public policy disputes. These include controversies in the news, such as the U.S. Census population undercount, the disappearing middle class, and the number of homeless individuals in the United States. A second criterion for including a controversy was its usefulness as a revealing illustration of a statistical issue. For example, while the rating of individual cities as the best places to live or the lists of the nation’s largest corporations are not particularly critical policy questions, debates about these numbers teach important lessons on the use and misuse of ranking in social statistics.

All the controversies discussed by the authors obviously predate the publication of this fourth edition of The Data Game. But readers should resist the temptation to reject the examples from past years as out of date. Almost all the debates are ongoing, perhaps with different individuals or institutions, but still involving the same issues. As long as the underlying social and economic system remains the same, controversies based on fundamental measurement problems will stay with us.

Each chapter concludes with “Case Study Questions,” which instructors may assign to students for further investigation and the stimulation of thought about the issues raised in each chapter. In most cases there is no

26

single correct answer to the questions; instead, they pose problems frequently encountered by researchers. In many instances citations related to the questions are provided as an aid to those readers interested in exploring underlying issues in greater depth.

Finally, readers should not overlook the “Notes” section at the end of each chapter. The works cited for each subject area are highly recommended guides to data sources, including both official government handbooks and privately published works. References tied to “Controversies” include popular presentations in magazines and newspapers, which are often the most accessible sources and are worth consulting to evaluate how the topic was generally understood—or misunderstood. In addition, there are references to summary reviews of each public policy debate; such reviews typically appear in academic journals. Citations for the key technical articles for each controversy are listed as well.

27

2

Demography □□□□

Demography, the scientific study of population, provides researchers with some of the most fundamental social statistics. This chapter looks at demographic controversies surrounding the size of the population, the classification of individuals by race and ethnicity, household characteristics, immigration, and trends in cohabitation and divorce. These controversies have public policy implications for congressional representation, social security financing, affirmative action, and family law. In addition, because demographic data are used in so many areas of social science, the potential problems described here have implications for research outside the field of demography itself.

In the United States, the major source of demographic data is the U.S. Census, an attempt made every 10 years to count each individual, citizen or noncitizen, with or without legal documentation, who resides in the country. Less well known but similarly comprehensive are U.S. Vital Statistics that tabulate most births, deaths, marriages, and divorces. Researchers accustomed to surveys and the nagging problem of sampling error might wonder how there can be controversies about statistics based on complete data. This chapter identifies four major problems: (1) despite valiant efforts to be inclusive, not everyone is counted; (2) the categories used to classify race, ethnicity, and type of household are arbitrary and therefore subject to debate; (3) there may be multiple ways to define and measure variables such as the divorce rate, so care must be taken to match the definition to the question of interest; and (4) predictions about trends in demographic data are often based on assumptions that may not hold in the future.

Where the Numbers Come From

28

Organizations Data sources URL Bureau of the Census, U.S. Department of Commerce

U.S. Census, American Community Survey

www.census.gov

National Center for Health Statistics, U.S. Department of Health and Human Services

U.S. Vital Statistics www.cdc.gov/nchs/

Office of Immigration Statistics, U.S. Department of Homeland Security

Records of admissions and naturalizations

www.dhs.gov

Data Sources

U.S. Census

Collected every 10 years since 1790, the U.S. Census is the longest- running consecutive data set in the world. It is also the world’s largest data set, compiling information about the sex, age, marital status, and race of nearly every individual residing in the United States. A number of surveys sponsored by the U.S. government use the Census as a statistical base, most notably the Current Population Survey (see Chapter 9).

Because of limited space, only a few questions can be asked on the short form filled out by all households. It was reduced in the 2000 Census to only eight questions and characterized as “the most efficient, cost-effective census in the nation’s history” by the then-acting director of the Census Bureau, James Holmes. In 2010, two additional administrative questions were added, but at 10 questions, it was still far shorter than any questionnaire sent prior to 2000. Nonetheless, every Census comes under fire for prying into individual privacy, a charge that often leads the Bureau to defend the importance of this data source.

Data Sample: In the 2010 U.S. Census, 77.3 percent of Hawaii residents reported a race other than non-Hispanic White, the highest percentage in the United States; the lowest minority percentage was 5.6 percent in Maine.

American Community Survey

For each decennial U.S. Census through 2000, about one in six households also received a “long form” asking 46 additional questions on such diverse

29

matters as occupation and level of education. The long form was abandoned in 2010 and replaced with the American Community Survey (ACS). First administered in 2005, the ACS is a national survey that, on an annual basis, provides the same demographic, housing, social, and economic data that was previously collected only once per decade; three- and five-year average estimates are also provided for researchers. The ACS is a continuous survey, with questionnaires sent out every month, and the sample for annual estimates is much smaller than the decennial Census long form; note, however, that the ACS sample for the five-year estimates is not that different from the long form (one in eight for the ACS rather than one in six for the long form).

Data Sample: According to the American Community Survey, there were approximately 594,000 same-sex couple households in 2010; 42,000 of those were in the six states (including Washington, DC) that permit samesex marriages.

Vital Statistics

Most countries have a system for recording births, deaths, marriages, and divorces—data commonly referred to as vital statistics. These data were among the first ever collected and thus were used by historians to estimate population for time periods before governments instituted national censuses. For the United States, vital statistics are more recent and, in some cases, are still incomplete. Data for U.S. Vital Statistics are collected by individual counties and states and are then assembled on a nationwide basis by the National Center for Health Statistics. Hospital records and doctors’ reports provide a nearly complete count of births and deaths (death rates are discussed in Chapter 4). In most cases, locally generated information captures a comprehensive count of marriages and divorces for the entire nation. However, because of inadequate records in some states, as of 2009 detailed statistics on national divorce rates omitted six states, including California, where a significant share of all divorces occur.

Data Sample: In 2009, U.S. Vital Statistics recorded 4,130,665 live births, of which 1,693,658 were born to unmarried women.

Controversies

Population Counts

30

How Many People Are There?

Since 1960, the U.S. Census has relied on self-enumeration, that is, voluntary completion of forms mailed to individual households. Finding those who fail to return the forms is the Census Bureau’s major expense, involving nearly half a million employees recruited for the 2010 Census alone. The homeless, migrants, and transients were counted in a less-than- precise “shelter and street night” visit to inexpensive hotels, shelters, parks, train stations, and abandoned buildings.

For all its efforts, the Census Bureau admits that it misses some individuals. In 1990 the shortfall was particularly problematic—between 3 and 5 million people, or roughly 1 to 2 percent of the U.S. population. Although this error might seem small at first glance, it had an especially large impact because it was divided unevenly across the country. The undercount was estimated to be four times as high for blacks as for whites and nine times as high for inner-city residents as for the general population. Research on the undercount suggests that most minority households were successfully contacted but that distrust of government officials caused some individuals in those households not to be counted. In one study, more than 12 percent of young black men were omitted from the Census.

After 1990, data collection methods improved, reducing the frequency of undercounting in the 2000 Census; still, the Census Bureau estimates that it missed around 3 million individuals, or 1.18 percent of the population, in 2000. Further changes were adopted for Census 2010, including making forms available in more languages, increasing the number of multilingual census takers, and opening assistance centers where people could pick up forms and get help filling them out. However, at this writing, it is unknown how large the 2010 undercount really was.

The undercount has important implications for congressional representation and dispersal of government funds. According to one study, the 2000 undercount meant a loss to states of over $4 billion in federal funds between 2002 and 2012, with shortfalls most pronounced in metropolitan counties. And although it is unlikely that the undercount alone is responsible for New York losing two congressional seats after the 2010 Census, New York City submitted a challenge to its census count, arguing that census takers missed at least 50,000 people.

No population count will ever be completely accurate, but when the

31

undercount is higher for some parts of the population than for others, the problem can present a major challenge for researchers. official population data may be misleading, especially for research that focuses on minority groups or innercity residents, who are most likely to be undercounted. For example, in 1987 sociologists Reynolds Farley and Walter R. Allen recomputed Census data to take into account the maximum effect the undercount might have had on important social and economic variables for black men aged 20 to 34. This change eliminated the apparent shortage of men relative to women and, according to Farley and Allen, disproved the thesis that black women remain single because there are relatively few black men in their age group.

Farley and Allen found that other important social statistics were affected by the undercount, albeit in a less dramatic manner. The pay gap between young white men and young black men, measured at 36 percent with traditional statistics, increased only slightly to 39 percent, assuming uncounted black men had extremely low earnings. Similarly, the difference in unemployment rates between black and white men, measured at 5.6 percent, increased to 7.1 percent, assuming the uncounted have very high unemployment rates. Thus, correcting for the undercount caused only a small increase in the measurement of the already severe economic deprivation among young black men.

Box 2.1 Undercount in History

Concern about Census undercount dates back to 1790, when 3,929,326 individuals were counted, but, as Thomas Jefferson wrote to George Washington, the omissions were “very great,” and “we are certainly above four million.”

Following the 1870 Census, New York City and Philadelphia successfully demanded recounts, increasing their official population by just over 2 percent each. Indianapolis, home of the powerful Senator Oliver Morton, obtained a recount based on land annexed after the census date so that the city could reach the prestigious 50,000 level. Subsequent research suggests that the biggest error in the 1870 Census—an undercount by 10 percent of recently freed southern blacks—went uncorrected.

32

Sources: Jefferson to Washington in Robey, “Two Hundred Years,” p. 35; 1870 Census in Anderson, American Census, p. 89.

The lesson for researchers is that population data should be analyzed for the effect of the potential population undercount. Even if the correction is relatively minor, as was the case with income and unemployment, such adjustment adds credibility to research results. Farley and Allen calculated their own estimates of the potential undercount. The availability of official Census Bureau adjustments will simplify the task for researchers, although debate will likely continue about the accuracy of the revised numbers as well.

Who Gets the Prisoners?

The undercounting in urban areas is exacerbated by the way the Census treats incarcerated individuals. For purposes of the Census, prison inmates are considered residents of the prisons, and many prisons are located in rural, mostly white towns. However, a large majority of those inmates came from, and upon release will return to, urban communities elsewhere in the state. Thus, counties with prisons appear to have larger populations; according to one study by the Prison Policy Initiative, there are 21counties that have at least 21 percent of their population in prison. The higher population count in prison counties means additional funding and political representation that they would not have if prisoners were counted as residents of their original jurisdictions.

Aside from the political implications, researchers need to be aware of the implications for demographics. Because more than 90 percent of prisoners are men, and nearly all prisoners are over the age of 18, the gender and age distributions in both the prison counties and the home counties of the prisoners will appear skewed. When prisons open (or close), it can give a misleading impression of growth (or decline) in the county population. In addition, it becomes difficult to interpret changes in black and Hispanic populations: The 2000 Census showed many Hispanics moving to more rural communities, but it is unclear whether that shift occurred by choice or because of higher rates of incarceration. Researchers should be careful to consider these issues when dealing with data that includes prison counties.

33

How Did 1.3 Million People Disappear in One Year?

Even if census counts were exactly correct, a full census is conducted only once every 10 years. For the nine years in between, the Census Bureau creates population estimates by using administrative data (such as births and deaths) to estimate annual changes in population that are then applied to the baseline decennial counts. Specifically, county-level records are used to compute overall population growth rates, which are applied to the previous year’s count. Then housing records are used to allocate the population growth among the individual cities within the counties. Thus, the census year count is the only true count; the other nine years are estimates based on population changes and housing records.

Because the inter-census year numbers are only estimates, there can be large gaps between the decennial headcount and the count from the previous year: in 2010, for instance, the national population headcount was roughly 3 million higher than the 2009 estimate. In contrast, the 2010 count for the country’s 50 largest cities was 1.3 million lower than the 2009 estimate. According to a report from the Pew Center on the States, Atlanta, Georgia was a particularly egregious example, with a 2009 estimated population of 540,000, dropping over 20 percent to 420,000 in the 2010 Census. Part of the discrepancy is related to the undercount issue in the decennial census discussed above, which is more likely to affect urban areas. Also contributing to the problem are errors in the between- census estimates. For some cities, including Atlanta, mid-decade adjustments to housing data led to overestimates of population in the central city.

As with an undercount in general, these cities are faced with a significant loss of federal funding and political representation. The fact that the 2010 estimate is so different from expected makes the potential loss even more difficult for municipalities to accept. For researchers, these year-to-year gaps can make it hard to correctly identify trends and make demographic comparisons. It is particularly important for social science researchers to recognize that the between-census population counts are only estimates, subject to revision, and there may be wide variation in the accuracy of those estimates across geographic locations.

Privacy in the Census: Double-Edged Sword?

One reason people may fail to return their U.S. Census form relates to privacy concerns. What is the government really going to do with that

34

information? For the most part, such concerns are unwarranted. In the Census population counts, individuals are never identified. In the individual-level sub-sample of the data that the government makes available to researchers (known as the public use micro sample or PUMS), several steps are taken to ensure the privacy of respondents, such as rounding off income values, averaging together outliers, and swapping the characteristics of some respondents (for example, a 62-year-old single woman in City A is moved and appears in the data as if she lives in City B while a similar woman in City B appears in the data as if she lives in City A).

These steps create additional noise in the data, but researchers have generally assumed that the errors do not systematically change patterns in the underlying data, so analysis with the manipulated data can still be assumed to apply to the general population. However, while investigating marriage rates using the 2000 Census data, University of Pennsylvania economist Betsey Stevenson discovered some strange anomalies, such as an implausibly large drop in the marriage rate for elderly women. Further investigation revealed several problems with the data that stemmed from a glitch in the program used by the Census Bureau to protect individual identities. In addition, similar errors were found in several other data sets, including the American Community Survey and the Current Population Survey. Particularly troublesome for researchers is that the coding problem led to a disproportionate number of errors in the data on older individuals so that the correlations between age and other variables cannot be trusted. Although the Census Bureau is aware of the problem, corrected data sets have not been released and even if they are, it is unclear how many studies that used the erroneous data would need to be redone. These findings suggest that any researcher using these data sets should be extra cautious and thoroughly investigate whether these errors have the potential to affect the analysis.

Some critics have questioned whether the Census Bureau’s attempts to protect privacy are even worth the time and effort, since these days, anyone with ill intentions can track down personal information in many easier ways than by digging through census files. But there are other reasons why privacy protections in the Census are important. In a 2001Social Research article titled “The Dark Side of Numbers: The Role of Population Data Systems in Human Rights Abuses,” demographer William Seltzer and historian Margo Anderson review how population data have been used to help governments target vulnerable populations,

35

such as Jews during the Nazi Holocaust in Germany and Japanese Americans during World War II in the United States. Obviously, one way to prevent such use is not to collect the data in the first place, but given that the Census is constitutionally mandated, alternative safeguards must be considered. Laws and regulations about the sharing of data across government agencies are important, though in times of war, such laws can be circumvented, as was the case with the Japanese American internment. The purposeful introduction of errors by the Census Bureau may be a more effective way of ensuring that the data are less useful for any nefarious purposes.

Box 2.2 Interstate Migration

According the U.S. Census Bureau, annual interstate migration dropped significantly between 2000 and 2010, to nearly half the level it had been in the 1990s, with a particularly large drop in 2006. This led some analysts to speculate that the decrease in mobility was associated with the housing crash and would slow the nation’s recovery from the recent recession. But Federal Reserve economists Greg Kaplan and Sam Schulhofer-Wohl argue that much of the drop in the rate between 2005 and 2006 can be explained by an unreported change in the procedure used by the Census Bureau to account for missing data. Prior to 2006, the bureau imputed missing data in a way that inflated the estimated migration rate; when that problem was fixed the resulting estimates dropped, but no attempt was made to go back and revise the prior estimates. Kaplan and Schulhofer-Wohl find that once this technical issue is corrected, a downward trend over the decade remains—but it is smooth, with no jump in 2006—and the entire change over the decade is only half the magnitude.

Sources: Greg Kaplan and Sam Schulhofer-Wohl, “Interstate Migration Has Fallen Less Than You Think: Consequences of Hot Deck Imputation in the Current Population Survey,” Federal Reserve Bank of Minneapolis Working Paper 681, November 7, 2010; Andrew Gelman, “Much of the Recent Reported Drop in Interstate Migration Is a Statistical Artifact,” Statistical Modeling, Causal Inference, and Social Science (blog), November 9, 2010, http://andrewgelman.com/2010/11/much_of_the_rec/.

36

Will There Be a Population Boom?

A recent report from the United Nations (UN) made headlines with an estimate that the world’s population may reach 10.1 billion by the year 2100. Previous estimates had predicted population would peak around 9 billion sometime in the middle of the century and then stabilize or decline, so the revised estimates set off concerns about the global impact of a continually growing population. But why did the estimates change by a billion people? A key part of the explanation is that population projections hinge on assumptions about fertility and mortality; when those underlying assumptions are off, even small differences can balloon when they are then extrapolated out over decades. In this case, the UN reported that in many poor countries, particularly in Africa, fertility is not declining as quickly as anticipated, as family planning programs have become less available. At the same time, mortality rates are somewhat lower than anticipated; in earlier decades, researchers predicted that the AIDS epidemic would be more devastating than it turned out to be.

It is not unusual for population projections to be far off from reality; past experience demonstrates the difficulty in making such projections. For example, after World War II, most demographers failed to predict the U.S. baby boom. In 1945, the U.S. Census Bureau underestimated by more than 25 percent the population growth for the following 25 years. Census Bureau predictions continue to be disputed today. Population experts Dennis A. Ahlburg and James W. Vaupel argue that the Census Bureau’s highest projections for the year 2080 are 300 million too low because they discount the possibility of another baby boom, greater longevity, and more immigration.

Demographer Nathan Keyfitz, who studied the error in 1,000 population growth predictions using modern methods, concludes: “We know virtually nothing about the population fifty years from now. We could not risk better than two to one odds on any range narrower than 285 million to 380 million for the year 2030.” Keyfitz advises researchers to be modest. Rather than attempt to predict population size far into the future, he would like to see social scientists study problems for the population already born. For example, schools need assistance in planning for enrollments, which can be easily anticipated on the basis of the current birthrate but— surprisingly—it is seldom used in education planning.

Undocumented Immigrants

37

Illegal immigration has become a huge issue in American politics, stirring strong emotions on both sides. Immigration hawks argue that illegal immigrants benefit from government programs subsidized by taxpayers, and supporters point out that America is a country of immigrants and policies like the Development, Relief, and Education for Alien Minors (DREAM) Act will strengthen our competitive position in the world. We discuss here two data controversies that are central to the larger debate: How many unauthorized immigrants are there, and do they pay more or less in taxes than the cost of the government services they use?

How Many Unauthorized Immigrants Are There?

How many immigrants enter the United States without legal status? No one knows for certain. In 2008 the Pew Hispanic Center estimated there were 11.9 million illegal immigrants in the country; the Department of Homeland Security put the number at 10.8 million in 2009. Others believe the numbers are much higher. Determining how many residents are undocumented aliens is difficult because many are reluctant to report their status to survey takers. Most researchers instead use a residual approach to calculate a reasonable number. First, estimates of legal immigrants are generated based on immigration records, deaths, and outmigration. The resulting number is then subtracted from the entire foreign-born population in the United States. Most studies assume that the official surveys miss about 10 percent of illegal immigrants, but this number is based on a relatively small study of Mexican-born residents of Los Angeles, California, conducted in 2001, which may or may not be applicable to current and larger samples.

An alternative method—one that results in much higher estimates—is based on arrests along the U.S.-Mexico border, which number well over a million each year. But these statistics are deceptive, because they include multiple arrests of the same individuals who cross repeatedly until gaining successful entry. Also, most illegal border crossers plan to work only temporarily in the United States; the large return flow to Mexico goes unnoticed. The majority of undocumented residents who remain in the United States originally crossed the border legally with visas and then failed to leave when the visas expired.

Given the difficulties with data collection on this issue, it is unlikely that researchers will ever be able to pin down more accurate numbers. But at stake is political representation in areas with large numbers of

38

undocumented immigrant populations: Even though noncitizens cannot vote, they are counted in the U.S. Census and therefore contribute to apportionment of congressional representatives. In addition, estimates of the illegal immigrant population are routinely cited in policy debates over border controls and immigration policies, but analysts should be aware of the many underlying assumptions that cannot be tested empirically.

Costly Immigrants?

Even more controversial than the basic count of illegal immigrants is whether they are a net cost or a net benefit to the country. Critics argue that unauthorized immigrants receive publicly provided services like education and health care but do not pay their share of the taxes that fund these programs. One study by the Federation for American Immigration Reform (FAIR), a group that advocates for immigration reduction, estimated illegal immigrants cost American taxpayers $113 billion. But PolitiFact.com, a fact-checking project of the Tampa Bay Times, points out that the FAIR analysis makes strong assumptions about the actual number of illegal immigrants in the United States (see previous section) and uses large underestimates of the taxes paid by immigrants, such as sales taxes or payroll taxes; although some undocumented workers are paid under the table, many are paid via payroll checks that withhold taxes.

A 2007 Congressional Budget Office report on the impact of unauthorized immigrants on state and local governments makes a number of important points about any attempt to quantify these effects. One is that whether undocumented immigrants represent a net cost or a net benefit depends first on which level of government one is talking about: At the federal level, they likely pay more in taxes than they receive in services, given that they cannot receive benefits from Social Security, Medicaid, and other welfare programs. But state and local governments are required to provide certain services to everyone, regardless of immigration status or ability to pay, so most of the public services that immigrants receive are funded by state and local governments. At the same time, although undocumented workers pay sales taxes and some income taxes, they are less likely to pay local property taxes directly (though they may pay them indirectly through rents), and their lower average income means that the amount of taxes paid to local governments is relatively low. Thus, the taxes paid by unauthorized immigrants generally do not offset the cost of services provided by state and local governments.

39

However, the CBO study also concludes that any attempt to measure the size of those costs is fraught with unanswered (and in some cases, unanswerable) questions: What is the size of the unauthorized immigrant population? What is the extent to which this population pays taxes and consumes services? Are all costs and revenues captured? What is the appropriate time period to analyze, particularly when some services (such as education) might be considered investments that may not show benefits for several years? Because so little information is available on these issues, there is no way to know for sure the magnitude of the costs or the benefits that might accrue from them.

Race and Ethnicity

The sixth question in the 2000 U.S. Census asked: “What is this person’s race? Mark one or more races to indicate what this person considers himself/herself to be.”

White

Black, African American, or Negro

American Indian or Alaska Native—Print name of enrolled or principal tribe

Asian Indian

Chinese

Filipino

Vietnamese

Other Asian—Print race

Native Hawaiian

Guamanian or Chamorro

Samoan

Other Pacific Islander—Print race

Some other race—Print race

The 2000 Census marked the first time in the history of American census taking that respondents were allowed to self-identify with more than one race, a practice that was continued with the 2010 Census. Many

40

respondents answered without difficulty; the question is common in student surveys, affirmative action programs, and other questionnaires where policymakers want to assess racial and ethnic populations. However, in a debate that is already two centuries old, social scientists remain uncertain about how to classify race and ethnicity. As the United States becomes an increasingly multiethnic nation, such classification will be even more difficult in the future. For the 2000 and 2010 Census, the President’s Office of Management and Budget (OMB) issued directives about who is multiracial, who is black, who is Asian, and who is Hispanic. In each case OMB attempted to resolve long-standing disputes that involved important policy implications.

Multiracial Backgrounds

The most contentious decision for the 2000 Census was how to count the increasing number of individuals from multiracial backgrounds. The nearly doubling of intermarriage between groups during the 1980s created a large group of children whose parents were dissatisfied with the single race checkoff. The Association of Multiethnic Americans and Project RACE (Reclassify All Children Equally) lobbied for a multiracial category, and a September 1997 article, “What Race Am I?” in Mademoiselle magazine urged readers to contact the federal government about the issue.

An interagency governmental committee cochaired by the Bureau of the Census and the Bureau of Labor Statistics met during the mid-1990s to con-sider the issue of multiracial backgrounds as well as other proposed changes for racial and ethnic categories. The committee decided against the use of a multiracial category, endorsing instead a system that would allow respondents to select more than one of the Census’s racial categories —a procedure accepted by OMB for the 2000 Census. This marked a significant change from previous Census questions that required respondents to select only one race. However, it is still unclear how those multiple responses should be tabulated. With the multiple check-off system, there are 63 total possible categories, an unwieldy number for readable demographic reports. As result, decisions must be made about how to collapse the categories: Should those who checked off a combination of white and black be counted as black? Or should there be a new “multiracial” category that includes everyone who checked off more than one category? As an indication of how difficult it can be to resolve

41

this issue, sociologists Carolyn Liebler and Andrew Halpern-Manners discuss multiple “bridging methods” (to convert multiple categories to one), none of which is ideal and all of which would result in different estimates of racial counts.

Who Is Black?

Enumeration by race dates back to the first U.S. Census in 1790, when black slaves were counted as three-fifths of a person in congressional apportionment. For free individuals, tabulation by race presented problems that continue to the present day. The rate of self-reported black identification was not expected to change significantly with the addition of the multiracial check-off in 2000. Although the NAACP estimated that as many as 70 percent of African Americans may have a mixed-race background, samples taken prior to 2000 suggested that only about 3 percent of blacks were likely to check more than one race. Generally, African Americans followed American social convention, which classifies someone as black if one parent is black; or, pushing the definition back one generation, if one of four grandparents is black. This practice is also used in official U.S. Vital Statistics, where instructions read: “When the husband is white and wife is not, the child is assigned the wife’s race. When the husband is not white, the child is assigned to the husband’s race.”

Until recently, official U.S. birth statistics followed social convention. Instructions stated that if only one parent is nonwhite, then the child is assigned the race of the nonwhite parent; if both parents are nonwhite, then the child is assigned the father’s race. Beginning in 1989, the procedure was simplified to assign the child’s race based solely on the mother’s race. Even though this new procedure differs from standard social practice in which a child is considered nonwhite if either parent is nonwhite, government statisticians believe it will reduce error and confusion. A study by the U.S. Centers for Disease Control and Prevention (CDC) found that the old procedure understated the infant mortality rate for nonwhites. Using the new race-of-mother guideline, infant mortality rates were 0.6 percent higher for blacks, 0.9 percent higher for Hispanics, and 3.7 percent higher for Filipinos. According to the CDC study, the “one nonwhite parent” rule caused some infants to be designated nonwhite even though society categorized them as white. Conversely, the determination of race at death, often made by a funeral director based on a visual examination of

42

the body, led to the erroneous designation of many individuals as white. Consequently, nonwhite births were overestimated and nonwhite deaths were underestimated, and official statistics continued to understate the deplorably high infant mortality rate for nonwhites (see Chapter 4).

Box 2.3 Black “Insanity”: An Argument for Slavery

When the 1840 U.S. Census counted the “insane and idiots,” they measured an extraordinarily high rate of 1 in 162 for northern free blacks (and 1 in 6.7 in Maine), compared to only 1 in 1,558 in the South. Pro-slavery advocates cited these Census results to argue that “free negroes of the northern states are the most vicious persons on this continent.” This curious statistical result was not fully explained until the 1980s, when historian Patricia Cline Cohen looked at the original Census forms. Apparently many census takers miscoded older senile whites, who were considered idiots in common parlance, as black. The mistake was easy to make because the two items were situated close together on an unwieldy 80-column form.

Sources: Patricia Cohen, A Calculating People, pp. 194–204; see also Anderson, American Census, pp. 29–31.

Who Is Asian?

Since 1870, the U.S. Census question about race has included choices of countries of origins for Asian Americans. Originally only Chinese or Indian, the choice was expanded by 1930 to include Japanese, Filipino, Hindu, Korean, and “other.” When later immigration included significant numbers from other countries as well, the Census Bureau intended respondents to use the “other” category. Instead, some Thais and Cambodians wrote in their background, usually in place of Vietnamese, causing the computer to misread their forms. In addition, some respondents used the “other” category to report themselves as Taiwanese (instead of Chinese, as the Census intended) or Greek (instead of white, as the Census intended).

43

Box 2.4 Second Largest “Ethnic” Group: “No Response”

Students taking the Scholastic Aptitude Test (SAT) are asked to self- identify their race or ethnicity. But 12 to 14 percent fail to respond— that’s more than the number of blacks, Asian Americans, or any other single minority group. Researcher Howard Wainer at the Educational Testing Service points out that the no-response group introduces an uncertainty into comparisons of scores by different ethnic groups that is greater than the measured change in minority-group test scores over a five-year period. In other words, media focus on test-score differences by ethnicity, sometimes as little as two or three points in a single year, may occur entirely because of our uncertainty about racial identification.

Sources: Howard Wainer, “How Accurately Can We Assess Changes in Minority Performance on the SAT?” American Psychologist 43, no. 10 (October 1988): 774–78.

To avoid these problems, and to save space in the form, the Census Bureau proposed that all Asian Americans write in their background for the 1990 Census. Census officials were eager to abandon the Asian “race question” because it took so much space in proportion to the number of respondents affected. Representatives of Asian communities protested that the procedure would cause a serious undercount because Asian Americans with poor English-language skills would be unable to write in their background. At stake were social programs such as English education that are sometimes allocated based on Census data. Ultimately, the check-off boxes for Chinese, Japanese, Filipino, Korean, and Vietnamese were retained and all other Asians were supposed to write in their specific background. The 2010 Census added specific examples (“Hmong, Laotian, Thai, Pakistani, Cambodian, and so on”) to the “print race” instructions for the “Other Asian” category.

Who Is Hispanic?

Debate about how to count U.S. residents from Spanish-speaking countries (except Spain) is now several decades old—but it too remains unresolved.

44

The government has changed how it counts Hispanics in nearly every census since 1930, beginning with the addition of answers to the race question: “other nonwhite” (1930), “persons of Spanish mother tongue” (1940), “white persons of Spanish surname” (1950 and 1960), “persons of both Spanish surname and Spanish mother tongue“ (1970), and finally, differentiation of Hispanics in a separate question (1980 and 1990).

Each of these designations created confusion for respondents. Re- interviews following the 1970 Census showed that more than 20 percent of those with Spanish backgrounds changed their answer to the same question, with an approximately equal number shifting themselves from “non-Spanish” to “Spanish” as from “Spanish” to “non-Spanish.” And, aside from errors in the original data, researchers do not have a consistently defined population to compare over time.

Dissatisfaction with such uncertain data, coupled with growing political power among Hispanics, resulted in a new question for the 1980 Census, modified only slightly in the 2000 Census to:

Is this person Spanish/Hispanic/Latino?

No, not Spanish/Hispanic/Latino

Yes, Mexican, Mexican American, Chicano

Yes, Puerto Rican

Yes, Cuban

Yes, other Spanish/Hispanic/Latino—Print group

The only significant change from the 1990 Census was the addition of “Latino,” which the Office of Management and Budget maintained might increase the response rate from the western portion of the United States, where the term Latino may be preferred to Hispanic. A May 1995 study for the Current Population Survey found that Latino was the preferred designation by only 11.7 percent of respondents, far behind Hispanic at 57.9 percent and “of Spanish origin” at 12.3 percent. In 2010 instructions were added to distinguish ethnicity from race (“For this census, Hispanic origins are not races”) and the question wording was modified slightly to “Is the person of Hispanic, Latino, or Spanish origin?”

To be consistent with the new multiple response option for the “race” question, the Census should allow respondents from mixed Hispanic

45

backgrounds to check both “of Hispanic origin” and “not of Hispanic origin.” In this way respondents would not be forced to choose between one of their parents’ ethnic heritages. In sheer numbers, the most common marriage between U.S. ethnicities occurs between Hispanics and other groups; because such an option has not yet been researched, OMB has so far rejected it.

Although separation of the race and Hispanic ethnicity questions eliminates one source of ambiguity in the Census, other problems remain. For example, Jamaicans and other West Indians are still often confused about how to classify themselves. As non-Hispanics and non-native blacks, they often list their background in the “other” category. These uncertainties, coupled with nonresponse to the Census by illegal immigrants, have led some researchers to conclude that the Hispanic count is far too low. The Census Bureau admits an undercount as high as 5 million, but some researchers claim there are as many as 16 million uncounted Hispanics— double the official number. At stake are programs to increase Hispanic political representation, particularly in western states where the Hispanic population is large enough to have a significant impact on how legislative districts are drawn.

Implications

The way in which the U.S. Census counts race and ethnicity shows how politics, science, and research needs intersect in data collection. The decennial adjustments to the race and ethnic categories clearly follow shifting political tides, from the racist enumeration of slaves as three-fifths of a person to the first opportunity to select multiple racial backgrounds in the 2000 count. As a result, the racial and ethnic questions are a hodgepodge of categories indefensible in scientific terms, conflating biological attributes such as skin color with nonbiological attributes such as national origin and native language.

Even if it has no scientific basis, self-enumeration by race has legitimate social validity. In other words, if people describe themselves as black or white, then for most research it makes sense to classify them in that category. The addition of multiple answers to the race question eliminates one set of problems for those individuals who define themselves as multiracial; however, it leaves an ambiguous situation for many individuals, particularly African Americans, whom the society at large identifies as black no matter what their multiracial background. Clearly,

46

researchers need to decide how to count those individuals who check off multiple categories. Should the categories follow social convention and previous Census practice by counting as black anyone who checked off black as one of his or her races? Alternatively, there could be entirely new multiracial categories that will complicate comparisons over time. Moreover, decisions need to be made about how many new categories to use, since there are 63 possible racial combinations.

Some official statistics will be affected by the new racial categories in confusing and unpredictable ways. For example, in official death rates, the numerator (race of deceased) is determined by an observing mortician or physician, while the denominator (race in the overall population) comes from the Census. Multiracial categories reported in the Census may not be recognized by the mortician or physician.

Future censuses will need to resolve the problem of counting multiracial answers. Moreover, there will be lobbying to extend the option to Hispanics and to allow for more accurate counting of groups such as those with Chinese ancestry who had been living in Malaysia. Finally, pressure is mounting from Arabs and Middle Eastern Americans for their own ethnicity category. The OMB has recommended further research to determine the best way to improve data, guaranteeing further debate.

Box 2.5 If It’s Tuesday, I’m Swedish

Researchers need to be aware that answers to ethnic ancestry are highly inconsistent; as many as one-half of respondents change their answer from one survey to another. One reason is that the increased number of mixed-ancestry marriages enables individuals to report a variety of ethnic backgrounds. Changing popularity trends for particular ethnic backgrounds may cause individuals to report different ethnicity at different times.

Sources: Reynolds Farley, “The New Census Question About Ancestry: What Did It Tell Us?” Demography 28, no. 3 (August1991): 421–23.

How Big Is the Gay Population?

47

Just how many lesbian, gay, or bisexual people are there in the United States? For several decades, the common answer was “one in ten,” based on Alfred Kinsey’s 1948 study of male prison inmates. Although Kinsey’s study has been criticized for being based on an unrepresentative sample, for many years there was little data to directly refute his calculation. More recent surveys have included questions about sexual orientation and produced estimates much lower than Kinsey’s number. For example, a 1993 Yankelovich Partners consumer survey found 5.7 percent of respondents described themselves as “gay, homosexual, or lesbian.” Demographer Gary Gates averaged the results of five surveys conducted since 2000 and concluded that about 3.5 percent of poll respondents self- identify as lesbian, gay, or bisexual; however, his results have been criticized for failing to consider the context of the individual surveys included in that average. For example, one of the larger studies was only of Californians, another (with the highest percentage identifying as gay) was administered online.

Within this debate are several data issues. Surveys conducted only in specific states may not be representative of the nation. Survey respondents may answer sensitive questions, such as those about sexual orientation, differently when being interviewed on the phone versus online. There are also issues with how questions are worded: The National Center for Health Statistics found that many people do not understand the words heterosexual or bisexual, leading to errors in the estimates.

Analysts should also be careful about equating sexual behavior with sexuality. Surveys generally find higher rates of same-sex sexual behavior than rates of homosexual self-identification. For example, data gathered by the National Opinion Research Center in 1992 found that 2.8 percent of men identified themselves as homosexuals, whereas 6 percent said they were attracted to other men and 9 percent of men reported having had at least one homosexual experience since puberty. Thus, whether one focuses on questions of behavior or identity can affect the resulting estimates.

Households and Families

What Is a Household? What Is a Family?

The U.S. Census Bureau uses the terms household and family differently from everyday usage. A household is defined by the housing unit, not by the social or biological relationship of the individuals who live there. Thus,

48

two families sharing a single home are counted as a single household, as are a family and a lodger or a group of individuals sharing a home. Households are subdivided into family households, groups of two or more persons related by birth, marriage, or adoption, and nonfamily households, including individuals living alone and unrelated individuals residing together (see Table 2.1).

As recently as 1950, nearly 80 percent of all households were married- couple families. By 2010 only half of all households fell into that category. The trend has been driven not only by increases in the number of individuals living alone but also by more unmarried couples living together. One group that may contribute to future changes in the share of family households is gay and lesbian couples: as of 2010 gay partners were not generally considered families by the Census definition since they could not legally marry in most states (and it was unclear whether the federal government would recognize such marriages even in states where it was legal); however, as gay marriage becomes more accepted and is legalized in more states, the number of gay family households may increase.

Unfortunately, many research projects are limited only to family households, thereby leaving out a significant proportion of the population. For example, income analyses that use family income may understate the well-being of U.S. households by leaving out the growing number of young, affluent individuals who live on their own (see Chapter 8) or those who are not legally married. Even less likely to be included in research projects are individuals who not do not live in households as defined by the Census Bureau. For example, the Current Population Survey, the source for many data on housing, education, income, and employment (see Chapters 3, 5, 8, and 9), leaves out most of those who live in group-quarter populations such as military barracks, prisons, hospitals, and nursing homes (college dormitory residents are counted with the parents’ families).

Table 2.1

Types of Households

Percentage of total persons in household type

Family households 67

49

Married-couple family 50 Male householder, no spouse present 5 Female householder, no spouse present 13 Nonfamily households 33 Living alone 27 Not living alone 6

Sources: Statistical Abstract, table 61: Households and Persons per Household by Type of Household: 1990 to 2010, http://www.census.gov/compendia/statab/2012/tables/12s0063.pd- f, based on data from U.S. Census Bureau, America’s Families and Living Arrangements, Current Population Reports, P20–537, 2001, and earlier reports. See also http://www.censu- s.gov/population/www/socdemo/hh-fam.html.

Box 2.6 Head of Household

Until 1980, the “head of household” was automatically assigned to a man—even if a woman respondent coded herself as the household head. The procedure was abandoned in part because it perpetuated sexist stereotyping but also because the increasing number of nontraditional family arrangements made it difficult to identify the head of the household. Now, one individual of either sex can be designated the “householder,” a reference person in whose name the housing unit is owned or rented and to whom the relationship of all other household members is recorded; over 7,043,827 married women out of over 54 million married couples identified themselves as householders in 2000.

Sources: Sweet and Bumpass, American Families, pp. 336–37, and in Tavia Simmons and Martin O’Connell, “Married-Couple and Unmarried- Partner Households: 2000,” CENSR-5 (Washington, DC: U.S. Bureau of the Census, 2003), http://www.census.gov/prod/2003pubs/censr-5.pdf.

What Is the Role of Cohabitation?

As mentioned in the previous section, at least part of the decline in traditional married-couple households is due to an increase in couples who live together without getting married. Economists Betsey Stevenson and

50

Justin Wolfers point out that forming a household with another person is increasingly separate from entering into a marriage with that person. However, in the past, most analysts have treated marriage and household formation as virtually the same decision, so the increase in cohabitation raises a host of questions about the similarities and differences between married households and cohabitating households.

Unfortunately, answering any of those questions is hampered by lack of good data, particularly for analyses over longer periods of time. The first direct questions about cohabitation appeared in the decennial 1990 Census and in the Current Population Survey in 1995. For earlier periods, the commonly used proxy was the “number of persons of the opposite sex sharing living quarters.” This indirect measure is often criticized for producing quite imprecise counts compared to direct estimates. A number of family surveys, such as the National Survey of Family Growth and the National Survey of Families and Households, began asking about cohabitation in the 1980s, but changes in the social acceptance of cohabitation may muddy comparisons in cohabitation rates across decades. Even more problematic is the finding of demographers Sarah R. Hayford and S. Philip Morgan that, in these surveys, retrospective data on cohabitation likely underestimate cohabitation in more distant past periods (i.e., prior to the time of the survey). Research has shown that people have a harder time with accurate recall of more distant events in general, and survey respondents may also differ in their definition of when cohabitating relationships officially began or ended.

Even if we believe the survey data, the interpretation of cohabitation data and the ability to predict the trends into the future is difficult because of the rapidly changing nature of the institution. Surveys suggest that cohabitation traditionally has been part of the path to marriage: Most cohabiters say they expect their relationship to segue into marriage, and the majority of marriages are preceded by cohabitation. However, in 2002 more than 20 percent of cohabitating couples had been living together for five years or more, which could indicate cohabitation is a more permanent state for at least some of these couples.

How Many Divorces Are There?

Data on divorce are less accurate than most other family statistics for several reasons. The usually comprehensive Vital Statistics are incomplete for divorce, leaving out several states, including California, home to a

51

large share of all divorces. To make matters worse, statistics from reporting states lack data on race in more than 25 percent of all divorce reports, and on age in over 10 percent of divorce reports. Alternative data from the U.S. Census Bureau are unsatisfactory because some respondents do not accurately report their marital history: Men report about 10 percent fewer divorces than women do (although the number should be nearly precisely equal). Even the more accurate women’s answers measure a divorce rate substantially lower than must have occurred based on legal records counted in U.S. Vital Statistics.

If we put aside these problems with the statistics, an important question remains: What does the divorce rate really measure? In some reports, the divorce rate is calculated as the number of divorces per 1,000 people; by this measure, the rate has been steadily declining since 1981. But while this per capita measure controls for changes in population, it does not take into account how many marriages there are. The number of marriages per 1,000 people has also been declining since the early 1980s, so is the falling number of divorces per capita attributable to fewer marriages failing or simply to fewer marriages?

What researchers are interested in capturing is how many marriages ever end in divorce, or how many people are divorced. Longitudinal surveys such as the Survey of Income and Program Participation (see Chapter 8) provide the data to estimate the percent ever-divorced among individuals ever-married, and the percent of first marriages that reach various anniversaries. For marriages occurring in the 1970s, the 25-year survival rate was 48 percent, the source of the frequent media observation that one- half of marriages end in divorce. However, 1970s marriages appear to be exceptional. Since then declining number of divorces indicate more stable marriages, not just fewer of them. Only about 25 percent of 1990s marriages had ended in divorce within the first 10 years, lower than the 10- year divorce rate for 1970s marriages.

Few social science studies cause such furor as estimates for the trends in marriage and divorce. Questions concerning family appear to raise doubts about both our individual futures and our society’s well-being. On the individual level, it is important to apply overall statistics with caution. The chance of marriage and divorce depends as much on individual circumstances as on social averages. For example, the divorce rate for couples who marry late in life is far lower than for those who marry young. It is also important to be aware of which measures are being used and recognize that different measures are answering different questions.

52

Divorces per 1,000 people may give an indication of the total number of divorces but may not be an accurate measure of how many marriages end in divorce or of how many people are divorced. In addition, there are likely to be problems with the underlying data itself that analysts should keep in mind.

Summary

The demographic controversies described in this chapter serve as an introduction to problems researchers find in all social science statistics.

First, we never have complete data. Even when our source attempts universal sampling, as in the U.S. Census and Vital Statistics, some of the population will be missed. There is no secret to this lack of completeness. Census Bureau researchers are at the forefront in identifying the extent of the undercount and developing methods to correct it. Similarly, shortcomings in the Vital Statistics, primarily nonreporting states, are a matter of public record. Responsibility lies with researchers to recognize explicitly the possible implications of less-than-complete data.

Second, all studies are limited by the categories in which the data are classified. This chapter summarizes U.S. experience in racial and ethnic categories. Generally accepted classifications of race and ethnicity from the past are considered hopelessly naive today. Likewise, researchers need to remember that categories based on the social biases of the early twenty- first century will change with new developments in race and ethnic relations. Official use of the terms household and family also demonstrate the importance of paying careful attention to classification. In this case, not only are official definitions different from everyday usage, but the changing characteristics of U.S. living situations mean that research on “traditional family households” will leave out increasing numbers of individuals.

Third, studies of population growth and family structure demonstrate the hazard of predicting the future. Even when our knowledge about past trends is relatively accurate, estimates for the future must involve debatable assumptions. We do not know if past trends will continue or if new trends will arise, as has been the case in world population projections. Although such uncertainties plague all predictive social science research, popular media coverage rarely warns readers about the problem.

Finally, researchers should be clear about how they are defining and measuring variables. A comparison of costs and benefits, such as with

53

1.

2.

3.

4.

undocumented immigrants, can change dramatically depending on what one includes as a cost or a benefit. Interpretation of trends in divorce can depend on whether one is comparing divorces per capita or divorces per marriage. Such issues mean that social science analysts need to think carefully to ensure the variables being used are appropriate to the research question.

The overall lesson for researchers is to use caution as they proceed in their studies. We have much to learn from demography. It would be a mistake to throw up our hands and reject all research using demographic statistics just because no data are perfect. Instead, the controversies imply a constant struggle to understand the world. By learning about limitations in the data and biases on the part of data users, we can better evaluate the policy implications of current research.

Case Study Questions

In a survey using Census Bureau race and ethnic categories, 20 percent of California college students left the question blank. How might these omissions affect research projects on enrollment, financial aid, or graduation rates?

Among Americans who first got married in the early 2000s, over half had lived with their future spouse before marriage, a significant increase over earlier time periods. How does cohabitation affect the marriage rate? The divorce rate?

There are conflicting measures of the U.S. divorce rate. In the U.S. Census, men report 10 percent fewer divorces than women. Vital Statistics measure more divorces than reported by women or men to the Census. Explain these discrepancies.

The Development, Relief, and Education for Alien Minors (DREAM) Act would have provided a path to citizenship for undocumented immigrants who met certain requirements. In particular, it would have required individuals to go to college or join the military to be eligible. Supporters of the DREAM Act argued that it would lead to higher incomes, and thus more tax revenue, as well as savings on public services; opponents argued that the cost of subsidizing education for immigrants was much higher than potential benefits. What data would be needed to determine which argument is correct?

54

5.

6.

The 1924 National Origins Act severely restricted immigration based on the proportion of national backgrounds already in the current population. Estimates of ethnicity were derived from surveys of family names in which “Mueller” might indicate German background, “O’Leary” might indicate Irish background, and “Miller” or “Leary” might represent English background. German and Irish lobbyists challenged the numbers and gained an increase in their quotas. Why?

Beginning with the 2000 Census, each respondent was able to check off more than one racial classification. Researchers have several options for using the resulting data. One possibility is to expand the number of categories to include all of the 63 possible racial combinations so that each individual will fit into exactly one group. As an alternative, the number of individuals checking off a racial category could be counted so that an individual with a multiracial background would be counted two or more times. Describe two research projects, one that would appropriately use each of these data sets.

References

Data Sources (8)

U.S. Census (8)

History of U.S. Census in Patricia Cline Cohen, A Calculating People (Chicago, IL: University of Chicago Press, 1982); Margo J. Anderson, The American Census: A Social History (New Haven, CT: Yale University Press, 1988); U.S. Census methods in Bryant Robey, “Two Hundred Years and Counting: The 1990 Census,” Population Bulletin 44, no. 1 (April 1989); James Holmes quoted in U.S. Bureau of the Census, “Census 2000 Questions Fewest in 180 Years” (press release, March 30, 1998); Hawaii census data in Karen R. Humes et al., Overview of Race and Hispanic Origin: 2010, U.S. Bureau of the Census, 2010 Census Briefs, C2010BR- 02, 2011.

American Community Survey (8)

History of long form and ACS in Mark Mather et al., “The American Community Survey,” Population Bulletin 60, no. 3 (2005); Nancy K. Torrieri, “America Is Changing, and So Is the Census: The American Community Survey,” Statistics 61, no. 1 (February 8, 2010): 16–21; ACS data from Daphne Lofquist, Same-Sex Couple Households: 2010, U.S. Bureau of the Census, ACSBR/10–03, 2011.

55

Vital Statistics (9)

Marriage and divorce statistics described in Betzaida Tejada-Vera and Paul D. Sutton, “Births, Marriages, Divorces, and Deaths: Provisional Data for 2009,” National Vital Statistics Reports 58, no. 25 (Hyattsville, MD: National Center for Health Statistics, 2010); Larry L. Bumpass and James A. Sweet, American Families and Households (New York: Russell Sage Foundation, 1987), chs. 2 and 5; data sample in J.A. Martin et al., “Births: Final Data for 2009,” National Vital Statistics Reports 60, no. 1 (Hyattsville, MD: National Center for Health Statistics. 2011).

Controversies (9)

Population Counts (9)

How Many People Are There? (9): In-person enumeration in Robey, “Two Hundred Years,” pp. 10–14. Undercount controversy in “Census Subject to Possible Correction,” Population Today 17, no. 9 (September 1989): 3ff; Dudley Kirk, “Politics of Demography,” Society 18 (January-February 1981): 22–25; “Census Mired in Dispute over Counting the Hidden,” Los Angeles Times, March 15, 1989, p. 1; Barry Edmonston, The Undercount in the 2000 Census, report prepared for the Annie E. Casey Foundation and The Population Reference Bureau, 2002, http://www.aecf.org/upload/publicationfiles/undercount_in2000cen- sus.pdf; 2010 changes and improvements in Robert Longley, “U.S. Census: The Cost of Not Being Counted—Census 2000 a Case in Point,” About.com: U.S. Government Info, updated February 1, 2010, http://usgovinfo.about.com/od/censu- sandstatistics/a/undercountcost.htm. New York challenge in “NYC Alleges Major U.S. Census Undercount,” UPI.com, March 28, 2011, http://www.upi.com/Top_N- ews/US/2011/03/28/NYC-alleges-major-US-Census-undercount/UPI-5280130132- 6737/. Effect of undercount in Reynolds Farley and Walter R. Allen, The Color Line and the Quality of Life in America (New York: Oxford University Press, 1987), pp. 420–38.

Who Gets the Prisoners? (11): David Sommerstein, “Urban, Rural Areas Battle for Census Prison Populace,” All Things Considered, NPR.org, February 15, 2010, http://www.npr.org/templates/story/story.ph- p?storyId=123663462; changes in Hispanic populations in Rose Heyer and Peter Wagner, “Too Big to Ignore: How Counting People in Prisons Distorted Census 2000,” Prison Policy Initiative, April 2004, http://ww- w.prisonersofthecensus.org/toobig/.

How Did 1.3 Million People Disappear in a Year? (12): Josh Goodman, “Low Census Counts Will Cost Big Cities Precious Funds,” Pew Center on the States Online, May 3, 2011, http://www.pewstates.org/projects/stateline/headlines/low-c-

56

ensus-counts-will-cost-big-cities-precious-funds-85899375075; “Two Population Counts by Census Bureau, Two Different Numbers: The Gap Between City 2009 Estimates and 2010 Census,” Teaching with Data (blog), May 3, 2011, http://teac- hingwithdata.blogspot.com/2011/05/every-year-census-bureau-conducts.html.

Privacy in the Census (13)

Privacy measures in Carl Bialik, “Census Bureau Obscured Personal Data—Too Well, Some Say,” Wall Street Journal, February 6, 2010, http://online.wsj.com/art- icle/SB20001424052748704533204575047241321811712.html; “Can You Trust Census Data?” Freakonomics (blog), February 2, 2010, http://www.freakonomic- s.com/2010/02/02/can-you-trust-census-data/. J. Trent Alexander, Michael Davern, and Betsey Stevenson, “Inaccurate Age and Sex Data in the Census PUMS Files: Evidence and Implications,” National Bureau of Economic Research Working Paper 15703, January 2010; William Seltzer and Margo Anderson, “The Dark Side of Numbers: The Role of Population Data Systems in Human Rights Abuses,” Social Research 68, no. 2 (Summer 2001).

Will There Be a Population Boom? (14)

UN projections at United Nations, Department of Economic and Social Affairs, Population Division, Population Estimates and Projections Section, updated October 20, 2011, http://esa.un.org/unpd/wpp/index.htm; Justin Gillis and Celia W. Dugger, “UN Forecasts 10.1 Billion People by Century’s End,” New York Times, May 3, 2011, http://www.nytimes.com/2011/05/04/world/04population.html; Jocelyn Kaiser, “10 Billion Plus: Why World Population Projections Were Too Low,” ScienceInsider, May 4, 2011, http://news.sciencemag.org/scienceinsider/20- 11/05/10-billion-plus-why-world-population.html; see also Brian C. O’Neill et al., “A Guide to Global Population Projections,” Demographic Research 4 (June 13, 2001): 203–288, http://www.demographic-research.org/volumes/vol4/8/. Criticism of census projections in Dennis A. Ahlburg and James W. Vaupel, “Alternative Projections of the U.S. Population,” Demography 27, no. 4 (November 1990): 641–50. Problems of prediction in “Census Bureau Demographer’s Unqualified Prediction,” New York Times, February 5, 1989, p. 30; 1945 estimate in Tony Kaye, “The Birth Dearth,” New Republic, January 19, 1987, p. 21. Nathan Keyfitz, “The Social and Political Context of Population Forecasting,” in The Politics of Numbers, ed. William Alonso and Paul Starr (New York: Russell Sage Foundation, 1987), pp. 256–58.

Undocumented Immigrants (15)

How Many Unauthorized Immigrants Are There? (16): Estimates of immigrant counts in Michael Hoefer, Nancy Rytina, and Bryan C. Baker, “Estimates of the Unauthorized Immigrant Population Residing in the

57

United States: January 2009,” Population Estimates, Office of Immigration Statistics, U.S. Department of Homeland Security (DHS), 2010, http://www.dhs.gov/xlibrary/assets/statistics/publications/ois_ill_pe- _2009.pdf; Carl Bialik, “In Counting Illegal Immigrants, Certain Assumptions Apply,” Wall Street Journal, May 7, 2010, http://online.ws- j.com/article/SB10001424052748704370704575228432695989918.html; and Bialik, “The Pitfalls of Counting Illegal Immigrants,” The Numbers Guy (blog), May 7, 2010, http://blogs.wsj.com/numbersguy/the-pitfalls-o- f-counting-illegal-immigrants-937/. Border crossing and yearly immigration in “U.S.-Mexico Study Sees Exaggeration of Migration Data,” New York Times, August 31, 1997, p. A1. See also Katharine M. Donato, Jorge Durand, and Douglas S. Massey, “Stemming the Tide? Assessing the Deterrent Effects of the Immigration Reform and Control Act,” Demography 29, no. 2 (May 1992): 139–57.

Costly Immigrants? (16): FAIR study at “The Fiscal Burden of Illegal Immigration on U.S. Taxpayers,” Federation for American Immigration Reform Online, updated February 2011, http://www.fairus.org/publications/the-fiscal-burd- en-of-illegal-immi-gration-on-u-s-taxpayers. PolitiFact.com, “Vern Buchanan Says Illegal Immigration Costs Florida Taxpayers $4 Billion a Year,” December 6, 2010, http://www.politifact.com/florida/statements/2010/dec/06/vern-buchanan/ve- rn-buchanan-says-illegal-immi-gration-costs-flori/, and “Workman Misquotes Immigration Report,” June 30, 2010, http://www.politifact.com/florida/statements/- 2010/jun/30/ritch-workman/workman-misquotes-immigration-report/. Edward Schumacher-Matos, “Illegal Immigration: What’s the Real Cost to Taxpayers?” Washington Post, September 9, 2010, http://www.washingtonpost.com/wp-dyn/co- ntent/article/2010/09/09/AR2010090902874.html; “The Impact of Unauthorized Immigrants on the Budgets of State and Local Governments” (Washington, DC: Congressional Budget Office, December 6, 2007), http://www.cbo.gov/ftpdocs/87- xx/doc8711/12-6-Immigration.pdf.

Race and Ethnicity (17)

Multiracial Backgrounds (18): On OMB decision see Office of Management and Budget, Executive Office of the President, “Revisions to the Standards for Classification of Federal Data on Race and Ethnicity,” Federal Register 62 (October 30, 1997): 58781–58790. BLS report in Bureau of Labor Statistics, U.S. Department of Labor, “A Test of Methods for Collecting Racial and Ethnic Information” (news release, October 26, 1995); response to multiracial category in “More Than Identity Rides on a New Racial Category,” New York Times, July 6, 1996, p. A1. Arti V. Finn, “What Race Am I?” Mademoiselle (September 1997). See also Humes et al., Overview of Race and Hispanic Origin: 2010, C2010BR-02. Carolyn A. Liebler and Andrew Halpern-Manners, “A Practical Approach to Using

58

Multiple-Race Response Data: A Bridging Method for Public-Use Microdata,” Demography 45, no. 1 (2008): 143–55.

Who Is Black? (20): History of census questions in Ira S. Lowry, The Science and Politics of Ethnic Enumeration (Santa Monica, CA: The Rand Corporation, 1980), pp. 8–9; and William Petersen, “Politics and the Measurement of Ethnicity,” in The Politics of Numbers, ed. William Alonso and Paul Starr (New York: Russell Sage Foundation, 1987), pp. 208–9. NAACP estimate in “New Rules for Marking Racial Identity, San Francisco Chronicle, December 26, 1997, p. A1; change in race designation and infant mortality in Robert A. Hahn, Joseph Mulinare, and Steven M. Teutsch, ”Inconsistencies in Coding of Race and Ethnicity Between Birth and Death in U.S. Infants,“ Journal of the American Medical Association 267, no. 2 (January 8,1992): 259–63.

Who Is Asian? (20): History of Asian enumeration in Lowry, The Science and Politics of Ethnic Enumeration, pp. 7–10. Problem answers in “Simpler 1990 Census Form Upsets Asian Americans,” Los Angeles Times, April 12, 1988, I, p. 3; “Concerns Raised on the ‘90 Census,” New York Times, April 17, 1988, p. 31; “Census Won’t List Various Asian Groups,” Wall Street Journal, May 23, 1988, p. 19. Limited space in census in “Scrambling to Be Counted in Census,” New York Times, December 3, 1989, p. A17. New form in Humes et al., Overview of Race and Hispanic Origin: 2010, C2010BR-02.

Who Is Hispanic? (21): Change in census method in Joan Moore and Harry Pachon, Hispanics in the United States (Englewood Cliffs, NJ: Prentice Hall, 1985), p. 3; see also Nancy A. Denton and Douglas S. Massey, “Racial Identity Among Caribbean Hispanics,” American Sociological Review 54 (October 1989): 790–94. Current Population Survey in Bureau of Labor Statistics, U.S. Department of Labor, “A Test of Methods for Collecting Racial and Ethnic Information” (news release, October 26, 1995). Respondent confusion in Lowry, The Science and Politics of Ethnic Enumeration, p. 13. Los Angeles City Council in “L.A. Cases Seek Hispanic Gain,” New York Times, July 10, 1989, p. A15.

Implications (23): History of concept of race in J.C. King, The Biology of Race (Berkeley: University of California Press, 1981); see also Stephen J. Gould, The Mismeasure of Man (New York: Norton, 1981); and Ashley Montagu, Man’s Most Dangerous Myth: The Fallacy of Race (New York: Oxford University Press, 1974). On future changes, see Office of Management and Budget, “Revisions to the Standards for Classification of Federal Data on Race and Ethnicity,” Federal Register 62 (October 30, 1997): 58781–790.

How Big Is the Gay Population? (24)

Yankelovich survey in “A Sharper View of Gay Consumers,” New York Times,

59

June 9,1994, p. C1; “Sex Surveys: Does Anyone Tell the Truth?” American Demographics, July 1993, p. 9; “Sex Survey of American Men Finds 1% Are Gay,” New York Times, April 15, 1993, p. A1; Priscilla Painton, “The Shrinking Ten Percent,” Time, April 26, 1993, pp. 27–29; “Polling on Sexual Issues Has Its Drawbacks,” New York Times, April 25, 1993, p. A23; NORC study in “Sex in America: Faithfulness in Marriage Thrives After All,” New York Times, October 7, 1994, p. A1. Gates study in Carl Bialik, “Reliable Tally of Gay Population Proves Elusive,” Wall Street Journal, April 16, 2011, http://online.wsj.com/article/SB100- 01424052748704116404576263383778476752.html; NCHS study in Bialik, “Sexual Stats in the Post-Kinsey Age,” The Numbers Guy (blog), April 15, 2011, http://blogs.wsj.com/numbersguy/sexual-stats-in-the-post-kinsey-age-1051/.

Households and Families (25)

What Is a Household? What Is a Family? (25): Changing household characteristics in Nancy Folbre, A Field Guide to the U.S. Economy (New York: Pantheon, 1987), p. 3.9; and U.S. Department of Agriculture, “Living Arrangements and Marital Status of Households and Families,” Family Economics Review 2, no. 3 (July 1989): 16. Christopher Jencks, “The Politics of Income Measurement,” in The Politics of Numbers, ed. Alonso and Starr, pp. 92–105.

What Is the Role of Cohabitation? (27): Betsey Stevenson and Justin Wolfers, “Marriage and Divorce: Changes and Their Driving Forces,” Journal of Economic Perspectives 21, no. 2 (2007): 27–52. Sarah R. Hayford and S. Philip Morgan, “The Quality of Retrospective Data on Cohabitation,” Demography 45, no. 1 (2008): 12941; Richard Fry and D’Vera Cohn, “Living Together: The Economics of Cohabitation,” Pew Research Center: Social and Demographic Trends, June 27, 2011, http://www.pewsocialtrends.org/2011/06/27/living-together-the-economics-- of-cohabitation/.

How Many Divorces Are There? (27): Problems with divorce data in Sweet and Bumpass, American Families, ch. 5; Bogue, Population of the United States, pp. 194–97. Men’s versus women’s responses in Sweet and Bumpass, American Families, p. 210. How divorce rate is calculated in Dan Hurley, “Divorce Rate: It’s Not as High as You Think,” New York Times, April 19, 2005, http://www.nytime- s.com/2005/04/19/health/19divo.html; Stevenson and Wolfers, “Marriage and Divorce,” 27–52; Pamela Paul, “Millennial Myths,” American Demographics, December 2001, p. 20.

Case Study Questions (30)

1. Harold Orlans, “The Politics of Minority Statistics,” Society 26 (May-June 1989): 25.

2. Betsey Stevenson and Justin Wolfers, “Marriage and Divorce: Changes and Their Driving Forces,” Journal of Economic Perspectives 21, no. 2 (2007): 27–52.

60

3. Larry L. Bumpass and James A. Sweet, American Families and Households (New York: Russell Sage Foundation, 1987), p. 210.

4. DREAM Act Portal, http://dreamact.info/_and National Immigration Law Center, http://www.nilc.org/DREAMact.html.

5. Margo J. Anderson, The American Census: A Social History (New Haven, CT: Yale University Press, 1988), pp. 144–49.

6. Office of Management and Budget, “Revisions to the Standards for Classification of Federal Data on Race and Ethnicity,” Federal Register 62 (October 30, 1997): 58781–790.

61

3

Housing □□□□

Housing is the largest single component of U.S. household budgets, comprising roughly one-third of expenditures as measured by the Bureau of Labor Statistics (see Chapter 9). In national income accounts, residential construction is the largest component of investment, averaging about $250 billion per year. Thus, housing statistics warrant separate and detailed attention.

This chapter first looks at two major debates related to the housing bubble and collapse: Just how big was the bubble? And could it have been foreseen? These questions will likely be debated for some time, but here we highlight the role of data choices made by analysts in studying them. We also review several debates about homeownership, including (1) whether racial gaps in homeownership really are as large as they appear, (2) who benefits from tax policy designed to encourage homeownership, and (3) how much discrimination there is in mortgage markets. The range of estimates that exist for the number of homeless, an interesting case study of matching the right statistic to the right policy question, is discussed as well. The final section of this chapter looks at geographic divisions used in social science research, most of which are based on place of residence. Researchers need to know how changes in geographic units affect the definition of urban, metropolitan, and rural areas. Two issues based on geographic divisions are summarized: the trend in racial segregation, and the desirability of different urban areas.

Where the Numbers Come From

Organizations Data sources URL Bureau of the Census, U.S. U.S. Census, American www.census.gov

62

Department of Commerce Community Survey, Housing Vacancy Survey, Survey of Market Absorption

U.S. Department of Housing and Urban Development; Bureau of the Census, U.S. Department of Commerce

American Housing Survey www.huduser.org

Bureau of Labor Statistics, U.S. Department of Labor

Consumer price indices for rent (Shelter Index)

www.bls.gov

Federal Housing Finance Agency

House Price Index www.fhfa.gov

National Association of Realtors

Price and sales statistics, Affordability Index

www.realtor.org

Standard & Poor’s Case-Shiller Home Price Index

www.standardandpoor- s.com

Data Sources

U.S. Census

Although the U.S. Census is best known as a population count, it is officially a “Census of Population and Housing.” Impetus for a national housing survey came during the Great Depression of the 1930s in order to determine the degree of inadequate housing and assess how new housing construction might stimulate the economy. When these surveys proved successful, the U.S. Census added housing to its 1940 population count. By 1990, the housing section of the Census had grown to six out of 14 questions on the short form (administered to all households) and 19 out of 59 questions on the long form (replaced in 2010 with the American Community Survey; see Chapter 2). Census housing data provide the most comprehensive statistics. The major drawback to the census is timeliness; many statistics of interest are not made available until years after the Census is taken.

Data Sample: In both the 2000 Census and the 2010 Census, West Virginia and Minnesota had the highest rates of homeownership, at 73.4 percent and 73.0 percent in 2010, respectively. New York, with 53.3 percent owner-occupied housing, had the lowest rate in both

63

decades.

American Housing Survey

The American Housing Survey (AHS), called the Annual Housing Survey until 1984, provides data on a speedier and more frequent basis than the Census. Conducted by the Census Bureau for the Department of Housing and Urban Development every year since 1973, the AHS samples households across the country on a staggered basis, providing data on the size and quality of housing, neighborhood characteristics, home financing, and recently moved households.

Data Sample: In 2009 the Housing Survey estimated that in 77,400 housing units in Detroit, Michigan, water from the primary water source was not safe to drink and 129,400 units did not have a working smoke detector.

Other Industry Data

In addition to the AHS, the Census Bureau conducts several economic surveys that provide data on the housing industry. Most closely watched are housing starts, a key measure of the economy’s overall health and one component of the Index of Leading Economic Indicators. The Census Bureau collects residential and nonresidential data from local permit- issuing offices. More detailed information is available commercially from McGraw-Hill Construction’s Dodge data, based on correspondent reports gathered directly by the construction industry.

Box 3.1 American Community Survey Versus American Housing Survey

One might expect that the American Community Survey (the ACS, which replaced the U.S. Census long form) and the American Housing Survey (AHS) would yield similar results. But in 2005 about 140,000 more units were counted in the ACS, and homeownership was approximately two percentage points lower than the AHS measure. The ACS also reports fewer houses built prior to 1939 and a greater proportion of homeowners with a mortgage. Several factors are thought to account for these differences:

64

Sampling differences—the two surveys are drawn from slightly different populations: the ACS draws from a continually updated address list that is also used for the decennial census, while the AHS sample is based on the 1980 Census, updated to account for new construction, with weights applied to replicate the complete housing stock.

Data collection methods—the AHS uses telephone and in- person interviews; the ACS starts with mail questionnaires and then follows up with telephone calls or in-person interviews.

Questionnaire design—the questionnaires for the two surveys are dissimilar, with different wording and order of questions. The AHS survey is also more detailed and requires more time to complete.

Sources: Differences described in Adams, Housing America in the 1980s, pp. 34–37; Frederick Eggers, “Comparison of Housing Information from the American Housing Survey and the American Community Survey” (U.S. Department of Housing and Urban Development, Office of Policy Development and Research, September 2007), http://www.huduser.org/porta- l/publications/polleg/compari-son_hsg.html.

Data Sample: According to the Survey of Market Absorption of Apartments (SoMA), there were 90,500 unfurnished rental apartments completed in 2010, 90 percent of which had air conditioning.

Price Data

The U.S. Census Bureau economic surveys include extensive local data on vacancies, mortgages, and rents, often used by the housing industry for planning purposes. The U.S. Department of Labor’s Bureau of Labor Statistics monitors housing costs in the Shelter Index, a part of the Consumer Price Index (see Chapter 11). The Shelter Index provides a single statistic for the cost of rents, new home prices, mortgage rates, and home upkeep for the United States and selected geographic areas. The Federal Housing Finance Agency tracks housing prices using data on mortgages. other widely reported statistics on new and existing home prices are assembled by the National Association of Realtors based on a

65

combination of census data and their own surveys. Finally, one of the most widely used residential housing price indices is Standard and Poor’s Case- Shiller Home Price Index. Developed by economists Karl Case and Robert Shiller in the 1980s, the index relies on data for repeat sales of homes, thus controlling for housing quality and characteristics.

Data Sample: The National Association of Realtors Affordability Index estimates that in 2010 the median family had 174 percent of the income necessary to qualify for a conventional mortgage on a median-priced single-family home. This was a 36 percentage point increase over 2008.

Controversies

Housing Crisis

Surely one of the biggest news stories of the 2000s was the housing bubble, the huge run-up in housing prices followed by a spectacular crash that triggered chaos in the financial system and a deep recession. There are several data controversies related to the housing crisis, including the size of the bubble, how long it lasted, and whether it could have been foreseen. As we will see, analysis of these issues varies, depending on how one looks at the data.

How Big Was the Bubble and Is It Over Yet?

One way to measure the magnitude of the bubble, and to assess whether the market is back to “normal,” is simply to examine the overall trend in housing prices. But which measure of housing prices should be used? Standard & Poor’s Case-Shiller Index is the most-often cited; according to that index, housing prices increased nationally by over 240 percent between 1997 and 2006, then dropped 120 percent between 2006 and 2009. In contrast, an index produced by the Federal Housing Finance Agency (FHFA), based on mortgage data from Fannie Mae and Freddie Mac, shows a much smaller increase and a correspondingly slower decline. A major source of the difference is that the Case-Shiller Index is based on data from county recorder and assessor offices and reflects all sales transactions, including those financed with subprime mortgages, while the FHFA’s index is based only on transactions involving conventional conforming mortgages. Thus, the Case-Shiller is more comprehensive but

66

also more volatile. The discrepancies between the indices are even more apparent for smaller regions: When the bubble was growing, Case-Shiller showed bigger increases in prices in areas with a large number of the riskiest mortgages; after the bubble popped, it showed larger declines in areas with large numbers of distressed sales such as short sales and foreclosures.

Source: Federal Housing Finance Agency price index “Housing Price Index,” http://www.fhfa.gov/Default.aspx?Page=14; Case-Shiller index in Standard and Poor’s, “S&P/Case-Shiller Home Price Indices,” http://ww- w.standardandpoors.com/indices/sp-caseshiller-home-price-indices/en/u- s/?indexId=spusa-cashpidff--/----p-us/. Figure 3.1 Case-Schiller and FHFA Indices, 2000–2011

For homeowners who want to know whether the housing market has stabilized, and the direction in which the value of their own home is likely to move, both the Case-Shiller and the FHFA’s indices suffer serious limitations. One is that the smallest regions covered by the indices are metropolitan areas; however, there can be significant variation in trends within these areas (see later section on geographic units). Some private companies estimate their own local indices, but the smaller the area, the fewer the sales, and estimates may be unduly influenced by one or two outliers.

Another problem with both indices is that they are based on sales and

67

therefore can accurately reflect the overall market only if the houses that sell are representative of the houses that are not sold. While this may be a safe assumption in normal times, it is not clear that this was the case in the aftermath of the bubble, when distressed sales became a disproportionately large share of transactions. With distressed sales, the price is generally lower than a comparable house could attract in a normal sales negotiation without the time pressure. FHFA does produce indices based on appraisals used in refinancing, and those show smaller declines than the more general housing index; however, those appraisal-based indices may overestimate average value because only those houses that have held their value would be eligible for refinancing.

Could the Bubble Have Been Foreseen?

Looking back at the run-up in housing prices, it is easy to say that it was a bubble—a situation where rising prices are unsustainable, driven by investor expectations and easy credit. Under these circumstances, prices are destined to fall back to more reasonable levels. By contrast, if prices had increased because of a reduction in housing supply or a sustainable increase in demand—what economists call market fundamentals—then price levels could conceivably stay at their high levels.

But should it have been obvious at the time that what we were experiencing was a bubble? Alan Greenspan, former chairman of the Federal Reserve, has famously argued that no one could have foreseen the bubble, but there certainly were some analysts that did. one was Dean Baker, who was arguing as early as 2002 that the increasing housing prices were not consistent with market fundamentals. Baker compared the trend in housing prices to the trend in several variables that have traditionally been correlated with housing prices, such as inflation, rental costs, vacancies, and demographics. Noting the divergence in trends during the late 1990s and early 2000s, he concluded that a bubble was the most plausible explanation for the increasing housing prices.

In contrast, economists Charles Himmelberg, Christopher Mayer, and Todd Sinai argued that conventional measures, such as the price-to-rent ratio, were not really accurate reflections of housing costs. Instead, they estimated the imputed annual rental cost of owning a home, which attempts to capture the true “user cost” of housing by incorporating long- run costs such as interest rates and tax rates. By that measure, 2004 housing prices in most metropolitan areas did not appear significantly

68

above their fundamental values.

The key source of disagreement is which variables were used as measures of market fundamentals for comparison with housing prices: Baker focused on traditional variables like inflation and vacancies; Himmelberg, Mayer, and Sinai proposed that their calculation of user cost was more appropriate. Because either measure is plausible, it is up to the analysts to make the case for one over the other. Unfortunately, when looking at the housing market, policymakers and most financial analysts focused on the wrong measures. That includes Federal Reserve chairman Ben Bernanke, who said in March 2006, “I agree with most of the commentary that the strong fundamentals support a relatively soft landing in housing.” Two years after those comments, the collapse in the housing market triggered a massive financial crisis and the nation fell into a severe recession.

Box 3.2 Housing Price Indices and Housing Quality

One difficulty with measuring changes in housing prices over time is that the housing stock itself changes. Of course, newly built houses are valued differently than older houses, but over longer stretches of time, the characteristics of new houses have been changing. For example, the average home size in 2009 was 2,700 square feet, almost double what it was in 1970; thus, it should not be surprising for average home prices also to be much higher.

Some measures of housing prices, such as median sale price, do not attempt to control for characteristics of houses or any aspects of the housing market. Such measures may be of limited value for comparisons across long periods of time or across geographic areas where there are likely to be large differences in housing stock. An alternative is price indices based on repeated sales, such as the Case- Shiller or the FHFA indices, which are constructed by measuring the change in price when a given property is sold more than once, with adjustments for the time between sales. This approach accounts for housing characteristics like size, since it is comparing two prices for the identical house.

In times when distressed sales comprise a large share of overall sales, even repeated sales indices may not capture the true change in

69

the value of a home for two reasons: First, a short sale typically would result in a lower price for a home than a regular sale would. Second, homes subject to foreclosure may be in poor condition; if owners have vacated and the house has been left to deteriorate, that will further depress the sale price. Simply comparing repeat sales prices for a home will not capture the lower quality that is contributing to the lower price. In both cases, repeat sales indices will appear to fall more than the decrease in the “true” value of homes.

Sources: Floyd Norris, “Studying Housing Through Distorted Indexes,” New York Times, May 28, 2011, p. B3, http://www.nytimes.com/2011/05/2- 8/business/28charts.html; Carl Bialik, “Only One Person Knows a Home’s Value: Its Buyer,” Wall Street Journal, November 21, 2008, http://online.ws- j.com/article/SB122722235538745845.html.

Homeownership

Is There a Homeownership Gap?

Although the rise in housing prices in the 1990s and first half of the 2000s may have been driven by unrealistic expectations, they also reflected a boom in hom-eownership. Homeownership rates increased steadily over those decades, reaching an all-time high of just over 69 percent in 2006. In addition, that trend spanned all racial and ethnic groups so that “gaps” in homeownership rates between white and nonwhite households began to shrink. Asians, in particular, have relatively high rates of homeownership, even though many are recent immigrants. However, some analysts have suggested that the trends are misleading because of the way homeownership rates are calculated. Specifically, demographers Zhou Yu and Dowell Myers point out that typically homeownership is reported as a function of households. That is, the homeownership rate is the number of owner-occupied households divided by the total number of households (both owner-and renter-occupied). Calculated in this way, the ownership rate will increase when renters become owners (more owner-occupied households) or when renters leave the market (fewer households overall), such as when adult children move back in with parents or individuals living alone start living with roommates. According to Yu and Myers, the rise in homeownership rates over the last two decades is due more to falling household formation than to conversion of renters to owners.

70

Furthermore, the “success” of Asians in attaining homeownership is due almost entirely to their very low rates of household formation; conversely, blacks and Latinos are much more likely to form renter households, which hold down homeownership rates. on a per capita basis, Asians, blacks, and Latinos all have a similar number of homeowners; in all cases the numbers are still far lower than they are for whites, but the gaps are not quite as large as observed rates would suggest.

This controversy highlights the importance of interpreting changes or differences in percentage rates carefully. When both the numerator and the denominator are potentially variable across time or categories, analysts should not jump to conclusions about which one is changing.

Is the Mortgage Deduction a Middle-Class Tax Break?

The importance of homeownership is reflected in the U.S. tax code, which allows homeowners to deduct their mortgage interest payments from their taxable income; in fact, the mortgage interest deduction is the largest federal subsidy for owner-occupied housing, estimated at $131 billion in 2012. Although politicians like to talk about the mortgage interest deduction as a major tax break for the middle class, some analysts argue that the real benefits go to wealthier upper-income households. The controversy hinges largely on how one defines middle class and how the relative benefits from the deduction are measured.

According to a recent report by the Tax Policy Center, the mortgage interest deduction disproportionately benefits those at the top of the income distribution because (1) many lower-income households do not itemize their deductions, (2) the value of the deduction is worth more to those in higher marginal tax brackets, and (3) higher-income households are more likely to buy higher-priced homes and pay more in mortgage interest.

Table 3.1 shows the impact if the mortgage interest deduction were to be eliminated. Among those in the fortieth to sixtieth percentile of income, the literal middle of the distribution, about 22 percent would face a higher tax bill without the deduction, and the average value of the deduction for those households is about $215 or 0.49 percent of income. As income goes up, so does both the percentage of households benefiting from the deduction and the size of that benefit in dollar and percentage terms. And below the fortieth percentile, fewer than 6 percent of households receive any benefit. Thus, one has to have a fairly expansive definition of middle

71

class to say the mortgage interest deduction is a “middle-class tax break.”

On the other hand, the averages in the table include the many households that do not take the deduction at all. If we restrict our attention only to those households that receive the deduction, data from the Joint Center on Taxation suggest that the benefit to those households represents a similar percentage of income across the distribution. For example, among households making $40,000 to $50,000, only 22.9 percent claim the mortgage interest deduction at all; however, among those who do, the average benefit is $797 or approximately 1.8 percent of income. Among households making $100,000 to $200,000, a much larger percentage (64 percent) claim the deduction; among those households, the average benefit is $2,856 or roughly 1.9 percent of income. So although the dollar value is higher, the percentage of income is similar. Many middle-class households may reasonably perceive the deduction as a clear benefit to them, although upper-income households gain from the deduction more often and by a larger dollar amount.

Racial Discrimination by Banks

Do banks discriminate based on race when approving mortgage loans? This seemingly simple question has generated a heated, decades-long research war between two parties: (1) those who want government regulators to force bank to grant loans to more deserving minorities, and (2) bank defenders who argue against government intervention. Early studies appeared to prove the existence of redlining—that is, a banking policy that limits or refuses loans to potential homeowners in predetermined geographic neighborhoods, often those with predominantly minority residents. Consequently, U.S. Congress passed the 1975 Home Mortgage Disclosure Act (HMDA), requiring banks to make records available of the number of loans made in specific census tracts. This new easy-to-use data source prompted newspaper studies of lending practices, often showing discrimination against minorities, including a 1988Atlanta Journal Constitution Pulitzer Prize-winning series called the “The Color of Money.”

Table 3.1

If Mortgage Tax Deduction Is Eliminated (table 2a in TPC report):

72

Income percentile

Percent with higher taxes

Percentage change in aftertax income

Average change in

federal tax bill ($)

Average after- tax income

All units 23.5 −0.93 559 60,371 0–20 0.6 −0.01 2 11,067 20–40 5.5 −0.12 32 25,893 40–60 21.5 −0.49 215 43,678 60–80 45.2 −0.96 689 71,839 80–90 68.5 −1.59 1723 108,418 90–95 74.2 −1.75 2643 151,680 95–99 70.4 −1.63 4234 259,935 Top 1 60.3 −0.41 5393 1,302,188

Housing researchers recognized that the HMDA data used in isolation might not tell the full story about discrimination. Most important, the data did not take into account differences in wealth and income that might readily explain a low loan acceptance rate in minority neighborhoods. In 1992 researchers at the Boston Federal Reserve Bank attempted to overcome this crucial drawback with a new study on all 1990 Boston-area minority loan applications including data on income, credit history, debts, and other financial characteristics. Even after adjusting for the effect of such factors, race remained a significant factor in determining the probability of getting a mortgage.

Paul Craig Roberts, an official in the Ronald Reagan administration, called the Fed study “consciously fraudulent.” The controversy contributed to President Bill Clinton’s withdrawal of Alicia Munnell’s nomination to the powerful Federal Reserve Board of Governors because she had coauthored the Fed study. Concurrent with this political flurry was a substantive statistical debate, including lengthy scholarly rebuttals by the American Banking Association. In the end, the Boston Fed researchers responded successfully to two main criticisms: First, bank defenders pointed out that the Boston Fed study may have suffered from omitted variables. If characteristics leading to loan denial had been overlooked, then their effect might mistakenly be attributed to race. Boston Fed researchers responded that their study included more than 60 variables and that no one had been able to identify a variable to replace race as the

73

explanation for loan denial.

A second criticism of the Boston Fed study pointed out that default rates were similar in minority and nonminority neighborhoods. If minorities had been treated unfairly, critics argued that they should have a lower default rate because minorities were being held to a higher standard. Here the Boston Fed rejoinder was quite technical and therefore not as widely reported as the initial criticism. The Boston Fed authors pointed out that one would not expect discrimination to cause a lower default rate if many minority applicants were near the denial threshold. In this case discrimination would raise the standard for minorities to receive loans and would lower the default rate for minorities compared to what it might have been. But it would not result in a lower default rate compared to whites who, on average, were better off and therefore less likely to default.

In a 1998 review of this debate, Duke University economist Helen F. Ladd concluded, “The [Boston Fed] study has survived close scrutiny by a host of skeptical critics.” Lawrence B. Lindsey, a conservative Reagan- appointee to the Federal Reserve Board of Governors, writing in the foreword to an American Banking Association report, conceded that “discriminatory practices may occur in cases of marginally qualified applicants.” But he nonetheless warned that “caution should be used before jumping to any conclusions” and “discrimination will ultimately be eliminated not by government agencies.” Clearly, policymakers on each side were locked in their positions despite the strong evidence in the Boston study.

In more recent years, researchers have conducted audit studies in which pairs of applicants, identical in all respects except race, request loan information from the same lenders. These studies typically find that African Americans and Latinos receive less information, are quoted higher loan rates, and are offered fewer discounts on closing costs than their white counterparts. These studies provide evidence consistent with the Boston study but are subject to their own set of critiques; for example, auditors generally do not go beyond the pre-application phase, so it is unknown if the discriminatory behavior carries into real loans. Although the debate will likely continue to rage, researchers do have vastly improved data to use in their investigations.

How Many Homeless Are There?

Depending on which news story you read, there are anywhere from

74

600,000 to 1.5 million to more than 3 million people who experience homelessness in the United States. Why such a large range of estimates? Primarily because of differences in the methods that researchers use to measure homelessness.

The lowest estimates come from “point-in-time” counts: literal headcounts of individuals in emergency shelters, transitional housing, or on the street on a given day. The U.S. Department of Housing and Urban Development (HUD), which now provides annual reports on homelessness to Congress, documented 643,067 homeless people on a single night in the last week of January 2009. Roughly one-third were unsheltered, sleeping on streets, in cars or abandoned buildings, or other locations “not meant for human habitation.” Critics of pointin-time estimates argue that many people experience temporary episodes of homelessness; that is, there is a great deal of turnover in the homeless population as some people find housing and others lose housing as a result of social or economic situations such as losing a job or domestic violence. In a study of Philadelphia shelter users, social psychologist Dennis Culhane found that the most common length of stay in a shelter was one night because, as he explained, “anyone who has ever had to stay in a shelter involuntarily knows that all you think about is how to make sure you never come back.” On the other hand, 10 percent of shelter residents were what Culhane called “chronically homeless,” staying for long periods of time, and in his studies these individuals accounted for a large share of city social service and health care expenditures. Point-in-time estimates do not identify how long people have been homeless and thus may both underestimate the total number of people who ever experience homelessness and overestimate the number of “chronically homeless.”

An alternative is to count the number of homeless over a longer period of time. HUD reports that over the one-year period between October 1, 2008 and September 30, 2009, approximately 1.56 million people spent at least one night in an emergency shelter or transitional housing. However, the one-year estimate does not include any measure of the unsheltered homeless and therefore likely underestimates the total number of homeless.

The highest estimates of homelessness come from extrapolations conducted by such organizations as the National Law Center on Homelessness and Poverty and the Urban Institute. Those organizations use point-in-time estimates and convert those numbers to a percentage of the population living in poverty (with a high estimate of 10 percent, based

75

on data from 1996). That percentage is then applied to the poverty population in later years and results in an estimate of 3.5 million. One problem with this approach is that there is no way to gauge whether the percentage of people in poverty who experience homelessness is stable over time.

Finally, some critics argue that any of these estimates are underestimates because even when unsheltered individuals are included, there are many people who will not be counted because researchers simply can’t find them. Many homeless do not want to be found, so they specifically look for shelter in places that are out of sight.

While media stories may tend to focus on the larger, more sensational homelessness numbers, these varying estimates actually provide different types of information that are useful for answering different questions. For researchers studying homelessness, it is therefore important to identify the reason for their work. As a ballpark estimate, roughly 600,000 is a reasonable estimate for those individuals in shelters or on the street on a given night. And as Over the Edge author Martha Burt points out, such estimates are useful if the purpose of the data is to find out how many shelter beds are needed on a particular night. However, if the purpose is to help all those without adequate homes, then the relevant count would be the higher estimates, taken over time, that include those episodically without homes. Thus, the meaning of “homeless” may depend on the question being asked, which in turn will determine the estimate that is most relevant.

Geographic Units

Housing location is used to designate geographic area, a common research variable. Although seemingly easy to define, geographic designation often presents research headaches and, in some cases, requires intervention by the U.S. president’s Office of Management and Budget (OMB).

The smallest geographic classifications are census blocks. For the 2010 Census, the United States was divided into about 11.1 million of these units, corresponding in urban areas to city blocks. They were subdivisions of approximately 65,000census tracts, the statistical unit for many research projects such as the studies of segregation described later.

What is a city? What is a rural area? These are obviously subjective questions, complicated by the growth of neither-urban-nor-rural suburbia. Not surprisingly, the division between city and country involves arbitrary

76

classification and categories that change over time. Beginning in 1910, the Census defined any incorporated place with more than 2,500 residents as urban. The standard shifted slowly so that by 1980, urban areas required a population of 50,000 and “built-up” characteristics. Since 1949, the Bureau of the Budget (now the OMB) has also attempted to create useful boundaries for urban areas called Metropolitan Statistical Areas (MSAs). Over the years, the idea of defining areas that are economically and socially linked has expanded so that now a more general term —metropolitan area or MA—is used to refer collectively to metropolitan statistical areas (MSAs), consolidated metropolitan statistical areas (CMSAs), and primary metropolitan statistical areas (PMSAs). In 2000 the term core based statistical area (CBSA) was created to refer to both metropolitan and micropolitan areas, where the latter have at least one urban population cluster of between 10,000 and 50,000 people.

There are over 360 metropolitan statistical areas and 560 micropolitan statistical areas in the United States; however, those numbers are in flux, as the standards for defining metropolitan areas have changed periodically, often coinciding with the decennial census. This can create problems for researchers when certain MSAs are eliminated and others are newly created. For example, Rapid City, South Dakota, lost its metropolitan status in 1980 because of population loss, while 35 new areas were established in 1981 based on the previous year’s census.

The increasing number and complexity of metropolitan-area terms is due largely to a corresponding complexity in the relationships of the areas that are captured. For example, prior to the introduction of PMSAs and CMSAs, Nassau and Suffolk counties of New York’s Long Island were designated a separate MSA, even though both counties are closely tied to the New York MSA, which included New York City and counties to the north. Nassau and Suffolk are now designated a PMSA within the larger New York-Northern New Jersey-Long Island CMSA.

Even with these various types of area designations, area boundaries can present research problems. Some MSAs include entire counties (or cities in New England) that have close social and economic relationships with the central urban area. Because some western counties are so large, the resulting MSA is geographically huge, stretching more than 50 miles into the Cascade Mountains in the case of Seattle, Washington, while Los Angeles County extends 25 miles into the Mojave Desert. In both cases the outer reaches of the MSA are entirely unpopulated, and a huge variation in political, social, and economic characteristics exists within the area.

77

What can researchers do about these problems? First, consult the data source. Government publications typically comment in detail about geographic issues, including any changes in the definition of urban areas and MSAs. If a researcher needs continuity of geographic units, he or she may need to consult a second set of statistics based on older boundaries. It is also important to consider the appropriate geographic area for a given analysis; although MSAs are still the most commonly used unit, there are many applications where the much smaller census tracts, or larger CMSAs or CBSAs, might be more informative. Finally, the U.S. Census Lookup (see “Where the Numbers Come From”) includes the capability to figure calculations for various geographical units.

Segregation

In 1968 the National Advisory Commission on Civil Disorders warned: “Our Nation is moving toward two societies, one black, one white— separate and unequal.” Has this prediction come true? Data on segregation were considered so sensitive by the Richard M. Nixon administration that studies by the Census Bureau were blocked during the 1970s out of fear that evidence of continued segregation would be politically explosive. Many more studies have been conducted since the 1980s, sometimes with contradictory findings because of the difficulties in measuring segregation.

The first problem confronting researchers is the relevant geographic scale. On the level of MSAs, the United States became more integrated during past decades in the sense that many cities had more diverse populations. But segregation is still present because, within those cities, racial groups live in separate neighborhoods. Consequently, most research on segregation looks at census tracts, for which total counts by racial group are available in each census (see Chapter 2 on problems in defining racial groups). But even on this small-scale level, there are different types of segregation indexes, discussed in a vast literature that seems to point to contradictory trends. For example, in a study released by the Manhattan Institute titled “The End of the Segregated Century,” economists Edward Glaeser and Jacob Vigdor calculate “dissimilarity indices,” which measure the proportion of a group that would have to move in order to achieve perfectly integrated neighborhoods. By this measure, segregation declined significantly between 1970 and 2010.

An alternative way to measure segregation is with an index of “exposure” of groups to other groups. In a rebuttal of the Manhattan

78

Institute study, Richard Rothstein of the Economic Policy Institute highlights the fact that in 2010, the average black American lived in a neighborhood that was only 35 percent white, a percentage that is actually lower than it was in the 1940s and that has not changed significantly in the last several decades.

Both measures mask significant variation across neighborhoods; some African Americans are living in neighborhoods with almost no white residents, while others live in neighborhoods that are quite integrated. The Manhattan Institute study also notes that there are important differences in the trends in various parts of the country, with Sun Belt cities showing larger reductions in the dissimilarity index than elsewhere. Sociologist Douglas Massey points out that the declines in dissimilarity indices have been largest in cities with relatively small black populations, while the changes in cities with large black populations are much more modest—if they exist at all. He also notes that in many locations, dissimilarity indices have fallen because of the movement of Asians and Hispanics, rather than African Americans.

One lesson for researchers is to be aware of the differences that can arise when there are multiple ways to measure a variable. The claim that segregation has ended appears extreme, particularly when multiple measures of segregation are considered.

Is Your City the Best Place to Live?

What is the overall desirability of a city or, more technically, an MSA? Obviously, such a statistic must combine a great variety of data. Two contrasting methods for making MSA comparisons illustrate the relative advantages of different summary statistics—one quite simple, the other complex.

Ratings

It seems like every few months, a new report comes out that names city X as “the most livable city” or “the best place to live” or “the safest city” or “the best city for the elderly,” etc. All of these rankings should be regarded with caution because they are created with methods that involve arbitrary choices. Most people don’t realize that small changes in methodology can sometimes lead to completely different rankings.

Most “best city” indices are based on several underlying attributes such

79

as climate, crime, housing, culture, education, access to public transportation, and so on. But different indices rely on different attributes. For example, Money Magazine’s “Best Places to Live” are screened first for income, diversity, education, and crime, and then the list is whittled down with rankings on “job growth, home affordability, safety, school quality, health care, arts and leisure, diversity, and several ease-of-living criteria”; the list is further reduced by factoring in economic data, which is weighted most heavily; journalists then visit cities in the final group to assess traffic, parks, and intangibles like “community spirit.” In contrast, Places Rated Almanac bases their rankings on nine categories (cost of living, economy, weather, transportation, health care, education, recreation, location, and safety), where a city’s rating in each area is simply summed up to one number. In both cases the rankings depend on which variables are included (Money Magazine includes fiscal strength of local and state governments in their measure of the economy while the Almanac does not), how variables are measured (Money Magazine measures housing costs with median home prices while the Almanac also incorporates utility costs), and how variables are combined.

The issue of how variables are combined is particularly contentious because it tends to be the least transparent aspect of the analysis. Most media reports of city rankings will mention what factors go into the index but rarely explain how those factors are weighted, that is, how much importance each variable is given. Even if two indices use the same variables, measured in the same way, different weights can lead to very different rankings. As noted, the Almanac simply adds up each city’s ranking in attributes ranging from climate to recreation. Researchers at AT&T’s Bell Laboratories reanalyzed the Places Rated data to show that with an alternative weighting scheme, 134 cities could claim the number one spot, or by reversing the process, 150 different cities could be ranked last (see Table 3.2). For example, with a large weight on crime, safe Beaver County, Pennsylvania, reaches number one, but a large weight on its economic outlook pushes Beaver County to the bottom of the list.

Table 3.2 shows how an unscrupulous user can manipulate ratings outcomes, as was true in the case of a Chicago publicist who asked the Bell researchers for a way to make Chicago come out first. Indeed, weighting on transportation and recreation lifts Chicago to the number one spot. In fairness to the authors of the Almanac, we should note that they caution against the use of the overall ratings altogether. Even though press reports typically focus on precisely this aspect of their list, a short list of

80

winners appears late in the book and the losers are not highlighted at all.

Hedonic Index

A more sophisticated estimate of the quality of life measures how willing individuals are to accept higher housing costs and lower wages in exchange for amenities such as good weather and good schools. Using this method, called hedonic pricing, economists Glenn C. Blomquist, Mark C. Berger, and John P. Hoehn estimated that in the 1980s, residents of Pueblo, Colorado, “traded” $3,289 in higher costs to live in this highest- ranked city, while residents of St. Louis, Missouri, enjoyed $1,857 compensation for their lowest-ranked city. Although this technique avoids the subjective judgments of ranking indices, it assumes there is a fluid marketplace for housing. If some desirable cities are inaccessible because of job and family ties, then the hedonic method will not accurately measure a city’s “value.”

For many social science topics there are relatively simple statistics such as the Places Rated Almanac, as well as complex statistics such as the he- donic index. A comparison of the two methods for rating cities illustrates drawbacks for each approach. Simple indexes are arbitrary in the sense that another researcher might attach different importance to each variable, for example, counting rain as a greater disadvantage for Seattle. And the hedonic index and many similar measurements assume a well-functioning economic market in which individuals freely make informed judgments. Because of the complexity of statistics such as the hedonic index, users may not be aware of the underlying assumptions it requires. Researchers can learn from both approaches, borrowing the easy-to-use characteristics of Places Rated with the more subtle insights of the hedonic index.

Table 3.2

The First Shall Be Last

Depending on the weighting of attributes in the Places Rated Almanac, the same city can be rated first (#1) or last (#329).

City Attributes weighted highly to cause #1 ranking

Attributes weighted highly to cause #329 ranking

81

Atlantic City, NJ Economy, recreation Housing, crime Beaver County, PA Crime Economy Detroit, MI Health care Crime, economy Duluth, MN Housing, crime Climate, economy Fort Myers, FL Recreation, economy Climate housing, art Honolulu, HI Crime, climate, recreation Housing, education Minneapolis, MN Art, health care Climate New York, NY Health care, art Crime, housing Salem, OR Climate, education Economy, housing Salt Lake City, UT Recreation Education, housing Washington, DC Education, transportation Housing, crime

Summary

The issues reviewed in this chapter underscore the ambiguity of social statistics. Attempts to measure the magnitude of the housing crisis, and debates over whether the bubble could have been foreseen, illustrate the discrepancies that can arise when multiple measures can be used to capture a given concept. Based on the Case-Shiller Index, the volatility in the housing market looks much worse than if one uses the FHFA index; similarly, the housing bubble seems much more apparent if one compares housing prices to certain other measures of the market. One lesson for researchers is to consider multiple measures and to be upfront about the pros and cons of each one.

The discussion of homeownership rates and so-called best city ratings systems highlight the challenges of analyzing variables that are themselves a combination of other variables. In the case of homeownership rates, variation in the percentage may be due to variation in either the numerator or the denominator, and isolating which one is changing has important policy implications. With city ratings that attempt to combine several location attributes into a single number, the method used to combine those attributes can have significant effects on which cities come out on top. Analysts who must work with these sorts of compound measures should make sure to investigate how their results change if different assumptions are made about the underlying variables.

Researchers should also be careful about how they define terms that might be politically loaded, such as middle class and homeless. Although news media use these terms with great frequency and little if any

82

1.

2.

explanation, a closer look at the data often reveals that how one defines middle class and homeless will determine the statistics that are subsequently reported. Similarly, the media often oversimplify complex issues, such as the sophisticated statistics used in the Boston Fed’s effort to measure bank mortgage racial discrimination. Both consumers and producers of social science statistics should be vigilant that such oversimplification does not distort findings in the process.

Finally, the choice of geographic units is a common step in research projects. Although the federal government attempts to standardize its data carefully along common geographic boundaries, the definition of urban versus rural areas and the designation of metropolitan areas can pose problems for some research projects. For example, analysis of racial segregation requires attention to population shifts within neighborhoods as well as within larger-scale units such as MSAs, and analyses of trends in housing markets at the MSA level may not capture within-area variation that is more meaningful for individual homeowners. The work of many social science researchers is constrained by data that has been collected and organized by others; when data are available for multiple geographic levels, however, it is important to select the appropriate level for the research question.

Case Study Questions

Between 1999 and 2009, the Annual Housing Survey measured the numbers of dwellings with deficiencies, as shown in the following table. During this time period, the number of occupied structures increased by about 9 percent. How might these data on deficiencies be used to measure improved, stable, and declining housing quality?

Condition present in neighborhood

1999 2009 Signs of rats in last 3 months 892,000 613,000 Streets in need of major repair 5,317,000 6,604,000 Bothersome odors present 3,983,000 5,434,000

Just as there is an underground economy (see Chapter 7), there are unreported and often illegal housing units. In particular, illegal housing conversions are estimated to account for the

83

3.

4.

5.

6.

largest proportion of new housing units for some locales, but they are often missed in the census and the AHS. How might these missing housing units affect measurement of housing quality?

The number of new-home sales changes dramatically from month to month. For example, December 2010 sales rose by 15.3 percent over the previous month, January 2011 sales fell by 6.4 percent, and February 2011 sales fell by 9.4 percent. These data are adjusted for seasonal variation—that is, they take into account typical month-by-month changes during past years. What other factors might cause such erratic month-to-month variation?

Which is the most-visited U.S. city? After Orlando predicted 48.6 million visitors in 2010, New York counted 48.7 million, only to be topped by Orlando’s 51.5 million final count. List the reasons why it is difficult to measure the number of tourists visiting a city.

Median home prices in Orange County, California, are generally two and half to three times as high as in Louisville, Kentucky. Nonetheless, housing experts believe this statistic underestimates the actual difference in home prices between the two locations. Why is this likely to be true?

From the list that follows, explain which criteria should be used to define homelessness: (1) individuals visible at 2:00 a.m. on a street location known to be used as a sleeping location; (2) residents of emergency shelters; (3) hotel residents using social service agency vouchers; (4) residents of shelters for abused women and runaway youth; (5) individuals in jail who claim no residence; (6) residents of sober living homes; (7) families doubled up with friends?

References

Data Sources (37)

U.S. Census (37)

History of housing surveys in John S. Adams, Housing America in the 1980s (New York: Russell Sage Foundation, 1987), pp. 4, 29–30; see also Joseph W. Duncan

84

and William C. Shelton, Revolution in United States Government Statistics (Washington, DC: U.S. Government Printing Office, 1978), p. 39. On census form, see Bryant Robey, “Two Hundred Years and Counting: The 1990 Census,” Population Bulletin 44, no. 1 (April 1989): 15–24. For advice on using census housing data, see U.S. Conference of Mayors, Assessing Elderly Housing (Washington, DC: American Association of Retired Persons, 1986). Data sample in Christopher Mazur and Ellen Wilson, Housing Characteristics: 2010, U.S. Census Bureau, 2010 Census Briefs, C2010BR-07, October 2011.

American Housing Survey (37)

History in Adams, Housing America in the 1980s, pp. 33–36. Data sample in U.S. Department of Housing and Urban Development (HUD) and U.S. Department of Commerce, American Housing Survey for Selected Metropolitan Areas, 2009:Current Housing Reports, H170/09, July 2011, http://www.census.gov/housi- ng/ahs/data/metro.html.

Other Industry Data (37)

On housing starts, see Norman Frumkin, Guide to Economic Indicators (Armonk, NY: M.E. Sharpe, 1990), pp. 127–31; “New Residential Construction,” Bureau of the Census, http://www.census.gov/construction/nrc/. On construction and housing surveys, see “Housing,” Bureau of the Census, http://www.census.gov/housing/. Dodge data in Cynthia Bansak and Anne Toohey, “Comparing Dodge’s Construction Potentials Data and the Census Bureau’s Building Permits Series,” Economic Review, March/April 1994, pp. 23–37. Data sample in “Market Absorption of Apartments, Characteristics Report,” Social, Economic, and Housing Statistics Division of the U.S. Bureau of the Census, H130/10-C, April 2011, http://www.census.gov/hhes/www/housing/soma/char10/char10txt.html.

Price Data (38)

Shelter index in U.S. Department of Labor, Bureau of Labor Statistics, “Consumer Price Indexes for Rent and Rental Equivalence,” http://www.bls.gov/cpi/cpifact- 6.htm. Federal Housing Finance Agency, “Housing Price Index,” http://www.fhf- a.gov/Default.aspx?Page=14; National Association of Realtors statistics in “Real Estate Sales Statistics: Existing-Home Sales and Pending Home Sales,” http://ww- w.realtor.org/research/research/ehspage; Case-Shiller Index in Standard and Poor’s, “S&P/ Case-Shiller Home Price Indices,” http://www.standardandpoors.co- m/indices/sp-case-shiller-home-price-indices/en/us/?indexId=spusa-cashpidff--p-u- s----. Data sample from National Association of Realtors, “Affordable Housing Real Estate Resource: Housing Affordability Index,” http://www.realtor.org/resear- ch/research/housinginx.

85

Controversies (39)

Housing Crisis (39)

How Big Was the Bubble and Is It Over Yet? (39): Differences in indices in Carl Bialik, “Behind the Home-Price Indexes,” The Numbers Guy (blog), November 20, 2008, http://blogs.wsj.com/numbersguy/behind-the-home-price-indexes-460/; Bialik, “Only One Person Knows a Home’s Value: Its Buyer,” Wall Street Journal, November 21, 2008, http://online.wsj.com/article/SB122722235538745845.html; Federal Housing Finance Agency, “The Impact of Distressed Sales on Repeat- Transaction House Price Indexes,” May 27, 2009, http://www.fhfa.gov/webfiles/2- 916/research-paper_distress%5b1%5d.pdf. Floyd Norris, “Studying Housing Through Distorted Indexes,” New York Times, May 28, 2011, p. B3, http://www.n- ytimes.com/2011/05/28/business/28charts.html.

Could the Bubble Have Been Foreseen? (41): Paul Rosenberg, “Who Could Have Foreseen the Housing Bubble Collapse? Dean Baker, That’s Who—in 2002,” Open Left (blog), August 29, 2009, http://www.openleft.com/diary/14839/who-co- uld-have-foreseen-the-housing-bubble-collapse-dean-baker-thats-whoin-2002; Dean Baker, “The Run-Up in Home Prices: Is it Real or Is it Another Bubble?” Center for Economic and Policy Research, August 2002, http://www.cepr.net/inde- x.php/publications/re-ports/the-run-up-in-home-prices-is-it-real-or-is-it-another-bu- bble/; Dean Baker, “The Housing Bubble and the Financial Crisis,” Real-World Economics Review 46 (May 20, 2008): 73–81, http://www.paecon.net/PAERevie- w/issue46/Baker46.pdf. Charles Himmelberg, Christopher Mayer, and Todd Sinai, “Assessing High House Prices: Bubbles, Fundamentals and Misperceptions,” Journal of Economic Perspectives 19, no. 4 (2005): 67–92. Bernanke quote in Martin Crutsinger, “Documents Show How Fed Missed Housing Bust,” Associated Press, January 13, 2012, http://my.news.yahoo.com/documents-show-fed-missed-- housing-bust-022039343.html.

Homeownership (43)

Is There a Homeownership Gap? (43): Zhou Yu and Dowell Myers, “Misleading Comparisons of Homeownership Rates When the Variable Effect of Household Formation Is Ignored: Explaining Rising Homeownership and the Homeownership Gap Between Blacks and Asians in the U.S.,” Urban Studies 47, no. 12 (2010): 2615–2640.

Is the Mortgage Deduction a Middle-Class Tax Break?(43): Benefits to top of distribution in Alexander C. Hart, “Is the Mortgage-Interest Deduction Really A Middle-Class Tax Break?” The New Republic, November 16, 2010, http://www.tn- r.com/blog/jonathan-cohn/79206/the-mortgage-interest-deduction-really-middle-cl- ass-tax-break. Eric Toder et al., “Reforming the Mortgage Interest Deduction” (Washington, DC: Urban-Brookings Tax Policy Center, April 2010), http://www.t-

86

axpolicycenter.org/uploadedpdf/412099-mortgage-deduction-reform.pdf; data from Joint Center on Taxation in Kevin Drum, “Why Everyone Loves the Mortgage Interest Deduction,” Mother Jones, July 13, 2011, http://motherjones.com/kevin-d- rum/2011/07/why-everyone-loves-mortgage-interest-deduction. See also Brian McCabe, “Despite Benefit Disparities, Middle Class Supports Mortgage Deduction,” FiveThirtyEight (blog), July 13, 2011, http://fivethirtyeight.blogs.nyti- mes.com/2011/07/13/despite-benefit-disparities-middle-class-supports-mortgage-- deduction/.

Racial Discrimination by Banks (44)

History of controversy from the banks’ point of view in Anthony M. Yezer, Fair Lending Analysis: A Compendium of Essays on the Use of Statistics (Washington, DC: American Bankers Association, 1995); “The Color of Money” series in Atlanta Journal-Constitution, May 1–4, 1988; Boston Fed Response in Lynn Elaine Browne and Geoffrey M.B. Tootell, “Mortgage Lending in Boston—A Response to the Critics,” New England Economic Review 19 (September-October 1995): 53–72. See also Susan Wachter, “Discrimination in Financial Services: What Do We Know?” Journal of Financial Services Research 11 (1997): 205–8, and George Benston, “Discrimination in Financial Services: What Do We Not Know?” Journal of Financial Services Research 11 (1997): 209–13; Paul Craig Roberts in Browne and Tootell, “Mortgage Lending,” p. 54; Alicia Munnell quoted in Peter Passell, “Race, Mortgages, and Statistics,” New York Times, May 10, 1996, p. C1; Helen F. Ladd, “Evidence on Discrimination in Mortgage Banking,” Journal of Economic Perspectives 12 (Spring 1998): 53. Audit studies discussed in Stephen Ross et al., “Mortgage Lending in Chicago and Los Angeles: A Paired Testing Study of the Pre-application Process,” Journal of Urban Economics 63, no. 3 (May 2008): 902–19.

How Many Homeless Are There? (46): HUD estimate in U.S. Department of Housing and Urban Development, The Annual Homeless Assessment Report to Congress, June 2010, http://www.huduser.org/portal/taxonomy/term/65. Comparison of estimates in “How Many People Experience Homelessness?” National Coalition for the Homeless, updated December 15, 2011, http://nationalh- omeless.org/factsheets/How_Many.html. Culhane quoted in Malcolm Gladwell, What the Dog Saw and Other Adventures (New York: Little, Brown, 2009); Martha Burt, Over the Edge: The Growth of Homelessness in the 1980s (New York: Russell Sage Foundation, 1992).

Geographic Units (49)

Origins of tracts in Duncan and Shelton, Revolution in United States Government Statistics, p. 209; defined in Robey, “Two Hundred Years,” p. 25. Guide to the Census in Bureau of the Census, “A Guide to State and Local Census Geography,”

87

CPH-I-18 (Washington, DC: U.S. Department of Commerce, 1993), and “Geographic Areas Reference Manual” (Washington, DC: U.S. Department of Commerce, November 1994). Definition of urban in James L. Newman and Gordon Matzke, Population: Patterns, Dynamics, and Prospects (Englewood Cliffs, NJ: Prentice Hall, 1984), pp. 50–51; definition of areas in Bureau of the Census, “Metropolitan and Micropolitan Statistical Areas,” http://www.census.go- v/population/metro/, and in Office of Management and Budget (OMB) Bulletins, http://www.census.gov/population/metro/data/omb.html.

Segregation (49)

“Our Nation,” National Advisory Commission on Civil Disorders, Report (Washington, DC: U.S. Government Printing Office, 1968), p. 1. Nixon administration blocked in “Middle-Class Black Housing Still Largely Segregated,” Washington Post, December 30, 1987, p. A4. On historical trend, see Reynolds Farley and Walter R. Allen, The Color Line and the Quality of Life in America (New York: Oxford University Press, 1987), pp. 139–57; see also Christine H. Rossell, “Does School Desegregation Policy Stimulate Residential Integration?” Urban Education 21, no. 4 (January 1987): 403. Overall trend in Douglas S. Massey and Nancy Denton, “Trends in the Residential Segregation of Blacks, Hispanics, and Asians: 1970–1980,” American Sociological Review 52 (December 1987): 802–25. Edward Glaeser and Jacob Vigdor, “The End of the Segregated Century: Racial Separation in America’s Neighborhoods, 1890–2010,” Manhattan Institute for Policy Research, Civic Report No. 66, January 2011, http://www.man- hattan-institute.org/html/cr_66.htm; Massey quoted in Sam Roberts, “Segregation Curtailed in U.S. Cities, Study Finds,” New York Times, January 30, 2012, http://- www.nytimes.com/2012/01/31/us/Segregation-Curtailed-in-US-Cities-Study-Find- s.html; Richard Rothstein, “The ‘End of the Segregated Century’?” Working Economics (blog), Economic Policy Institute, February 3, 2012, http://www.epi.or- g/blog/manhattan-institute-study-segregation/.

Is Your City the Best Place to Live? (51)

Ratings in David Savageau, Places Rated Almanac (Washington, DC: Places Rated Books, 2007), http://placesrated.expertchoice.com; “Best Places to Live,” Money Magazine, http://money.cnn.com/magazines/moneymag/bplive/2011/. See also Carl Bialik, “The Trouble with Rankings,” The Numbers Guy (blog), April 30, 2010, http://blogs.wsj.com/numbersguy/the-trouble-with-rankings-934/, and Bialik, “Why One Top-10 List’s Leader Is Another’s Also-Ran,” Wall Street Journal, May 1, 2010, http://online.wsj.com/article/SB10001424052748704302304575214- 791145457742.html. Discussion of Bell Labs study in Richard A. Becker, Lorraine Denby, Robert McGill, and Allan R. Wilks, “Analysis of Data from the Places Rated Almanac,” The American Statistician 41, no. 3 (August 1987): 169–86. Explanation of hedonic approach in G.C. Blomquist, M.C. Berger, and J.P. Hoehn,

88

“New Estimates of Quality of Life in Urban Areas,” American Economic Review 78, no. 1 (March 1988): 89–107.

Case Study Questions (54)

1. U.S. Department of Housing and Urban Development, American Housing Survey for the United States, H150/09, March 2009, http://www.census.gov/housi- ng/ahs/data/national.html.

2. John S. Adams, Housing America in the 1980s (New York: Russell Sage Foundation, 1987), p. 44.

3. U.S. Department of Housing and Urban Development, “New Residential Sales in February 2011” (press release, March 2011), http://www.census.gov/cons- truction/nrs/.

4. Carl Bialik, “Seeking Standards in Tourism Management,” The Numbers Guy (blog), June 3, 2011, http://blogs.wsj.com/numbersguy/seeking-standards-in-t- ourism-measurement-1063/.

5. “Orange County Tops U.S.,” Los Angeles Times, August 12, 1988, p. 22.

6. City of Pasadena, “The City of Pasadena Homeless Count,” September 23, 1992, p. 15, final report dated April 18, 1994, available at http://www.iurd.org/pas- a-denaresearch/documents/1992HomelessCount.html.

89

4

Health □□□□

Making sense of health statistics involves the expertise of many disciplines ranging from economics to medicine. This chapter reviews controversies about a cross-section of these statistics, including infant mortality, life expectancy, cancer, HIV, heat-related deaths, and traffic fatalities. These diverse examples were chosen because they offer real-world illustrations of conflicting statistics about health and safety. A final section examines benefit-cost analysis, a commonly used technique for evaluating public health policies that also yields contradictory results. As with other social science statistics, careful examination of the underlying data helps resolve the apparent statistical quandaries. Clearly, the task is made especially difficult by the varied sources used in health research, but an understanding of the data is a necessary first step for successful evaluation of health care issues.

Where the Numbers Come From

Organizations Data sources URL National Center for Health Statistics, U.S. Department of Health and Human Services

U.S. Vital Statistics; National Survey of Family Growth

www.cdc.gov/nchs/

Bureau of the Census, U.S. Department of Commerce

Current Population Survey www.census.gov

National Cancer Institute, National Institutes of Health

Surveillance, epidemiology, and end results

www.cancer.gov

World Health Organization Individual country reporting

www.who.int

90

Data Sources

U.S. Health Surveys

Some U.S. health statistics are based on complete counts of all relevant individuals. For example, U.S. Vital Statistics attempts to count all births and deaths (see Chapter 2), and the U.S. Centers for Disease Control and Prevention collect as complete a count as possible of many diseases and causes of death. But because it is not feasible to take a complete count for most other health statistics, survey data are used instead. Extensive surveys are compiled by the National Center for Health Statistics in the publications of the Vital and Health Statistics Series, which are extraordinarily specific, covering such detail as “Percent of persons 18 years of age and over who ate breakfast every day.”

For data on drug abuse, the National Survey on Drug Use and Health (NSDUH) by the U.S. Health and Human Services Department offers user demographics for a number of illicit drugs. The Monitoring the Future (MTF) survey, sponsored by the National Institute on Drug Abuse (NIDA; a branch of the U.S. National Institutes of Health) and conducted by the University of Michigan’s Survey Research Center, provides data on drug use by eighth-, tenth-, and twelfth-grade students.

Other Government Surveys

U.S. government surveys covering health-related topics include the U.S. Census Bureau’s Current Population Survey (see Chapter 9), and specialized data collected by the Social Security Administration, the Bureau of Indian Affairs, the National Highway Traffic and Safety Administration, and the Food and Drug Administration.

Private Surveys

Private-sector organizations also assemble health statistics. Examples include health care associations for physicians, nurses, and hospital administrators; voluntary organizations such as the American Heart Association, the American Cancer Society, and the Cystic Fibrosis Foundation; and the insurance industry, most notably the Metropolitan Life Insurance Company (MetLife) surveys of health and medical costs.

Data Sample: In 2010 Monitoring the Future surveyed hookah

91

smoking for the first time, finding that 17 percent of twelfth graders had used a hookah at least once in the past year. Males had only a slightly higher usage rate, 19 percent versus 15 percent for females.

Worldwide Data

For worldwide health statistics, research will likely begin with World Health Organization (WHO) publications, complete and detailed vital statistics, and comparative data on infectious diseases and health care usage. Researchers should be cautious in making comparisons between countries because reporting systems and standards vary widely. The WHO information service will assist researchers in assessing the comparability of data between countries.

Data Sample: In the World Health Statistics Annual, we learn that in El Salvador there were 20 cases of malaria and 126 cases of mumps in 2009.

Controversies

Infant Mortality

The infant mortality rate is a key health status indicator. Because it is easy to collect data on infant mortality, at least in countries where hospital births prevail, this statistic provides a potentially accurate comparison of health care in different places and at different times. In addition, the infant mortality rate responds quickly to improvements—or failures—in health care. For example, one of the first signals of problems in health programs in the Soviet Union was an increase in infant mortality during the 1970s.

Controversy about the U.S. infant mortality rate focuses on an apparent failure of the United States to reduce infant mortality to levels already achieved in other countries (see Table 4.1). Infant mortality traditionally is measured in deaths per 1,000 live births. A 1912 Federal Children’s Bureau survey estimated this death rate at more than 100 of every 1,000. Since then, improvement has been ongoing, and by 2011 the death rate was estimated at 5.98 per 1,000 (see Figure 4.1). Nonetheless, infant mortality in the United States exceeds that of most other industrialized countries. Japan, Singapore, Sweden, and Hong Kong have the lowest infant mortality of all, fewer than three per 1,000—that’s less than one- half the U.S. rate.

92

Bernadine Healy, former director of the National Institutes of Health, argues that we can’t fairly compare U.S. infant mortality rates with other countries. one reason is that the definition of infant mortality varies from country to country and causes U.S. rates to be misrepresented when compared with rates on a global scale. The measurement problems arise because in some countries, such as Switzerland, early births are counted as miscarriages and therefore are not calculated in the infant mortality rate. In the United States, however, the infant mortality rate measures the rate at which babies less than a year old die. Therefore, all births showing any signs of life figure into the value, no matter how long the pregnancy or how small the fetus. Even after correcting for different measures of infant mortality, though, the U.S. rate still appears high. A U.S. Centers for Disease Control and Prevention (CDC) study, taking into account reporting variations, concluded that the United States ranked behind seventeen European countries, followed by Slovakia, Poland, and Hungary.

By one measure the United States stands ahead of other countries: the infant mortality for pre-term infants is among the lowest in the world. Unfortunately, the United States also has far more of these early and far riskier births. The U.S. mortality rate for infants born after 37 weeks is only slightly higher than most European countries; it is the relatively large proportion of pre-term infants, born alive but die shortly thereafter, that pulls the overall infant mortality to the bottom of the list. In explaining the disproportionate number of U.S. pre-term births, some researchers point to the high rate of babies born to unmarried teenage mothers, for whom pre- term births are more common. Others attribute these early births and the somewhat higher mortality rate for full-terms births to the prevalence of poverty in the United States and limited access to prenatal care for low- income mothers.

Comparison of infant mortality rates across countries presents a challenge to researchers, and there is considerable controversy about the differences. First, the often-cited official infant mortality rates should be interpreted as suggestive, but not definitive, especially when comparing two countries for which the underlying definition of infant mortality may differ. Moreover, a full understanding of U.S. infant mortality also needs to confront the high overall death rate as well as the troubling number of pre-term babies—a stark contrast to the nation’s relative success in keeping pre-term babies alive.

93

Abortion

Abortion statistics are unusual in that the best data are available from a nongovernmental source. The Alan Guttmacher Institute, a private foundation, provides the most detailed abortion data based on a survey of all facilities known or expected to have provided abortion services, including hospitals, clinics, and physicians’ offices. In addition, abortion data have been collected by the CDC from state health departments since 1970, but the reported total number of abortions is incomplete and lower than the Guttmacher estimates by about 50 percent. In both surveys, abortions as a proportion of live births have been declining for many years, falling to 22 percent in 2008, or 1,210,000 abortions using the Guttmacher count.

Are We Living Longer?

It is difficult not to feel a little cheered when one reads about increases in average life expectancy. For example, in 2009 overall U.S. average life expectancy at birth was 78.2 years, a three and one-half year increase since 1984. But few readers understand the limitations of this statistic, especially when it is used for political purposes, as it often is in the debate about U.S. health care.

Sources: U.S. Center for Disease Control and Prevention, National Center for Health Statistics, table 17. Infant mortality rates (1950–2005); 2010 data from National Vital Statistics Reports 60, no. 4 (January 11, 2012), Deaths: Preliminary Data for 2010, p. 9.

94

Figure 4.1 U.S. Infant Mortality, 1970–2010

Table 4.1

Infant Mortality by Country (Number of deaths of infants under one year old per 1,000 live births, 2011)

1. Japan 2.21 17. S. Korea 4.08 2. Singapore 2.65 18. Slovenia 4.12 3. Sweden 2.74 19. Denmark 4.19 4. Hong Kong 2.90 20. Austria 4.26 5. Iceland 3.18 21. Belgium 4.28 6. Italy 3.36 22. Australia 4.55 7. France 3.37 23. United

Kingdom 4.56

8. Spain 3.37 24. Portugal 4.60 9. Finland 3.40 25. New Zealand 4.72 10. Norway 3.50 26. Cuba 4.83 11. Germany 3.51 27. Canada 4.85 12. Czech Republic

3.70 28. Greece 4.92

13. Netherlands 3.73 29. Taiwan 5.10 14. Ireland 3.81 30. Hungary 5.24 15. Switzerland 4.03 31. United States 5.98 16. Israel 4.07

Source: U.S. Central Intelligence Agency, The World Factbook, https://www.cia.gov/library/publications/the-world-factbook/.

In its most commonly used form, U.S. life expectancy is below that of nearly every European country as well as that of Australia, New Zealand, Canada, and even Jordan and Bosnia and Herzegovina. This low ranking is used to criticize the U.S. health care system in Michael Moore’s film Sicko. As an alternative statistic, economists writing for the American Enterprise Institute have proposed a measure adjusted for the greater

95

number of U.S. accidents, suicides, and homicides—deaths that in their view are not affected by the health care system—so that the remaining deaths are a better measure of U.S. medical care. Because accidents, suicides, and homicides often are early-age deaths, they greatly reduce overall U.S. life expectancy. Using the alternative statistic, the U.S. ranks number one in the world, erasing the nearly four-year advantage held by Japan using the traditional life expectancy statistic, and in the view of these conservative critics, providing evidence for the superiority of U.S. health care. Critics, however, have pointed out that accidents, suicides, and homicide rates are, to a degree, affected by the health care system in preventing death when an injured person is brought to the hospital.

Hispanic/Mexican American Life Expectancy: A Paradox?

Beginning with research by gerontologist Kyriakos S. Markides and anthropologist Jeannie Coreil, demographers have noted a relatively long life expectancy for U.S. Hispanics, especially immigrants born in Mexico. What is surprising is that these groups all share characteristics that typically reduce life expectancy such as lower incomes, less educational attainment, and infrequent insurance coverage. In 2004 researchers Alberto Palloni and Elizabeth Arias wrote Paradox Lost, proposing that a concept known as “return migration” by the unhealthy elderly reduces that measured death rate. According to the authors, the life expectancy measurement for U.S. Hispanics is raised artificially by the absence of immigrants who choose to return to Mexico to die. However, this theory does not account for the mortality advantage for other foreign-born Hispanics or for infants born to Mexican immigrants; in the latter group, a lower mortality rate was very unlikely to be affected by out-migration during the first week of life when most infant deaths occur. Other researchers suggests that perhaps immigrants are a self-selected, more resilient proportion of the home population. In a 2011 study, demographers Laura Blue and Andrew Fenelon proposed what they called a factor “hiding in plain sight”: lower smoking rates by Hispanic immigrants. More research is needed to identify whether the greater longevity of Mexican and other Hispanic Americans is a statistical artifact or the result of one of these or another as-yet-unidentified health factor.

Blacks’Life Expectancy

Life expectancy is not the same for all groups. There is a troublesome gap in length of life between blacks and whites that increased during the 1980s

96

— attributed primarily to high homicide rates—and closed slightly since then, but remained large, particularly for men. Black male life expectancy in 2007 was only 70.0 years, compared to 75.9 years for white males, while black women’s life expectancy was 76.8 years, compared to 80.8 for white females. It should be pointed out that the overall life expectancy data conceal important differences by age. By age 75, differences between whites and blacks are relatively smaller, less than one year for both men and women. The shorter overall black life span at birth is a result of a 2 percent chance of death for blacks before age 20, nearly double the approximately 1 percent rate for whites—a difference that continued even after the decline in homicides since 1993 (see Chapter 6).

The Very Elderly

Life expectancy data may also be unreliable for the very elderly because many individuals do not truthfully report their ages in the U.S. Census. Those under 21 years tend to exaggerate their age, and adults 21 through 70 report their ages relatively accurately, while the largest misreporting occurs for the extremely elderly, who tend to exaggerate their age. Thus there are too many people claiming to be of advanced age in the Census, but fewer who actually die at those ages, causing an exaggerated survival rate. According to researchers at the University of Pennsylvania in 1980, this problem caused significant underestimation of the death rate for nonwhites and misrepresentation of the trend in the cancer death rate for all persons over 85 years of age.

How to Measure Longevity

Mean and Median Life Expectancy

Another problem with longevity statistics is confusion about the meaning of mean and median life expectancy. In 1982, noted paleontologist Stephen Jay Gould was diagnosed with mesothelioma, a serious cancer, for which he learned the median life expectancy is only eight months after discovery. Gould recounts how his knowledge of statistics gave him hope, a sense of optimism he credits with helping him survive until his death in 2002. What Gould knew was that life expectancy for mesothelioma was pulled down by a large number of patients who die soon after the cancer is discovered—that is, one-half are dead within eight months, but an equal number of individuals survive longer, many far longer. Consequently, the mean, or average, life expectancy is much greater than the median. Thus,

97

those who survive beyond eight months need not expect to die momentarily, having already lived longer than most who have the cancer. Some individuals will live for years, increasing the value of the mean, but leaving the median at eight months. Being young and receiving the best treatment, Gould thought he would fall in this longer-lived group, an outlook that itself may have helped him fight the disease.

Box 4.1 Are Americans “Food Insecure”?

The U.S. Department of Agriculture monitors how much Americans have to eat through supplementary questions in the Census Bureau’s Current Population Survey (see Chapter 9). In 2008 and 2009, the department reported over 17 million U.S. households to be “food insecure,” the highest since measurement began in 1995. Newspaper headlines transcribed the bureaucratic-sounding “food insecure” to “17 million hungry Americans,” prompting Heritage Foundation fellow Robert Rector to criticize claims of “widespread hunger” as “far from accurate” because “only one child in 150 will miss at least one meal in a given month due to food shortages.”

Not surprisingly, measuring the degree of hunger is imprecise. The Department of Agriculture survey classifies households as “food insecure” if they report (1) being unable to afford balanced meals, (2) cutting the size of meals, or (3) being hungry because of too little money for food. Although the Heritage Foundation critics were correct in stating that few families experienced this food insecurity at any single point in time, as James Weill, president of the Food Research and Action Center pointed out, the survey was indicative of the growing need for food pantries and government programs such as food stamps for struggling households.

Sources: 2009 data in Mark Nord et al., “Household Food Security in the United States,” ERS Report Summary, U.S. Department of Agriculture, November 2010; Robert Rector in “Top Chef v. Heritage Policy Analysts: Stirring Up the Debate on Federal Welfare,” The Foundry: Conservative Policy News Blog from The Heritage Foundation, http://blog.heritage.org/2010/07/07/top-chef-v-heritage-policy-analyst- stirring-up-the-debate-on-federal-welfare/; James Weill in Jason DeParle, “49 Million Americans Report a Lack of Food,” New York Times, November 17,

98

2009, p. A14.

The difference between the mean and the median applies to life expectancy for the population as a whole. For example, the 2011 U.S. mean life expectancy for men ranked fiftieth in the world at 78.4 years. But because high infant mortality and young male homicides cause so many early deaths, more than half of the population will survive longer than 78.4 years; that is, the median life expectancy is greater than the mean life expectancy. Thus, at age 65, most American men on average were expected to live past 82, putting the U.S. twelfth in the world. For health researchers, mean life expectancy provides a reasonable measure of longevity. For individuals who want to know a typical life span, however, leaving out the possibility of tragic early death, the higher median life expectancy is the most representative statistic.

Box 4.2 The Oldest Old

Although the number of elderly individuals is growing, reports of extreme ages often prove to be false. Tales of yogurt-eating Russian Caucasus villagers living to mid-100 years old lack documentation. Even cases with written records sometimes prove inaccurate, as in the celebrated example of American Charlie Smith, who claimed to be 137 years old based on records of his sale into slavery in 1854. A subsequently discovered marriage certificate showed he was a hearty but much younger 104 when he died. Similarly, another American claiming to be 130 was proved to have used his father’s documents to avoid army service in World War I. For many years the oldest human with a verified age was Madame Jeanne Calment of Arles, France, who died in 1997 at age 122.

Of more interest than the current age-record holder to health researchers is the number of supercentenarians, those living beyond their 110th birthday. Recent studies by U.S. Social Security Administration demographers have attempted to identify such extremely old individuals. Beginning with Medicare records, they corroborated birth years from birth certificates or U.S. Census

99

records, counting over 325 persons who lived to 110 and died in the two decades prior to 2003. Such careful documentation is unusual and provided data on characteristics of supercentenarians: 90 percent were female, 58 percent worked primarily as a homemaker, 12 percent never married, and about one-half completed high school. A disproportionate number were born in the U.S. Southeast, a factor attributed to a greater degree of erroneous records in those states.

Source: May Berstein, “France’s Doyenne of Humanity,” U.S. News & World Report 123, no. 7 (August 18, 1997): 9; Roy L. Walford, Maximum Life Span (New York: Norton, 1983), pp. 12–15. Supercentenarians in Bert Kestenbaum and B. Renee Ferguson, “Supercentenarians in the United States,” in H. Maier et al., Supercentenarians (Berlin, Germany: Springer- Verlag, 2010), pp. 43–58.

Age-Adjusted Rates

Many research projects adjust data to take into account varying ages in comparison groups. For example, in order to compare the effect of pollution on heart disease in two geographic areas, it is important to adjust for age differences that would cause one group’s overall older population to have more heart problems regardless of pollution levels. The most commonly used standardization method, prepared by the U.S. Department of Health and Human Services, was revised based on the 2000 U.S. Census; it replaced an earlier one that used the 1940 U.S. population as a starting point and relied on updates from several government agencies that could not be compared easily.

The new age-adjustment caused abrupt changes in commonly used data, making it extremely important for researchers to pay close attention to their application. For example, the new 2000 Census standard caused New Jersey’s overall death rate to be nearly twice as high when compared with the previous age-adjustments. Of course, death rates had not actually doubled. Instead, the age-adjusted death rates according to the old standard were artificially low because the then current population appeared “unusually” old. The new age-adjusted death rate assumed that the current much older population was “normal” and no longer needed its death rate adjusted downward.

More substantive issues arose when the new standard was used to rank

100

cause of death. In New Jersey the new age adjustment revealed that cancer was the number one cause, replacing heart disease. Pneumonia/influenza jumped from the eighth most common cause to number five. Also the comparison of age-adjusted death rates for whites and blacks grew closer, but remained high at 30 percent higher for blacks instead of the prior 60 percent higher rate. In all these cases, the older standard had caused an unwarranted adjustment for age, thereby misrepresenting a population that actually was not unusual when compared to the new 2000 Census.

Cancer

In 1971 President Richard Nixon signed the National Cancer Act, calling for a war on cancer of the “same kind of concentrated effort that split the atom and took man to the moon.” Initially, hopes were raised that massive funding could find a cure for cancer by the 1976 U.S. bicentennial. Four decades and hundreds of billions of dollars later, cancer is now the number one cause of death, replacing heart disease, for which significant life- extending medical interventions have succeeded. Are we losing the war on cancer? One key statistical point is that cancer is actually a collection of several diseases; prevention and even cures have been made for some types of cancers in recent years, affecting different age groups and in some cases extending life, even if the death rate is unchanged.

Until about 2000, cancer statistics appeared to tell a bleaker story. Take, for example, a 1997 New England Journal of Medicine article titled “Cancer Undefeated,” by epidemiologist John Bailar. By 2011 Bailar had become far more optimistic, concluding that “mortality rates for cancer are going down.” The ups and downs in the war on cancer are in part a matter of statistical measures as well as actual medical progress.

Much of the increase in cancer mortality since 1950 is attributable to lung cancer. Thus, not surprisingly, the overall cancer death rate falls if lung cancer deaths are excluded, a statistic used by some National Cancer Institute (NCI) officials to argue that the war on cancer has been successful. In this view, increased lung cancer deaths (caused mainly by cigarette smoking) could not be prevented by medical research. Bailar pointed out that the same logic could be used to exclude stomach cancers, for which early detection, rather than NCI-sponsored research, reduced the death rate.

Advocates of the effectiveness of the war on cancer also point to improvements in cancer survival rates. Although this statistic receives

101

much publicity, epidemiologists question its validity. Better record keeping in recent years means that more cancer survivors are retained in the data files, thereby pushing up the measured survival rate. In addition, because cancer is typically detected earlier now than it was in the past, cancer patients appear to live longer.

Another statistic used to defend the war on cancer is a decline in the cancer death rate for those under 55 years old. Researcher Richard Doll concluded that “we are, for the most part, winning the fight against cancer” because the mortality rate is low for younger people. He and his collaborator, Richard Peto, in their influential study The Causes of Cancer, downplay the higher cancer rate for the elderly, in part because medical records are unreliable for this group and because the enactment of Medicare increased the number of cancers detected but not necessarily the underlying cancer rate. Critics maintain that those under 55 years old account for less than 25 percent of all cancers. Devra Lee Davis of the U.S. Department of Health and Human Services emphasizes that increased cancers among the elderly may be a harbinger of dangers in the environment; the elderly population’s longer exposure to carcinogens indicate a danger that might later affect the young. Doll, himself renowned for identifying environmental causes of cancer, worries that arguments such as this may detract attention from the fundamental issue—tobacco’s role in causing cancer.

The measured degree of progress against cancer depends in large part on the comparison year. Because overall cancer rates peaked in the early 1990s, about the time that reduced smoking first began to lower lung cancer rates, progress appears significant in the following two decades. Cancer deaths dropped about 22 percent for men and 14 percent for women between 1990 and 2007, equivalent to saving nearly 900,000 lives. However, comparisons between the present and years prior to 1990 show a much smaller decline because of the lower cancer rates reported before 1990. For example, today’s lung cancer rate is higher than it was in 1960 for men and is down only slightly from its 1990 high for women. Moderate progress has been made for other cancers, notably those related to breast, colon, and prostate tumors.

Box 4.3 Likelihood of Breast Cancer

102

Former U.S. Department of Health and Human Services Secretary Margaret M. Heckler asserted that “1 out of every 11 women in this country will develop breast cancer.” Although often cited, this statistic does not actually measure the average likelihood for breast cancer among women. The 1-out-of-every-11 cancer likelihood is derived from the sum of current age-specific rates for ages 1 to 85. Because many women will not live to age 85, and because future cancer incidence rates may change, this figure does not measure the actual likelihood of cancer (although the statistic is useful for epidemiologists in assessing various cancer risks). Physician Richard Love calculates that the actual risk of breast cancer is about 1 in 30 during the fifth and sixth decades of life—the time when breast cancer is most worrisome, and a period of life through which most women can expect to survive.

Sources: Richard R. Love, “The Risk of Breast Cancer in American Women,” Journal of the American Medical Association 257 (March 20, 1987): 1470.

Cancer Incidence

In addition to the cancer death rate, researchers are interested in the prevalence of cancer among the living, or what epidemiologists call the incidence rate. The problem is that this statistic changes not only with the number of people who have cancer but also with changing rates of cancer detection. For example, there was an unprecedented increase in reported breast cancer in 1974, when public disclosure of the disease by the wives of the U.S. president and vice-president prompted more women to have their breasts examined. As a result, more borderline lesions were discovered and reported to the NCI, although the cancer mortality rate was not affected in subsequent years.

Similarly, the male prostate cancer rate varies, depending on the level of examination. Autopsy uncovers prostate cancer in as many as 80 percent of the men over 80 years old. But since the overall lifetime risk of medically diagnosed prostate cancer is only about 10 percent for men over 50, it appears that lesions discovered during autopsy grow so slowly that they do not qualify according to traditional definitions of cancer. Also, prostate cancer detection is susceptible to the level of screening, so the prostate cancer incidence rate rose dramatically between 1989 and 1992,

103

when the PSA blood test was first widely used. Since then, both the incidence and death rates have fallen, and both are attributable to better detection and treatment. Similar measurement problems exist for skin cancer and in situ cervical cancer, which are relatively frequent cancers in need of treatment, but not commonly fatal. The inclusion of such cancers in incidence reports would cause the cancer rate to vary with the detection rate of these common but less severe cancers. Because of these problems, the overall cancer incidence rate often excludes superficial skin cancers. Most researchers prefer to look at the incidence rates for specific cancer sites (lung, breast, and so on), taking into account possible changes in cancer detection rates.

HIV/AIDS

Since the mid-1990s, when antiretroviral therapy helped HIV-positive individuals avoid a nearly predictable progression to AIDS, data on HIV has become more important; AIDS cases alone no longer provide a full picture of the epidemic. Beginning in 2004, all U.S. states have reported HIV diagnoses to the U.S. Centers for Disease Control and Prevention, and 33 states had in place a confidential name-based HIV reporting system for at least five years as of 2011, collecting sufficient data to monitor trends.

Using this data, the CDC is able to estimate the total HIV incidence. By comparing reported AIDS cases (likely to be a relatively full count) with previously reported HIV cases, the CDC calculates how many HIV infections have not been reported. As of July 2010, the CDC estimated that more than one million people in the U.S. carried HIV, of which 21 percent were unaware of their infection. However, the CDC notes that the data are subject to possible error because several “high morbidity” areas (namely, California, Illinois, Maryland, and the District of Columbia) were not included in the original data. The CDC estimates that about one-half of HIV infections occur in men who have sex with other men. Also disproportionately burdened with HIV infection are African Americans, who make up almost one-half of all new HIV infections, yet account for only about 12 percent of the U.S. population.

In 2008 about 16,600 Americans with AIDS died, adding to the more than one-half million who have died since the epidemic began. The occurrence of new HIV infections appears to be stable at about 50,000 new cases each year. However, the CDC suggests caution in interpreting this trend because more recent estimates include states not reporting earlier,

104

and, even more important, the data have been refined to remove duplicate cases. As a result, the earlier estimate for 2003 was reduced by about 12 percent, indicating that HIV infection increased in the following years and was not as stable as the unadjusted data would have indicated.

Worldwide, the numbers of HIV infections and AIDS cases are even more uncertain. As in the United States, HIV incidence can only be estimated because many cases go unreported. Particularly in resource-poor countries, these numbers are far less accurate than in the United States. Often HIV prevalence is estimated from health clinic surveys of women soon after childbirth, a group that studies use as a reasonable indicator of HIV infection for the entire population, once adjustments are made for differences in male versus female infection rates and underreporting from rural areas where women may not be served by clinics.

Data on deaths from AIDS are similarly less accurate in resource-poor countries. In part because people living with AIDS are often stigmatized, the World Health Organization estimates that up to 90 percent of AIDS cases go unreported in some countries. With these limitations in mind, researchers surmise that in 2009 about 2.6 million people were infected with HIV, over two-thirds of that number in sub-Saharan Africa. About 1.8 million individuals died from AIDS-related causes in 2009—that’s about a 19 percent reduction since 2004, primarily because of the expanded availability of antiretroviral therapy.

Are Americans Getting Fatter?

The 2007–2008 National Health and Nutrition Examination Survey found that the number of obese Americans increased to about 34 percent of the adult population, more than double the percentage in 1980. Maps featured at the U.S. Centers for Disease Control website show that in 1985, no state had more than 15 percent of its population in the obese range, whereas all states had rates over 20 percent by 2010, and 12 states had more than 30 percent of its adult population in the obese range.

However, as Rockefeller University obesity researcher Jeffrey Friedman points out, the apparent doubling of obesity occurred because individuals moved just over the threshold, gaining only a few pounds but sufficient to push many into the “obese” category, defined as a body mass index of 30 or higher. Moreover, the average gain, about one pound per year, obscured a very unequal distribution, with the lowest weight group adding no pounds over time while the very obese had substantial weight gains, more

105

than two pounds per year.

At issue in the debate about obesity is its impact on public health. Even though it has proved surprisingly difficult to pin down the precise number of deaths attributable to obesity, a long list of diseases have been tied to weight gain, including coronary heart disease, diabetes, cancer, hypertension, stroke, and osteoarthritis. In reviewing the debate, researchers should be aware that financial interests may affect data analysis, including funding of the American Obesity Association by weight-loss businesses as well as influence by the U.S. food industry in studies of changes in the typical American diet.

Box 4.4 How Much Did You Eat?

Nutrition researchers recognize that survey respondents typically underestimate their food intake. Validation studies suggest as much as a 33 percent undercount in the most commonly cited Women’s Health Initiative and the Nurses’ Study food consumption surveys. Food writer Michael Pollan suggests even greater error based on the nearly 100 percent difference between food produced for each American and the amount reported as eaten, a disparity partially due to waste but, according to Pollan, largely a result of respondent ignorance or outright misrepresentation. Pollan points out that diet surveys ask “impossible to answer” questions, requiring respondents to remember how much of a particular food they ate in the last three months and to know the oil used in the case of frying even though about half of food dollars are spent outside the home. Pollan observes, “I’m not sure Marcel Proust himself could recall his dietary intake over the last ninety days with the sort of precision demanded by the FFQ [Food Frequency Questionnaire].” Researchers attempt to correct for such shortcomings, adjusting reports upward to account for underestimated consumption and using 24-hour recall surveys to gather information in a more timely fashion. Epidemiologist Gladys Block, developer of the Women’s Health Initiative, advises researchers that FFQ data are best for ranking relative consumption, not for absolute measurement of caloric or nutritional intake.

Sources: Michael Pollan, In Defense of Food: An Eater’s Manifesto (New York: Penguin Press, 2008), pp. 74–78.

106

Chicago Heat Wave: What Caused the Tragedy?

During an extraordinary 1995 Chicago heat wave, 739 deaths were blamed on the high temperatures. Although poor and elderly individuals living alone were the most susceptible, mortality was uneven—more common in some neighborhoods than other equally disadvantaged neighborhoods. In his 2002 book, Heat Wave: A Social Autopsy of Disaster in Chicago, sociologist Eric Klinenberg maintains that degraded conditions in some neighborhoods were responsible for the heat wave deaths, contradicting explanations from the Chicago mayor’s office and newspaper accounts that blamed drug abuse or failure of friends and relatives to provide assistance.

Klinenberg’s conclusions have been challenged by other researchers. Writing in the American Sociological Review, Mitchell Duneier suggests that the ecological fallacy (see also Chapter 6) caused Klinenberg to misinterpret Chicago’s deaths. In Duneier’s view, group level correlation between neighborhood characteristics and mortality did not mean that on the individual level these neighborhood factors caused heat-related deaths. Based on interviews, Duneier found that individual characteristics such as drug abuse were more likely responsible. In a bitter response, Klinenberg faulted Duneier’s data, pointing out that it was collected nine years later based on interviews with current residents, some of whom had faulty or reconstructed memories of the tragedy that may have exaggerated the victims’ self-destructive behavior. Using his own data from medical examiner autopsy reports, Klinenberg maintained that drug or alcohol abuse was not a significant cause for excess deaths.

Researchers need to beware of the ecological fallacy and not draw unwarranted conclusions from grouped data. In U.S. voting statistics, for example (see Chapter 11), the apparent correlation between a state’s average income level and the state’s vote for the Democratic Party misrepresents the correlation between individual income and voting behavior. In this case, richer states tend more often to vote Democratic even though in all states rich individuals are more likely to vote Republican.

However, grouped data are not necessarily wrong. If Klinenberg’s data on Chicago’s heat wave death are correct, then both group and individual data correlate neighborhood characteristics with heat deaths. In this case,

107

the policy implications of such a conclusion would be that the deaths in certain neighborhoods could have been avoided, even for poor and elderly residents, if residents had felt safe enough to go outside and seek assistance rather than remaining locked in un-air-conditioned apartments.

What’s Unsafe on the Road? Speed, Texting, Teens, Motorcycles, or Alcohol?

The good news is that the total number of traffic fatalities declined significantly in recent years, down to 33,963 deaths in 2009, a reduction of 22 percent in only five years and the lowest total since the 1950s. Adjusted for the far greater number of miles driven, the decline is even more dramatic, falling by about 80 percent since 1950. However, researchers debate the reasons behind this favorable trend as well as the factors that continue to contribute to unsafe driving.

An unintended consequence of a 1974 government regulation to conserve fuel by lowering the nationwide speed limit to 55 miles per hour was an estimated annual reduction of 2,000 to 4,000 highway fatalities. When Congress agreed to raise the speed limit on rural interstate highways to 65 miles per hour in 1987, some safety experts feared that there would be an increase in fatalities. Specifically, the National Research Council estimated that the higher speed limit would result in an increase of 500 highway deaths per year. But the effect of this change has been difficult to measure, leading to considerable controversy about the relationship between speed limits and highway safety. The problem is that highway fatality data have not followed an easy-to-understand pattern. Data collected by the U.S. Transportation Department show that highway fatalities continued to drop long after the lower speed limit was introduced, even though average speeds increased well above the legal 55-miles-per- hour limit. One intriguing possibility put forward by economist Charles A. Lave points to variance in speeds as the critical factor. In other words, different speeds on the same highway may contribute to serious accidents.

A more recent debate targeted texting while driving as a cause of fatal accidents. Overall the U.S. Transportation Department estimated distracted driving caused 16 percent of all traffic deaths in 2008, an increase from 10 percent in 2005. However, the estimate is linked to accident reports in which drivers have an incentive not to admit behavior (such as texting while driving) that may have contributed to the event and about which police officers usually have no independent information. A study in the

108

American Journal of Public Health correlating cell phone use with traffic fatalities ascribed over 2,000 additional annual traffic deaths to increased texting. However, without better data, experts disagree whether hands-free phone use is any better, or whether other distractions such as eating, drinking, or holding intense conversations are equally likely to cause accidents.

Another growing concern in fatal accidents is the role of teenage drivers, leading many states to impose additional restrictions on their driving rights. Overall, the statistics are quite startling for teenagers. Traffic fatalities are the leading cause of teenage deaths, and the annual fatality rate per teenage driver is more than double the rate for adults—and more than triple if measured per mile driven.

Emphasis in the media on speed limits, texting while driving, and teenage drivers misses two trends that contribute more significantly to fatalities. First, while overall deaths are down, motorcycle rider deaths increased to more than 5,000 per year in 2008. Traffic researchers blame the increase on the greater number of inexperienced middle-aged male cycle drivers. The second major cause of accidental death is alcohol, implicated in 32 percent of traffic fatalities. In this case the data are relatively complete because of laws in most states that allow testing for blood alcohol.

The lesson for researchers is that an overall decline, in this case fewer traffic fatalities resulting from a number of factors including vehicle safety technology, can mask other trends. It is important to analyze traffic data by driver characteristics, vehicle type, and, when better data are available, by driver behavior.

Drug Use

The Office of National Drug Control Policy (ONDCP), the current coordinator of a “war on drugs” first declared by President Nixon in 1971, has been assigned ambitious goals to “stop drug use before it starts.” However, as analyzed by political scientists Renee Scherlen and Matthew B. Robinson, the ONDCP’s reports are textbook cases in the mishandling of data. For example, a “substantial decline” in marijuana use could be supported only by comparing recent years to the particularly high level measured in the 1979 National Household Survey on Drug Use data. In recent years, usage has settled down to a relatively narrow range. On the other hand, when drug use by high school students measured in the

109

Monitoring the Future survey showed a long-term increase, the ONCDP focused instead on a shorter time frame, noting that there were years “without significant changes,” even though illicit drug use was up by more than 60 percent in the previous decade.

Accurate understanding of drug use statistics requires that researchers look behind politically motivated summary statements made by government officials. Survey data are reported in government documents, even those featuring misleading summaries. Because overall illicit drug use has been remarkably stable since the 1970s, it is difficult to put a successful spin on the war on drugs. Nonetheless, there have been important changes in the types of drugs used, in particular a large decline in cocaine use after 1990.

Benefit-Cost Analysis

Is government regulation necessary to protect the nation’s health? In answering this question, policy makers often use benefit-cost analysis, a method developed by economists to measure the relative advantages (benefits) and disadvantages (costs) of a proposed program. Although seemingly straightforward, benefit-cost analysis is controversial because of problems in putting a dollar value on all benefits and costs.

Value of Human Life

Full accounting within the benefit-cost framework requires putting a dollar value on human lives. The idea often shocks the general public, but valuing human life is a well-established practice among economists, who maintain that tradeoffs between money and human life are implicit in everyday life. In this view, whenever we buy a less than perfectly safe car, accept a dangerous job, or even cross the street, we trade a risk of death— albeit slight—in return for measurable financial gain. The proportion between these risks and the gain (see Box 4.6) is called the willingness to pay by an individual for his or her own life.

Box 4.5 Did Depenalization Increase Drug Use?

The 1976 decision by the Netherlands to depenalize possession of small amounts of marijuana offered an important case study for

110

researchers on drug policies. However, popular accounts of marijuana use since this change often have been misinterpreted. As summarized in a review by Peter Mac-Coun of his own research, lifetime marijuana use by Dutch teens had either “fallen from 15 to 2 percent” or “risen from 5 to 14 percent.” The most common story was far greater drug use among Dutch teenagers than U.S. teens, an apparent indictment of the depenalization decision. However, Mac-Coun and his coresearcher, Peter Reuter, conclude that usage rates actually differed between the Netherlands and the United States by no more than 1 percent, well within the sampling error. The misinterpretations arose in part because journalists had made a number of errors. For example, some reports compared Dutch usage in 1996 with U.S. data for 1992, just before a significant increase in U.S. drug use by 1996. Also, journalists compared the lower U.S. usage by teens of all ages with drug use by Dutch 18 year-olds, not surprisingly higher because this older Dutch group had more years to try drugs and were at an age of greater experimentation.

Sources: Robert J. MacCoun, “American Distortion of Dutch Drug Statistics,” Society 38, no. 3 (March/April 2001): 23–26; Robert MacCoun and Peter Reuter, “Interpreting Dutch Cannabis Policy: Reasoning by Analogy in the Legalization Debate,” Science 278 (October 3, 1997): 47–52.

A major problem with the willingness-to-pay technique is that it yields widely varying estimates for the value of life. Risky jobs such as mining and elephant-keeping imply values as low as $300,000 per life, whereas surveys about the willingness to pay for environmental safeguards measure values at more than $8 million per life. As an alternative method, some researchers advocate the use of estimates of lost earnings. The advantage to this method is that it yields relatively consistent estimates, but it suffers from the shortcoming of severe inequity. For example, in one study, an 85- year-old black woman was “valued” at $128, clearly an unacceptable measure. The variation in values for human lives makes it difficult to derive policy advice from benefit-cost analysis. Projects judged wasteful with one measure may be quite worthwhile if human lives are more highly valued. For example, in opposing a safety standard for construction workers who handle concrete, the U.S. Office of Management and Budget advocated valuing these workers at $1 million per life rather than the $3.5

111

million proposed by the Occupational Safety and Health Administration.

One technique that avoids the problems of valuing human life is cost effectiveness. For example, in evaluating different methods of reducing infant mortality, researchers measure the cost per infant life and then compare which technique saves the most lives without putting a dollar value on the lives saved. In a similar manner, consumer advocate Ralph Nader’s associate, Mark Green, supports the idea of separate accounts for human lives and dollar benefits. By this method, Green calculates that many terminated health and safety programs had been value effective in terms of decreased medical care, lost work time, and other benefits measurable in money—even without placing a dollar value on human life.

Risk Assessment

A second controversial technique in benefit-cost analysis is the measurement of the probability of unlikely events, a calculation called risk assessment. Some policymakers argue that these risks are far lower than is commonly assumed. In this view, the public consistently overestimates the risk associated with nuclear power plants, acid rain, and other environmental hazards targeted by proregulation activists. For example, in a series of advertisements, the Mobil Oil Corporation observed that most people accept the potential risks of lawn mowers, vacuum cleaners, bathtubs, stairs, and other everyday necessities that are responsible for over a million yearly accidents because the benefits of their use outweigh the risks. Mobil would like the same logic applied to the risks of oil exploration compared to the benefits derived from the use of oil.

Critics of risk assessment agree that why people worry about certain risks more than others is an interesting puzzle; still, such apparent confusion about risk may not be irrational. Sociologist William R. Freudenberg points out that the experts are often wrong, as demonstrated by the explosion of the space shuttle Challenger. NASA rejected its own studies indicating a failure rate as frequent as 1 in 70 space shuttle boosters, favoring instead a study measuring only a 1 in 100,000 risk. The shuttle exploded on its twenty-fifth launch. The “uninformed” consensus of the public, evaluating the risk of new technologies such as nuclear power and off-shore oil drilling, may be more accurate than expert estimates, which traditionally have proved too optimistic.

Measurement problems for the value of human life and risk assessment seriously undermine the usefulness of benefit-cost accounting for

112

policymaking purposes. As illustrated by the space shuttle tragedy, erroneous assumptions can produce seemingly precise but wildly erroneous results. Similarly, different values for human life can cause completely different evaluations of public policies. At a minimum, it is good research practice to use “sensitivity analysis” for all assumptions— that is, the use of different values for human life and risk assessment to see how such substitution alters the results. By demonstrating that results do not depend on arbitrary assumptions in any of these controversial measures, researchers can broaden the influence of their findings beyond the already convinced.

Box 4.6 Cost of Tamper-Proof Closures

Mathematically, the value of human life is calculated by dividing society’s willingness to pay for a reduction in the risk of mortality by the likelihood of death. For example, Paul MacAvoy, a member of President Reagan’s Council of Economic Advisers, calculated that tamper-proof closures, introduced after the 1982 Tylenol poisonings, implied the value of human life was equal to $2 million: $.02 (the cost of the closures) divided by 1 in 100,000,000 (the chance of a single bottle containing poison based on historical record). MacAvoy opposed the tamper-proof closures on the grounds that $2 million per human life was greater than the “actual” value of a human life measured in willingness-to-pay studies. Neither drug companies nor government regulators followed MacAvoy’s advice.

Sources: Paul W. MacAvoy, “FDA Regulation—At What Price?” New York Times, November 21, 1982, p. III–3.

Alternatively, researchers should consider whether benefit-cost analysis and risk assessment are the most appropriate techniques to use. As described earlier, cost-effectiveness measurement may provide useful policy advice that does not encumber the analysis with unwarranted assumptions about the value of life. In the case of risk assessment, policy adviser Langdon Winner warns that traditional technique biases the analysis in favor of the status quo by imposing an unnecessary burden of

113

1.

proof on those who would like to see alternatives to unsafe technology. For instance, in evaluating genetic engineering, Winner believes that risk assessment is nearly impossible and that attempts to make this tricky estimate have caused policymakers to sidestep the more critical moral debate about direct control of human evolution. In such instances, discussion of a wider scope may provide better guidance, even if it does not appear as mathematically precise as the calculations in benefit-cost analysis.

Summary

Each of the controversies reviewed in this chapter is fueled by statistical ambiguity. There is uncertainty about the trend in infant mortality and the reason for its remarkably high level in the United States. There is disagreement about the best way to measure life expectancy as well as uncertainty about data showing longevity for some immigrant groups. There are conflicting interpretations of the effect of the 55-miles-per-hour speed limit and a bitter debate about data on the cause of death in the Chicago heat wave. Even attempts to state explicitly the costs and benefits of government policies are controversial because of assumptions about the value of life and risk assessment.

Nevertheless, especially in the area of health policy, there is little opportunity for social scientists to take a wait-and-see attitude. Decisions must be made about the allocation of health care resources, even if the statistics or decision-making techniques are less than perfect. At the very least, however, these decisions must be made with a full understanding of the data’s limitations. For example, in the debate about the war on cancer, critics have been able to warn us that some statistics overstate the success of existing health policies. Similarly, careful analysis of life expectancy statistics has provided insights about the uncertain effect of modern medicine and potential shortcomings in health care for black Americans. The analysis of benefit-cost analysis and risk assessment shows that researchers must look critically at official studies that “prove” the efficacy or failure of government regulations. Overall, these examples prove the usefulness of health statistics--but only with careful attention to their limitations.

Case Study Questions

In the U.S. Department of Agriculture food security survey, 11

114

2.

3.

4.

5.

percent of households reported difficulty at some time during the year providing enough food for all members due to a lack of money and other resources. Of these, most obtained enough food to avoid hunger by eating less varied diets, participating in federal food programs, or getting emergency food from community programs. About 3.5 percent of households had members who were hungry at some time during the year. One- half of 1 percent of children were hungry, while 3.8 percent of households had adult members who were hungry. How could these data be used to support the position that the United States has a hunger problem? How could the data be used to support the position that the U.S hunger problem is insignificant?

In recent years, more teenagers died in traffic accidents than by any other cause, accounting for about one-third of teenage deaths. How might this statistic be misleading about the safety of teenage drivers?

The state of Utah has an extremely low death rate from heart disease, a benefit sometimes ascribed to Mormon abstinence from tea, coffee, alcohol, and cigarettes. But, based on age- adjusted death rates—that is, taking into account the relatively young age of Utah’s overall population—Utah’s heart disease death rate is higher than its neighboring mountain states. On the other hand, Utah’s liver cirrhosis and lung cancer rates are lower, even in the age-adjusted data. What implications can you infer about the benefits of lifestyles prevalent in Utah?

The Federal Aviation Administration (FAA) used benefit-cost analysis to analyze the need for safety seats on airlines for children under 2 years old. Because young children usually ride in their parents’ laps, the primary cost of providing such safety is the additional regular airline seats, estimated to cost $56 million per year. On average, five infants die per year in U.S. commercial airline accidents. The FAA valued infant lives at $500,000 each. Based on benefit-cost analysis, what did the FAA decide? How would you evaluate this decision-making method?

In effort to gather more accurate statistics on abortion, researchers have worked to increase the number of women willing to participate in surveys that discuss this highly personal

115

issue. However, efforts to include initially recalcitrant women respondents led to less accurate estimates for the number of women who have had an abortion. Why do you think this occurred?

References

Data Sources (60)

National Center for Health Statistics (NCHS) at http://www.cdc.gov/nchs/. On Current Population Survey, see Chapter 9. Information on the National Survey on Drug Use and Health is available at http://www.oas.samhsa.gov/nhsda.htm. See Monitoring the Future (MTF) survey at http://monitoringthefuture.org. Hookah use data sample in Lloyd D. Johnston et al., Monitoring the Future: National Results on Adolescent Drug Use— Overview of Key Findings, 2010 (Ann Arbor: University of Michigan Institute for Social Research, 2011), p. 40. El Salvador data sample in World Health Organization, Global Health Observatory Data Repository, http://apps.who.int/ghodata/?theme=country.

Controversies (62)

Infant Mortality (62)

As a health indicator in C. Arden Miller, “Infant Mortality in the U.S.,” Scientific American 253 (July 1985): 31–37; Soviet Union in “Getting Russia Well Again,” The Economist, November 21, 1987, pp. 51–52; historical data in Miller, “Infant Mortality in the U.S.,” pp. 31–32; international rankings in Marian F. MacDorman and T.J. Mathews, “Behind International Rankings of Infant Mortality,” NCHS Data Brief, no. 23 (Hyattsville, MD: National Center for Health Statistics, November 2009). Bernadine Healy, “Behind the Baby Count,” U.S. News and World Report, September 24, 2006, http://health.usnews.com/usnews/health/articles/060924/2healy.htm. Problems in measurement in Nicholas Eberstadt, “America’s Infant-Mortality Puzzle,” Public Interest, no. 105 (Fall 1991): 30–47.

Abortion (63)

Centers for Disease Control and Prevention (CDC) data in the Guttmacher Institute’s “Abortion Reporting Requirements,” State Policies in Brief, February 1, 2011; see also 2008 data in the Institute’s “Facts on Induced Abortion in the United States,” In Brief (fact sheet), January 2011, www.guttmacher.org/pubs/fb_induced_abortion.html.

116

Are We Living Longer? (63)

Life expectancy in “Life Expectancy Remains at Record Level,” Statistical Bulletin 70, no. 3 (July-September 1989): 26–30. Comparison with other countries in U.S. Central Intelligence Agency, World Factbook, https://www.cia.gov/library/publications/the-world- factbook/rankorder/2102rank.html. Sicko at http://sickothemovie.com. American Enterprise Institute (AEI) in Robert L. Ohsfeldt and John E. Schneider, The Business of Health (Washington, DC: AEI Press, 2006). Criticism in Carl Bialik, “Does the U.S. Lead in Life Expectancy?” Wall Street Journal, November 12, 2007, http://blogs.wsj.com/numbersguy/does-the-us-lead-in-life-expectancy-223/.

Hispanic/Mexican American Life Expectancy: A Paradox? (65): S. Markides Kyriakos and J. Coreil, “The Health of Hispanics in the Southwestern United States: An Epidemiologic Paradox,” Public Health Reports 101 (1986): 253; Alberto Palloni and Elizabeth Arias, “Paradox Lost: Explaining the Hispanic Adult Mortality Advantage,” Demography 41, no. 3 (2004): 385–415; Robert Hummer et al., “Paradox Found (Again): Infant Mortality Among the Mexican-Origin Population in the United States,” Demography, 44, no. 3 (August 2007): 441–57; Laura Blue and Andrew Fenelon, “Explaining Low Mortality Among U.S. Immigrants Relative to Native-Born Americans: The Role of Smoking,” International Journal of Epidemiology, February 15, 2011, http://ije.oxfordjournals.org/content/40/3/786; and Laura Blue, “The Ethnic Advantage,” Scientific American, October 2011, p. 30.

Blacks’ Life Expectancy (66): National Vital Statistics Reports, “Table 22: Life Expectancy at Birth, at 65 Years of Age, and at 75 Years of Age, by Race and Sex,” www.cdc.gov/nchs/data/hus/2010/022.pdf; “Differences by Race,” Child Trends Data Bank, www.childtrendsdatabank.org.

The Very Elderly (66): Unreliable rates in Ira Rosenwaike, Nurit Yaffe, and Philip C. Sagi, “The Recent Decline in Mortality of the Extreme Aged,” American Journal of Public Health 70 (October 1980): 1074–80. Who misrepresents their age, in Kenneth C.W. Kammeyer and Helen L. Ginn, An Introduction to Population (Chicago, IL: Dorsey Press, 1986), p. 71.

How to Measure Longevity (66)

Mean and Median Life Expectancy (66): Stephen Jay Gould, “The Median Isn’t the Message,” Discover, June 1985, pp. 40–42; recent data in U.S. Central Intelligence Agency, World Factbook, www.cia.gov/library/publications/the-world-fact- book/rankorder/2102rank.html.

Age-Adjusted Rates (69): U.S. Department of Health and Human Services adjustment in “Notice to Readers: New Population Standard for Age-Adjusting Death Rates,” MMWR Weekly 48, no. 69 (February 19, 1999): 126–27. New Jersey

117

in Rose Marie Martin, “Age Standardization of Death Rates in New Jersey: Implications of a Change in the Standard Population,” New Jersey Department of Health and Senior Services, Center for Health Statistics, www.state.nj.us/health/chs/topicaa.htm.

Cancer (69)

“Same kind of concentrated effort” in Ralph W. Moss, The Cancer Syndrome (New York: Grove Press, 1980), p. 16; see also Richard M. Nixon, “Acting Against Cancer,” Saturday Evening Post, July/August 1986, pp. 67–69. J.C. Bailar and Heather L. Gornik, “Cancer Undefeated,” New England Journal of Medicine 336 (1997):1560–1574; J.C. Bailar III and Elaine M. Smith, “Progress Against Cancer?” New England Journal of Medicine 314, no. 19 (May 8, 1986): 1226; R. Doll in Robert N. Proctor, Cancer Wars (New York: Basic Books, 1996), p. 254; see also D. Davis and R. Proctor in Cancer Wars, p. 268; John Bailar in Carl Bialik, “Bleak Cancer Reports Mask Major Advances,” Wall Street Journal, September 2, 2009, http://online.wsj.com/article/SB125185000130377889.html. R. Doll and R. Peto, “The Causes of Cancer: Quantitative Estimates of Avoidable Risks of Cancer in the United States Today,” Journal of the National Cancer Institute 66, no. 6 (1981). Reduction in deaths in Stacy Simon, “Annual Report: U.S. Cancer Death Rates Decline, but Disparities Remain,” American Cancer Society, June 17, 2011, http://www.cancer.org/Cancer/news/News/annualreport- u.s-cancer-death-rates-decline-but-disparities-remain.

Cancer Incidence (71): Cancer incidence in Bailar and Smith, “Progress Against Cancer?” pp. 1229–30. Prostate cancer in Richard Martin, “Commentary: Prostate Cancer Is Omnipresent, but Should We Screen for It?” International Journal of Epidemiology, 36 (2007): 278–81.

Overall rate in American Cancer Society, Cancer Facts& Figures, 2012, http://www.cancer.org/Research/CancerFactsFigures/CancerFactsFigures/cancer- facts-figures-2012.

HIV/AIDS (72)

U.S. reporting in CDC, “HIV/AIDS Statistics and Surveillance,” http://www.cdc.gov/hiv/topics/surveillance/index.htm; U.S. data in CDC, “HIV in the United States at a Glance,” http://www.cdc.gov/hiv/resources/factsheets/PDF/HIV_at_a_glance.pdf. Worldwide AIDS deaths in AVERT, “Number of People Infected During 2009, and the Number of Deaths,” http://www.avert.org/worlstatinfo.htm. HIV prevalence estimates in AVERT, “Understanding HIV and AIDS Statistics,” http://www.avert.org/statistics.htm. Problems with AIDS reporting in World Health Organization, Global Alert and Response (GAR), “WHO Report on Global Surveillance of Epidemic-Prone Infectious Diseases—(HIV/ AIDS),”

118

http://www.who.int/csr/resources/CSR_ISR_2000_1hiv/en/index.html.

Are Americans Getting Fatter? (73)

National Health and Nutrition Examination Survey (NHANES) in Cynthia L. Ogden and Margaret D. Carroll, “Prevalence of Overweight, Obesity and Extreme Obesity Among Adults,” June 2010, http://www.cdc.gov/nchs/data/hestat/obesity_adult_07_08/obesity_adult_07_08.pdf. Maps in J. Eric Oliver, “The Politics of Pathology: How Obesity Became an Epidemic Disease,” Perspectives in Biology and Medicine 49, no. 4 (Autumn 2006): 611–27. Jeffrey Friedman in Gina Kolata, “The Fat Epidemic: He Says It’s an Illusion,” New York Times, June 8, 2004, p. F5. Funding of the American Obesity Association in Oliver, “The Politics of Pathology.”

Chicago Heat Wave: What Caused the Tragedy? (74)

Eric Klinenberg, Heat Wave: A Social Autopsy of Disaster in Chicago (Chicago: University of Chicago Press, 2003); Christopher Browning et al., “Neighborhood Social Processes, Physical Conditions, and Disaster-Related Mortality: The Case of the 1995 Chicago Heat Wave,” American Sociological Review 71, no. 4 (August 2006): 661–78; Mitchell Duneier, “Ethnography, the Ecological Fallacy, and the 1995 Chicago Heat Wave,” American Sociological Review 71, no. 4 (August 2006): 679–88; Eric Klinenberg, “Blaming the Victims: Hearsay, Labeling, and the Hazards of Quick-Hit Disaster Ethnography,” American Sociological Review 71, no. 4 (August 2006): 689–98.

What’s Unsafe on the Road? Speed, Texting, Teens, Motorcycles or Alcohol? (75)

Reduction in traffic deaths in Joseph B. White, “Deaths in Crashes Decline Amid Gains in Car Safety,” Wall Street Journal, September 20, 2010, http://online.wsj.com/article/SB10001424052748704644404575481603221075946.html. Drop in 1974 fatalities in Transportation Research Board, Managing Speed: Review of Current Practices for Setting and Enforcing Speed Limits National Academy Press, Special Report 254, p. 278, http://www.nap.edu/catalog.php? record_id=11387. Predicted increase in fatalities in “Does Speed Kill?” Newsweek, July 21, 1986, p. 16. Increase in fatalities for 1987, “65 MPH Not Costing Lives,” New York Times, May 4, 1988, p. A18; critics’ answer in “U.S. Issues Fatality Data on 65 MPH,” New York Times, May 7, 1988, p. 36. Varying speeds in “Official Report Says,” p. 29; Charles A. Lave, “Speeding, Coordination, and the 55 MPH Limit,” American Economic Review 75, no. 5 (December 1985): 1159–64. Comments by Peter Asch, David T. Levy, Richard Fowles, Peter D. Loeb, Donald W. Snyder, and Charles A. Lave in American Economic Review 79, no. 4 (September 1989): 913–31. Distracted driving in Fernando Wilson, “Trends in

119

Fatalities from Distracted Driving in the United States, 1999 to 2008,” American Journal of Public Health 100, no. 11 (November 2010): 2213–19. Teen drivers in Carl Bialik “The Dangers of Teen Driving,” Wall Street Journal, September 26, 2008, http://blogs.wsj.com/numbersguy/the-dangers-of-teen-driving-418/. Alcohol and motorcycle accidents in Joseph B. White, “New Puzzle: Why Fewer Are Killed in Car Crashes,” Wall Street Journal, December 15, 2010, http://online.wsj.com/article/SB10001424052748703734204576019602118693930.html.

Drug Use (77)

“Stop drug use,” at Office of National Drug Control Policy http://www.whitehouse.gov/ondcp; Matthew B. Robinson and Renee Scherlen, Lies, Damned Lies, and Drug War Statistics: A Critical Analysis of Claims Made by the Office of National Drug Control Policy (Albany: State University of New York), 2007; 1979 National Household Survey on Drug Use at http://www.icpsr.umich.edu/icpsrweb/SAMHDA/studies/6843.

Benefit-Cost Analysis (77)

For description of benefit-cost analysis, see James T. Campen, Benefit, Cost and Beyond (Cambridge, MA: Ballinger, 1986), or any of the numerous textbooks on the topic such as Edward M. Gramlich, Benefit-Cost Analysis of Government Programs (Englewood Cliffs, NJ: Prentice Hall, 1981).

Value of Human Life (77): Criticisms in Mark Green and Norman Waitzman, “Cost, Benefit, and Class,” Working Papers for a New Society, May/June 1980, pp. 39–51. Dollar estimate of $300,000 to $8 million in “What Is the Audited Value of Life?” New York Times, October 26, 1984, p. 24; and $128 in Green and Waitzman, “Cost, Benefit, and Class,” p. 43. Also, $400,000 and construction workers in “What Is the Audited Value of Life?” and “Cost, Benefit, and Class,” pp. 48–49. Advocates of risk assessment in Mary Douglas and Aaron Wildavsky, Risk and Culture: The Selection of Technical and Environmental Dangers (Berkeley: University of California Press, 1982); see also Council of Economic Advisers, Economic Report of the President, 1987 (Washington, DC: U.S. Government Printing Office, 1987), pp. 179–207; Henry Fairlie, “Fear of Living,” New Republic, January 23, 1989, pp. 16–18.

Risk Assessment (79): Criticism in William R. Freudenberg, “Perceived Risk, Real Risk: Social Science and the Art of Probabilistic Risk Assessment,” Science 242 (October 7, 1988): 44–49; and Langdon Winner, “On Not Hitting the Tar- Baby: Risk Assessment and Conservatism,” in Mary Gibson, To Breathe Freely (Totowa, NJ: Rowman and Allanheld, 1985), pp. 269–84. Lawn mowers versus nuclear power, ibid., p. 277. Chernobyl in “Life’s Risks: Balancing Fear Against Reality of Statistics,” New York Times, May 8, 1989, p. A1; “Genuinely puzzling” in Winner, “On Not Hitting the Tar-Baby,” p. 275. William R. Freudenberg in

120

“Perceived Risk,” p. 47. Summary of the problem in Benefit, Cost and Beyond, pp. 52–55, and Winner in “On Not Hitting the Tar-Baby,” pp. 280–82.

Case Study Questions (81)

1. See Mark Nord, “Measuring U.S. Household Food Insecurity,” Amber Waves, 3, no. 2 (April 2005): 10–11.

2. U.S. Centers for Disease Control and Prevention, “Teen Drivers Fact Sheet,” http://www.cdc.gov/Motorvehiclesafety/Teen_Drivers/teendrivers_factsheet.html.

3. Utah in Myron Johnston, “Young and Alive,” American Demographics, December 1986, p. 7.

4. “Babies’ Seats Are Air Safety Issue,” New York Times, August 18, 1985, sec. 4, p. 23; and “Tighter Safety Rules Planned for Young Children in Planes,” New York Times, November 4, 1989, p. A7.

5. Andy Peytchev et al., “Measurement Error, Unit Nonresponse, and Self- Reports of Abortion Experiences,” Public Opinion Quarterly, 74, no. 2 (Summer 2010): 319.

121

5

Education □□□□

Education is a critical issue with global ramifications. In order to develop sound education policies in the United States, we need a wide variety of data about schools, students, and educational outcomes. over the last decade, a stronger focus on school and district accountability has led to improvements in the quality and amount of educational data that is collected. However, in comparison to the extensive statistical programs for other social science disciplines, education data are still far from ideal, particularly in regard to timeliness of release. This shortcoming is the first controversy discussed here.

In addition, this chapter summarizes debates about such educational issues as high school and college dropout rates, charter schools, teacher effectiveness, international rankings, and standardized test scores. Although newspapers and magazines cover these issues extensively, there has been little assessment of the weaknesses and strengths in the underlying education data used by the media.

Where the Numbers Come From

Organizations Data sources URL National Center for Education Statistics, Institute of Education Sciences, U.S. Department of Education

Common Core of Data, National Assessment of Educational Progress, Integrated Postsecondary Education Data System; surveys of teachers, students, schools, colleges, recent graduates

www.nces.ed.gov

Bureau of the Census, U.S. Current Population Survey www.census.gov

122

Department of Commerce Education Week Quality Counts state

reports www.edweek.org

National Education Association

Compiles state and federal data

www.nea.org

Data Sources

National Center for Education Statistics

The U.S. Department of Education’s National Center for Education Statistics (NCES) is the conduit for most U.S. education data. State and local agencies report a variety of information to NCES, including data on enrollment, number of teachers, expenditures, and characteristics of public elementary and secondary schools. Data on private schools and higher education are gathered through separate NCES surveys. In addition, the NCES conducts several national surveys to collect more detailed data on specific areas of interest, such as teachers (Schools and Staffing Survey), early education (Early Childhood Longitudinal Study), and high school and post-college transitions (High School and Beyond; Baccalaureate and Beyond). Many of these are longitudinal; that is, they study individuals over periods of time.

Data Sample: NCES survey data show that in 2009–10, approximately 19 percent of public schools required students to wear uniforms and 57 percent of public schools enforced a strict dress code.

U.S. Census Bureau

Data on school enrollment and educational attainment are correlated with other individual characteristics every 10 years in the U.S. Census (see Chapter 2), annually in the Current Population Survey (see Chapter 9) and American Community Survey (see Chapter 2), and occasionally in the Survey of Income and Program Participation (see Chapter 8).

Data Sample: In the 2010 Current Population Survey, 29.9 percent of the U.S. population 25 years and over had completed a bachelor’s degree or higher level of education, an almost sixfold increase from 4.6 percent in 1940, when the U.S. Census began asking for the number of years of school completed by each individual.

123

Other Surveys

Much debate about school desegregation is based on data collected by the Office of Civil Rights in the U.S. Education Department. In the private sector, the National Education Association, an organization of teachers and school administrators, conducts its own survey of state education data, often published in advance of NCES data. Education Week’s annual Quality Counts reports, begun in the late 1990s, provide state-level at-a- glance statistics on student achievement, school finance, teachers, accountability, and other state education policies. Finally, there are several sources of longitudinal data, such as the National Longitudinal Survey and several university-sponsored surveys.

Data Sample: In the 2008 National Longitudinal Survey, 8 percent of 23-year-olds had earned a GED credential and were not enrolled in college.

State Data

In each individual state, the Department of Education is likely to collect a wide range of administrative data on students, schools, and districts. Although there is wide variation in the exact data collected, and whether or how it is made publicly available, states may have data on student demographics, performance on state assessments, staffing, and district finances.

Data Sample: In 2010, according to the California Department of Education’s DataQuest system, 71 economically disadvantaged students in the San Diego Unified School District took the High School Exit Exam in math and 18 percent passed, while 80 economically disadvantaged students took the Exit Exam in English and 8 percent passed.

Controversies

Poor Data

Although the U.S. government has collected educational data since 1869, their quality has lagged far behind comparable statistics for other areas of study. Before the 1950s each state reported information to the federal government as it saw fit, and sometimes not at all. The office of

124

Education’s handbooks, first published in 1954, attempted to remedy some of these obvious flaws, but truly national data were not collected until the 1977–78 school year in what is called the Common Core of Data. originally the Common Core was an ambitious project in which educational statistics would be gathered in a speedy and comprehensive manner, similar to the compilation of data by the National Center for Health Statistics or the Bureau of Labor Statistics. But lack of funding and resistance by state officials and other federal agencies to relinquish their power over educational data resulted in a much-scaled-back project focusing on enrollment, attendance, revenues, and expenditures. During the 1980s the National Center for Education Statistics suffered reductions in budgets more than triple the cutbacks for other federal statistical agencies. Forty percent of the items not dealing with funding were eliminated from the federal database between 1981 and 1983, including questions about critical policy issues such as school busing and the gender mix of teachers. However, the Common Core today is still the most widely used national source for basic information about schools and districts.

The biggest complaint about education data is the delay in obtaining it. The NCES data often are published so long after collection that they cannot be used for policymaking. In order to maintain a reasonably current public record, the U.S. Education Department’s chief statistics publication, Digest of Education Statistics, relies partially on data collected by private organizations, such as the National Education Association, because their data are available before the department’s own official figures.

There are also issues with standardization across states. The No Child Left Behind (NCLB) legislation, the 2001 reauthorization of the federal Elementary and Secondary Education Act, required all states to track student performance (as measured by annual test scores and graduation rates) for specific subgroups of students as well as the overall population. However, each state was largely allowed to set its own standards and determine its own accountability systems. As a result, it can be difficult for researchers to compare data between states. For example, variations in how states adjust for absent students make it difficult to compare public school enrollment, and performance on the standardized tests used in accountability programs are generally not comparable across states. Recent efforts to create uniform reporting procedures for educational statistics have made some headway (see discussion of dropouts) but there is still much room for further improvement. Even if such upgrades are made, the stinginess of usable information from previous decades will hamper

125

research efforts because there is no long-standing database for making historical comparisons.

Educational Attainment

How Many Students Complete High School?

One example of an educational statistic that is poorly measured is the high school dropout rate, or its complement, the graduation rate. Based on the Common Core of Data, the national dropout rate is often reported as around one-third, and as high as 50 percent for blacks and Hispanics. However, researchers at the Economic Policy Institute have argued this is an overstatement compared to better data from Department of Education surveys and census household surveys, which indicate the overall high school graduation rate is closer to 80 percent and in the range of 70 percent for blacks and Hispanics. In addition, some researchers have computed state-level graduation rates that differ sharply from the rates reported by the states themselves; in one analysis, almost half of the states reported rates that were ten or more percentage points above those calculated by independent researchers.

A big reason for the differing estimates is that calculating dropout or graduation rates requires the analyst to make several decisions, none of which has a clear “right” choice. As a ratio, it seems obvious that the numerator of a graduation rate should be the number of students who obtain a diploma, but should diplomas by equivalency exam be included? Should we count students who take longer than the traditional four years to graduate? What about those who are allowed to take modified coursework (such as students with disabilities)? And what is the comparison group to use in the denominator? Should it be the number of students who started high school four years earlier, which could include students who repeat grade 9? Or should we use the size of the senior class in the year of graduation, which would miss anyone who drops out prior to grade 12?

Other problems stem from the use of school-level administrative data: When a student leaves one school, it is not always clear whether they have transferred to another school or dropped out. Even if a student is known to transfer, does that student count as part of the estimate for the sending school or the receiving school? And what to do with students who are expelled or die?

Answering these questions is complicated by the fact that researchers

126

and policymakers may be interested in different measures depending on the context. For policymakers who want a general idea of how many high school or college-educated workers there are in the population, it may not matter how those degrees were achieved or how long it took to earn them; if the concern is whether specific schools are getting better (or worse) at successfully graduating students in four years, then a different measure is needed.

Accurate measurement of high school completion is becoming even more important as state and federal accountability systems are increasingly including dropout or graduation rates among the mix of student outcome measures that determine whether individual schools are considered successful or not. And it is not just the reputations of individual schools that are at stake; these measures can also affect larger policy decisions about where and how to target limited educational dollars.

For at least some purposes, researchers will soon have data that is comparable across states. As part of the NCLB legislation, states are required to report graduation rates as part of their accountability system, but it was not until 2008 that the Department of Education issued guidelines for a standard definition of that rate. The 2010–11 school year was the first year in which all states were required to use the official measure, which defines a “four-year adjusted cohort graduation rate” as “the number of students who graduate in four years with a regular high school diploma divided by the number of students who entered high school four years earlier (adjusting for transfers in and out, émigrés, and deceased students).” Although some may quibble with the exclusion of modified diplomas or diplomas by exam, the standardized definition will at least allow for more consistent comparisons going forward and across geographic regions.

Box 5.1 Schools Are Only as Bad as We Think They Are

The top problems in U.S. schools are drug abuse, alcohol abuse, pregnancy, suicide, and rape. In the 1940s, the top five offenses were talking, chewing gum, making noise, running in the halls, and getting out of turn in line. Between 1985 and 1993, these lists were cited by Newsweek, Harper’s, CBS News, former secretary of education William Bennett, Ross Perot, Senator John Glenn, and columnists

127

Carl Rowan and Anna Quindlen-even though both lists were complete fabrications!

Because so many reputable sources used the lists, often referencing one another, it took detective work by Yale professor Barry O’Neill to track down the origin of the comparison. O’Neill found its creator, T. Cullen Davis, a fundamentalist Christian fighting against sex education and the teaching of creationism in Texas schools. Davis admitted he made up the list. “How did I know what the offenses in schools were in the 1940s? I was there. How do I know what they are now? I read the newspapers.”

The unscientific origins of the list are less upsetting than its repetition as fact by so many reputable sources, none of whom bothered to check it for accuracy. O’Neill traced 250 versions of the list, noting how it changed to suit public concern. For example, Davis placed rape as the number one problem, inadvertently remembering rape as the most serious, but not necessarily the most common crime on a standard reporting form. The California Department of Education reproduced the list, copied from Harper’s, but put drugs at the top of the list, probably to confirm public perception that drugs were a problem. The lesson for researchers is that experts quoting experts can lead to national folklore based more on our fears than on factual research.

Source: Barry O’Neill, “The History of a Hoax,” New York Times Magazine, March 6, 1994, pp. 46–49.

How Many Students Complete College?

Pinning down college completion rates is also a problem for education researchers. According to the Department of Education’s Integrated Postsecondary Education Data System (IPEDS), among bachelor’s degree- seeking students who began college in 2002, 57.2 percent graduated within six years. That puts the United States second to last among countries in the Organisation for Economic Co-operation and Development (OECD). However, the U.S. ranks first in the world for share of population earning bachelor’s degrees. At least part of this discrepancy comes from the fact that IPEDS only measures graduation from the school of entry, so anyone who transferred institutions at any time before graduation is simply not

128

counted. Schools that send many students to other colleges lower their own graduation rates, even if those students eventually graduate elsewhere; at the same time, when those students do graduate, they are not counted as part of the rate at the new school either. thus, schools that happen to send or accept large numbers of transfer students will have lower graduation rates, regardless of the job they are doing helping students complete their degrees. in contrast, every other country reports graduation rates within systems. in the united States, that would correspond to tracking students who start at any four-year college and graduate from any four-year college. By that measure, the national rate would increase to 63 percent, putting the united States much closer to the top.

Another reason why the united States looks worse in international comparisons is that countries seldom use a standard time frame. For example, norway’s completion rate of 67 percent is for students who finish within ten years, not six. Even within the united States, critics of the IPEDS measure point out that the six-year time frame can create an appearance of poor quality for colleges that serve low-income and nontraditional populations who are more likely to go to school part-time, to require remedial coursework, or to take semesters off—all of which can increase the time to graduation. John Bassett, president of Heritage university in toppenish, Washington, which serves a largely low-income, local rural population, points out that Heritage may have six-year IPEDS graduation rates in the teens, but for students who start as traditional freshmen—having completed any necessary remedial work, and not including those who transfer—the eight-year graduation rate is over 40 percent. among students who make it through their first year, the graduation rate jumps to 81 percent. Bassett uses these differences to highlight the need for a completion measure that accounts for transfer patterns and for the different needs of students from varying demographic backgrounds.

Given these data issues, researchers should take care in making comparisons between colleges in the united States and those abroad, and even among schools within the united States. in particular, any policy prescriptions based on IPEDS data should consider how time frame and demographics affect the observed data.

Testing

Does the United States Rank Last?

129

For years now, students in the united States have trailed students in other countries on a number of different standardized tests (see table 5.1). Each new round of results sparks headlines bemoaning the terrible state of american schools or more generally, America’s fall from being a world leader in science and innovation. But do the tests really justify these dreary pronouncements? Educational experts are not so sure.

Box 5.2 Are Women Crowding Men Out of College?

In 2008 women made up over 56 percent of college undergraduates and over 60 percent of graduate students. Since 1980, when the number of women in college first surpassed the number of men, the gender gap in higher education has grown steadily. Some policymakers have interpreted this trend as a worsening situation for men, or that the increase in female enrollments is somehow the cause of the decline in male college enrollments. But are men really worse off? A closer look at the data shows that male enrollments are at an all-time high. Men continue to attend college at increasing rates; the gender gap is growing because female enrollments are simply growing faster. Thus, women are not replacing men in higher education but are joining them at higher rates than in the past. While some may still consider a lopsided gender share a problem, it is a fundamentally different problem than if the rate of enrollment were actually falling for men.

Sources: Marcus Weaver-Hightower, “Where the Guys Are: Males in Higher Education,” Change, May-June 2010, http://www.changemag.org/Archives/- Back%20Issues/May-June%202010/where-guys-full.html.

130

Source: U.S. Department of Education, National Center for Education Statistics, Digest of Education Statistics 2010, “Table 213: Total Undergraduate Fall Enrollment in Degree-Granting Institutions, by Attendance Status, Sex of Student, and Control of Institution: 1967 Through 2009,” http://nces.ed.gov/programs/digest/d10/tables/dt10_213.asp?referrer- =list.

Figure 5.1 College Enrollment by Gender, 1970–2009

Table 5.1

U.S. Performance on International Tests

Test U.S. Ranking Trends in International Math and Science Study (TIMSS) 2007: 4th-grade math 11 out of 36 8th-grade math 9 out of 48 4th-grade science 8 out of 36 8th-grade science 11 out of 48 Progress in International Reading Literacy Study (PIRLS) 2006 18 out of 45 Programme for International Student Assessment (PISA) 2009: Reading 12 out of 34 Math 25 out of 34

131

Science 17 out of 34

Source: U.S. Department of Education, National Center for Education Statistics, International Activities Program, Assessments and Surveys, http://nces.ed.gov/surveys/international/assessments.asp.

One of the most publicized tests is the Trends in International Mathematics and Science Study (TIMSS), first conducted in 1995 by the International Association for the Evaluation of Educational Achievement (IEA) and administered every four years to fourth and eighth graders around the world. The IEA also began assessing fourth-grade reading achievement in 2001 with the Progress in International Reading Literacy Study (PIRLS), repeated every five years. Another well-publicized set of tests come from the OECD’s Programme for International Student Assessment (PISA), given every three years to 15-year-olds and assessing reading, mathematical, and scientific literacy.

The United States has not been among the top-scoring countries on any of these assessments since the early 1990s but researchers at the Urban Institute point out that the implications are not nearly as dire as the media would have us believe. Although the United States consistently trails Asian countries on the TIMSS, comparisons to other countries are less consistent because there is significant movement up and down in the rankings in different years and on different tests. For example, in 2003, the United States ranked twelfth in fourth-grade math, while Lithuania ranked eighth and Hungary ranked eleventh. By 2007, the United States moved up one to eleventh while Lithuania moved down to tenth and Hungary ranked fifteenth. Slight differences in rankings are also often not statistically significant; for example, in 2007, American eighth graders ranked ninth in math (out of 48 countries) but were not statistically different from Hungary (#6), England (#7) or the Russian Federation (#8). And even if the test scores do suggest there is room for improvement, they have not translated into a decline or shortage of science and engineering graduates.

Another limitation of international comparisons is that they may compare different populations of students. For example, during the 1960s IEA studies compared U.S. high schoolers with students in countries where as little as 10 percent of the population continued to the secondary level. When the tests apply only to an educational elite, it is not surprising that students scored better than the 70 percent of U.S. students remaining in secondary school. An obvious solution to this is to limit comparisons to

132

lower grades, where almost all children are still in school, or only include the growing number of countries for which education through the equivalent of U.S. high school is the norm. Even with these restrictions, some analysts note that the population in the United States is far more diverse and that heterogeneity poses challenges that other countries do not face, so the samples are not really comparable. On the other hand, economist Eric Hanushek compared math outcomes for all students in other countries to just white students or students with college-educated parents in individual American states and found that these more-privileged students still do not perform as well as those in many other countries, though there is a lot of variation across states.

One problem for researchers is that international comparisons do not tell us why some countries do well, so it is difficult to design educational policies that will achieve higher performance. Education researcher Richard Jaeger concludes that with the exception of time spent on homework, factors such as class size and amount of time exposed to instruction are poor predictors of variation in test scores between countries. If we are to move beyond the frightening headlines and begin to understand why students perform differently in other countries, we need improved international data, looking at more subjects and educational indicators in addition to standardized test results.

Are Students Learning Less?

A similar dispute examines the trend in U.S. education: are schools doing a better or worse job than in the past? On the one hand, we have more data to answer this question for the last decade than in previous eras, thanks to the testing requirements of No Child Left Behind, which have led to a huge increase in the amount of testing data collected at the state level, disaggregated by racial and economic subgroups. On the other hand, because every state administers its own tests and sets its own standards about what scores denote proficient performance, characterizing the trend in U.S. scores has also become much more complicated.

Overall, it seems that students are learning more; according to state tests, the percentage of students performing at a proficient level in both reading and math has increased steadily since 1992 in all but a small handful of states. This is mirrored in the National Assessment of Educational Progress (NAEP), the one consistent national source of trend data on student achievement. Administered by the Educational Testing

133

Service (the same organization that conducts the SAT) to more than 200,000 students, the NAEP, also called “The Nation’s Report Card,” is a test series that dates back to 1969 and is the only ongoing measurement of achievement for youths in public and private schools in grades 4 through 12. It is recognized by education experts as a relatively accurate evaluation instrument because it involves tight security and relatively little “teaching to the test” by teachers to inflate student scores.

Although both the NAEP and state tests show an overall upward trend, the magnitude of the gains on the two sets of tests appear quite inconsistent. Recall that NCLB mandated that all students be tested, but it left it up to the states to determine their own standards; thus, what it means to be “proficient” can vary widely across states and need not bear any relationship to the federal standard of proficiency measured on the NAEP. According to researchers at the University of California, Berkeley, large gaps exist between the reported percent proficient on state tests and the percent proficient on the NAEP; for example, in Texas, there is a gap of 52 percentage points in math and 56 percentage points in reading. Overall, for the sample of states in that study, the gaps ranged from 1 percent (in math in Massachusetts) to 60 percent (in math in Oklahoma); none of the state tests showed lower percent proficient than the NAEP. Along similar lines, the researchers found that improvements over time were typically higher when measured with state tests than the NAEP.

The difference between state and national tests is particularly important when one considers that even though the trend has been positive, the NAEP still indicates performance levels that are disturbingly low. For example, on the 2009 NAEP, only 33 percent of fourth graders performed at the proficient level in reading and only 26 percent of twelfth graders were proficient in math. Although these statistics look better when we use state tests, there are still large numbers of students who are performing below grade level across the country. A study sponsored by the American Federation of Teachers found the actual test questions answered correctly by “average” European high school students “far exceeds what American students have achieved by that grade.” For example, German students wrote a 120-word essay in English and French students wrote for one hour on “The Causes of the First World War.” Education professors David C. Berliner and Bruce J. Biddle defend U.S. education, pointing out that U.S. schools downplay the “stress on the subservient conformity that generates high levels of subject-matter achievement” in order to promote creativity, social responsibility and friendliness—hard-to-measure attributes of U.S.

134

students.

Box 5.3 Are Boys Better at Math and Science?

In 2005 then president of Harvard University Larry Summers set off a firestorm with remarks that many interpreted as suggesting men are innately better at math and science than women are. Part of Summers’s argument was based on studies that show boys perform better than girls on math tests. In particular, although average math scores for boys are only slightly higher than average scores for girls, the gap widens considerably at the top end of the distribution. That is, boys are much more likely than girls to excel on math tests. Although this empirical pattern has been well documented, there is still much debate about why this difference exists. Summers’s implication that the difference is rooted in nature versus nurture ultimately contributed to his stepping down from the Harvard presidency, but the debate rages on.

In a symposium on “Tests and Gender” in the Journal of Economic Perspectives, two sets of economists provided some evidence for the nurture side of the debate. First, Devin Pope and Justin Sydnor examined how the gender gap among high-performers varies across states. They found that the pervasive gap seen in national statistics is not observed equally across the country; instead, some states are consistently more “gender-equal” than others. Their work highlights the pitfalls of relying on national statistics that can mask within- country variation.

Economists Muriel Niederle and Lise Vesterlund focus on the tests themselves. They argue that at least part of the gender gap at higher percentiles of the distribution can be explained, not by true differences in math skills but by differences in how the genders respond to the competitive environment in which the tests are taken. Prior studies have shown that women perform less well in more competitive environments, a distortion that is likely to be particularly large on math tests. One implication is that any education researcher using testing instruments that are administered in competitive environments should take care in interpreting gender differences in outcomes.

135

Sources: Ruth Marcus, “Summers Storm,” Washington Post, January 22, 2005, p. A17, http://www.washingtonpost.com/wp-dyn/articles/A27819– 2005Jan21.html; Tamar Lewin, “Math Scores Show No Gap for Girls, Study Finds,” New York Times, July 25, 2008, http://www.nytimes.com/2008/07/25/education/25math.html; Devin G. Pope and Justin R. Sydnor, “Geographic Variation in the Gender Differences in Test Scores,” Journal of Economic Perspectives, 24, no. 2 (2010): 95–108; Muriel Niederle and Lise Vesterlund, “Explaining the Gender Gap in Math Test Scores: The Role of Competition,” Journal of Economic Perspectives, 24, no. 2 (2010): 129–44.

Are We Closing Achievement Gaps?

NCLB and the accompanying accountability reforms have focused largely on the differences in performance between white students from higher- income, better-educated families and nonwhite students from households of lower socioeconomic status. States are now required to report test scores for racial and economic subgroups so policymakers and researchers can better track any existing gaps over time. Unfortunately, although most subgroups follow the general upward trend discussed in the previous section, there appears to have been little closure of the gaps between groups. That is, black, Hispanic, and low-income students all show gains over time, but the size of those gains is roughly the same as the gains for white students.

One factor that complicates subgroup comparisons is the treatment of English-language learners and students with disabilities. For much of the history of the NAEP, states have been able to exclude English learners and students with disabilities, and there have been wide disparities in how many of these students were included in NAEP testing. States also have varying policies about how these students are treated in the administration of state tests. These disparities may be particularly problematic for comparisons between states with very different demographics; for example, the Hispanic and Asian subgroups in California are more likely to include large shares of English learners than those same subgroups in Massachusetts. Researchers should take care to consider how these issues might affect their empirical analysis.

Why Do SAT Scores Keep Falling?

136

SAT scores are the one of the most highly publicized measures of educational performance. Every year, newspapers report the one-or two- point change in the math and verbal portion of this test taken primarily by high school seniors for admittance to college. In 2011 the Washington Post’s headline was “SAT reading scores drop to lowest point in decades,” reflecting that the average score of 497 was down three points from the previous year and 33 points from 1972.

On closer examination, however, the picture is more complicated. 2011 was also the first year in which more than half of all high school graduates took the exam; additionally, test takers were more diverse than ever, with 44 percent minority students. It is not surprising that the 1.65 million SAT takers today perform worse than the 11,000 mostly Ivy League-bound students who first took the SAT in 1941. The College Board has argued that the decrease in test scores for the entire population is because of an increase in the number of test takers who came from lower-scoring, mostly lower-income, groups. SAT scores have been plagued in many years by Simpson’s Paradox (see Box 5.4), in which trends for the whole population diverge from trends for subpopulations because of changes in the overall population’s composition.

Source: College Board, “Total Group Profile Report,” 2010, http://professionals.collegeboard.com/profdownload/2010-total-group-profile- report-cbs.pdf.

137

Figure 5.2 SAT Scores, 1972–2009

On the other hand, it is unclear how much of the recent drop in SAT scores can really be blamed on composition changes. Although scores increased for nearly every subgroup during the 1990s, subgroup scores have stagnated since 2000; only Asian scores have increased while white, black, and Hispanic averages have been flat or fallen. So, similar to the NAEP results discussed earlier, achievement gaps between white students and black or Hispanic students appear to be holding relatively steady. What remains unanswered is why the trend has changed in the last decade.

Box 5.4 Simpson’s Paradox

Named for the British statistician who first formally described it, Simpson’s Paradox arises when a relationship that appears in aggregated data disappears or is reversed when the data are broken down into subgroups. It typically occurs when the subgroups are different sizes. For example, scores on the National Assessment of Educational Progress are reported by state for all students as well as by particular subgroups. In 2009—in both reading and math, for both 4th-graders and 8th-graders—Wisconsin had higher overall scores than Texas. But on each of those tests, the scores for individual subgroups by race were higher in Texas; that is, white students in Texas outperformed white students in Wisconsin, black students in Texas outperformed black students in Wisconsin, and Hispanic students in Texas outperformed Hispanic students in Wisconsin. The paradox exists because Texas has many more black and Hispanic students, who had lower average scores, and that pulled down the average for the state. So even though scores for each group were higher, the average for the total population appears lower in Texas than in Wisconsin.

Source: David Burge, “Longhorns 17, Badgers 1,” Iowahawk (blog), March 2, 2011, http://iowahawk.typepad.com/iowahawk/2011/03/longhorns- 17-badgers-1.html.

138

Charter Schools: Are They More Effective Than Regular Public Schools?

Since 2000, there has been a huge increase in the number of charter schools in the United States. Charter schools are technically public schools, but they are exempt from many of the rules and regulations that govern traditional public schools. In exchange for this freedom, charter schools are expected to produce specific results in terms of student achievement; if they do not, they risk losing their charter and being forced to close. An obvious question is whether charters are more effective than regular public schools, but that has not been an easy question to answer definitively. A big problem is that students who choose to attend charter schools may be quite different from students who choose not to, so simply comparing outcomes at charter schools to outcomes at regular public schools would not be comparing apples to apples. To address that issue, some researchers have focused on charters that select their students using a random lottery; students who participate in the lottery but do not win a spot can then be compared to those who do win and enter the charter. In those studies, charters generally perform well; critics, however, point out that any school popular enough to require a lottery (which means it is oversubscribed in the first place) is likely better than other charter schools, so the results should not be generalized to all charters. An alternative approach— one used by researchers at Stanford University—is to match charter school students with students in regular schools by matching students on things like demographics and participation in special programs (such as English learners or special education). Using those comparisons, the Stanford group found that only 17 percent of charters were outperforming traditional schools, and just over one-third of charters actually performed worse. Critics of that study argue that the groups of students being compared, even though matched on several characteristics, still differ in important ways—such as prior achievement or parental involvement—that are not captured in the analysis.

One thing all these charter studies have in common is finding tremendous variation in charter quality. For example, the Stanford study found that in New York, over half the charter schools were more effective in math than their traditional counterparts—much higher than for the national sample. This kind of variation makes it even more difficult to make sweeping generalizations about charters versus traditional schools. It is not difficult for the media to find success stories among charters, but that does not mean all charters are doing as well. The lessons for

139

researchers are that no single method is perfect: These sorts of comparisons often require a nuanced approach that recognizes the large variation within types of schools.

Teacher Compensation

One education issue where there is solid consensus among education researchers is the importance of teachers. Although family background variables such as income and parental education are consistently the best predictors of student achievement, researchers have found that having a good teacher can make a significant difference as well. This has led to intense policy debates about how to attract and retain high-quality teachers, with much of the debate centered on teacher compensation. Some observers argue that all members of the teaching profession should be paid more; others argue that the fundamental structure of teacher pay needs to be changed so that compensation is better tied to student performance. Both of these lines of argument are plagued by data controversies.

Are Teachers Underpaid?

Conventional wisdom tells us that no one becomes a teacher for the money, but are teachers in the United States actually underpaid relative to other professional workers with similar training? Michael Podgursky, an economist at the University of Missouri, says that the answer is no, teachers’ average weekly pay is actually greater than that of comparable professionals, once one accounts for teachers’ shorter work year. But Lawrence Mishel, Sean Corcoran, and Sylvia Allegretto, researchers with the Economic Policy Institute, counter that Podgursky’s measure overstates the amount of time off that teachers have and that teachers actually do earn less than comparable workers.

The so-called time-off factor is a key point in discussions about teacher pay. Podgursky considers summers as time not spent teaching and subtracts that time before converting teachers’ annual salaries to a weekly rate; the EPI team argues that many teachers are required to spend their summers engaged in professional development to maintain their credentials, or to prepare class materials, so Podgursky’s measure will overestimate weekly pay. A similar difference arises in debates about how many hours teachers work per week, since comparisons of weekly or annual wages between teachers to other workers depend on the assumption that weekly hours are comparable. Some researchers assume that only time

140

spent in the classroom or directly related to the classroom (such as grading) is considered working, so teachers will therefore appear to work fewer hours than other occupations; others argue that teachers spend a good deal of “off-campus” time still working, making teachers’ weekly hours comparable to or even higher than hours logged by other professionals.

Another open question is which workers are most comparable to teachers? The EPI group examines the specific skills required for various occupations and selects sixteen professions that require skills similar to teaching, including accountants, nurses, counselors, and computer programmers. Podgursky uses a much broader group—all occupations that require a college degree— arguing that any attempt to identify specific comparable occupations will be fraught with arbitrary decisions. Both approaches have merit but result in very different conclusions about whether teachers are underpaid. One lesson for researchers is to be clear about what assumptions underlie the numbers and to discuss how those assumptions affect the resulting analysis.

Who Is an Effective Teacher?

In addition to questions about whether the level of teacher pay should be higher, there is a vigorous debate among education researchers and policymakers over the structure of teacher pay. For decades, teachers in almost every district have been paid according to a step-and-column salary schedule, where teachers earn increasing pay by moving down the steps with each additional year of experience and across the columns with additional educational credits. Thus, within a district, any two teachers with identical experience and education will earn the identical salary (not including any additional for-pay duties such as coaching or administrative roles). However, researchers have found no evidence that additional experience, beyond the first few years, nor additional education (whether that means advanced degrees or accumulated credits) are significantly related to student outcomes. This has led to numerous proposals for performance pay, where teacher salary is more closely tied to measures of teacher effectiveness. These proposals range from one-time bonuses awarded on top of the salary schedule to complete replacement of the step- and-column schedule with a performance-based system. In all such proposals, a key sticking point is how one measures teacher effectiveness. While many parents will say they know who the good (and bad) teachers are in their child’s school, defining an effective teacher for the purposes of

141

compensation and promotion is a contentious and often politically charged issue.

One option is to define an effective teacher as one who increases student learning as measured by student performance on standardized tests. In 2010 the Los Angeles Times itself became news when it published rankings of teachers in Los Angeles Unified School District based on value-added test scores. “Value-added” measures attempt to capture the change in student test scores that can be attributed to a particular teacher; by focusing on the change in scores, value-added measures attempt to control for the many other factors that can affect test scores, such as student characteristics. The LA Times analysis set off a huge debate in the education community about the use of such measures. Several critics emphasize problems with the measures themselves, particularly the large variability in scores from one year to the next; for example, a teacher could be ranked at the top in one year and at the bottom in the next. others focus on issues with the underlying tests, arguing that the tests were not designed for the purpose of evaluating teaching and there is more to good teaching than just higher test scores. There are also concerns that tying teacher compensation entirely to test-related measures will create perverse incentives for teachers to cheat and selectively ignore materials and concepts that do not typically appear on standardized tests. Partly because of these concerns, most education analysts prefer using multiple measures; for example, combining value-added scores with classroom observations and evaluations from administrators or peers.

Implications

In debates about educational policy, partisans are able to select evidence showing that U.S. schools are doing well and that they are doing poorly. But researchers are less definitive about whether we should praise or condemn the U.S. educational system. We have seen improvements since 2000 but a key question is whether those improvements, and the current trajectory, are sufficient, or whether we need significant reform. Unfavorable international comparisons and falling standardized test scores are cited most often by those who favor such reforms as increased choice and performance-based teacher compensation. In this view, only free- market competition will shake up the dismal failure of public school systems; however, it remains unclear whether such competition really leads to more effective schools.

142

1.

Other scholars maintain that massive school reform deflects resources from where they are needed. Harold Hodgkinson, director of the Center for Demographic Policy of the Institute for Education Leadership, points out, “We know where the bad schools are—in the middle of our largest cities.” In this view, the data show that many, mostly suburban, U.S. schools work well, arguably among the best in the world. Instead of blaming the overall structure of U.S. schools, we should focus attention on inner cities, which need not only better schools but also improved health care, jobs, and safe neighborhoods.

Summary

A fundamental problem with education statistics is poor data. More timely and standardized data will assist researchers, as will increased willingness by the National Center for Education Statistics to tackle controversial issues. In addition, better-quality longitudinal data will provide more accurate information on enrollment rates, dropout rates, and the effect of different schools on students’ future careers.

For several of the issues covered in this chapter, including dropouts, charters, testing, and international comparisons of math and science achievement, relatively adequate data exist. Controversies arise because of different interpretation of the numbers. How many students complete high school and college? Are charters more effective than traditional public schools? Do U.S. students perform as well as students in other countries? Is achievement in U.S. schools falling? Media accounts of these disputes rarely explain the origins of the statistical discrepancies, focusing instead on the different policies that researchers recommended. In each case, closer examination of the statistics reveals important information about education. For example, hidden within the debate about the precise dropout rate is the more critical issue about what level of education is required for succeeding in the United States. And, although experts debate whether or not there has been a decline in U.S. student achievement levels, the overriding issue is the unchanged low level of achievement for too many students. Although many educational data remain uncollected and many research questions remain unanswered, there is still much to be learned from careful attention to the data.

Case Study Questions

According to a 1988 study of state education departments, all

143

2.

3.

4.

claimed their students performed above the national average. This mathematical impossibility was termed the “Lake Wobegon effect,” after Garrison Keillor’s mythical Minnesota town where “all the children are above average.” Technically, the national average is a median, meaning one-half of tested students score below it and one-half score above it, but that is not the source of the paradox. The standard was set in the early 1980s. Why were all the children reported to be “above average”?

In 1990 the U.S. Census Bureau changed its question about education attainment from “What is the highest grade (or year) … attended?” to “What is the highest level completed or degree received?” Why do you think this change was necessary?

In a 2010 report on state-level education measures, the Heartland Institute ranked states on student achievement, education expenditures, and learning standards. The report does not control for demographics or other student or state characteristics. Each state (plus the District of Columbia) was assigned a letter grade, and the grades were uniformly distributed so that 10 states received A’s, ten received B’s, and so on. What effect do you think it would have on the rankings if student demographics were included in the analysis? How might this grading scheme create a misleading picture of failing states?

According to a 2009 study, alternatively certified teachers (teachers who receive certification through special programs) are just as successful as teachers who are traditionally certified. In order to compare teachers in similar environments, the study matched alternatively certified and traditionally certified teachers working in the same school. Thus, the sample focused on schools that routinely hire a relatively large number of alternatively certified teachers. How might this choice of sample have affected the study results?

References

Data Sources (87)

National Center for Education Statistics (87)

Description of NCES data in U.S. Department of Education, National Center for

144

Education Statistics, From Data to Information: New Directions for the National Center for Education Statistics, NCES 96–901 (Washington, DC: U.S. Government Printing Office, 1996); surveys in “Surveys and Programs,” NCES online, http://nces.ed.gov/surveys/. Data sample in Indicators of School Crime and Safety: 2011, “Table 20.1: Percentage of Public Schools That Used Safety and Security Measures: Various School Years, 1999–2000 through 2009–10,” NCES 2012–002, February 2012, http://nces.ed.gov/programs/crimeindicators/crimeindicators2011/tables/table_20_1.asp.

U.S. Census Bureau (88)

Current Population Survey data sample in Digest of Education Statistics: 2010, “Table 8: Percentage of Persons Age 25 and Over and 25 to 29, by Race/Ethnicity, Years of School Completed, and Sex: Selected Years, 1910 through 2010,” NCES 2011–015, April 2011, http://nces.ed.gov/programs/digest/d10/tables/dt10_008.asp?referrer=list.

Other Surveys (88)

Civil rights data in Civil Rights Data Collection (formerly the Elementary and Secondary School Survey [E&S Survey]), U.S. Department of Education, Office for Civil Rights, http://ocrdata.ed.gov. National Education Association data described in National Education Association, “Rankings and Estimates: Rankings of the States 2011 and Estimates of School Statistics 2012,” December 2011, http://www.nea.org/assets/docs/NEA_Rankings_And_Estimates_FINAL_20120209.pdf. Quality Counts reports in Education Week, http://www.edweek.org/ew/qc/index.html. National Longitudinal Survey data sample in U.S. Bureau of Labor Statistics, “America’s Young Adults at 23: School Enrollment, Training, and Employment Transitions Between Ages 22 and 23” (news release, February 9, 2011), http://www.bls.gov/news.release/archives/nlsyth_02092011.pdf.

State Data (89)

State Department of Education contacts available from U.S. Department of Education, http://www2.ed.gov/about/contacts/state/index.html. California data sample in Data Quest, California Department of Education, http://dq.cde.ca.gov/dataquest/.

Controversies (89)

Poor Data (89)

Summary of NCES problems in Charles Cooke, Alan Ginsburg, and Marshall

145

Smith, “The Sorry State of Education Statistics,” Education Digest 51, no. 4 (December 1985: 28–30; see also Janet A. Weiss and Judith Gruber, “The Managed Irrelevance of Federal Education Statistics,” in William Alonso and Paul Starr, The Politics of Numbers (New York: Russell Sage Foundation, 1987), and Anne C. Lewis, “New Data Collection System Raises New Questions,” Phi Delta Kappan 76, no. 10 (June 1986, pp. 699–700. Standardization discussed in U.S. Department of Education, “Common Education Data Standards: CEDS 101,” http://ceds.ed.gov/pdf/ceds-101.pdf. See also Sarah D. Sparks, “Education Department Overhauls Data Website,” Inside School Research (blog), August 25, 2011, http://blogs.edweek.org/edweek/inside-school- research/2011/08/education_department_overhauls.html.

Educational Attainment (90)

How Many Students Complete High School? (90): Lawrence Mishel and Joydeep Roy, “Education Week’s Graduation Rate Estimates Are ‘Exceedingly Inaccurate,’ Experts Say,” Economic Policy Institute online, June 4, 2008, http://www.epi.org/publications/entry/webfeatures_viewpoints_20080604_gradrates/; Mishel and Roy, Rethinking High School Graduation Rates and Trends (Washington, DC: Economic Policy Institute, 2006); measurement issues discussed in High School Dropout, Graduation, and Complete Rates: Better Data, Better Measures, Better Decisions, ed. Robert M. Hauser and Judith Anderson Koenig (Washington, DC: National Academies Press, 2011), http://www.nap.edu/catalog.php?record_id=13035. Standardized definition of graduation rate in U.S. Department of Education, “A Uniform, Comparable Graduation Rate,” October 2008, http://www2.ed.gov/policy/elsec/reg/proposal/uniform-grad-rate.pdf. See also Kathryn Baron, “Grad Rates Trending Up—or Down?” Thoughts on Public Education (blog), June 8, 2011, http://toped.svefoundation.org/2011/06/08/graduation-rates-trending-up-or-maybe- down/; Sarah D. Sparks, “Scholars Call for More Nuanced Graduation Measures,” Inside School Research (blog), October 28, 2010, http://blogs.edweek.org/edweek/inside-schoolre- search/2010/10/scholars_call_for_fine-grained.html.

How Many Students Complete College? (92): IPEDS data in U.S. Department of Education, National Center for Education Statistics, Digest of Education Statistics: 2010, “Table 341: Graduation Rates of First-Time Postsecondary Students Who Started as Full-Time Degree-Seeking Students, by Sex, Race/Ethnicity, Time Between Starting and Graduating, and Level and Control of Institution Where Student Started: Selected Cohort Entry Years, 1996 Through 2005,” http://nces.ed.gov/programs/digest/d10/tables/dt10_341.asp. Clifford Adelman, “The Spaces Between Numbers: Getting International Data on Higher Education Straight” (research report), Institute for Higher Education Policy, November 2009,

146

http://www.ihep.org/Publications/publications-detail.cfm?id=131; John Bassett, “An Alternative to Graduation Rates,” Inside Higher Ed, April 1, 2011, http://www.insidehighered.com/views/2011/04/01/bassett_essay_on_improving_federal_measures_of_college_student_success.

Testing (93)

Does the United States Rank Last? (93): IAE and IAEP description and scores in National Center for Education Statistics, International Mathematics and Science Assessments: What Have We Learned? 92–011 (Washington, DC: U.S. Government Printing Office, 1992). B. Lindsay Lowell and Harold Salzman, “Into the Eye of the Storm: Assessing the Evidence on Science and Engineering Education, Quality, and Workforce Demand” (Washington, DC: Urban Institute, 2007), http://www.urban.org/publications/411562.html; Carl Bialik, “Whose Fourth Graders Are Smartest?” The Numbers Guy (blog), October 31, 2007, http://blogs.wsj.com/numbersguy/whose-fourth-graders-are-smartest-215/. Early problems with IEA in Richard Wolf, “The NAEP and International Comparisons,” Phi Delta Kappan, April 1988, pp. 580–82. Hanushek in Amanda Ripley, “Your Child Left Behind,” The Atlantic, December 2010, http://www.theatlantic.com/magazine/archive/2010/12/your-child-left- behind/8310/; Jeremy Kilpatrick, “International Math Study Doesn’t Add Up,” National Education Policy Center online, January 11, 2011, http://nepc.colorado.edu/newsletter/2011/01/international-math-study- doesn%E2%80%99t-add. Richard Jaeger, “Weak Measurement Serving Presumptive Policy,” Phi Delta Kappan, October 1992, pp. 118–27.

Are Students Learning Less? (96): NAEP versus state tests in Bruce Fuller et al., “Gauging Growth: How to Judge No Child Left Behind?” Educational Researcher 36, no. 5 (2006): 268–78. NAEP trend in Lawrence C. Stedman, “The Sandia Report and U.S. Achievement: An Assessment,” Journal of Educational Research, 87, no. 3 (January/February 1994): 137–40; American Federation of Teachers study in “Helping Students in the Middle,” American Educator 19 (Winter 1995– 1996): 2–19; “stress on” in David C. Berliner and Bruce J. Biddle, The Manufactured Crisis: Myths, Fraud, and the Attack on America’s Public Schools (Reading, MA: Addison-Wesley, 1995), pp. 52–53.

Charter Schools: Are They More Effective Than Regular Public Schools (101)

Stanford study in Multiple Choice: Charter School Performance in 16 States (Stanford, CA: Stanford University Center for Research on Education Outcomes, June 2009), http://credo.stanford.edu/research-reports.html; Carl Bialik, “The Conflicting Charter-School Numbers,” The Numbers Guy (blog), November 19, 2010, http://blogs.wsj.com/numbersguy/the-conflicting-charter-school-numbers- 1014/; and Bialik, “Studies That Grade Charter Schools Rely on Imperfect Math,”

147

Wall Street Journal, November 20, 2010, http://online.wsj.com/article/SB10001424052748704170404575624562978485450.html. See also Sean Reardon, “Review of How New York City’s Charter Schools Affect Achievement,” National Education Policy Center online, November 12, 2009, http://nepc.colorado.edu/thinktank/review-how-New-York-City-Charter.

Teacher Compensation (102)

Are Teachers Underpaid? (102): Michael Podgursky, “Fringe Benefits,” Education Next 3, no. 3 (Summer 2003): 71–76, http://educationnext.org/fringebenefits/; Sylvia A. Allegretto, Sean P. Corcoran, and Lawrence Mishel, How Does Teacher Pay Compare? Methodological Challenges and Answers (Washington, DC: Economic Policy Institute, 2004), http://www.epi.org/publication/books_teacher_pay/. See also “NCTQ Square-Off: Are Teachers Underpaid? Two Economists Tackle an Intractable Controversy,” National Council on Teacher Quality, July 2005, http://www.nctq.org/p/publications/docs/nctq_square_off_20071202080402.pdf.

Who Is an Effective Teacher? (103): Susan Moore Johnson and John P. Papay, Redesigning Teacher Pay (Washington, DC: Economic Policy Institute, 2009). Los Angeles controversy in Jason Felch, Jason Song, and Doug Smith, “Who’s Teaching L.A.’s Kids?” Los Angeles Times, August 14, 2010, http://articles.latimes.com/2010/aug/14/local/la-me-teachers-value-20100815; see also Andrew A. Rotherham, “Rating Teachers: The Trouble with Value-Added Data,” Time Magazine, September 23, 2010, http://www.time.com/time/nation/article/0,8599,2020867,00.html; Robert Manwaring, “LA Times Value-Added Release—Problems and Solutions,” The Quick and the Ed (blog), August 18, 2010, http://www.quickanded.com/2010/08/la-times-value-added-release-%E2%80%93- problems-and-solutions.html.

Implications (104): “We know where” in Harold Hodgkinson, Letter to the Editor, New York Times, November 26, 1991, p. A20; see also Hodgkinson, “American Education: The Good, the Bad, and the Task,” Phi Delta Kappan, April 1993, pp. 619–23.

Case Study Questions (106)

1. “The Misleading Concept of ‘Average,’” New York Times, July 12, 1989, p. B7.

2. “Making the Grade,” American Demographics, May 1987, p. 8.

3. Herbert Walberg and Marc Oestreich, “2010 State School Report Card,” Heartland Policy Study, Heartland Institute, October 25, 2010.

4. Sean Corcoran and Jennifer Jennings, Review of “An Evaluation of Teachers

148

Trained Through Different Routes to Certification: Final Report,” National Education Policy Center online, March 10, 2009, http://nepc.colorado.edu/thinktank/review-evaluation-of-teachers.

149

6

Crime □□□□

Researchers new to the criminal justice field may be surprised at the amount of data available, including nearly complete records of arrests and crimes reported to police and a survey of crime victims almost as large in sample size as the Current Population Survey (see Chapter 9). These data are the source of a number of controversies about crime, each with important public policy consequences. This chapter reviews several debates in criminology, including the accuracy of crime statistics, why the U.S. crime rate fell, where crimes occur, the effect of poverty on crime, who commits crimes, the role of gun control in crime rates, the deterrent effect of capital punishment, and the impact of white-collar crime.

Where the Numbers Come From

Organizations Data sources URL Federal Bureau of Investigation, U.S. Department of Justice

Uniform Crime Reports; National Incident-Based Reporting System

www.fbi.gov

Statistics Division, U.S. Department of Justice

National Crime Victimization Survey

www.bjs.gov

Data Sources

There are two major sources of U.S. crime data: (1) the Uniform Crime Reports (UCR), a compilation of police reports by the U.S. Justice Department’s Federal Bureau of Investigation (FBI), and (2) the National Crime Victimization Survey (NCVS), a survey of households conducted by the U.S. Census Bureau for the U.S. Justice Department. The UCR is

150

slowly being replaced by the National Incident-Based Reporting System (NIBRS), which makes new data available for some locations but is not yet used for national statistics. In addition, the FBI collects data on bank crime, violence affecting institutions of higher education, mass marketing fraud, mortgage fraud, Internet crime, drug threats, gang threats, and terrorism incidents. For easily accessible archived data, including examples of studies already completed using the NIBRS, consult the Inter- University Consortium for Political and Social Research website (www.psr.umic.edu).

Data Sample: In the Internet Crime Complaint Center 2010 report, the District of Columbia topped the ranking for rate of Internet crime perpetrators at 833 per 100,000 population, more than 10 times that of Montana, the next highest rate. However, the Center warns that D.C.’s numbers may be inflated by scams impersonating the FBI in which “complainants believe the incident has taken place in D.C., even though often the perpetrator is not based there.”

Uniform Crime Reports

When we hear that crime is up or down, the figures usually come from the Uniform Crime Reports (UCR). Since 1930 the Federal Bureau of Investigation (FBI) has collected reports from nearly 17,000 police departments across the country. These data include type of crime committed; time of occurrence; locality; and age, sex, and race of the offender. Even though overall coverage in the UCR has improved, significant problems remain with reporting from individual city police departments. The lowest level crimes are the most easily manipulated—for example, downgrading stolen goods values to $49, just under the threshold for grand larceny. In 1997 a former Atlanta, Georgia, police chief was accused of authorizing officers to reclassify some violent crimes to lesser crimes and to classify unsolved crimes as baseless reports. This misrepresentation contributed to Atlanta’s apparently improved crime situation in 1996, but it returned to its prior position in 1997 as the second most violent U.S. city. Between 1996 and 1998, Philadelphia, Pennsylvania, had to withdraw its crime figures from the national system because of underreporting and downgrading of crimes. More recently, a survey by criminologists found that New York City retired officers were aware of “ethically inappropriate” changes to crime complaints. One way to detect downgrading of crimes is to compare the major crime indicators

151

with the number of lesser crimes; an actual reduction in crime should cause both statistics to fall, while manipulation of data would cause fewer major crimes but more total crimes. Unfortunately, unlike nearly every other New York State police agency, New York City did not disclose data on lower level crimes.

Prior to 2004, the FBI published a total of seven main offenses called the “Crime Index,” often used to compare different localities, yet highly misleading because it created a bias against places with more minor thefts even if there were fewer serious crimes. Since then, the UCR program aggregates data by violent and property crime, adding a “pop-up” caution when entering the website: The pop-up warns that simplistic comparisons may create “misleading perceptions adversely affecting communities and their residents” because “rough rankings provide no insight into the numerous variables that mold crime in a particular town, city, county, state or region.” Such explicit warnings are unusual from data collection agencies, demonstrating the awareness by the FBI of the problems in crime data described in this chapter, as well as the likelihood that a location might be undeservedly labeled as crime-ridden or crime-free.

Data Sample: The 2009 UCR lists one violent crime and 171 property crimes at Brown University.

National Crime Victimization Survey

The second major data source on crime in the United States comes from crime victims in the National Crime Victimization Survey. These data are published annually by the U.S. Justice Department’s Bureau of Justice based on a special Census Bureau survey of about 40,000 households every six months for three years. The survey includes crimes that were and crimes that were not reported to the police, but excludes murder, commercial burglary and robbery, victimless crimes, prostitution, and white-collar crimes such as fraud and embezzlement. As a supplement to the victimization survey, every three years since 1996, the Bureau of Justice has surveyed “Police-Public Contact,” an effort, in their summary, “to estimate the likelihood of a driver being pulled over in a traffic stop and the percentage of all contacts that involve the use of force by police.”

The National Crime Victimization Survey can be inaccurate when, for example, one respondent answers for everyone in the household—a practice that leads to underreporting of crimes for other household members but exaggeration of minor crimes committed against the

152

respondent. To make matters worse, when a person has been the victim of several crimes, he or she tends not to report every instance of victimization. Apparently, respondents feel they have been cooperative by reporting a few crimes and do not want to be bothered to recollect the entire number of incidents. Finally, respondents tend to telescope past events, erroneously attributing crimes experienced more than a year previously to the “past year.” These errors tend to be greater for some social groups, especially those in low-income households who have suffered repeat victimization but report to survey takers only the most serious or most recent crime.

Data Sample: In the 2008 National Crime Victimization Survey, individuals in families with income under $15,000 said they reported 35 percent of property crimes to police, while those with incomes over $75,000 reported 43 percent of property crimes.

National Incident-Based Reporting System

Studies initiated in the late 1970s by law enforcement executives, criminologists, and the FBI recommended new guidelines for the eventual replacement of the UCR. More than 30 years later, a new program called the National Incident-Based Reporting System (NIBRS) became the adopted model, but it is still in its implementation stage. Nationally, the new program covered only 25 percent of the U.S. population in 2007 (and 44 percent of reporting agencies in 2009); while 32 states were certified for NIBRS participation, only 10 submitted their data exclusively in the new format. Implementation began with smaller states and smaller cities, where changeover to a new computer system was easier than in larger police departments.

NIBRS was designed to remedy recognized problems with UCR data. Most important, the new system is incident-based, recording data on the offense, offender, victim, property, and arrests for individual events. Consequently, the underlying data source is far richer than the UCR, which included only the most serious offense for each incident (called the hierarchy rule) and had little information about the victim.

Overall, the NIBRS expands UCR data to collect relevant data on 46 specific major crimes, including terrorism, white-collar crime, missing children, hate crimes, spousal abuse, and organized crime, thus greatly increasing data-based research possibilities. Also, a new category, “crimes against society”—such as gambling, prostitution, and weapons violations

153

—supplements the UCR’s “violent” and “property” crime categories.

The NIBRS will greatly expand analysis of crime victimization from the NCVS. For instance, hospital emergency room data show that many gunshot injuries are missed in the NCVS but are more likely to be included in reports to the police used in the NIBRS. Also, victims missed in the NCVS household survey because they are homeless or temporarily sheltered will be counted. The NIBRS will separate resident versus nonresident victimization, key for understanding crime in cities associated with large influxes of suburban commuters to urban areas such as Washington, DC, or visitors to touristdestination cities such as Las Vegas. Finally, because NIBRS data are based on local police records that, at least in theory, offer complete tabulations of reported crime, researchers will be able to better understand crime victimization differences between different cities. As an example, criminologist Michael Maxfield suggests that we could explore demographic characteristics that explain why auto theft rates are far higher in Jersey City, New Jersey, than in neighboring Hoboken.

Data Sample: Using NIBRS data, Illinois State University researchers found that the abusers of the elderly were most often acquaintances, followed by children, spouses, and other family members, rather than strangers.

Controversies

UCR, NCVS, or NIBRS?

Because of drawbacks to the Uniform Crime Reports, the National Crime Victimization Survey, and the National Incident-Based Reporting System, the challenge for researchers is to choose the data for which inaccuracies are least likely to affect the issue being studied. For some purposes the choice is readily apparent: the UCR more accurately designates types of crime because it uses standard legal definitions of specific offenses, while the NCVS gives a better estimate of the total number of crimes, including those not reported to police. The NIBRS combines the superior attributes of the UCR and NCVS, but as of 2012 it was not available for the entire country.

For measuring the overall crime rate, the choice between crime measures is less clear-cut and, unfortunately, may be based on political convenience. Criminologists Albert Biderman and James P. Lynch conducted an exhaustive comparison of the UCR and NCVS and pointed

154

out numerous reasons why researchers cannot use the two crime measures interchangeably. Different denominators—population in the UCR and households in the NCVS—by themselves account for a 13.9 percent divergence in the two series. Moreover, year-to-year variations in the crime rate may occur solely because of errors in measurement and thus do not reflect any changes in crime policy for which political leaders can take blame or credit. The lesson for researchers is to be aware of both sets of crime statistics and to wait for longer-term data to determine the upward or downward trend.

Given the slow progress in implementing NIBRS, national statistics or comparisons between most localities will continue to rely on the UCR. The FBI points out that it will continue to use the traditional format until it receives the majority of data via the NIBRS. Nonetheless, researchers can take heart that improved data is on the horizon and already available for studies of some crimes and for some localities. The transition from UCR to NIBRS will introduce new problems of data consistency over time. Already some localities complain that crime numbers are higher with NIBRS because it counts all parts of a single event so that, say, battery accompanying robbery is listed as two crimes, whereas robbery took hierarchical precedence in the UCR. The Bureau of Justice Statistics responded to these concerns by declaring that “multiple offense crimes account for a very small number of offenses and should not have a great impact on the overall crime statistics.”

Crime Is Down—And We Don’t Know Why

The numbers are stunning: between 1991 and 2000, the United States saw a 44 percent drop in homicides; a 47 percent decline in robberies; and 42 percent fewer burglaries. The numbers held true across the board for most locations and most perpetrator age groups. These rates stabilized from 2001 through 2008, but then dropped once again in 2009 and 2010.

The experts don’t know why. In fact, prior to the 1990s, many predicted a coming crime upswing fed by the growth in the number of young men. Quite unexpectedly, crime by youths fell beginning in 1993, despite predictions by eminent social scientists such as James Wilson, who anticipated more “muggers, killers and thieves than we have now,” and John DiIulio, who predicted “approximately 270,000 [more] juvenile superpredators on the streets” by 2010.

Social scientists suggest several explanations for the sudden crime

155

decline, but none fully explains the crime drop. Steven Levitt (author of Freakonomics) and collaborator John Donohue put forward the controversial and widely reported claim that the availability of abortions beginning one generation earlier may have reduced the number of unwanted children born to low-income mothers. But the timing of additional abortions does not quite fit the later crime dip. Criminologist Alfred Blumstein points to the close correlation between youth violence and the rise and fall of crack cocaine. However, the theory cannot explain the drop in property crime, nor the decline in crime by adults that began before crack was introduced. Similarly, theories about the effect of bad economic conditions on crime are undercut by the fall-off in adult crime prior to the 1990s economic boom and are entirely confounded by reductions in robberies and all violent crime in 2009 and 2010, despite the most severe U.S. recession since the 1930s. Even the number of people incarcerated did not have the widely expected effect on crime. Crime increased in the late 1980s, when the number of convicted criminals behind bars increased the most, and then crime fell in 2009 and 2010 as the percentage of convicts in prison declined.

The most often-cited explanation for the drop in crime is policing efforts. New York City is a key player in this particular explanation: Policing practice was changed and crime dropped precipitously—overall double the national decline—so that homicide, robbery, and burglary all fell by more than 80 percent. As a result, New York City’s crime rates dropped far below the levels of typical U.S. cities. However, cities without new policing practices also experienced less crime, just not as much of a decline as in New York City. Furthermore, if New York City policing did have an effect, criminologists have not been able to disentangle the effect of more police versus new strategies such as low tolerance for minor crimes. Nor are criminologists certain if New York City policies would be applicable to other cities that didn’t start with extraordinarily high crime rates, or to cities without New York’s population density and strict gun laws.

In social science research it is rare to find a situation in which change is so sudden and unambiguous as in the lower U.S. crime rate after 1995. More typically, researchers must try to tease out cause and effect from tiny changes over time. In this case, however, crime fell by nearly one-half across a wide variety of offenses, although the decline was asymmetric in terms of the offender age, falling first for adults and only later for youth. Tempting single answers—abortion availability, demographic changes,

156

policing innovations, new drugs, or altered economic conditions—work only if data is used selectively, ignoring other time periods, locations, and age groups that don’t fit the hypothesis. Researchers need to make certain that analyses are robust across these variables. It is likely that the reasons for widespread and long-lived changes in crime will require multiple explanations.

Are There More Female Criminals?

Media reports of “women gone wild” grossly exaggerate crimes perpetrated by women, which remain a small fraction of crimes committed by men. For example, women commit fewer than one-tenth of U.S. murders. Nonetheless, the gender crime gap has changed in ways that demonstrate the complexity of crime statistics. The ratio of offenses committed by women compared to those committed by men has increased, while at the same time the number of women offenders has decreased since the mid-1990s—just not as quickly as the decline in the number of male offenders.

A 2009 study by criminologists Lauritsen, Heimer, and Lynch using NCVS found more women were involved in single-offender incidents, suggesting that the increase was not simply increased female collaboration as secondary participants in male-led crimes. Also, the increase was found for robberies and aggravated assaults, indicating the trend applied to serious as well as minor crimes.

The limited understanding on the male-female crime gap has not prevented researchers from suggesting underlying causes. One hypothesis is that more equal gender roles has given women the freedom to commit more crimes. However, data show that the greatest relative increase in female crime occurred well after the greatest changes in women’s rights and increased female labor participation. As an alternative, other researchers attribute increased crime by women to the post-1990 reduction in welfare assistance that placed greater economic pressures on women. In any case, much more research is needed in order to test these hypotheses.

Box 6.1 Is New York City Telling the Truth?

If allegations are true that New York City police officers fudged

157

reports so that crime appeared lower in their precinct, can we believe the claims for such dramatic decline in overall crime? County health records keeping track of all deaths by cause show that the murder rate indeed fell just as the police reports show. Auto theft also allows for an independent check through insurance company records. Criminologist Franklin E. Zimring found that two separate insurance industry data bureaus measured auto theft claims fell by nearly precisely the same 94 percent decline from 1990 to 2010. Thus it appears that New York’s crime drop was real.

Sources: Franklin E. Zimring, “How New York Beat Crime,” Scientific American, August 2011, p. 77.

Human Trafficking: How Often Does It Occur?

U.S. government efforts to fight trafficking in persons, often called modern-day slavery, increased dramatically after a 1999 CIA report that an estimated 45,000 to 50,000 women and children are trafficked annually to the United States. The number was based on a CIA survey of news clippings, extrapolated to estimate the number of victims. Such estimating methods are necessary in the absence of direct counts. But in this case, the estimates were so uncertain that the CIA actually lowered its estimate in 2004 to 14,500 to 17,500 each year, a number that former U.S. Attorney General Alberto Gonzales admitted might still be an overstatement. Meanwhile, based on the initial estimates, Congress funded new programs to encourage law enforcement to take a more proactive stance against human trafficking on the grounds that victims were unlikely to come forward voluntarily. However, efforts to identify actual victims were relatively unsuccessful, and government grants to outreach agencies yielded few results. In 2007 there were 148 cases filed by the Justice Department’s Civil Rights Division and 1,362 total victims identified, calling into question the originally frightening statistics.

At the same time, the New York Times, USA Today, CNN, and CSPAN reported that there were 100,000 to 300,000 U.S. child prostitutes, estimates that proved even more unwarranted than the human trafficking numbers. The U.S. child prostitute number was based entirely on a study by University of Pennsylvania School of Social Work researchers who counted children at risk for prostitution because they were runaways, traveling to Mexico, gang members, or transgender. The publications and

158

the Hollywood celebrities campaigning against child prostitution failed to include the key words “at risk,” implying instead that many hundred thousand children actually were prostitutes. A follow-up study by Village Voice reporters counted 827 child prostitute arrests in the 37 largest U.S. cities over a ten year period, and U.S. social service programs identified 248 children involved in sex trafficking between January 2008 and June 2010. While certainly not a full count of child prostitutes, these lower numbers suggest that the estimate of 100,000 to 300,000 child prostitutes cannot be substantiated.

Box 6.2 Misleading Boundaries

Using boundaries for the estimate of human trafficking at 45,000 to 50,000 in 1999 and 14,500 to 17,500 in 2004, the Central Intelligence Agency (CIA) gave an unwarranted sense of reliability to the numbers. When such ranges are reported, it usually means that a statistician used a random sample for which we should expect the reported variable to be within the reported range 90 or 95 percent of the time. By contrast, for the human trafficking data based on extrapolation from newspaper accounts, there was no standard random sampling error, so the ranges do not mean that the actual value likely lies between the two numbers. Indeed, subsequent downward revisions indicate that the range was quite misleading and had not been reached through standard statistical practice.

Sources: CIA report in April Riegler, “Missing the Mark: Why the Trafficking Victims Projection Act Fails to Project Sex Trafficking Victims in the United States,” Harvard Journal of Law and Gender 30 (2007): 233.

Reports of human trafficking and child prostitution prompted a frightened public to take note of what everyone agreed were egregious crimes. However, by exaggerating the problem, researchers may have caused policymakers to divert funds from preventing more prevalent crimes

Where Is Crime the Worst?

159

Despite cautions from the FBI that its national crime data should not be used to compare the effectiveness of law enforcement in different localities, crime rankings continue to be published. According to the website Crime in America (crimeinamerica.net), the term crime rankings is one of the most popular Internet crime-related searches. The most commonly cited rankings are published by CQ Press, justified by them as useful for law enforcement agencies, city governments, grant funding, and the media to compare crime rates across cities and years. Topping the list in CQ Press’s City Crime Rankings 2009–2010: Crime in Metropolitan America were heavily urbanized Camden, New Jersey; St. Louis, Missouri; Oakland, California; and Detroit, Michigan. In the book’s purpose statement, the publisher notes that crime rates are higher in these central cities because they have more victims and targets than the surrounding metropolitan area.

In direct response to such rankings, criminologists working for the Improving Crime Data project have ranked city homicide adjusted for socioeconomic factors. As they explain, “the model produces a more meaningful comparison of city homicide levels, especially for proving insight into the effectiveness of criminal justice policies and programs.” With these adjustments, some rankings nearly reverse: Detroit and Cleveland, numbers 2 and 8 respectively in the unadjusted rank, fell to 53 and 56 out of 60 cities. At the same time, Albuquerque, New Mexico, rose from 44 to 11, and Colorado Springs, Colorado, went from 52 to 22.

Similar adjustments are made to school test scores (see Chapter 5), an attempt to take into account social and economic conditions that cause children to score differently and then to measure changes in these adjusted scores over time. As in the case of test scores, there is no single best way to adjust crime statistics for social and economic conditions; the result depends always on the weight assigned to each explanatory variable. The Improving Crime Data researchers corrected for poverty, income, unemployment, race, and female-headed households. A different set of factors, or these same ones weighted in a different manner, would have yielded different results (see similar effects in Chapter 3 on ranking the best place to live).

Investigators need to choose adjusted or unadjusted statistics depending on their research purpose. In order to determine the impact of crime on people’s lives or the need for intervention to reduce crime, the raw data may be more instructive. However, in order to estimate the effectiveness of law enforcement agencies, the adjusted data are helpful because they

160

take into account social and economic factors that cause one agency to face a more difficult task. In either case, researchers should be aware of one additional variable that affects both statistics: Has the underlying population been counted accurately? As described in Chapter 2, the U.S. Census is adjusted for undercounts, particularly in urban areas. Sometimes this adjustment overcorrects, as in the case of New Orleans when the 2010 Census showed the 2008 adjustment to be almost 100,000 too high. Without the overcorrection, the homicide rate proved to be substantially higher—the highest, in fact, among large U.S. cities. With the overcorrection, the homicide rate appeared much lower, masking New Orleans’s problem with an overestimated population, the denominator in the homicide rate.

Rape

Most social scientists agree that rape is a difficult crime to measure, both because of ambiguity in what constitutes rape and because of the difficulty in gathering data from rape victims. The UCR statistics indicate that rape has followed national violent crime rates: falling after 1993 through 1999, then rising slightly and falling again in 2008 and 2009. However, these rates omit rapes not reported to the police and include only male against female vaginal rape, a limitation changed beginning in 2012.

During the 1990s estimates for the level of college student rape ranged from 15 percent of the female student population by psychologist Mary Koss to as low as 0.1 percent by social welfare professor Neil Gilbert, who complained that Koss and others tilted their numbers to fit preconceived “feminist prescriptions.” A 2000 U.S. Department of Justice study found a middle estimate by interviewing 4,446 college women, measuring sexual victimization based on information about specific behaviors for a specific time period. By using graphically worded questions, the study revealed a 2.8 percent victimization rate in the previous seven months, 11 times higher than in the NCVS for completed rape and six times higher for attempted rape. Significantly, only 46.5 percent of the interviewed women answered “yes” to the question: “Do you consider this incident to be a rape?” Furthermore, fewer than 5 percent of the incidents were reported to police, even though rape was defined as “unwanted completed penetration by force or threat of force.”

In January 2012 the FBI announced that the UCR would change the definition of “forcible rape” to include attacks on men as well as women,

161

expand what is counted as “penetration,” and designate an act as “nonconsensual” even if no overt physical force was involved. Although some police departments had already used an expanded rape definition, it will take years before all state and local law enforcement agencies comply with the FBI reporting requirement. As an aid to researchers during the transition, the FBI promises to ask for dual reporting using both definitions. However, omissions have already occurred, as in 2010 when Chicago data were excluded from the FBI’s 2010 Uniform Crime Report because they used an expanded definition of rape before its official 2012 adoption.

Even with the new definition, approved statistics all undercount sexual crimes significantly, making it important for researchers to recognize that a serious problem is being understated. A National Institute of Justice and Centers for Disease Control and Prevention study known as the National Violence Against Women Survey found that 17.6 percent of women reported they had been victim of a completed or attempted rape at some point in their lives, 21.6 percent of which occurred before the victim reached the age of 12.

Does Poverty Cause Crime?

It seems obvious that crime is associated with poverty. Basic street sense tells us that poor neighborhoods are more dangerous than wealthy neighborhoods. Superficially, at least, the data support such generalizations. Compton, California, a high poverty section of inner-city Los Angeles, suffered 36 murders in 2009 according to the UCR, while the upscale Los Angeles suburb of Mission Viejo, with a similar population, had no reported murders. Robberies were also disproportionate: 509 in Compton; 35 in Mission Viejo.

Nonetheless, some social scientists are uncertain about how poverty is related to crime. The problem is that grouped data such as the comparison between Compton and Mission Viejo do not necessarily imply individual differences in the propensity to commit crimes by residents in those neighborhoods. In other words, just because crime is correlated with poverty at the group level (crime is high in poor neighborhoods), it does not necessarily follow that crime is correlated with poverty at the individual level (poor individuals commit more crimes).

Sociologist C.S. Tittle and other researchers argue that low-income Compton residents may be no more likely to be criminals than high-

162

income residents of Mission Viejo. Tittle rejects grouped data studies for perpetuating the “myth of social class and criminality” and the “prejudice that lower class people are characterized by pejorative traits such as immorality, inferiority and criminality.” In place of grouped data, Tittle suggests we look at self-report studies in which sociologists ask individuals about their criminal past. According to Tittle, self-report studies show practically no association between social class and criminal activity.

If economic circumstances are not responsible for crime, then income programs are unlikely to lower the crime rate. But those who support these economic solutions to crime take issue with Tittle’s reliance on self-report studies. The rebuttal does not focus on the obvious drawback of obtaining accurate self-incriminating data; instead, self-report studies are criticized for failing to take into account the seriousness of crimes. It is only by including very minor thefts as crimes that middle-class and upper-class youths have as high a crime rate as lower-class youths. Or, as sociologist and criminology professor Elliott Currie mocked, we learned that “American youths of all backgrounds sometimes acted up.”

A second debate looks at how changes in poverty affect crime. The continued crime decline during the 2007–2009 recession surprised some criminologists, such as Richard Rosenfeld, who had predicted property crime increases on the heels of the worst economic downturn since the Great Depression. Manhattan Institute attorney Heather MacDonald argues that this new evidence demolishes “the idea that the root cause of crime lies in income inequality and social injustice.” Rosenfeld admits that the data are a “real break in the pattern from past relationships between economic downturns and crime increases.”

Overall, this debate illustrates the limitation of research on changes as a way to understand the underlying relationship between two variables, in this case poverty and crime. The already complex connection between income levels and crime cannot be determined easily, because other factors occurring during the recession could account for the drop in crime —among them decreased mobility that leads to neighborhood stability and thus to less crime, or simply the decline in consumption that gives thieves less to steal.

Why Is the Black Crime Rate So High?

The Uniform Crime Reports tabulate criminal offenders by sex, age, and

163

race. The most obvious observation is that men make up the largest percentage of offenders, about 80 percent for violent crimes. Roughly 60 percent of all violent crime offenders are under the age of 30. However, it is the disproportionate number of young black men in the criminal justice system—blacks constitute more than 30 percent of those arrested—that has provoked heated research debate. Part of the difference in crime rates between racial groups depends on what we define as crime. Whites are arrested more frequently for white-collar offenses, a type of crime with ramifications that are larger in dollar amount than other crimes (see next section) but not included in the UCR. Also, for one of the most common crimes of all, drunk driving, the arrest rate for blacks is lower than for whites.

Overall, however, blacks constitute about 39 percent of those in prison, several times the percentage of blacks in the U.S. population. Sociologist Andrew Hacker ascribes part of this difference to the discriminatory access to crime; that is, black offenders often resort to street robberies, which almost always result in arrest, whereas white offenders are able to operate freely in better-off neighborhoods where they can commit burglaries, a far more profitable and less hazardous occupation. Bruce Wright, a former New York State Supreme Court judge, adds that police discrimination causes blacks to be arrested more frequently, and discrimination by the judicial system causes blacks to receive more frequent and longer prison sentences than nonblack offenders.

Criminologist Elliott Currie responds that outright bias explains only a small part of the difference in crime rates for blacks and for whites. As evidence, Currie cites homicide data available from the Public Health Service showing that black men from age 25 to 44 are eight times as likely to be murdered as their white age-mates. Because other studies show that most murders are intraracial (black against black), it seems likely that the higher homicide rate for blacks is real and not a result of police or judicial bias. Currie maintains that the eagerness by some researchers to dismiss race as an important factor in crime has distracted policymakers from the task of accounting for what Currie calls the “genuine social disaster wrought by extremes of economic inequality we have tolerated in the United States.”

The issue of race and crime has a long, controversial history. Statistics from the Justice Department’s Bureau of Justice Statistics show a grisly past in which blacks have constituted about one-half of those executed by the legal system, including many for crimes other than murder.

164

Researchers need to be aware of the issues raised by Hacker and Wright that cause bias in official statistics, as well as of the complex social factors that Currie urges us to address in understanding the relationship between race and crime.

Does Prison Pay?

It costs about $25,000 per year to keep a prisoner behind bars. Does society save an equal amount in crimes not committed because perpetrators are incarcerated? At issue is the dramatic increase in U.S. incarceration rates beginning around 1980, so that the number of U.S. men between the ages of 18 and 64 in prison or jail jumped to 2,009,512 in 2008 from 315,258 in 1980. As a result, in 2008 over 2 percent of all men between the ages of 18 and 64 were in prison or jail. Of the 30 most developed countries in the world, not one imprisons even one-third as many residents. Poland is second at 224 per 100,000 population; Iceland is lowest at 44 per 100,000, 17 times less than the U.S. rate.

National Institute of Justice economist Edwin Zedlewski argues that imprisonment is cost effective because the typical offender commits 187 crimes a year, costing $2,300 per crime, including property loss and victims’ pain and suffering, for a total of $430,000 in benefits by keeping the criminal in prison. His study was widely quoted by advocates of more prisons, such as those mentioned in the Reader’s Digest article “Why Don’t We Have the Prisons We Need?” which concluded that “to pen every serious offender will cost billions, but it’s money well spent.”

Criminologists Franklin E. Zimring and Gordon Hawkins disagree with Zedlewski’s findings. Most criminals are arrested for only a small proportion of their crimes, so Zedlewski attempted to estimate the number of crimes per offender based on self-reports by prison inmates. According to Zimring and Hawkins, the actual number and cost of crimes per inmate is far lower than the estimate used by Zedlewski. Thieves and burglars commit a large number of less serious crimes so that the average number of crimes per inmate is high. But one-half of those imprisoned committed fewer than fifteen crimes, albeit perhaps ones of a more serious nature. Using this median number of crimes reduces the apparent savings of imprisonment from over $400,000 to about $10,000 per year.

Princeton University researchers John J. DiIulio and Anne Morrison Piehl take a middle position based on estimates for the cost of different crimes ranging from fraud, forgery, and petty theft at only $110 per

165

incident to robbery at $12,060 per incident. In their view, imprisonment saves society $46,000 for the typical prisoner, less than $2,000 for the 10 percent of prisoners who committed the fewest and pettiest crimes, and nearly $2 million for one exceptional individual who committed 151 robberies.

In addition to the number of crimes, other assumptions in these cost estimates affect the disputed savings from imprisonment. The social cost of each crime is based on jury awards, providing only a crude estimate of the cost of victims’ pain and suffering. And most estimates omit drug sales, deeming them “victimless crimes” that are nearly impossible to value but that cause researchers to leave out about 90 percent of crimes reported by prisoners. Because of the varying estimates, partisans can find statistics to support both cost effectiveness or, alternatively, the wastefulness of imprisonment. Careful users will be alert to the underlying data problems that complicate measurement of this important social issue.

Hate Crimes

As defined by the U.S. Congress, a hate crime is a “criminal offense against a person or property motivated in whole or in part by an offender’s bias against a race, religion, disability, ethnic origin or sexual orientation.” As might be expected from such a vague description, there is an eighteen- fold difference between hate crimes as measured in the UCR, about 8,000 reported to police in 2009, versus nearly 150,000 counted in the National Crime Victimization Survey, clearly presenting a challenge to researchers and highlighting the difficulties in determining what constitutes a hate crime. The higher NCVS hate number results from two factors. First, about one-half of hate victims do not report the incident to police because it was “handled another way” (32 percent), “not important enough” (19 percent), or “the police would not help” (15 percent). Of those incidents in which the victim claimed to have reported the hate aspect to police, only 15 percent were confirmed by police as hate crimes. The main reason is that nearly all victims, 99 percent of violent hate crimes, cited “language” as evidence that the crime was motivated by hate, often insufficient for police to confirm and report it officially as a hate crime. As the Department of Justice explains from the police perspective: “It is sometimes difficult to know with certainty whether a crime resulted from the offender’s bias,” and the “presence of bias alone does not mean that a crime can be considered a hate crime.”

166

Despite the far greater inclusiveness of victim-based NCVS numbers, researchers commonly use the new police report-based NIBRS data to study hate crimes because they include offenses against schools or religious institutions not counted in the NCVS. With the NIBRS, it is now possible to analyze differences in hate crimes over time and between cities, a breakdown that was previously impossible to determine with UCR or NCVS data. However, even with this improved source, researchers need to be careful: Local level data are not broken out by type of crime, so more serious crime trends cannot be separated from reports of relatively minor vandalism or intimidation, about one-half of hate crimes.

Data Sample: Of the hate crime victims reported in the UCR between 2003 and 2009, 52 percent were targeted because of race, 17 percent because of religion, 16 percent because of sexual orientation, 14 percent because of ethnicity, and 1 percent because of a disability. Despite its different database, the NCVS found quite similar results: 58 percent motivated by race, 15 percent by sexual orientation, 30 percent by ethnicity, 12 percent by religion, and 10 percent by disability.

Does Capital Punishment Deter Murder?

Much recent criminology research focuses on punishment as a deterrent to crime; at the center of the debate is the efficacy of the death penalty. Since the restoration of U.S. capital punishment in 1976, there have been two rounds of controversy, each time raising identical and as-yet-unresolved statistical issues. During the 1970s, economist Isaac Ehrlich estimated that eight murders were prevented by each legal execution, evidence used in a 1976 brief by the solicitor general to the U.S. Supreme Court in favor of capital punishment. In competing Supreme Court testimony, critics took issue with Ehrlich’s results, arguing that his analysis depended on several arbitrary assumptions. Any one of the following changes reduces the deterrence effect measured by Ehrlich: (1) if Vital Statistics replace Ehrlich’s FBI crime data; (2) if raw numbers are substituted for the logarithms used by Ehrlich; (3) if the years 1963 through 1969 are removed, and (4) if states that never had a penalty are studied separately.

Box 6.3 Do Sex Offenders Repeat Their Crime?

167

A recent debate between Canadian criminologists illustrates the difficulty in measuring criminal recidivism. At issue was the likelihood for sex offenders to repeat their crime. A widely distributed 2004 study found an 88.3 percent recidivism rate, whereas other researchers estimate rates as low as 24 percent.

The variation was due in part to the difficulty of defining recidivism: Is it repeating a crime more than once (as in the study finding high recidivism), or is it repeating the crime after receiving a prison sentence or treatment (as in the study finding low recidivism). Also, researchers needed to decide what constituted a repeated crime. Is it conviction once again for a sex offense, or it simply conviction for a crime such as home invasion or assault, often a plea-bargained admission when in fact the defendant intended a sexual crime as well? Finally, how should the studies treat missing records? Complete records are more likely available for those with subsequent criminal convictions, thus raising the apparent recidivism rate. Many criminologists weighed in on the debate with the consensus that recidivism is difficult to measure but is likely far higher than the 24 percent re-conviction rate might indicate.

Sources: Marnie Rice and Grant Harris, “What Population and What Question?” Canadian Journal of Criminology and Criminal Justice 48, no. 1 (January 2006): 95–101; R. Karl Hanson et al., “Long-Term Follow Up Studies Are Difficult,” Canadian Journal of Criminology and Criminal Justice 48, no. 1 (January 2006): 103–17.

An almost identical debate resurfaced in the 2000s. Data on the now more prevalent death penalty was used to find a deterrence effect ranging from three to 32 fewer murders for every execution. As in the case of Ehrlich’s study, these results have been used in legal briefs to overturn moratoria on the death penalty and expand its use in more states. Emory University law professor Joanna Shepherd testified before Congress that “studies are unanimous,” and there is a “strong consensus among economists that capital punishment deters crime.”

Because the death penalty was not reinstated in all states and in some states was used intermittently, such variation provided researchers with natural “experiments” to measure its effect. One group, primarily economists, maintained that the new studies offer overwhelming evidence

168

of deterrence, while criminologists and many legal theorists remain unconvinced; one survey of top criminologists found that 88 percent disagreed with the statement, “Do you feel that the death penalty acts as a deterrent to the commitment to murder?” Critics charge that the economists’ results showing a death penalty deterrent were not robust, since they depended on such factors as the variables included or excluded and the impact of one state, Texas, where one-third of all executions occurred. In addition, anti-death penalty advocates pointed out that even if we cannot rule out a deterrent effect, the difficulty in proving that one exists suggests that the deterrence effect is quite small compared with other factors affecting the crime rate. Other social policies, especially programs to provide jobs and boost income, have also been demonstrated to affect the crime rate. But such social policies are less well studied; as economist Richard McGahey points out, there are exhaustive studies of the death penalty, but no corresponding “spate of articles on … ‘Murder and Poverty,’ or ‘The Preventative Effect of Higher Incomes.’”

Box 6.4 Wrongful Incarceration

In February 2010 the Innocence Project noted the 250th U.S. prisoner exonerated on the basis of DNA evidence. Overall, how many convictions were wrong? For murder, studies estimate that about 1 percent of those convicted were officially exonerated, a number likely to rise as the Innocence Project continually finds new cases of wrongful imprisonment. Surprisingly, the best comprehensive data on wrongful incarceration come from inmate self-reports, which one might expect to be self-serving and exaggerate prisoners’ innocence. However, data collected by the RAND corporation found that inmates usually revealed more arrests than in official records. Using this data, Tony G. Poveda found that about 15 percent of these inmates claimed they did not commit the crime for which had been convicted (or a similar one). In a caution, however, Poveda notes that inmates’ perception of crime may be less strict than the legal definition and that the offender may also reshape his memory in unconscious ways, thus exaggerating unwarranted imprisonment.

Sources: Innocence Project at http://www.innocenceproject.org/Content/I- n_250th_DNA_Exoneration_Nationwide_New_YorkMan_Is_Proven_Inno--

169

cent_33_Years_After_Wrongful_Conviction_for_Rape.php; Tony G. Poveda, “Estimating Wrongful Convictions,” Justice Quarterly 18, no. 3 (September 2001): 689–708.

More Guns/More Crime or Less Crime?

In theory, gun ownership could deter criminals if they fear potential victims possess a gun that could be used defensively. Alternatively, gun ownership could increase crime if disputes are more likely to involve a lethal weapon and therefore escalate to cause injury or death. Social science debate over these competing hypotheses has been intense, fueled by strong feelings about the appropriateness of gun control laws.

One problem is the absence of good data on the number of guns and who owns them. Overall, there is a consensus that U.S. households own nearly 300 million guns, a number that has grown in recent years. (Gun ownership is likely becoming more concentrated, so some households own many guns while about one-half own no guns at all.) However, in the absence of gun registration in most localities, combined with illegal ownership, the data on gun ownership is not sufficiently detailed to correlate it with crime.

Instead, researchers rely on indirect measurement. During the 1990s debate focused on survey data in which individuals were asked about “defensive gun use (DGU).” A study by Gary Kleck and Marc Gertz used telephone surveys to estimate more than 2 million DGUs annually, many times higher than the 80,000 uses reported through the NCVS. According to Kleck and Gertz, defensive gun use has “saved lives, prevented injuries, thwarted rape attempts, driven off burglars and helped victims retain their property.” Critics charged that the telephone survey was faulty because respondents may have telescoped prior events inappropriately into the time period, thus overstating the number of gun uses. Also, respondents may have counted gun usage in response to minor transgressions such as trespassing, thus exaggerating the number of crimes prevented.

More recently, the debate has shifted to research on the correlation between crime prevention and legislation allowing gun owners to carry concealed weapons (CCW). In an influential study titled More Guns, Less Crime by economist John R. Lott, Jr., such legislation appeared to cause a significant decline in violent crime. In direct response, economist Mark Duggan wrote “More Guns, More Crime,” in which he measured gun use

170

by sales data for Guns&Ammo magazine, an indirect, but in his view accurate, measure of gun ownership. He found that the number of murders with guns was significantly increased by more gun ownership at the state and county level. Lott disputed Duggan’s results on the basis that the Guns&Ammo publisher tried to boost sales by distributing free copies in areas where crime rates were rising, thus explaining the apparent relationship between murder and magazine circulations.

Other researchers have disputed Lott’s correlation of legalized concealed weapons and the crime decline on other grounds. For example, the results were too strong: Why would small increases in concealed gun permits such as 2 percent in Florida cause a dramatic 8 percent drop in murders? Also, robberies, the crime one might expect to be deterred most often by defensive gun use, was less affected than murder or rape. Even if Lott’s data could stand up to these criticisms—and he has written three editions of More Guns, Less Crime responding to them—his policy conclusion is limited to the relation-ship between new concealed weapon laws and crime. Lott’s results may apply only if one accepts the status quo of a gun-prevalent society. It may also be true that reducing the number of guns would be an even more effective policy than allowing concealed weapons. Nonetheless, Lott’s research had a major impact on gun laws, so that the number of U.S. states with right-to-carry weapon laws increased to 39 in 2010, up from 18 in 1992.

Box 6.5 Are You More Likely to Be Murdered by an Acquaintance Than a Stranger?

Are murders committed most often by friends and acquaintances? Uniform Crime Reports data suggest that only 13 percent of murders were committed by complete strangers, 18 percent by family members, 40 percent by someone “known” to the victim (and 30 percent of the cases had an attacker with undetermined relationship to the victim). These data are used to discredit the defensive role of guns and to show that fewer guns would reduce the opportunity for crimes of passion by family members or friends. In arguing for less gun control, John Lott points out that victims “known” to the assailant include gang members and drug sale participants, for whom gun murders were more likely premeditated rather than passion

171

induced. Others maintain that greater numbers of guns and easier availability are also factors in homicides by gang and drug sale “acquaintances,” even if they were premeditated.

Source: John R. Lott, More Guns, Less Crime 3d ed. (Chicago: University of Chicago Press, 2010), pp. 2–5, 152–54.

Box 6.6 Who Pays for Research on Guns?

Some critics of John Lott’s conclusions about guns and crime prevention, charged a conflict of interest in Lott’s position as Olin Fellow at the University of Chicago. Funded by the John M. Olin Foundation, the post had been created many years earlier by the founder of a chemical and munitions manufacturer. Direct influence on Lott’s research is unlikely; the Olin Foundation supports a number of other conservative-leaning researchers without regard to guns or munitions. More troubling is pressure on U.S. government researchers by the National Rifle Association (NRA) that led to the cutoff for funding needed to answer basic questions about the role of guns in the United States. When the U.S. Centers for Disease Control and Prevention’s (CDC) National Center for Injury Control and Prevention planned to fund firearms research, the NRA promoted language in the center’s congressional appropriation that funds could not be used to “advocate or promote gun control.” Even though the prohibition appeared to apply to congressional lobbying, something already disallowed, it was sufficient to dissuade CDC requests for firearms research.

Source: John R. Lott, More Guns, Less Crime, 3d ed. (Chicago: University of Chicago Press, 2010), pp. 127–28, 202–203; NRA influence in Michael Luo, “Sway of NRA Blocks Studies, Scientists Say,” New York Times, January 26, 2011, p. A1.

What About White-Collar Crime?

172

In 1939 Edwin H. Sutherland used the occasion of his American Sociological Society presidential address to urge study of what he termed white-collar crime. Sutherland suggested that criminologists previously ignored “upper-world” crime committed in the course of an occupation, offenses he claimed were as important as traditionally defined crime. Although sociologists sometimes define white-collar crime by the status of the offender, for most research purposes it is defined by the type of offense. But which types should be included, and what can be done about white-collar crimes for which little data is available?

Although the FBI defines white-collar crime as “illegal acts of deceit, concealment or violation of trust … which are not dependent upon the application of threat of physical force of violence,” UCR data are in fact quite limited in scope, counting only fraud, forgery, counterfeiting, and embezzlement in the white-collar category. This shortcoming will be somewhat improved by the NIBRS, which includes bribery and bad checks, and also breaks down fraud into categories such as credit card scams and impersonation. Moreover, the NIBRS has more information about the offender and victim, including whether or not a computer was used in the commission of a crime. Even with improved NIBRS data, most white-collar crime is still not included. A 2005 survey by the National White Collar Crime Center found that almost half of households were victims of white-collar crime in the previous year, but fewer than one in 10 reported the crime to police.

Far more difficult to measure are crimes involving violation of environmental, financial, antitrust, or other regulatory laws. It is rare for criminal charges to be filed in such cases; instead, they are handled as civil cases or enforced by regulatory agencies or professional associations. Even if there are criminal charges, they are likely counted in the UCR and NIBRS “all other offenses” category, making them indistinguishable from non-white-collar crimes.

The dollar value of white-collar crime is extremely difficult to quantify. In NIBRS reports, white-collar crime averages a loss of only $10,000, while the median (one-half of the cases above a certain point and one-half below) is about $200. The data are skewed by a few crimes with substantial losses, whereas most losses are low (the mode is $100) and many have no measured loss because they involved items such as yet- unused credit cards with no “fair market value.”

Using a broader definition of white-collar crime to include unprosecuted

173

employee theft and other incidents not counted in the NIBRS, costs estimates range from $250 billion to over $1 trillion. If white-collar crime is defined even more broadly to include malfeasance such as unlawful acts during construction, then the truly extraordinary multi-hundred billion- dollar devastation caused by Hurricane Katrina could count as a white- collar crime. More useful for research purposes than these arbitrary and wildly varying efforts to measure total white-collar crime are estimates for the cost of particular problems. For example, one study by Joseph Eaton and David Eaton on price fixing and kickbacks in the mortgage insurance industry measured the cost to homebuyers at several billion dollars.

At stake in the definition of white-collar crime is the allocation of crime prevention resources. Criminologists Harold Pepinsky and Paul Jesilow point out that even the lowest estimates for the value of white-collar crime are still about ten times as high as the total property loss in traditionally defined crime. But crime prevention funds are allocated in just the reverse proportions, with most money spent dealing with non-white-collar criminals. Philosopher David Reiman argues that rational allocation would increase policing of operating rooms and dangerous workplaces, where four times more people are killed than in traditionally defined murders. Unfortunately, federal support for research on white-collar crime has not always been strong, declining in particular during conservative political administrations.

Summary

Statistical research on crime is relatively complete and is typically published in a timely way. U.S. crime statistics have changed dramatically and in a positive direction. The crime drop, nearly one-half for homicides and over 90 percent for auto thefts, cries out for explanation. Nonetheless, this chapter has illustrated several major problems in interpreting these crime statistics and their implications for policymakers.

First, crime statistics are unusual in that the same organizations— individual police departments and the FBI—are responsible for carrying out public policy as well as collecting and publishing the most important source of crime data, the Uniform Crime Reports. This conflict of interest sometimes causes inaccuracies in the data, as in the case of police department misrepresentations to the UCR. A different sort of survey error affects the National Crime Victimization Survey, in which respondents are not necessarily forthcoming with accurate answers about their experience

174

with crime. Overreporting of crime occurs when respondents exaggerate their own experiences or telescope distant events into the past year being surveyed. Underreporting appears to be the more serious problem, including a severe undercount of the number of rapes, as well as underreporting for other crimes caused by noncooperation or lack of knowledge about crimes involving other household members. All of these measurement problems are well studied. Even if there is no method to “correct” the underlying data for under-or overreporting, the criminal justice literature provides ample evidence of the direction in which survey errors are likely to lie.

Second, there is disagreement in the study of crime about the fundamental question, “What is crime?” The most commonly used crime data, the UCR, include only crimes reported to police, a number that most certainly is much lower than the number of crimes committed. A more complete count of crime is available in the National Crime Victimization Survey, but it omits categories such as victimless crime and white-collar crime. This last category is probably the most controversial of all. There is no standard definition for what constitutes white-collar crime, and there are only imprecise estimates about its dollar value. Similarly, the crime of rape is subject to varying definitions, which has led to a bitter dispute about the prevalence of rape victimization. Because alternative definitions are, by necessity, somewhat arbitrary, most research projects rely on the “official” crime definitions. Nonetheless, careful researchers will note these measurement issues and speculate about the effect they may have on the study of crime. For example, conclusions about the cost of crime and the relationship between race and crime depend critically on how crime is defined.

Third, as with other social statistics, crime statistics are only as meaningful as the skill of the researcher who uses them. More problematic is the use of crime statistics in a selective manner. The issue arose in several critical matters of social policy, including the relationship between race and crime, the impact of the death penalty, and the role of gun control. In all cases, critics charged that much standard crime analysis focuses too narrowly on race and the death penalty in isolation, ignoring other social and economic variables that are more important for understanding crime. The choice of variables for analysis crosses over into theoretical issues of criminology that are beyond the scope of this book. At a minimum, however, researchers must be clear about how they have made such choices, making sure their selection of variables has a sound

175

1.

2.

3.

4.

theoretical basis.

Case Study Questions

In 2004 the U.S. Department of Justice dropped its “Crime Index,” published since 1960 and a widely used measure of crime based on a total of eight of the most serious criminal offenses. Why was this a potentially misleading guide to the severity of crime?

Mr. Smith, who profits from “inside” knowledge about the stock market, is robbed by three youths. How many crimes have been committed according to the Uniform Crime Reports? According to the National Crime Victimization Survey? In the National Incident-Based Reporting System?

Most criminologists believe that apprehension has a deterrence effect, especially for “amateur” crimes. Yet a study of shoplifting by high school students found that apprehended offenders were more likely to repeat the crime than those who eluded detection. Why was there no apparent deterrence effect? (Hint: Remember that there is wide variation in the frequency of shoplifting by individuals.)

Comparison of the National Crime Victimization Survey with police data on nonfatal gunshot injuries suggests that the survey undercounts these injuries by a factor of three. Why might this occur?

References

Data Sources (111)

Data sample in 2010 Internet Crime Report, National White Collar Crime Center, http://www.ic3.gov/media/annualreport/2010_IC3Report.pdf p. 23.

Uniform Crime Reports (112)

Uniform Crime Reports (UCR) described in U.S. Department of Justice, Federal Bureau of Investigation, “Uniform Crime Reports,” http://www.fbi.gov/about- us/cjis/ucr/ucr. Atlanta in “Manipulation of Crime Figures Alleged,” Atlanta Constitution, May 21, 1998, p. E1; Philadelphia in “As Crime Falls, Pressure Rises to Alter Data,” New York Times, August 3, 1992, p. A1; New York City in Ray

176

Rivera and Al Baker, “As Rates for Major Crimes Fall, Data on Minor Ones Stay Under Radar,” New York Times, November 2, 2010, p. A22. Pop-up at http://www.fbi.gov/about-us/cjis/ucr/crime-in-the-u.s/2010/crime-in-the- u.s.-2010/caution-against-ranking. Data sample in U.S. Department of Justice, Federal Bureau of Investigation, “Uniform Crime Reports,” Crime in the United States: 2009,http://www2.fbi.gov/ucr/cius2009/data/table_09_ri.html.

National Crime Victimization Survey (113)

NCVS described in U.S. Department of Justice, Bureau of Justice Statistics, http:// bjs.ojp.usdoj.gov/index.cfm?ty=dcdetail&iid=245. “Police-Public Contact” in U.S. Department of Justice, Bureau of Justice Statistics, http://bjs.ojp.usdoj.gov/index.cf m?ty=dcdetail&iid=251#Methodology. Problems with NCVS in A.D. Biderman and J.P. Lynch, Understanding Crime Incidence Statistics (New York: Springer- Verlag, 1991). Data sample in U.S. Department of Justice, “Criminal Victimization in the United States,” 2008 Statistical Tables, Table 99; ibid., p. 107.

National Incident-Based Reporting System (114)

NIBRS described in Inter-University Consortium for Political and Social Research, “National Incident-Based Reporting System Resource Guide,” http://www.icpsr.umich.edu/icpsrweb/NACJD/NIBRS/; see also Federal Bureau of Investigation, “NIBRS General FAQs,” http://www.fbi.gov/about- us/cjis/ucr/frequently-asked-questions/nibrs_faqs. NCVS missed gun shootings in Michael G. Maxfield, “The National Incident-Based Reporting System: Research and Policy Applications,” Journal of Quantitative Criminology 15, no. 2 (1999): 130. Hoboken v. Jersey City auto theft in Maxfield, pp. 126–27. Data sample in Jessie L. Krienert, Jeffrey A. Walsh, and Moriah Turner, “Elderly in America: A Descriptive Study of Elder Abuse Examining National Incident-Based Reporting System (NIBRS) Data, 2000–2005,” Journal of Elder Abuse and Neglect 21 (2009): 325–45.

Controversies (115)

UCR, NCVS, or NIBRS? (115)

Comparison of UCR and National Crime Survey in A.D. Biderman and J.P. Lynch, Understanding Crime Incidence Statistics (New York: Springer-Verlag, 1991). UCR reporting improvements in Christopher Jencks, “Is Violent Crime Increasing?” American Prospect, Winter 1991, pp. 98–109. Misrepresentation of crime in New York City in Marvin E. Wolfgang, “Uniform Crime Reports: A Critical Reappraisal,” in Crime in America, ed. Bruce J. Cohen (Itasca, IL: F.E. Peacock, 1985), p. 41. Washington, DC, in James P. Levine, Michael C. Musheno, and Dennis J. Palumbo, Criminal Justice in America (New York: Wiley, 1986), p.

177

99; Atlanta in “Manipulation of Crime Figures Alleged,” Atlanta Constitution, May 21, 1998, p. E1; Philadelphia in “As Crime Falls, Pressure Rises to Alter Data,” New York Times, August 3, 1992, p. A1. Future reporting systems in “Agencies Are Trying Better Crime-Data System,” William G. Lopez, Letter to the Editor, New York Times, August 10, 1998, p. A18. See also Biderman and Lynch, Understanding Crime Incidence Statistics. FBI using traditional format and possible rise in number of reported offenses in “NIBRS General FAQs,” http://www.fbi.gov/about-us/cjis/ucr/frequently-asked-questions/nibrs_faqs. Higher numbers in NIBRS in Bill Novaki, “New Crime Reporting System More Detailed,” Wisconsin State Journal, December 21, 2010, p. 5.

Crime Is Down—And We Don’t Know Why (116)

Drop in crime in Richard Rosenfeld, “The Case of the Unsolved Crime Decline,” Scientific American, February 2004, pp. 80–89; New York City in Franklin E. Zimring, “How New York Beat Crime,” Scientific American, August 2011, pp. 75– 79.

Are There More Female Criminals? (117)

Janet Lauritsen, Karen Heimer, and James P. Lynch, “Trends in the Gender Gap in Violent Offending: New Evidence from the National Crime Victimization Survey,” Criminology 47, no. 2 (May 2009): 361–99; “Women gone wild,” ibid.; more equal roles and reduction in welfare assistance, ibid., 385–92. See also Jennifer Schwartz, Darrell J. Steffensmeier, and Ben Feldmeyer, “Assessing Trends in Women’s Violence via Data Triangulation: Arrests, Convictions, Incarcerations, and Victim Reports,” Social Problems 56, no. 3 (2009): 494–525.

Human Trafficking: How Often Does It Occur? (118)

CIA report in April Rieger, “Missing the Mark: Why the Trafficking Victims Protection Act Fails to Project Sex Trafficking Victims in the United States,” Harvard Journal of Law and Gender 30 (2007): 233. Criticism of estimates in Martin Cizmar, Ellis Conklin, and Kristen Hinman, “Real Men Get Their Facts Straight: Ashton and Demi and Sex Trafficking,” The Village Voice, June 29, 2011, http://www.vil-lagevoice.com/content/printVersion/2651144/. Attorney General Gonzales in “U.S. Estimates Thousands of Victims, Efforts to Find Them Fall Short,” September 24, 2007, humantrafficking.org. Campaign against child prostitution in Cizmar, Conklin, and Hinman, “Real Men.”

Where Is Crime the Worst? (119)

FBI cautions at http://www.fbi.gov/about-us/cjis/ucr/crime-in-the-u.s/2010/crime- in-the-u.s.-2010/caution-against-ranking. Internet searches at

178

http://crimeinamerica. net/2010/02/09/crime-rankings-for-cities-a-fair-comparison- crime-statistics-2/. Crime rankings in Kathleen O’Leary Morgan and Scott Morgan, eds., City Crime Rankings 2009–2010: Crime in Metropolitan America (Washington, DC: CQ Press, 2009); publisher’s press release dated November 23, 2009, available at http://os.cqpress.com/citycrime/CQPress_CityCrime2009_PressRelease.pdf. Criticism in Carl Bialik, “In Crime Lists, Nuance Is a Victim,” Wall Street Journal, December 4, 2010, p. A2. Improving Crime Data (ICD) project in “City Homicide Rankings Adjusted for Differences in Socio-Economic Factors,” Georgia State University, http://www.cjgsu.net/initiatives/HomRates-PR-2009–03-16.htm. New Orleans adjustment in Mark J. Van Landingham, “Making Murder Count,” New York Times, July 16, 2011, p. A19.

Rape (121)

UCR rape statistics in Mary Koss, “Rape on Campus: Facts and Measures,” Planning for Higher Education 2 (1992): 21–28; Mary Koss, Christine Gidycz, and Nadine Wisniewski, “The Scope of Rape,” Journal of Consulting and Criminal Psychology 55 no. 2 (1987): 162–70. Neil Gilbert, “Miscounting Social Ills,” Society, March/April 1994, p. 26, and Gilbert, “The Phantom Epidemic of Sexual Assault,” The Public Interest, Spring 1991, pp. 54–66. Department of Justice data in Bonnie S. Fisher, Francis T. Cullen, and Michael G. Turner, “The Sexual Victimization of College Women” (research report), U.S. Department of Justice, National Institute of Justice, Bureau of Justice Statistics, December 2000, https://www.ncjrs.gov/pdffiles1/nij/182369.pdf. UCR changes in Charles Savage, “U.S. to Expand Its Definition of Rape in Statistics,” New York Times, January 7, 2012, p. A10. National Violence Against Women Survey in Patricia Tjaden and Nancy Thoennes, “Full Report of the Prevalence, Incidence, and Consequences of Violence Against Women,” National Institute of Justice/Centers for Disease Control and Prevention, November 2000, https://www.ncjrs.gov/pdffiles1/nij/183781.pdf.

Does Poverty Cause Crime? (122)

Crime rate in Compton and Mission Viejo in U.S. Department of Justice, Federal Bureau of Investigation, “Uniform Crime Reports,” http://www.fbi.gov/about- us/cjis/ucr/crime-in-the-u.s/2009/crime2009. Research in Charles R. Tittle, Wayne J. Villemez, and Douglas A. Smith, “The Myth of Social Class and Criminality: An Empirical Assessment of the Empirical Evidence,” American Sociological Review 43 (1978): 643–56. Criticisms in Gary Kleck, “On the Use of Self-Report Data to Determine the Class Distribution of Criminal and Delinquent Behavior,” American Sociological Review 47 (June 1982): 427–33. “American youths of all backgrounds” in Currie, Elliott Currie, Confronting Crime: An American Challenge (New York: Pantheon): 157. Behavioral Research Institute study in

179

ibid., pp. 158–59; Rosenfeld in Patrik Jonsson, “Poverty Rate Paradox,” Christian Science Monitor, September 13, 2010, http://www.csmoni- tor.com/USA/2010/0913/Poverty-rate-paradox-Poverty-rises-but-FBI-crime-rate- falls; MacDonald in Bradford Plumer, “Crime Conundrum,” The New Republic, December 22, 2010, http://www.tnr.com/article/80316/relationship-poverty-crime- rates-economic-conditions. “Real break in the pattern” in Jonsson, “Poverty Rate Paradox.”

Why Is the Black Crime Rate So High? (123)

Crime by gender, age, race in U.S. Department of Justice, Federal Bureau of Investigation, “Uniform Crime Reports,” Crime in the United States: 2010, tables 33, 38, 43a; Andrew Hacker, “Black Crime, White Racism,” New York Review of Books, March 3, 1988, p. 36. Bruce Wright, Black Robes, White Justice (New York: Carol Publishing, 1987). Currie, Confronting Crime, pp. 152–59; “genuine social disaster,” ibid., p. 160.

Does Prison Pay? (124)

Cost of imprisonment and incarceration rates in John Schmitt, Kris Warner, and Sarika Gupta, “The High Budgetary Cost of Incarceration,” Washington, DC: Center for Economic and Policy Research, June 2000; Poland v. U.S. in ibid., pp. 2–3. Edwin Zedlewski, “When Have We Punished Enough?” Public Administration Review 45 (November 1985): 771–79. Franklin E. Zimring and Gordon Hawkins, “The New Mathematics of Imprisonment,” Crime and Delinquency 34 (October 1988): 425–36; Edwin Zedlewski, “New Mathematics of Imprisonment: A Reply to Zimring and Hawkins,” Crime and Delinquency 35 (January 1989): 169–173. “To pen every” in Eugene H. Methvin, “Why Don’t We Have the Prisons We Need?” Reader’s Digest, November 1990, p. 71. Middle position in John J. DiIulio, Jr., and Anne Morrison Piehl, “Does Prison Pay?” The Brookings Review, Fall 1991, pp. 28–35, and Anne Morrison Piehl and John J. DiIulio, Jr., “Does Prison Pay? Revisited,” The Brookings Review, Winter 1995, pp. 21–25.

Hate Crimes (125)

Congressional definition in http://www.fbi.gov/about- us/investigate/civilrights/hate_crimes/overview. UCR data on bias motivation, U.S. Department of Justice, Federal Bureau of Investigation, Hate Crime Statistics, 2009, http://www2.fbi.gov/ucr/hc2009/data/table_04.html. National Crime Victimization Survey data in Lynn Langton and Michael Planty, “Hate Crime, 2003–2009” (special report), NCJ 234085, Washington, DC: U.S. Department of Justice, Office of Justice Programs, Bureau of Justice Statistics, June 2011, p. 1. NCVS higher in Langton, p. 2. “It is sometimes difficult to know” in Federal

180

Bureau of Investigation, Criminal Justice Information Services Division, “Hate Crime Statistics, 2009 Methodology,” p. 1. Data sample: NCVS data in Langton and Planty, “Hate Crime, 2003–2009,” p. 4; UCR in Hate Crime Statistics, 2009, http://www.2.fbi.gov/ucr/hc2009/data/table_04.html.

Does Capital Punishment Deter Murder? (126)

Isaac Ehrlich time series study in Isaac Ehrlich, “The Deterrent Effect of Capital Punishment: A Question of Life and Death,” American Economic Review, June 1975, pp. 397–417. Use before Supreme Court in Richard M. McGahey, “Dr. Ehrlich’s Magic Bullet: Economic Theory, Econometrics, and the Death Penalty,” Crime and Delinquency, October 1980, p. 485. Criticisms of Ehrlich summarized in ibid., pp. 485–502; see also Jan Palmer, “Economic Analyses of the Deterrent Effect of Punishment: A Review,” Journal of Research in Crime and Delinquency, January 1977, pp. 4–21; Peter Passell and John B. Taylor, “The Deterrent Effect of Capital Punishment: Another View,” American Economic Review, June 1977, p. 445, and Yale Law Journal symposium, December 1975. Isaac Ehrlich cross- sectional study in “Capital Punishment and Deterrence: Some Further Thoughts and Additional Evidence,” Journal of Political Economy, August 1977, pp. 741– 88. Criticisms in McGahey, “Dr. Ehrlich’s Magic Bullet,” pp. 496–98; Franklin E. Zimring and Gordon Hawkins, Capital Punishment and the American Agenda (New York: Cambridge University Press, 1987), pp. 178–84. Very small deterrence in ibid., pp. 180–81. Range in 2000s studies in Jeffrey Fagan, “Death and Deterrence Redux: Science, Law, and Causal Reasoning on Capital Punishment,” Ohio State Journal of Criminal Law 4 (2006): 255–319, and Criminal Justice Legal Foundation, “Articles on Death Penalty Deterrence,” www.cjlf.org/deathpenalty/DPDeterence.html; Shepherd’s testimony that “studies are unanimous” in Fagan, “Death and Deterrence Redux,” p. 259. Survey of top criminologists in Michael L. Radelet and Traci L. Lacock, “Do Executions Lower Homicide Rates? The Views of Leading Criminologists,” Journal of Criminal Law and Criminology 99, no. 2 (Spring 2009): 489–519. Texas data in Kenneth C. Land, Raymond H.C. Teske, Jr., and Hui Zheng, “The Short-Term Effects of Executions on Homicides,” Criminology 47, no. 4 (2009): 1009–43. No corresponding “spate of articles” remark in McGahey, “Dr. Ehrlich’s Magic Bullet,” p. 501.

More Guns/More Crime or Less Crime? (128)

Number of guns in “Americans and Their Guns,” Wall Street Journal (blog), September 6, 2007, http://blogs.wsj.com/numbersguy/amerians-and-their-guns- 183/. Gary Kleck and Marc Gertz, “Armed Resistance to Crime: The Prevalence and Nature of Self-Defense with a Gun,” Journal of Criminal Law and Criminology 86 (Fall 1995): 150–87; Tom W. Smith, “A Call for a Truce in the DGU War,” Journal of Criminal Law and Criminology 87 (Summer 1997): 1462–

181

69; David Hemenway, “Survey Research and Self-Defense,” Journal of Criminal Law and Criminology 87 (Summer 1997): 1430. Gary Kleck and Marc Gertz, “The Illegitimacy of One-Sided Speculation,” Journal of Criminal Law and Criminology 87 (Summer 1997): 1446–61. John R. Lott, More Guns, Less Crime: Understanding Crime and Gun Control Laws, 3rd ed. (Chicago: University of Chicago Press, 2010); Mark Duggan, “More Guns, More Crime,” National Bureau of Economic Research Working Paper No. 7967, October 2000; impact on laws in Lott’s preface; Florida in Duggan, p. 133; violent crime and rape in Duggan, p. 137.

What About White-Collar Crime? (131)

Sutherland on white-collar crime quoted in David O. Friedrichs, Trusted Criminals: White Collar Crime in Contemporary Society (Belmont, CA: Wadsworth, 2010). FBI white-collar crime data in “U.S. Reports 18% Rise in ‘85 in White-Collar Convictions,” New York Times, September 29, 1987, p. A24. FBI definition in Cynthia Barnett, “The Measurement of White-Collar Crime Using Uniform Crime Reporting (UCR) Data,” U.S. Department of Justice, Federal Bureau of Investigation, Criminal Justice Information Services Division, http://www.fbi.gov/about-us/cjis/ucr/nibrs/nibrs_wcc.pdf; NIBRS estimates and improved data in ibid. National White Collar Crime Center survey in Friedrichs, Trusted Criminals, pp. 45, 47; NIBRS estimates in Barnett, “The Measurement of White-Collar Crime,” p. 4; broader definition in Friedrichs, p. 50. Joseph W. Eaton and David J. Eaton, The American Title Insurance Industry (New York: New York University Press, 2007). Harold E. Pepinsky and Paul Jesilow, Myths That Cause Crime (Santa Ana, CA: Seven Locks Press), 58–65; Reiman on white-collar crime in Pepinsky and Jesilow, Myths That Cause Crime, p. 33.

Case Study Questions (133)

1. See http://www.fbi.gov/about-us/cjis/ucr/frequently-asked- questions/ucr_faqs.

2. See references on accuracy of UCR, NCVS, and NIBRS.

3. Richard A. Wright, In Defense of Prisons (Westport, CT: Greenwood Press, 1994), pp. 97–98.

4. See Philip J. Cook, “The Case of the Missing Victims: Gunshot Woundings in the National Crime Survey,” Journal of Quantitative Criminology 1, no. 1 (1985): 91–102.

182

7

The National Economy □□□□

This chapter looks at statistical controversies for a variety of national economic statistics, including gross domestic product (GDP), productivity measures, savings rates, imports and exports, and domestic investments. These statistics are very much in the news, but often in contradictory terms, portraying at the same time both the good and the bad health of the U.S. economy. In some instances, these discrepancies occur because of problems with the underlying data; for example, data on imports and exports is difficult to collect, and consequently some researchers question official trade statistics. More often, however, U.S. national economic data are considered a model of survey technique, refined over several decades of collection with well-understood limitations. In these cases, contradictory statistics arise because of different methods for interpreting the data. Such examples provide constructive case studies of the intersection between economic theory and the construction of economic statistics.

Where the Numbers Come From

Organizations Data sources URL Bureau of Economic Analysis, U.S. Department of Commerce

National income and product accounts

www.bea.gov

Bureau of the Census, U.S. Department of Commerce

Economic censuses www.census.gov

Bureau of Labor Statistics, U.S. Department of Labor

Productivity computations www.bls.gov

Board of Governors, U.S. Flow of Funds data www.federalreserve

183

Federal Reserve gov U.S. Small Business Administration

Statistics of U.S. Businesses, Business Dynamics Statistics and Business Employment Dynamics

www.sba.gov

World Bank Country data data.worldbank.org Penn World Tables Country data pwt.econ.upenn.edu United Nations Country data data.un.org U.S. Central Intelligence Agency

World Factbook www.cia.gov

Data Sources

U.S. Commerce Department

National Income and Product Accounts

The U.S. Department of Commerce’s Bureau of Economic Analysis (BEA) is the single most important agency for statistics on the entire U.S. economy. It is the conduit for data collected throughout the government and consolidated in National Income and Product Accounts, better known by its most comprehensive statistic, gross domestic product (GDP).* These data were first collected systematically during the 1930s under the leadership of economist Simon Kuznets, who later won the Nobel Prize in economics for his efforts. At that time, new categories were created that gained worldwide acceptance for the measurement of national economic activity.

In brief, total economic transactions are added up twice. on one side of the ledger, total incomes, including wages, salaries, rents, profits, interest, and taxes, are computed from data collected by the Internal Revenue Service, the Bureau of Labor Statistics, the Social Security Administration, and other government agencies. on the other side of the ledger, total expenditures are computed based primarily on Census Bureau economic surveys.

Data Sample: In 2009 U.S. private investment in medical equipment and instruments totaled $73.1 billion, an increase from $43.8 billion in 2002.

184

Industry Statistics

Frequently researchers require statistics on parts of the economy—such as on a single industry or a particular geographic area. Such data are collected and made available every five years in U.S. Economic Censuses covering 7 million businesses, supplemented by annual surveys of manufacturing, services, construction capital expenditures, county business patterns, public employment, and retail trade. The Annual Survey of Manufactures (ASM) is especially expansive, based on a survey of 50,000 U.S. establishments, providing data by North American Industry Classification System (NAICS) code (see chapter 10) for shipments, costs, employment, and payroll. A useful source of local information is the annual County Business Patterns, with data on employment, payroll, and type of business for every U.S. state and county.

Data Sample: The Survey of Manufactures estimated 16,700 employees in the 2008 dog and cat food industry (NAICS code 311111), down from 16,931 employees in 2007.

Trade Statistics

The BEA publishes many international statistics based on data collected by other government agencies. Imported and exported goods are measured from customs data and oil and gas pipelines that cross national boundaries, while imported and exported services are estimated through business survey data. Military sales are derived from U.S. Department of Defense reports. Foreign investment income comes from Treasury Department data. Travel expenditures are estimated from U.S. Customs data and surveys of air travelers.

Data Sample: U.S. Census Bureau trade statistics report imported fish and other marine products from Iceland valued at $100,811,000 in 2010.

U.S. Labor Department

Productivity

U.S. Commerce Department data are used by the Labor Department’s Bureau of Labor Statistics (BLS) to calculate productivity—that is, the relationship between national output and inputs. Most often, productivity

185

is measured in terms of real output per hour of labor, although the BLS has other productivity measures, including a multifactor productivity statistic that attempts to take into account changes in all inputs.

Data Sample: According to the Bureau of Labor Statistics, U.S. photofin-ishing output per hour grew by 24.7 percent in 2009.

U.S. Federal Reserve Board

The nation’s central government bank, the Federal Reserve, tracks the economy in its own statistical series. Most important for researchers are flow-of-funds data, a quarterly report on the flow of money through the economy. This is a fundamental source on banking, credit, investments, and international capital transactions. In addition, the Federal Reserve Board compiles statistics on industrial production, including the frequently consulted index of capacity utilization, a measure of how different sectors of the economy are using their existing resources. These data are used for Federal Reserve monetary policy decisions.

Data Sample: The Federal Reserve Board of Governors reported $800.2 billion in consumer revolving credit in 2010.

U.S. Small Business Administration

Economy-wide data—classified by employment and establishment size of firms and gathered from other U.S. government agencies—is available from the U.S. Small Business Administration’s office of Advocacy. The data include information on the size of firms by industry, as well as business births, deaths, and growth.

Data Sample: For 2007 the Small Business Administration reported 301,068 real estate firms employing 2,224,175 workers, of which one-third worked for the 1,239 firms with 500 or more employees.

International Statistics

World Bank data sets, once available only by subscription, now can be accessed without charge at the World Bank website. They include various specialized data sets compiled for World Bank research activities and sorted by topic, focusing on development finance and economic development used in the World Bank’s World Development Indicators.

186

Penn World Tables provide purchasing power parity and national income accounts converted to international prices for 189 countries/territories from 1950 through 2009.

United Nations compilations report on countries’ economic, social, environmental, and demographic data, including data used in the United Nations Human Development Index.

The U.S. Central Intelligence Agency (CIA) World Factbook summarizes data from other sources on the history, people, government, economy, geography, communications, transportation, and military for 266 world entities.

Controversies

Which GDP?

over the years GDP reporting has improved from Simon Kuznets’s first estimate in 1934, when it took two years to compile the data (too late for timely policymaking). Now, the Commerce Department estimates GDP statistics with as little delay as possible. The first available data, called “advance GDP,” are published one month after the end of the quarter-year being measured. The day before the GDP data are published, they are calculated on a stand-alone computer to which only a select few officials have access. Such secrecy is necessary because an unexpected rise or fall in GDP can affect stocks and other financial markets, creating the possibility of personal profit by those who know GDP data before the general public does. In 1985 two Commerce Department employees allegedly invested in bonds in anticipation of not-yet-released GNP data that would show less than expected growth and thus would raise future bond prices. Since then, security has been improved.

Advance GDP estimates are almost always revised as more data become available (see Table 7.1). Actual survey data is available for less than one- half of GDP in the advance estimate; other components are extrapolated based on partial data and past experience. For example, consumer spending on electricity and gas are estimated using recent temperatures compared to past heating and cooling expenditures. Experts recommend that researchers wait for final revisions a year later or, even more preferably, look at annual rates of change rather than the volatile and potentially misleading quarterly reports. Hasty use of preliminary data can cause erroneous economic policy.

187

These changing estimates underscore the tentative nature of GDP accounts; there will never be a final “correct” statistic. For researchers, the revisions can be an inconvenience, as for example when results must be recalculated because of newly released data. Fortunately, national income account data are compiled frequently and entered on the Commerce Department website along with comprehensive discussion of revisions to the original numbers.

Adjusting GDP Growth for Inflation

GDP can rise simply because of inflation, so economists usually replace nominal GDP, the dollar value of GDP, with real GDP adjusted for price changes. However, measuring inflation is a tricky issue because of changing spending patterns and the introduction of new products. In 1996 the U.S. Department of Commerce’s Bureau of Economic Analysis (BEA) introduced a new method for dealing with this issue that caused a revision of past and subsequent GDP statistics with important consequences for researchers.

Prior to 1996, the BEA used a fixed weight method to calculate economic growth: It updated expenditure choices every five years and revised GDP growth estimates based on the more recent expenditure distribution. The fixed weight system was relatively accurate every five years when the weighting was updated, but then deteriorated over time, causing economic growth to be overstated. This occurred because the method assumed continued spending on high inflation items and projected too little on low inflation items. In addition, the corrections every five years imposed a more recent expenditure pattern on earlier periods. As a result, past economic ups and downs appeared to smooth out, and the overall growth rate was lowered artificially.

Table 7.1

GDP Revisions (Real Annual GDP Growth for January–March 2010)

Advance estimate (May 2010) 3.2% Second estimate (June 2010) 3.0% Third estimate (July 2010) 2.7% Revision (August 2010) 3.7%

188

Annual revision (August 2011) 3.9%

Source: U.S. Department of Commerce, Survey of Current Business, May 2010, June 2010, July 2010, August 2010, August 2011. http://www.bea.gov/scb/date_guide.asp.

Beginning in 1996, BEA introduced what is called a chain-type inflation adjustment (see chapter 11). Essentially the new method takes an average of the expenditure distribution over two years so that changes in spending are smoothed out or “chained” from year to year. Experts consider the chain-weight calculation a superior, although more mathematically complex, method; it has been adopted nearly universally by other industrial nations.

For researchers, the new system means an entirely new set of “real” GDP growth statistics revised back to 1929. For example, the drop in GDP following World War II had been measured at 25 percent using the old fixed weights based on 1987 prices. The new chain weight statistics show only a 13 percent decrease, based on more realistic (lower) postwar defense equipment prices. Also, with the chain-weight index, recovery from the 2001 recession was shown to have a slow growth of 2.7 percent, not the 4.3 percent growth that would have been measured with the old method.

Problems with GDP

Gross domestic product statistics, both preliminary numbers and subsequent revisions, are estimated according to a set of accounting rules developed by the Commerce Department in consultation with the economics profession. But the standard accounting method is not universally accepted. In recent years several potential problems have led researchers to propose alternative statistics to those officially published.

In designing the original GDP accounts, Simon Kuznets deliberately chose to include all production—regardless of moral or aesthetic considerations. Thus it was considered better to add together all production, including cigarettes and the health costs they incur, rather than create different GDP accounts based on the subjective values assigned by each economist. Increased environmental awareness and concerns about resource sustainability have prompted many economists to reassess GDP as a measure of actual national progress. As an alternative statistic, economists William Nordhaus and James Tobin created a Measure of Economic Welfare, a concept updated by the public policy institute

189

Redefining Progress in its Genuine Progress Indicator. This measure showed that while official GDP per capita, adjusted for inflation, more than doubled between 1974 and 2004; however, once corrected for environmental degradation, inequality, costs of crime, and loss of leisure time as well as benefits such as public infrastructure, volunteer activities, and housework, “genuine progress” was essentially zero.

A 2010 study, Mismeasuring Our Lives: Why GDP Doesn’t Add Up, commissioned by the former French president Nicholas Sarkozy and led by Nobel Prize-winning economists Joseph Stiglitz and Amartya Sen, applauded efforts such as the Measure of Economic Welfare and the Genuine Progress Indicator. However, these authors also urged that an adjusted GDP take into account sustainability—that is, the availability of resources for future generations. For this purpose, the report recommends a “dashboard” of indicators measuring the available stock of resources and the environmental impact of depleting them.

Researchers compiling alternative measures to GDP recognize major measurement problems, most notably putting dollar values on items that are not directly bought or sold in the market—items such as a degraded environment. As a result, these alternative GDP statistics remain on the fringe of economic research, and when they may affect public policy, critics are quick to suppress their impact. For example, in 2007 the Chinese government canceled its once-heralded “Green GDP” project when it found that 20 percent of China’s GDP growth was counterbalanced by natural resource depletion and environmental degradation.

Underground Economy

Some economic transactions are not reported to the government, either because the activity is illegal or because those involved want to avoid taxation. Together these transactions constitute the underground economy, a difficult-to-measure entity that may cause underestimation of total economic activity and, according to some economists, overestimation of the poverty rate (see chapter 8), the unemployment rate (see chapter 9), and productivity (this chapter).

Because the underground economy is, by definition, unreported, it needs to be measured indirectly, and several methods have been developed to accomplish this, among them (1) surveys asking individuals about underground activity; (2) comparisons of national income versus national

190

expenditures to reveal unreported income; (3) analysis of labor force data to show “missing” participants; (4) comparisons of electricity usage with reported production; and (5) analysis of the money supply or currency distribution to reveal excess transactions. Using these varied methods, U.S. estimates of the underground economy’s proportion of GDP vary from 6.2 percent to 19.4 percent, compared to estimates of over 60 percent of GDP in Nigeria, Thailand, and Egypt.

These underground economy numbers are simply best guesses. For example, based on the enormous volume of U.S. currency in circulation (almost $3,000 per U.S. household), economist Edgar Feige puts the 2010 U.S. underground economy at over $2 trillion, implying an uncollected tax gap as high as $475 billion. In order to make this estimate, Feige needed first to calculate how much currency is used for non-underground activity and how much is circulating overseas. Then he had to make assumptions about how much underground activity that excess currency could support, since each dollar is circulated and reused several times per year. Finally, the question remains as to what percentage of these cash transactions that go unreported actually should be counted in the GDP. The Commerce Department has pointed out that many underground transactions, such as fencing stolen goods or failing to report the full value of a used car transaction, are simply transfers of money that do not involve production and thus would not be counted in GDP. Research by economist Friedrich Schneider at the International Monetary Fund suggests a U.S. underground economy smaller than Feige’s estimate by about one-third.

Implications

Although professional economists recognize each of these problems in GDP accounting, few corrections are ever made in actual research practice. The major difficulty is that even if adjustments can be justified on logical grounds, the alternative data require many poorly understood assumptions. Researchers do not want their results to depend on data that are not widely credible. Nonetheless, researchers and other users of national economic statistics should be alert to situations where traditional data can be misleading, as in comparing the national economies of different countries.

Intercountry Comparisons

Gross domestic product is the most common variable for comparing the economic size of nations. The absolute level of GDP measures a country’s

191

overall economic size, while GDP per capita (per person) measures a country’s production level, corrected for population size. on the basis of these numbers, countries are compared in terms of economic development. The World Bank, for example, classifies 40 economies as “low income,” with the Democratic Republic of the Congo at $347 per capita in 2010. But GDP per capita may not give an accurate comparison of living standards for several reasons. Unreported transactions are especially significant in developing countries (including the underground activity described earlier) and in nonmarket production such as homegrown food, family help in medical care, and community construction of homes. These activities, largely unreported in GDP accounts, make it possible for a resident of the Democratic Republic of the Congo to live on $347 per year, an income on which survival would be impossible in the United States. To complicate matters further, some countries’ GDP measures are subject to revisions even greater than the multibillion-dollar revisions noted earlier for the United States. In countries without sophisticated data networks, new products or services may go unrecognized for years, requiring substantial changes in existing statistics.

Box 7.1 What’s the Cost of a Holiday?

The 2011 British royal wedding of Prince William and Catherine Middleton sparked debate about its cost to the country—not only the direct cost of the Westminster event, but rather the total effect on the British economy. Of particular concern was the near-shutdown of many workplaces on that day, leading to media estimates that the wedding cost the economy as much as $50 billion. However, experts point out that one-day events such as the wedding usually have a much smaller impact, typically 0.1 percent of GDP, or about $2 billion for the United Kingdom. The main reason is the forgone economic activity is simply shifted to a later date. In addition, the event itself can prompt new sales such as wedding memorabilia, and for some tourist and retail locations, a holiday actually increases consumption.

Even more difficult to measure, but important for policy decisions, are the economic impacts on localities of special events such as the Olympics. As with holidays, heavily attended sports events slow

192

down some normal economic activity while at the same time increasing restaurant, hotel, and other such sales. Overall, researchers find little evidence of positive economic impact from major sporting events. Even health benefits prompted by a greater increase in sports have not shown up in studies of areas where the Olympics took place.

Sources: “How Much Did Royal Wedding Cost Britain?” Wall Street Journal, April 20, 2011, http://blogs.wsj.com/numbersguy/how-much-did- royal-wedding-cost-britain; Olympics in Jeffrey G. Owen, “Estimating the Cost and Benefit of Hosting Olympic Games,” The Industrial Geographer 3, no. 1 (Fall 2005): 1–18; “Benefits of Hosting Olympics Unproven—Report,” Reuters, May 21, 2010, http://uk.reuters.com/article/2010/05/21/uk-britain- olympics-bmj-idUK-TRE64J7AX20100521.

Uncounted items may cause a bias in highly developed countries as well. For example, residents of many Western European countries enjoy superior vacation and holiday benefits (up to seven weeks total per year in Sweden; see chapter 9). Such benefits contribute in a major way to well- being, but they are not accounted for in traditional GDP accounting. In summary, although there are few alternatives to GDP for comparisons between countries, researchers need to be aware of the many factors that cause GDP to be an imperfect measure of economic development.

Has China Caught Up with the United States?

In the 1890s the United States passed Great Britain as the world’s largest economy. In the not-too-distant future, China will supplant the United States as the largest world economy—perhaps as early as 2016, according to an International Monetary Fund report. However, the precise size of China’s economy and its true growth rate are subject to debate.

In order to compare the size of two economies, it is necessary to convert GDP to a common currency. If the Chinese yuan is converted to U.S. dollars at the market rate available at banks, the Chinese economy appears comparatively quite small ($5.878 trillion for China versus $14.66 trillion for the United States in 2010). However, the market conversion rate is not recommended by economists because this method does not measure actual buying power in each country. As an alternative, economists developed the purchasing power parity (PPP) method that adjusts currency values for the

193

amount of goods and services a given unit of currency can buy. A comparison of China and the United States in this regard illustrates the difficulties in using purchasing power parity.

The most commonly used PPP-adjusted GDP statistics are available in the Penn World Tables based on the United Nations International Comparison Program (ICP), a worldwide statistical partnership to collect comparative price data in collaboration with the World Bank and the organisation for Economic Co-operation and Development (oECD). The ICP data are used not only to compare contemporary economies but also for important historical research because they allow for long-run studies of economic development going back even before the industrial revolution. The ICP revises its PPP retroactively as new data become available. Thus China’s 2005 per capita GDP, first measured at $6,757, was revised a few years later to $4,088. However, even these adjusted numbers have come under scrutiny by economists Angus Deaton and Alan Heston, who suggest that the ICP underestimates China’s GDP because their data came from urban areas where prices were higher than in the country as a whole.

Before official Chinese statistics are converted from yuan to dollars, there are concerns about the basic numbers, specifically their degree of accuracy. For many years China has reported extraordinary economic growth rates, averaging over 10 percent each year from 1992 to 2010, far above the average in more developed industrial nations. However, these numbers were challenged in 1998 by British economist Angus Maddison, who concluded that the official Chinese economic growth rate was more than 2 percent too high because of inaccurate valuation of services—tricky measurements even in more developed countries with more transparent statistical agencies (see section on measuring productivity). Hong Kong University economist Carsten Holz, while not implying that China’s official data were correct, offers evidence that Maddison’s adjustment was no better, calling into question the widely used Penn World Tables that incorporated Maddison’s revisions. One indication of problems in Chinese economic statistics appears when the GDP is projected backward. The calculation shows the average Chinese existing on an implausibly low $279 per year in 1952, a level not thought possible to sustain a population for more than a short period. As Deaton points out, on this basis “it is not possible that both the current PPP estimate of Chinese GDP and the official growth rates of the economy can be correct.”

The lesson is that we cannot expect precision for the size of the Chinese economy, nor for its exact growth. These statistical discrepancies aside,

194

China’s economy will soon match the size of the U.S. economy. With a much larger population, though, China’s GDP per capita—$7,600 in 2010 using the purchasing power parity method—was low compared to the United States at $47,200 per capita.

Measuring Productivity

Economists were surprised by an unexpected decade-long resurgence in U.S. productivity after 1995, reversing a slow and sometimes negative trend in U.S. productivity growth beginning in the mid-1970s. Since about 2005, productivity resumed its more erratic course, jumping from small increases, as in 2008, to a major increase in 2010 (see Figure 7.1).

Which Years?

Clearly, a major problem is identifying a trend in a statistic that jumps from year to year in an apparently erratic manner. As shown in Figure 7.1, annual productivity statistics vary widely, so a selective choice of years can project a misleading picture of productivity. For example, a widely distributed U.S. Chamber of Commerce program supporting former President Ronald Reagan’s economic policies used productivity data for 1966, 1973, 1978, and 1979, conveniently selected to show a steady decline in productivity prior to his terms in office. other years—1969, 1973, and 1976, for example—would have indicated increasing productivity.

Box 7.2 Comparing Apples and Oranges

one problem that makes it tricky to measure economies is that countries differ in the goods and services they produce. To take an extreme example, we can’t compare the relative economic well-being of a rural Thai agricultural laborer living on rice with an Ethiopian farmer raising the cereal called teff. Neither product is eaten in the other country, nor is it available except at exorbitant cost. Thus there is no way to convert rice prices into teff prices and therefore no unambiguous way to compare production of foodstuffs in these two countries. Although the rice-teff difference is extraordinary, all cultures consume varying goods and services, and these items are not

195

always traded across borders. As a result, purchasing power parity (PPP) requires arbitrary assumptions about what would be the value of goods and services not bought or sold in a particular country.

Source: Angus Deaton and Alan Heston, “Understanding PPPs and PPP- Based National Accounts,” American Economic Journal: Macroeconomics 2, no. 4 (october 2010): 1–35.

These erratic year-to-year productivity changes are caused by fluctuations in the business cycle—that is, alternating periods of growth and stagnation in the economy. When sales begin to falter, employers cut back on output, thereby causing measured productivity to fall until workers are laid off. Conversely, when the economy starts to pick up after a recession, output at first increases faster than new hiring, causing an apparent increase in productivity. Accurate use of productivity data requires correction for the business cycle, which includes looking at productivity changes between peak growth years.

Even after adjusting for the business cycle, there is considerable uncertainty in measuring productivity. Technically, it can be measured in many ways, but the most common statistic—and the one regarded as most reliable—is the U.S. Labor Department’s Bureau of Labor Statistics measure of output per hour of labor. Problems arise in the denominator— how many hours are worked?—and most important in the numerator— how is the output measured?

Hours Worked

In order to measure labor productivity, accurate data are needed on the number of hours put in by workers. With the rise of portable information appliances allowing employees to work nearly anywhere, official statistics of time on the job may miss work done elsewhere, potentially overestimating productivity by not counting all work hours. However, studies by the Department of Labor concluded that although there has been an increase in unmeasured work at home, this additional labor grew at virtually the same rate as measured hours, therefore leading to no bias in productivity growth statistics. In other words, even though U.S. workers were putting in more hours at home, they were putting in an equal increase on the job, so that productivity, measured by the percentage change in

196

labor input, would be unaffected. on the other hand, BLS economists suggest a 0.1 percent productivity overestimate because of outsourcing and offshore production that reduces costs but does not mean a firm’s workers are more productive.

Source: U.S. Department of Labor, Bureau of Labor Statistics, “Major Sector Productivity and Costs, Labor Productivity ( output Per Hour).” Data extracted January 1, 2012.

Figure 7.1 U.S. Productivity, 1970–2010

Changing Products

A major problem for productivity statistics is that today’s output includes items that once did not exist. For example, how can we measure productivity changes in the manufacture of digital recording devices or smart phones that did not exist in previous years? Also, some products change so quickly that meaningful comparisons are difficult. A typical home computer produced in 2007 is entirely obsolete and no longer made in 2012, making it difficult to measure computer manufacturing productivity between these two years.

The U.S. Department of Commerce adjusts GDP accounts for new and higher quality products, but doing so does not assure accurate productivity measures. For example, in analyzing the post-1995 productivity boom, economists were perplexed to find that much of the increase occurred in

197

U.S. retail sales. Productivity in this sector appeared to grow at a phenomenal 17 percent annual rate between 1995 and 2002, especially in electronic and appliance stores. The problem lay in the adjustment for quality. When applied to retail sales in electronics and appliance stores, workers appeared to be more productive simply because they were selling better computers. Researchers call this the “inside the box” effect, because workers were selling the same boxes, but being inappropriately credited for higher value “inside the box.” Correcting for the effect reduces the apparent productivity growth rate by about one-third.

Services

The most vexing challenge is measuring the productivity of services not associated with a physical product. In the goods-producing sector, productivity is based on the number of items produced. Although this task is complicated by the aforementioned changing quality of goods, at least government statisticians have data, for example, on the number of cars, computers, and shirts produced. But what is the productivity of a physician, a schoolteacher, or a police officer? Is the physician more productive for seeing twice as many patients in an hour, or the teacher more productive with a larger class size, or the police officer more productive for giving more tickets? Ironically, this issue is illustrated at the Bureau of Labor Statistics itself, where it appears productivity increased when statisticians attempted to measure it in 40 service-sector industries in 1992, an increase from 15 service industries in 1982, despite no increase in staff. one staff member explained, “I guess you could say we’re becoming more productive, or perhaps it would be more accurate to say we don’t have enough people to do a job that in any event is just about impossible.”

The difficulties in measuring service productivity were highlighted in the 1990s, when economists could not document the expected productivity improvement brought about by computerization in industries such as banking, finance, and airline ticketing. The reason was that better service output could not be measured easily in official statistics, so the reported higher business costs of introducing computers instead caused apparently reduced productivity. U.S. Labor Department officials report progress in adjusting for quality improvements, but as ever-new computer-related services become available, researchers need to be aware of ongoing productivity underestimates.

Implications

198

The problem of year-to-year changes in productivity caused by the business cycle can be resolved with the adjustment suggested above. Still, the problem of changing products and services is a dilemma that requires far more caution in the use of productivity statistics than is commonly found in media reports or even some research projects. ongoing cooperation between the Department of Commerce and its critics may result in better productivity data, especially for fast-changing sectors such as the computer industry. But problems in measuring service-sector productivity are so fundamental that they will not be overcome easily. At issue are policy decisions in areas such as government services and health care, where it is alleged that costs have increased without comparable improvements in output. Researchers should interpret these numbers with extreme caution, realizing that some overall productivity statistics, such as “private business productivity,” include questionable productivity inferences for the service sector.

The Savings Rate

For many years in an unusual display of unanimity, both Democrats and Republicans have declared that the nation’s chief economic problem is a low savings rate. As evidence, policymakers point to the Commerce Department’s estimate of personal savings, which fell to less than 2 percent in 2005 from over 10 percent in the early 1980s. By 2010 the savings rate had bounced back to over 5 percent. However, there is still some debate about how best to measure the personal savings rate, and theoretical questions remain about its appropriateness for assessing the health of the national economy (see Figure 7.2).

Measurement Problems

The most commonly used savings rate comes from the Commerce Department’s GDP accounts. The rate is calculated by subtracting spending from total income to estimate how much is saved. The U.S. Federal Reserve also calculates a personal savings from its Flow of Funds data, calculating rates that closely track those measured by the Commerce Department. However, both data sets can be misleading from a household’s point of view if there is a change in assets such as real estate, stocks, or bonds. When these assets change in value, households consider themselves in better (or worse) financial condition; during the housing price bubble, for instance, the recorded savings rates plunged as consumers saved less based on their apparently growing wealth. Economist Dean

199

Baker maintains that the official Commerce Department numbers overstated the savings rate by as much as 2 percent because some gains from investments were mistakenly counted as income. The difference is significant because Baker’s revised lower savings rate during the housing bubble implies that consumers had to reduce spending considerably to get back to the more normal savings rate after 2008. In his view, this spending shortfall explains part of the unexpected sluggish economic recovery from 2009 through 2011.

Source: U.S. Department of Commerce, Bureau of Economic Analysis, table 2.1. Personal Income and Its Disposition (last revised on March 29, 2012).

Figure 7.2 U.S. Savings Rate, 1980–2010

To complicate matters further, many economists believe that the frequently cited personal savings statistics are misleading. Although popular newspaper and magazine articles suggest a link between declining personal savings rates and declining investment, the two trends may not be related because much investment comes from business savings. Some economists go further, maintaining that investment is unrelated to savings altogether. In this view, investment depends on the health of financial markets, which are largely unaffected by household or business savings. Yale economist William D. Nordhaus computes what he calls a “genuine national savings rate,” including all capital formation by households, business, and government. By this measure, “savings” are over 50 percent

200

of GDP and show no downward trend. Nordhaus admits that his calculation is illustrative; some might question his decision to include most health expenditures and defense spending as investment. Nonetheless, Nordhaus claims that his estimates are superior to “stone-age” traditional savings and investment measurements. By focusing on capital creation, which is the goal of saving, we find that there is no savings crisis and that common prescriptions to increase savings by reducing government programs actually would reduce “genuine” national savings because government purchases tend to have a higher capital component than do private household purchases.

Controversy about the savings rate illustrates the dangers of trying to oversimplify economic problems to a single statistic. The idea of “spendthrift” U.S. households hurting the U.S. economy carried a moral tone that gained attention in popular accounts. But when analyzed closely, savings are a complex concept; they are difficult to measure and tricky to conceptualize in terms of their effect on the economy. Researchers need to carefully assess the accuracy of the numbers they use, as well as the relevance of those statistics to the policy question at hand.

Consumer Confidence

Three nongovernmental organizations use polling data to gauge consumer confidence in the economy: the Conference Board’s Consumer Confidence Index, the University of Michigan’s Consumer Sentiment Index, and the Washington-ABC News Consumer Comfort Index. Despite different question wording and mode of interviewing, the three indices track one another closely, moving up and down with the economy. After reaching all-time highs in the late 1990s, they fell to all-time lows in 2009.

Researchers disagree about whether the confidence measures capture information not already available through other economic statistic s and whether they can be relied upon to predict future economic events. Economic theory suggests that consumer confidence should be important because consumer spending makes up such a large portion of GDP and because it is subject to jolts based on people’s states of mind, an otherwise difficult-to-measure variable. However, in practice, the confidence measures have a mixed record. ABC polling analyst Gary Langer notes that his company’s index raised an alarm about the oncoming 1991 recession before it was predicted by most economists—although he allows that the index does a poorer job of anticipating economic recovery. on the

201

other hand, economist Sydney Ludvigson, in a study of the consumer indices, concluded that they add “modest incremental information about the future path of spending.” one problem may be the focus on the overall index reported by each group rather than changes in answers to individual questions. Richard T. Curtin, director of the University of Michigan survey, suggests we are “better off using the details of the survey rather than the overall measure.” For example, income inequality affects the results, with low-income households most often overly optimistic about future incomes, while upper-income households are more volatile, overly pessimistic in downturns and optimistic when the economy recovers.

International Statistics

In recent years, the U.S. international economic situation has come under increased scrutiny. Headlines asked: Is the U.S. trade deficit too large? Are foreigners buying too much of the United States? Answers to these questions depend critically on the accuracy of statistics collected by the U.S. Commerce Department and the U.S. Federal Reserve.

Month-to-Month Volatility

Month-to-month changes in U.S. trade typically are reported as important news: “october’s trade deficit highest ever,” followed confusingly by “November trade deficit falls.” These sudden shifts in monthly trade figures do not usually signify a trend. often a single event unrelated to the economy’s health will cause trade figures to change, as they did in 1986 when the Japanese government bought more than $2 billion worth of gold for medals honoring Emperor Hirohito’s sixtieth anniversary. Because the gold was shipped through the United States (it was purchased in Europe), exports temporarily appeared to improve.

Most experts recommend that we look at quarterly or annual data instead of monthly data. Even so, the extent of the U.S. deficit is difficult to quantify because of measurement problems.

Poor Quality Data

one indication that worldwide trade data are inaccurate is the difference between all countries’ imports and all countries’ exports, which should match because every import is also another country’s export. But according to the OECD, there has been as much as an $80 billion excess in

202

total imports measured worldwide compared to total exports. The error also appears in reports of trade between countries. For example, the International Monetary Fund’s Direction of Trade Statistics Yearbook reported only $14.2 billion in U.S. exports to France, but $21.3 billion in French imports from the United States. Much of this error may occur because of the difficulty in measuring services. According to Harry L. Freeman, vice president of American Express, U.S. administration, legal advice, consulting, and other services provided in other countries are chronically underreported. In 1985, investigators found significant nonreporting by truckers carrying goods to Canada, the largest trading partner of the United States, and comical miscoding of country origins, such as “not available” data mistakenly ascribed to the southern African nation of Namibia, which is also abbreviated “NA.”

Globalization

Globalization of production further complicates trade statistics when products contain components manufactured in several countries, then are assembled somewhere else, before being shipped to the final consumer. A2007 study by economists Barry Bosworth, Susan M. Collins, and Gabriel Chodorow-Reich indicates that U.S. balance of payment statistics are distorted when American corporations shift their profits to low-tax jurisdictions. For example, the Christian Science Monitor cites Microsoft’s ability to shift apparent sales of its software to Ireland in order to pay taxes at the lower Irish rate. In addition to skewing data on trade, the practice deprives the U.S. Treasury of potentially billions in tax revenue.

Chinese authorities have used the trade statistic inaccuracies as a reason not to revalue their currency, a policy move that would seem to be indicated by China’s trade surplus. In their view, the surplus was actually several times smaller than officially reported and thus less indicative of a need to reduce the value of the Chinese currency in order to reduce exports and increase imports. The Wall Street Journal reported a study tracing an Apple iPhone, for which its $178.96 total wholesale cost was credited to China even though the assembly cost in China was $6.50. However, the discrepancy worked the other way as well, because an undetermined amount of the Apple suppliers in fact were in China making parts that were shipped elsewhere before being returned to China for final assembly. Thus, in this case, allocating the entire wholesale cost to China clearly is an error, but the actual Chinese contribution to iPhone production was difficult to determine. In 2010 a U.S.-China Joint Commission on

203

Commerce and Trade report concluded that the statistical discrepancies amounted to about $84 billion per year.

Box 7.3 Errors in Trade Data

one reason to look skeptically at the most recent monthly trade data is that sometimes they are wildly wrong. Two errors in the billion- dollar range illustrate the problem. In February 1997 U.S. Customs data caught up on prior unreported oil imports and included the numbers in that month’s official data. Unaware of this one-time adjustment, the U.S. Commerce Department reported a trade deficit of $11.6 billion—$1.2 billion too high. The mistake affected currency trades before the error could be reported.

A second, equally large error occurred in 2005, when U.S. exports to Canada were off by more than $1 billion because Canadian customs, the source for U.S. export data, failed to count an entire month’s data during a computer upgrade. When the correct numbers were released, measured U.S. economic growth—including exports —jumped by 0.5 percent, reversing expectations of a slowdown.

Sources: oil shipment error in Beth Belton, “Error Good News for Trade Deficit,” USA Today, April 18, 1997, p. 6; “Disappointing GDP? Blame Canada,” Toronto Star, February 1, 2005, p. 1.

Implications

These measurement problems in trade statistics pose different challenges to researchers. The misleading character of month-to-month trade data is widely recognized by economists, if not by newspaper and magazine reporters. Longer-term data are an easy solution. The issue of inaccurate or missing data is less controversial, but one that will have no solution until international trade statistics are improved. Finally, the single “country of origin” concept, once a reasonable simplifying assumption, no longer holds in a world in which iPhones and many other goods are made from parts manufactured in different countries and then often shipped from country to country for assembly before being sold in yet another country.

204

As a result, researchers need to realize that trade statistics may inaccurately represent where production actually takes place.

Summary

In this chapter we have seen debates about a wide range of national economic issues. one goal was to unravel the source of apparently contradictory findings that appear both to applaud the success of the U.S. economy and to measure apparent continued economic decline. Three major problems contributed to these contradictory statistics.

For productivity, it is necessary to take into account the business cycle that causes swings in the rate of productivity unrelated to the long-term historical trend that researchers want to measure. For GDP and trade statistics, the lesson is to use data covering a long enough time period to smooth out short-term fluctuations unrelated to the overall economic issues of interest. Both of these research precautions are relatively easy to follow, with well-documented procedures available in the economics literature.

Second, and less easily resolved, is the problem of unavailable information. Many areas of economic activity are not readily measured because there are no prices attached to products, or because the prices used do not represent actual value. Gross domestic product leaves out entire sectors of production, such as housework, because no prices are attached to those services—a bias that strongly affects comparisons between countries. The same problem also affects comparisons of countries with different levels of vacation benefits because leisure time is not a commodity with a price. Similarly, investment in the form of education and training is not counted as part of the investment category of GDP accounts because the increased value of human “investments” has no easily measured resale price. Commerce Department statisticians recognize these measurement problems and have commented extensively on them. Nonetheless, it is the responsibility of researchers to make sure that their use of the statistics is not unduly influenced by what is not included in the underlying data.

Third, the service sector is a problem area for economic statistics, complicating the measurement of productivity and international trade. This sector is growing and thus will provide an increasing challenge to researchers. In response to the issue, government researchers have developed new statistics, but whether a service-dominated economy can be

205

1.

2.

3.

4.

measured as accurately as one dominated by the production of tangible goods remains to be seen.

The challenge for researchers is to make discriminating judgments when competing statistics are available: When should a statistic be used and what are its shortcomings? For example, as this chapter demonstrates, there are objections to official GDP accounting methods—and economics textbooks list even more. Yet there are practically no widely accepted alternative statistics, so most researchers have little choice but to use official numbers. At a minimum, researchers can recognize and cite the importance of the assumptions that lie behind the traditional methods.

Competing measures also exist for the savings rate, the trade deficit, and international investment. In these cases, researchers can use economic theory to make informed choices between different statistics. This link between theory and the use of data should be clearly stated. Failure to do so is a major shortcoming in many research projects. Such lack of theoretical clarity contributes to the confusing state of affairs in which it seems that statistics can be chosen to fit any conclusion, when in fact it is the implicit economic theory that leads researchers to use particular statistics.

Case Study Questions

The median income is a better measure of well-being than the more-common GDP per capita. Why? The median income requires more data than a mean income statistic such as GDP per capita. Explain why.

The U.S. Bureau of Labor Statistics Consumer Expenditure Survey isn’t used to measure expenditures in GDP accounts because it underestimates what consumers spend on tobacco, liquor, gambling, and infrequently purchased goods and services. Explain why.

India has a substantial, informal, underground economy. Consequently, Indian national accounts makers rely on data from household surveys rather than from reports from enterprises (as is done in the United States). Why is this potentially a more accurate way to measure Indian output?

The Bureau of Labor Statistics has developed measures of productivity for some government services. For which types of

206

government employment would it be possible to measure productivity with reasonable accuracy? Which government services do not lend themselves to easy measurement of productivity?

References

Data Sources (141)

U.S. Commerce Department (141)

History of national income and produce accounts summarized in Joseph W. Duncan and William C. Shelton, Revolution in United States Government Statistics (Washington, DC: U.S. Government Printing office, 1978), pp. 74–107; and Mark Perl-man, “Political Purpose and the National Accounts,” in W. Alonso and P. Starr, The Politics ofNumbers (New York: Russell Sage Foundation, 1987), pp. 135–51. Private investment data sample in U.S. Department of Commerce, Bureau of Economic Analysis, National Income and Product Accounts, http://www.bea.gov/iTable/iTable.cfm?ReqID=9&step=1. Industry statistics in U.S. Bureau of the Census, “Annual Survey of Manufactures (ASM): How the Data Are Collected,” http://www.census. gov/manufacturing/asm/how_the_data_are_collected/. Trade statistics described in U.S. Bureau of the Census, “Foreign Trade,” http://www.census.gov/foreign-trade/. Dog and cat food data sample in http://factfinder2.census.gov/faces/tableservices/jsf/pages/productview.xhtml? pid=ASM_2008_31GS101 &prodType=table. Iceland data sample in U.S. Department of Commerce, “Foreign Trade: U.S. Imports from Iceland by 5-Digit End-Use Code 2002–2011,” http://www.census.gov/foreign- trade/statistics/product/enduse/imports/c4000.html. Small business data sample in U.S. Small Business Administration, “Major Industries by NAIC Codes,” http://archive.sba.gov/advo/research/us98_01_07n_mi.pdf.

U.S. Labor Department (142)

Data described in U.S. Department of Labor, Bureau of Labor Statistics, “Labor Productivity and Costs,” http://www.bls.gov/lpc/. Data sample in Bureau of Labor Statistics, “Table 1. Percent change in output per hour, output, hours, compensation, and unit labor costs, 2008–2009,” http://www.bls.gov/news.release/prin2.t01.htm.

U.S. Federal Reserve Board (143)

Data described in “Economic Research and Data,”

207

http://www.federalreserve.gov/econresdata/default.htm. Data sample in Board of Governors, Federal Reserve Board, Consumer Credit G-19.

U.S. Small Business Administration (143)

International Statistics (143)

Data from Office of Advocacy at http://www.sba.gov/advocacy; data sample in World Bank data at http://data.worldbank.org; Penn World Tables at http://pwt.econ.up-enn.edu/; United Nations data at http://www.un.org/en/development/; CIA World-Factbook at https://www.cia.gov/library/publications/the-world-factbook/.

Controversies (143)

Which GDP? (XX)

Revisions described in U.S. Department of Commerce, “Terminology for the Quarterly Estimates,” Survey of Current Business, October 1993, p. 30. On GDP secrecy, see Allan H. Young, “Evaluation of GNP Estimates,” Survey of Current Business, August 1987, p. 36; and “Sealed Lips, Locked Safes,” Business Week, May 23, 1988, pp. 96–101. Alleged insider trading in “The Washington Version of Insider Trading,” New York Times, June 15, 1986, http://www.nytimes.com/1986/06/15/weekinreview/the-nation-the-washington- version-of-insider-trading.html? scp=1&sq=department+of+commerce+gnp+employees+dismissed&st=nytl.

Adjusting GDP Growth for Inflation (144)

On chain weighting, see J. Steven Landefeld and Robert P. Parker, “BEA’s Chain Indexes, Time Series, and Measures of Long-Term Economic Growth,” Survey of Current Business, May 1997, pp. 58–67, and Robert P. Parker and Jack E. Triplett, “Chain-Type Measures of Real Output and Prices in the U.S. National Income and Product Accounts,” Business Economics 31 (October 1996): 37–44.

Problems with GDP (145)

Kuznets’s role is discussed in Duncan and Shelton, eds., Revolution in United States Government Statistics, p. 77. Measure of Economic Welfare in William Nordhaus and James Tobin, “Is Growth obsolete?” in National Bureau of Economic Research, Fiftieth Anniversary Colloquium V (New York: Columbia University Press, 1972). Genuine Progress Indicator in Clifford Cobb et al., “Why Bigger Is Not Better: The Genuine Progress Indicator,” Redefining Progress, http://www.nber.org/~rosenbla/econ302/lecture/GPI-GDP/gpi1999.pdf. Joseph E.

208

Stiglitz, Amartya Sen, and Jean-Paul Fitoussi, Mismeasuring Our Lives: Why GDP Doesn’t Add Up (New York: Free Press, 2010). Summary of issues in Robert Costanza et al., “Beyond GDP: The Need for New Measures of Progress,” The Pardee Papers, no. 4 (January 2009), Boston, MA: The Frederick S. Pardee Center for the Study of the Longer-Range Future, Boston University. China abandons Green GDP in Melinda Liu, “Where Poor Is a Poor Excuse,” in Newsweek, June 28, 2008, http://www.thedailybeast.com/newsweek/2008/06/28/where-poor-is-a- poor-excuse.html.

Underground Economy (146)

Estimates of size in Friedrich Schneider and Dominik H. Enste, “Shadow Economies: Size, Causes, and Consequences,” Journal of Economic Literature 38, no.1 (March 2000): 77–114; Edgar Feige, “New Estimates of U.S. Currency Abroad, the Domestic Money Supply, and the Unreported Economy,” Crime, Law, and Social Change, 57, no. 3 (April 2012): 239–63; already counted in GDP in Edward F. Denison, “Is U.S. Growth Understated Because of the Underground Economy?” Review ofEconomics and Income and Wealth, March 1982, pp. 1–16. Unemployment and income data accuracy in Richard J. McDonald, “The ‘Underground Economy’ and BLS Statistical Data,” Monthly Labor Review, 107, no. 1 (January 1984): 4–16.

Intercountry Comparisons (147)

Liberia and Congo in “GDP per capita, PPP (current international $),” The World Bank online, http://data.worldbank.org/indicator/NY.GDP.PCAP.PP.CD.

Has China Caught Up with the United States? (149)

Great Britain in Bill Jennings, “IMF Predicts China Economy Will Pass U.S. Sooner Than Expected,” NewsyType.com, April 25, 2011, http://www.newsytype.com/5690-imf-china-economy-us/. IMF prediction in Mark Weisbrot, “2016: When China overtakes the U.S.,” The Guardian, April 27, 2011, http://www.guardian.co.uk/commentisfree/cifamerica/2011/apr/27/china-imf- economy-2016; Angus Deaton and Alan Heston, “Understanding PPPs and PPP- Based National Accounts,” American Economic Journal: Macroeconomics, 2, no. 4 (2010): 1–35; Angus Deaton and olivier Dupriez, “Purchasing Power Parity Exchange Rates for the Global Poor,” American Economic Journal: Applied Economics 3, April 2011, pp. 137–66; Carsten A. Holz, “China’s Reform Period Economic Growth: Why Angus Maddison Got it Wrong and What That Means,“ http://ideas.repec.org/p/wpa/wuwpdc/0504012.html; Carsten A. Holz, ”China’s Reform Period Economic Growth: How Reliable Are Angus Mad-dison’s Estimates? Response to Angus Maddison’s Reply,“ Review of Income and Wealth, 52, no. 3 (September 2006): 471–75.

209

Measuring Productivity (150)

Which Years? (150): In Molly McUsic, “U.S. Manufacturing: Any Cause for Alarm?” New England Economic Review, January-February 1987, pp. 9–10. Misrepresentative choice of years in U.S. Chamber of Commerce, Supply Side Economics (Washington, DC: U.S. Government Printing Office, 1981), slide 10. Data showing increase in Council of Economic Advisers, Economic Report of the President 1988, p. 301.

Hours Worked (151): In Alan B. Krueger, “Economic Scene,” New York Times, June 21, 2001, p. C2; work at home in Lucy P. Eldridge and Sabrina Wulff Pabilonia, “Bringing Work Home: Implications for BLS Productivity Measures,” Monthly Labor Review, 133, no. 12 (December 2010): 18–35; Lucy Eldridge et al., “Alternative Measures of Supervisory Employee Hours and Productivity Growth,” Monthly Labor Review 127, no. 4 (April 2004): 9–28; Susan Houseman and Kenneth F. Ryder, “Measurement Issues Arising from the Growth of Globalization,” Survey of Current Business, 91, no. 2 (February 2011): 1–29; Lucy Eldridge and Michael J. Harper, “Effects of Imported Intermediate Inputs on Productivity,” Monthly Labor Review 133, no. 6 (June 2010): 3–15.

Changing Products (152): Marcel P. Timmer et al., “Alternative Output Measurement for the U.S. Retail Trade Sector,” Monthly Labor Review 128, no. 7 (July 2005): 39–45.

Services(153): In Jerome A. Mark, “Problems Encountered in Measuring Single- and Multifactor Productivity,” Monthly Labor Review, 109, no. 12 (December 1986): 4–6.

Implications(154): Problems are readily admitted by the BLS, see, for example, Mark, “Problems Encountered,” pp. 3–11; BLS in “Dubious Figures: Productivity Statistics for the Service Sector May Understate Gains,” Wall Street Journal, August 12, 1992, p. A1. Mandel in “Two Ways to Read the Numbers,” Business Week, December 1, 1997, pp. 34–35; Robert H. McGuckin and Kevin J. Stiroh, “Do Computers Make Output Harder to Measure?” The Conference Board, Economic Program Working Paper Series, E-0002–00-WP (April 2000), http://www.conference-board.org/topics/publicationdetail.cfm?publicationid=1135. Improvement in official data in Jack E. Triplett and Barry P. Bosworth, “Productivity Measurement Issues in Service Industries,” Economic Policy Review, Federal Reserve Bank of New York, vol. 9, no. 3 (September 2003): 23–33.

The Savings Rate (154)

Dean Baker, “Competent Economists Were Not Kept Awake Worrying About ‘A Collapse in the Value of the Dollar and of U.S. Government Securities,’” Center for Economic and Policy Research, July 9, 2010, http://www.cepr.net/index.php/blogs/beat-the-press/competent-economists-were-

210

not-kept-awake-worrying-about-qa-collapse-in-the-value-of-the-dollar-and-of-us- government-securitiesq. William D. Nordhaus, “Budget Deficits and National Saving,” Challenge, March-April 1996, pp. 45–49. See also Milka Kirova and Robert Lipsey, “Measuring Real Investment: Trends in the United States and International Comparisons,” National Bureau of Economic Research (NBER) Working Paper 6404, 1998.

Consumer Confidence (156)

Recent rates in Daniel M. Merkle, Gary Langer, and Dalia Sussman, “Consumer Confidence: Measurement and Meaning,” American Association for Public opinion Research, Nashville, TN, May 15–18, 2003, http://abcnews.go.com/images/PollingUnit/Consumer-Confidence061804.pdf; Gary Langer, “The Public Can Predict What Economists Can’t,” Wall Street Journal, May 1, 1991, p. 14; Sydney Ludvigson, “Consumer Confidence and Consumer Spending,” Journal of Economic Perspectives, 18, No. 2 (Spring 2004): 43, http://www.nytimes.com/2002/06/08/arts/consumer-confidence-index-goes- from-an-aha-to-a-hmm.html; Richard Curtin in Louis Uchitelle, “Consumer Confidence Index Goes from an Aha to a Hmm,” New York Times, January 8, 2002, p. B11.

International Statistics (157)

Month-to-Month Volatility (157): In Edwin A. Finn, Jr., “of Apples, oranges, and Toyotas,” Forbes, January 26, 1987, pp. 34–35; Taiwan and Japanese gold in “Trade Reports Sometimes Not What They Seem,” Los Angeles Times, July 4, 1988, IV, p. 1.

Poor Quality Data (157): on trade data see William A. Sherden, The Fortune Sellers: The Big Business of Buying and Selling Predictions (New York: John Wiley & Sons, 1998), p. 75; Freeman in “Measuring the Service Economy,” New York Times, october 27, 1985, p. F-4; customs data in “Canada Says Data on’86 U.S. Trade Err by $12 Billion,” Wall Street Journal, March 10, 1987, p.1. Globalization (158): In Barry Bosworth, Susan M. Collins, and Gabriel Chodorow- Reich, “Returns on FDI: Does the U.S. Really Do Better?” National Bureau of Economic Research (NBER) Working Paper 13313, August 2007; David Wessel, “Big Discrepancy Exists Between Data from U.S. and China on Trade Deficit,” Wall Street Journal, January 22, 1998, p. 2; Mei Xinyu, “Confusing Statistics Hide Sino-U.S. Trade Reality,” China Daily, November 23, 2005, p. 4; Wang Xing, “Trade Figures Present a Distorted Picture of U.S. Deficit” (advertisement), New York Times, May 9, 2011, p. A12; U.S. Department of Commerce, Economics and Statistics Administration, “Report on the Statistical Discrepancy of Merchandise Trade Between the United States and China,” September 16, 2010, http://www.esa.doc.gov/print/Reports/report-statistical-discrepancy-merchandise-

211

trade-between-united-states-and-china; Andrew Batson, “Not Really ‘Made in China,’” Wall Street Journal, December 15, 2010, http://online.wsj.com/ article/SB10001424052748704828104576021142902413796.html.

Case Study Questions (161)

1. See Joseph E. Stiglitz, Amartya Sen, and Jean-Paul Fitoussi, Mismeasuring Our Lives: Why GDP Doesn’t Add Up (New York: Free Press, 2010), p. 45.

2. See J. Steven Landefeld, Eugene P. Seskin, and Barbara M. Fraumeni, “Taking the Pulse of the Economy: Measuring GDP,” Journal of Economic Perspectives, 22, no. 2 (Spring 2008): 193–216.

3. Barry Bosworth and Susan M. Collins, “Accounting for Growth: Comparing China and India,” Journal of Economic Perspectives, 22, no. 1 (Winter 2008): 45– 66.

4. See Donald M. Fisk, “Measuring Productivity in State and Local Government,” Bulletin 2166, U.S. Department of Labor, Bureau of Labor Statistics (Washington, DC: U.S. Government Printing office, January 1984).

* Prior to 1992, the Department of Commerce reported GNP, or gross national product, including the profits earned by U.S. firms on production overseas, but not the profits earned by foreigners on production in the United States. For the United States, the difference between GDP and GNP is slight.

212

8

Wealth, Income, and Poverty □□□□

Economists differentiate between wealth, a stock of value such as a house, and income, a flow of value over time such as a salary or interest payment. Data on wealth and income typically come from different sources, each with its own set of problems. There is relatively little information on the distribution of wealth in the United States. one reason for the scarcity is the difficulty of measurement: Because many items of wealth have not been sold recently, they do not have a readily identifiable price. A second reason for the lack of data on wealth is its ownership by a relatively small group whose members are not eager to share information about how much wealth they own. As the controversies below illustrate, this second factor severely restricts how much we know about wealth.

Unlike wealth, incomes leave a “paper trail” of readily measured dollar amounts. As a result, income is well documented in a variety of government and private-sector sources. But there is still serious disagreement about fundamental findings, including the trend in well- being for the typical family and the extent of poverty in the United States.

Where the Numbers Come From

Organizations Data sources URL Board of Governors, U.S. Federal Reserve

Survey of Consumer Finances

www.federalreserve.gov

Bureau of the Census, U.S. Department of Commerce

Survey of Income and Program Participation, U.S. Census of Population, American Community Survey, Current Population

www.census.gov

213

Survey Bureau of Labor Statistics, U.S. Department of Labor

Establishment Survey www.bls.gov

Statistics of Income Division, Internal Revenue Service

Tax returns www.irs.gov/taxstats/

Forbes Lists of wealthiest individuals

www.forbes.com

Data Sources

Wealth

Federal Reserve Survey

The most comprehensive survey on wealth holdings is conducted by the U.S. Federal Reserve (the Fed), a quasi-independent government body that acts as the country’s central bank (see Chapter 11). At three-year intervals beginning in 1983, the Fed has sponsored its own independent Survey of Consumer Finances (SCF), conducted by the University of Chicago’s National opinion Research Center. This survey covers several thousand households, asking 100 pages of questions about each family’s financial status; there is also a supplemental sample of high-income households. Prior to 1983, the Survey of Consumer Finances included only a limited wealth survey; the only serious study of wealth until that point occurred in a 1962 special Federal Reserve Board survey.

Data Sample: The 2007 Survey of Consumer Finances found that 51.1 percent of families owned stocks directly or indirectly, with a median value of $35,000. Stock ownership ranged from 13.6 percent among families in the lowest quintile (or fifth) of income to 91 percent among families in the highest quintile of income.

Survey of Income and Program Participation

Wealth has been surveyed on a sporadic basis by the U.S. Census Bureau. Such data were collected on a regular basis only between 1850 and 1890 and then not again until 1984 in the Survey of Income and Program Participation (SIPP). For several years, measures of wealth were collected intermittently in the SIPP (in 1988, 1991, and 1993); since 1996, however,

214

the SIPP has collected data on real and financial assets annually. Wealth data also is available in the Panel Survey of Income Dynamics (discussed later). Because both surveys focus on income, though, they generally measure wealth inaccurately, leading to significant underestimates of wealth compared to the SCF.

Data Sample: In the 2009 Survey of Income and Program Participation, equity in homes constituted 55 percent of the wealth of Hispanic households, but only 38 percent of the wealth of white households.

Indirect Estimates from Tax Records

For historical data prior to 1983 (when the improved SCF was initiated), wealth distribution can be calculated using two methods. The estate- multiplier method pioneered by Robert J. Lampman of the National Bureau for Economic Research uses Internal Revenue Service (IRS) records for the 1 percent of estates subject to taxes. The Statistics of Income (SOI) Division of the IRS has used the estate-multiplier technique to produce wealth estimates since 1962, available in the estate tax data files (ETD). A related method, called “income capitalization,” works backward from tax reports of rent, dividends, and interest to estimate wealth from which these incomes are derived.

Data Sample: Lampman measured a decline in the share of wealth held by the top 1 percent from 36 percent in 1929 to 26 percent in 1956.

Direct Counts

The extraordinarily rich usually are well-known individuals about whom information can be obtained from public sources. Based on stock ownership records, media coverage, and independent investigation, Forbes magazine estimates wealth holdings for a select number of these individuals and compiles an annual list of the wealthiest individuals in the United States and the world.

Income

U.S. Census/American Community Survey

215

The long-form of the decennial Census of Population has included questions about income since 1940, providing researchers with a tremendously detailed data source, but one that is available only at ten- year intervals. In 2010 the Census long-form was replaced with the American Community Survey (ACS), administered annually (see Chapter 2).

Data Sample: Between the 2009 and 2010 American Community Surveys, no state experienced an increase in median household income; for 15 states and the District of Columbia, median income was statistically the same in both years.

Current Population Survey

By far the most frequently cited source of income data is the U.S. Census Bureau’s Current Population Survey (CPS), covering about 60,000 households (see Chapter 9 on the origins of the CPS). In 1947 the questions were simply, “How much did you earn in wages and salaries in this year?” and “How much income from all sources did you receive in this year?“ Now data is collected on 20 income sources. Especially detailed are questions on government benefits; in 1998 these questions were broadened in order to gather information on the variously named state programs that replaced Aid to Families with Dependent Children.

Data Sample: According to the Current Population Survey, the ratio of income at the ninetieth percentile to income at the tenth percentile grew from to 8.68 in 1968 to 11.67 in 2010.

Bureau of Labor Statistics’ Establishment Survey

Data collected in the Bureau of Labor Statistics’ (BLS) Establishment Survey are often quite accurate because the survey is based on employer records rather than on respondent recall. But the Establishment Survey measures earnings for individual jobs, not the total income for a worker, who may have more than one employer, or for a family with several separate sources of income.

Data Sample: In the August 2011 Establishment Survey, the lowest- paying industry was Leisure and Hospitality, with average hourly earnings of $11.47 for nonsupervisory production workers.

216

Panel Study of Income Dynamics

For comparisons over time, or longitudinal studies, researchers frequently turn to a private survey, the University of Michigan’s Panel Study of Income Dynamics (PSID). Begun with 5,000 households in 1968, PSID followed more than 9,000 households in 2009, and included files on more than 70,000 individuals, some tracked for as long as four decades. A sample of immigrant families was added in the late 1990s, and a supplemental child development survey was administered to a subsample of families between 1997 and 2008.

Controversies

Wealth

What Is Wealth?

In discussions of wealth, most analysts focus on net worth, defined as the current value of assets minus the current value of liabilities or debt. Using the Federal Reserve Board’s Survey of Consumer Finances, mean net worth in 2007 was $556,300 per family, while median net worth was only $120,300 (see Box 8.1 on the difference between mean and median). For nonwhite or Hispanic families, median net worth was only $27,800 in 2007, less than one-fifth of median net worth for white households.

For many decades, most household wealth was in housing, but as the stock market surged in the 1990s, more and more households began investing in stocks so that by 2001, over half of U.S. households were stockholders. The years since 2001 have seen dramatic swings in both the housing and stock markets, and housing as a share of total assets rose again between 2001 and 2007. For some research purposes, it may be preferable to exclude ownership of homes, vehicles, real estate, and business assets, in order to focus on ownership of wealth that is a source of economic power, such as stocks and bonds. This financial wealth is even more unequally distributed than total wealth: In 2007 the top 1 percent of households owned 43 percent of total non-home wealth (versus 35 percent of all wealth), and the top 20 percent owned 93 percent (versus 85 percent of all wealth). New York University economist Edward N. Wolff found that the inequality of non-home wealth has also followed a somewhat different trend than all wealth: While the inequality of total wealth stayed almost flat between 1989 and 2007, the inequality of non-home wealth

217

decreased slightly through the 1990s and then increased again after 2001.

Some researchers argue that rather than restrict the definition of wealth, we should expand it, particularly to include the value of pensions and social security. Although most people do not consider these benefits to be wealth, they are similar to bonds and other arrangements that promise a source of future income. There have also been significant changes over the last few decades in how retirement accounts are structured, with more and more individuals now having defined contribution plans such as IRAs and 401(k) plans, many of which are invested in stocks and bonds. The increase in these types of accounts explains much of the growth in stock ownership during the 1990s; however, whether this constitutes real growth in wealth depends on whether one includes other types of retirement accounts in prior periods: Because defined benefit pensions were not typically included in past calculations of wealth, the switch to defined contribution accounts (which are included) may appear as an increase in household wealth even if the real value of all retirement accounts has not changed.

Researchers need to choose carefully between these different measures of wealth. For analysis of the distribution of resources, total wealth measures are appropriate, perhaps also including the value of retirement funds. But if the research goal is to understand how wealth affects power, then it might be proper to use a narrower definition that excludes housing and retirement benefits.

Box 8.1 Mean, Median, and Mode

The median is the data point for which one-half the observations are above and one-half below; the mean is the average or, in this case, total income divided by the number in the sample; the mode is the data point with the most observations, a statistic rarely used in income analysis. Median income is the most common starting point for research on the “typical” household. A relatively small number of families with very high incomes skews the mean at more than 30 percent higher than the median. Thus, mean annual household income in the United States was $67,530 in 2010, while the median income was $49,445. The major drawback to using the median is that it is unchanged by redistribution of income above and below it. For

218

example, when the poor become poorer and the rich become richer, it is possible for the median to stay the same.

Sources: Data in Carmen DeNavas-Walt et al., “Income, Poverty, and Health Insurance: Coverage in the United States, 2010,” Current Population Reports, P60–239 (U.S. Bureau of the Census, September 2011), p. 34; Census Bureau use of mean and median in Christopher Jencks, “The Politics of Income Measurement,” in The Politics of Numbers, ed. William Alonso and Paul Starr (New York: Russell Sage Foundation, 1987), pp. 86–88.

How Wealthy Is Wealthy?

Those individuals who have high levels of wealth also tend to have high levels of income, and vice versa. Almost everyone in the top 1 percent of earners is in the top 5 percent of wealth. However, wealth and income are distributed very differently across the population. In particular, wealth is much more concentrated at the top than income: in 2010 the wealthiest 1 percent of Americans held about one-third of all wealth, while the highest earning 1 percent of earners received about one-fifth of all income. However, the true value of the wealth held by those at the top of the distribution is more difficult to measure than income (see Box 8.2).

Reports about wealth at the top of the distribution may also conflict when analysts use different data sources. For example, the concentration of wealth among the very wealthiest people appears bigger in the Federal Reserve’s Survey of Consumer Finance (SCF) compared to estate tax data collected by the Internal Revenue Service. In 2004 the SCF had the top 1 percent of households owning just over 34 percent of all wealth, while the tax data suggested that the top 1 percent of individuals owned a bit under 20 percent of all wealth. The reason for the difference is that the two data sources differ in coverage and reporting accuracy. Specifically, the unit of observation in the SCF is households, which may include multiple individuals, while the tax data is based on individuals only.

The tax data is also restricted to those with estates above the estate tax filing threshold, but even with that restriction, the pool of individuals in the sample is many times larger than the survey data; the 2004 sample includes over 25,000 returns for people with assets of $1.5 million or more, while the 2004 SCF has fewer than a thousand households at that level. In addition, because the IRS data comes directly from tax records, it

219

does not suffer from issues of self-reporting that may plague SCF data (see Box 8.3), though the reported values are also likely to be as conservative as (legally) possible. So, while the Federal Reserve’s Survey of Consumer Finances is the most often used source of wealth data, when it comes to studying the very wealthiest individuals, researchers often turn to the IRS summary of tax returns.

Income

Does Income of $250,000 Mean You Are Rich?

In the 2008 presidential campaign, then-candidate Barack Obama repeatedly argued for cutting taxes for families making less than $250,000 and raising taxes on those above that cutoff. After the election, President obama continued to use $250,000 as a dividing line, advocating for “taxing the rich” in debates over extending the Bush-era tax cuts. However, many people who fell just above that threshold argued that they are not “rich” and should not be lumped together with millionaires.

So does an income of $250,000 make you “rich”? It certainly would put you at the top of the income distribution: only 3 percent of American households have incomes above that level. But many households with that income do not feel rich. According to Gallup surveys conducted from 2005 to 2011, only 6 percent of people in households earning over $250,000 a year think their taxes are “too low,” but among that same group, 30 percent believe that “upper-income people” pay too little in taxes. In other words, quite a few upper-income people apparently do not think of themselves as upper-income. Part of this disconnect may stem from differences in costs of living; higher-income households are more likely to live in areas of the country with higher costs of living so their incomes do not seem to go as far. There are also large differences in income among those in the top few percentiles of the income distribution, with some highly visible millionaires and billionaires at the very top, so those relatively near the $250,000 mark may look at people who are even richer and feel their lifestyle is modest in comparison.

Are the Rich Getting Richer?

Although income is not as concentrated among the top tier of households as wealth is, the share of total income held by the top 1 percent of households has been increasing dramatically over time, while the share of

220

wealth held by the wealthiest households has remained relatively stable. Thus, whether “the rich are getting richer” depends critically on whether one focuses on income or wealth.

Box 8.2 Measuring Wealth

In 1983 the U.S. Federal Reserve changed its usual format for measuring wealth by adding 438 individuals already known to be wealthy. In theory, this commonly used method of “enriching” the data sample should have increased our knowledge about wealth holdings. Indeed, the survey indicated that the share of wealth going to the top 0.5 percent increased to 35 percent in 1983 from 25 percent in 1962. These results received a barrage of media attention after the Joint Economic Committee of Congress released its own interpretation of the Federal Reserve’s study using catchy labels —“super rich,” “very rich,” “rich,” and “everyone else”—to underscore the vast holdings of the wealthy that apparently had increased since 1962. Given the media attention, the Reagan administration asked the Federal Reserve to reexamine the survey and suddenly an error appeared: One of the 438 wealthy individuals was not as rich as he or she reported. Removing this individual caused the estimate for the share owned by the super rich to fall by almost 10 percent, wiping out the apparent increase since 1962.

The accuracy of the one data point is difficult to determine. A 1986 follow-up survey showing this person’s wealth to be only $2.3 million led the Fed to conclude that the reported 1983 wealth of $200 million was a coding error. Critics responded that the individual, known to own Texas oil and gas wells, indeed might have suffered a financial setback reported between 1983 and 1986. Whatever the actual circumstances of this one Texas magnate, the dispute shows how difficult it is to measure wealth. A basic research principle is to include a survey sample that is large enough so that errors for a single individual do not affect overall findings. Even the relatively large Federal Reserve survey—nearly 4,000 households—was still too small to measure changes in wealth holdings accurately because so few individuals owned so much.

Sources: Federal Reserve studies in Federal Reserve Bulletin, December

221

1984 and March 1986. Joint Economic Committee study in James D. Smith and Stephen D. Franklin, “The Concentration of Personal Wealth, 1922– 1969,” American Economic Review 64, no. 2 (1974): 162–67. Controversy in “Scandal at the Fed?” Dollars and Sense, April 1987, pp. 10–22.

The discrepancy between the trends in income and wealth is puzzling when one considers that people who earn higher incomes are typically more likely to use that income to acquire investments. So why haven’t increasing incomes for those at the top correlated with increasing wealth? It may be that wealth shares simply are not measured very well, or that the rich are spending all of their new income on consumption. It could also be that changes in the composition of wealth for many households has boosted wealth for those in the lower part of the distribution, thus reducing any gain in the share of wealth for the top. Specifically, economist Edward Wolff notes that over the last few decades, the retirement programs for most Americans have shifted from defined-benefit pensions to defined- contribution plans. The former were not counted as wealth assets for individuals; the latter are, thus increasing the wealth of many outside the top echelon. In addition, for many lower-income households, their largest asset is their house, and increases in housing prices throughout the first several years of the 2000s may have inflated wealth measures. Inequality of non-home wealth has increased more than all wealth, and more recent analyses show that even all-wealth inequality has increased sharply in the wake of the housing bust and ensuing recession.

Box 8.3 I’ve Got a Secret

In addition to small sample size and possible coding errors, data on wealth are subject to outright refusal of respondents to answer questions. The problem is particularly severe for the Census Bureau’s Survey of Income and Program Participation. Nearly 90 percent of respondents report the size of their checking accounts, but more than 40 percent refuse to tell the interviewer the value of stocks and mutual funds. The University of Michigan’s Survey of Consumer Finances does somewhat better, failing to gather data on stocks for just 25 percent of respondents. Research on this type of measurement

222

error suggests that those who do not answer are likely to be either poorer than average or far richer. The latter group introduces a possibly critical error into all wealth data: We are missing information on precisely those who own the most. Although statisticians who gather the data are aware of its shortcomings, the problem is rarely highlighted in official reports and almost never mentioned in press coverage about the concentration of wealth.

Sources: Eugene P. Ericksen, “Estimating the Concentration of Wealth in America,” Public Opinion Quarterly 52 (1988): 243–53.

Is Rags-to-Riches Just an American Myth?

Some conservatives maintain that the emphasis on gains for the rich are misplaced because the United States has so much economic mobility. Those who are at the top of the income distribution in one year are not necessarily the same as those who are at the top in another year. In this view, distribution of income at one point in time does not matter as much as the potential for large numbers of people to reach the upper echelon. It has long been an American ideal that every child has the opportunity to do better than her parents. But how often does that actually happen?

Box 8.4 Who Is the Wealthiest of Them All?

The top three wealthiest individuals in the 2011 Forbes survey are listed below. Forbes openly discusses problems in finding out about wealth; sometimes an individual will catapult nearly to the top when it is discovered that he or she is in fact quite wealthy. A previously uncounted Mars Candy Company benefactor, billionaire heiress Jacqueline Mars Vogel, was belatedly “discovered” and added to the Forbes list in 1987.

Wealth (in billions) as listed in the Forbes 400, 2011 (U.S. only)

1. William Henry Gates III $56 2. Warren Edward Buffett $50 3. Larry Ellison $39.5

223

One source of such discrepancies and omissions is the lack of information on privately held wealth. The Mars Company, for example, is one of a small number of large firms still owned entirely by a few individuals. Unlike publicly held corporations, there is no stock price to estimate the value of the company (see Chapter 10). Similarly, Forbes must estimate the value of property that has not been bought or sold in recent years. Finally, family trusts complicate wealth measurement by dispersing fortunes among various descendants. Forbes no longer includes such dispersed fortunes; however, if family wealth can be traced to one living person, they may add “& family” to that person’s entry on the list.

Sources: “The Forbes 400: The Richest People in America,” Forbes, http://www. forbes.com/forbes-400/. Discovery of Vogel in “The 400 Richest People in America,” Forbes, October 26, 1987, p. 106. Discussion of problems in measuring wealth in “The Forbes 400 Methodology,” Forbes, last modified September 21, 2011, http://www.forbes.com/sites/kerryadolan/2011/09/21/the-forbes-400- methodology.

One simple way to measure mobility is to look at who is at the top from one time period to another. In 2012 political scientist James Q. Wilson titled a Washington Post piece, “Angry About Inequality? Don’t Blame the Rich,” pointing out that a Federal Reserve study showed only one-half of the people in the top 1 percent of income were there 10 years later. However, such studies likely overstate the fall-off in riches because many of the top one-percenters haven’t fallen very far. In the Federal Reserve study, over 80 percent of the top group remained in the top 10 percent a decade later. Moreover, changes over a lifetime explain why individuals rise to the top during peak earning years, ages 45–65, and then fall from the top income brackets simply because they aren’t working as much.

A better measure of mobility compares income levels of the households into which a child is born with that person’s later earnings. One study estimates that about 65 percent of those born in the bottom fifth of incomes stay in the bottom two-fifths, and about 62 percent of those born in the top fifth stay in the top two-fifths. Similarly, economists Gary Solon and David J. Zimmerman independently analyzed the income data of fathers and sons and found correlations of at least 0.4, where 1.0 would

224

have meant that every son fell into the same category as his father. Solon concludes: “It’s not that you inherit the same position, but there’s substantial correlation.” More recent studies using data from the Social Security Administration (with longer earnings histories), along with PSID data that included women, found similar or even higher correlations.

The United States also has far lower mobility between income groups than other countries seem to have. Canadian economist Miles Corak, after reviewing several different studies of mobility, ranked the United States and Britain as the least mobile societies out of nine countries, while Canada, Norway, Finland, and Denmark were the most mobile (Sweden, Germany, and France fall in the middle). In Canada, only 16 percent of men born into the bottom tenth of incomes were still there as adults; in the United States, 22 percent were stuck.

But some argue that such cross-country comparisons are not appropriate. For one thing, the poor in the United States look demographically different than the poor in other countries in ways that make it harder to move into higher classes; for example, poor Americans are more likely to be raised in single-parent households. Perhaps even more relevant is that the income distribution in the United States is so much wider than in other countries that the population here only appears less mobile. By one calculation, it would take an additional $93,000 for an American family to move from the tenth percentile to the ninetieth percentile; a Danish family could make that same move with half as much additional income. Thus, it might be that on an absolute basis, lower-and middle-class American families are improving their fortunes by just as much as families in other countries, but those increased earnings simply do not move them into a different class. Indeed, many conservative critics prefer to focus on this “absolute mobility,” pointing out that most children do end up with higher incomes than their parents, even if that doesn’t move them into a different income class. One lesson for researchers is to take care when discussing income mobility over time, as measures of relative mobility (how people move within the income distribution) may paint a very different story than measures of absolute mobility (how people’s income levels move).

Are We Better Off?

The distinction between relative and absolute mobility is particularly important because they tell different stories about well-being over time.

225

That is, it appears that poorer and middle-class Americans are less likely to move up into the top tier of the income distribution, but perhaps this is less of a concern if almost everyone is still better off than their parents because the country as a whole is richer. But is the typical American really better off than 30 years ago? This simple question has provoked a broad spectrum of answers depending on the data source used. Between 1980 and 2009, Americans were somewhere between being no better or even possibly worse off when measured by average weekly earnings, and being more than 60 percent better off when measured by income per capita (see Figure 8.1).

Per Capita Income

The rosiest view of economic progress is based on per capita income, a statistic from the U.S. Commerce Department’s National Income and Product Accounts (see Chapter 7). Per capita income is simply the total of all wages, interest, rents, and other incomes divided by the number of people in the country. Critics charge that the statistic is misleading because the number of wage earners increased relative to the number of dependents. Or, to make the same point with a story: Assume a household of four people in 1990 with one person working outside the home earning $40,000 had a per capita income of $10,000. If, in 2000, the original worker’s income falls to $38,000, but a second family member now works outside the home, and earns $10,000, then per capita income increases to $12,000 ([$38,000 + $10,000]/4). In other words, total income increased, but only because more people were working, each of whom earned less. This was a particularly salient issue for growth between 1970 and 2000, when the number of two-income households increased dramatically.

Individual Earnings

The most pessimistic view of U.S. well-being comes from the data on average hourly earnings, which show a decline through most of the 1980s and early 1990s before slowly rising again, so that real earnings in 2009 are quite similar to what they were in 1980. However, one problem with wage statistics is accounting for fringe benefits, most importantly health care and pensions. Once benefits are included, real hourly compensation steadily rose by about 27 percent between 1980 and 2010. For some purposes, such as measuring the cost to employers, total compensation may be the relevant variable. But pay levels alone may be appropriate if the goal is to measure employee well-being. After all, from the point of

226

view of wage earners, high cost health care benefits to pay for high cost health care do not make them feel as well off as a comparable pay increase.

Sources: Council of Economic Advisers, Economic Report of the President 2011 (Washington, DC: U.S. Government Printing Office, 2011), http://www.gpoaccess.gov/eop/; U.S. Bureau of the Census, Historical Income Tables, table F-6, http://www.census.gov/hhes/www/income/data/historical/families/.

Figure 8.1 Are We Better Off?

The rise and fall of average wages also masks underlying inequality within the population. A study by the Congressional Budget office found that that between 1979 and 2009, the gap between low-wage (tenth percentile) and high-wage (ninetieth percentile) workers increased dramatically, driven largely by increases at the top of the distribution. There are also differences across gender: while wages for low-wage workers were stagnant for men and women, middle-wage women saw larger and more steady increases than middle-wage men. High-wage women also saw larger increases than high-wage men, although the level of women’s wages was lower than men’s wages across the distribution. Thus, we must take care in using average wages as a measure of whether Americans overall are better off.

227

Family Income

The middle view of Figure 8.1 combines falling individual earnings with the effect of additional workers in each household in the statistic of family income Median family income was relatively constant during most of the 1980s and early 1990s, but then rose steadily in the latter half of the 1990s and has held steady for most of the 2000s. From one perspective, this provides evidence that families are “better off” than three decades ago, but these data do not reflect changes in what constitutes a “family.”

It is widely recognized that the proportion of wives and mothers who work has increased dramatically; labor force participation of mothers with children under the age of 18 rose from 47 percent in 1973 to 71 percent in 2009. Thus many families were better off only because of additional working hours. Less often understood is the U.S. Census definition of a family as two or more related individuals living together, leaving out more than 38 million individuals in 2010 who lived in nonfamily households, either as single-person households or with unrelated individuals. Median income for families is more than double the median income for nonfamily individuals; therefore, looking only at families overstates the typical income of U.S. households. On the other hand, in recent years nonfamily household incomes increased faster than the average, so looking only at families underestimates improvement in overall well-being.

A second problem with family income data is that they do not take into account the decline in family size from 3.01 in 1973 to 2.59 in 2010. According to one measure of family well-being, these smaller families are 20 percent better off because there are fewer children to support. Other researchers object that these families might have fewer children precisely because of stagnating incomes. In this view, families are worse off because they could not afford to have as many children as they might have liked.

Because of these problems, researchers need to be careful in using the family income data; used by itself it does not tell the whole story about changing living standards.

Disappearing Middle Class

The disparity between sharply increasing per capita income and stagnant average earnings or relatively low growth in median family income is often seen as evidence of the growing inequality of the overall distribution. That is, per capita income may be increasing because the rich are getting

228

richer while middle-class households gain little ground. However, Terry Fitzgerald, a senior economist at the Minneapolis Federal Reserve, suggests that things are not as bad for the middle class as they might appear. First, per capita income is a more inclusive measure of income than either earnings or family income. Earnings exclude benefits as well as any unearned income, while the U.S. Census definition of income also excludes all nonmonetary sources of income such as employee benefits or in-kind government transfers such as food stamps or Medicare payments. Fitzgerald estimates that growth in measured income by the Bureau of the Census would be 8 percentage points higher between 1976 and 2006 if it included nonmonetary income.

Median income and average earnings would also be higher if a different adjustment were used to account for changes in prices. The most commonly used price adjustment is the Consumer Price Index (CPI), but many analysts argue that even revised versions of the CPI may overstate inflation (see Chapter 11 for a discussion of the issues involved in accurately measuring inflation). Incomes adjusted for inflation with the CPI will thus appear to grow much more slowly than when other inflation measures are used. Using the price index for personal consumption expenditures (PCE), Fitzgerald estimates income growth would be another 8 percentage points higher. The PCE index is based on the basket of goods and services consumed each year and is commonly used by macroeconomists.

After these adjustments, growth in census income is still lower than growth in per capita income—but the gap is much smaller, suggesting that while middle-class Americans may not be doing quite as well as those at the top of the distribution, neither are they falling dramatically behind. On the other hand, some analysts argue that even if the average American has more money, those dollars do not buy the same standard of living as they used to. In the report Middle Class in America, economists in the Department of Commerce note that it is not just income that defines “middle class.” In one survey, over 90 percent of respondents considered themselves “middle class” or “working class,” and many social scientists consider “middle class” to refer to a particular combination of values, expectations, and aspirations, not just income levels. Therefore, the researchers examine the cost of achieving a middle-class lifestyle, which they define as aspiring to homeownership, a car, college education, health and retirement security, and the occasional vacation. While recognizing that different families will weight these items differently and the costs will

229

vary across the country, the researchers determine that on average, this middle-class lifestyle is attainable for even relatively lower-income families (down to the 25th per-centile). However, just a few unplanned expenses can also put these aspirations out of reach, and many families will need to make regular sacrifices to ensure stability. Furthermore, because the cost of housing, health care, and college have risen (and continue to rise) faster than income, the researchers conclude that it is becoming harder to attain this middle-class lifestyle than it used to be.

Statistics for Every Theory?

At this point it may be tempting to question the usefulness of income statistics that tell such conflicting stories. If evidence can be found to support any position, what good are income statistics? The answer is that each income statistic measures a slightly different concept; they must be used together, with careful distinctions made about what is being measured in each specific case. For example, per capita income provides a broad measure of well-being, but it tells us little about the distribution of that well-being, which must be measured by other statistics. The trends in average earnings and family income reflect major changes that were occurring in the labor force and the family in addition to an underlying slowdown in the growth of income. Finally, the debate over whether middle-class Americans are better off illustrates the importance of identifying how one is defining such nebulous terms as middle class, and whether the comparison being made is relative or absolute. The lesson for researchers is to examine the robustness of results under different assumptions and to examine multiple measures before drawing conclusions.

Poverty

Research on poverty typically confronts a perplexing question: how do we define poverty? The easy solution is to rely on the official U.S. Census Bureau poverty line, the most commonly used in social science research. This official poverty line is also the basis of government programs ranging from Food Stamps to Foster Grandparents. But, as demonstrated by its origins, the Census Bureau poverty line is quite arbitrary and is therefore inadequate for some purposes.

The Official Poverty Line

230

During the early 1960s, Mollie orshansky, a Social Security Administration staff economist, was asked to develop a definition of poverty for the War on Poverty. By her own admission, this “orshansky” poverty line was a compromise between scientific justification and political expediency. She began with a U.S. Agriculture Department estimate for a nutritionally sound diet and then estimated the total poverty budget based on the proportion of food expenditures in a typical total family budget. orshansky was unhappy with these assumptions— especially the food budget, which was intended as a stopgap grocery list but required a sophistication of purchasing and food preparation techniques not likely to be practiced by poor families. But orshansky was under pressure from the Johnson administration to calculate a poverty line at about $3,000 for a family of four—low enough so that the War on Poverty could reasonably be expected to help all those designated as poor. She had more leeway in defining poverty for other-size families, so she set the poverty line on slightly more favorable terms for smaller-size households and families with more than four people.

Not surprisingly, these highly arbitrary poverty lines have been widely criticized. For example, at the request of the U.S. Congress Joint Economic Committee, the National Research Council (NRC) set up a panel to address concerns about the poverty measure. In its 1995 report, the study panel concluded that the official poverty measure was “outmoded and no longer accurately characterizes differences in poverty among population groups, across areas of the country, or over time.” However, researchers disagree about how to correct the poverty line, or even whether the current measure underestimates or overestimates the actual rate of poverty.

The Poverty Line Is Too Low

The official U.S. poverty line is an absolute standard, adjusted only for inflation. As discussed in Chapter 11, accounting for inflation is a controversial task in itself. Over the years since 1965, the standard of living in the United States has increased, so that on a relative basis, the poverty line has fallen. In 1965 it was nearly exactly one-half of the median after-tax family income; by 2009 the official poverty threshold was a little over one-quarter of median income. A relative poverty standard can be justified on the grounds that it matches the consensus view about what constitutes poverty. When experts and the public at large are asked to define poverty, the typical view is a level of poverty that rises when

231

overall well-being rises. If the poverty line measured poverty at a constant relative rate, then poverty would have increased between 1965 and 2009, instead of falling as it did in the official absolute standard.

The Poverty Line Is Too High

Others have criticized the Orshansky poverty line for overestimating the poverty line. For example, economist Rose Friedman points out that poor families spend a higher proportion of their income on food than the ratio used by Orshansky. Friedman recalculates Orshansky’s poverty level with new ratios, estimating U.S. poverty at one-half its official level.

A second correction that reduces the apparent level of poverty is to include the value of noncash government programs such as food stamps, school meals, housing subsidies, and medical care. One such effort measured government programs at their market value—that is, what it would cost the poor to buy services comparable to those provided by the government. By this estimate, the poverty rate falls by about one-third, which prompted President Ronald Reagan’s chief economic adviser, Martin Anderson, to conclude that the War on Poverty had been “won.” But many economists question the relevance of market values for adjusting the poverty rate. In particular, medical care is so expensive in the private sector that the market value of Medicare and Medicaid is sufficient by itself to lift—in theory—almost all the elderly poor out of poverty.

The National Research Council panel recommended a number of corrections for benefits, adding the income tax credit and in-kind benefits such as food stamps, and subtracting taxes, child care, and other work- related expenses. Health care benefits were not included because they vary so greatly from family to family and do not readily substitute for money. The council also recommended that the Survey of Income and Program Participation (SIPP) be used instead of the CPS because SIPP more completely measures nonmonetary benefits. In addition to more accurate data, the more frequently conducted SIPP would allow researchers to investigate episodes of poverty that may last for short periods of time.

Poverty Is Overestimated for Some, Underestimated for Others

The official orshansky poverty line has withstood the test of time; although the Census Bureau has used the panel’s recommendations to create experimental poverty measures, none has replaced the official measure. In 2009 an Interagency Technical Working Group on Developing a

232

Supplemental Poverty Measure was formed and tasked with developing a Supplemental Measure that defines income thresholds differently but is intended to complement, not supplant, the official measure. Retaining the original poverty line makes sense for many research projects because conclusions can be compared readily with other studies that also use this poverty line. Adopting any alternative also would require similarly arbitrary assumptions that make it easy to criticize the orshansky method.

Nonetheless, researchers should be aware of biases in the official number. When the Supplemental Poverty Measure was first officially released in 2011, it showed that the overall poverty rate for the U.S. was slightly higher than indicated by the orshansky measure (16 percent instead of 15.2 percent), but what garnered the most media attention were the differences for specific subgroups (see Table 8.1). Under the Supplemental Measure, which includes in-kind government programs, the poverty rate among children was about four percentage points lower than with the official measure, and about two percentage points lower among African Americans. on the other hand, the Supplemental Measure showed poverty among the elderly at almost seven percentage points higher than with the official measure. That jump has been attributed largely to the way the Supplemental Measure deducts medical expenses from income. At the same time, many analysts believe that both measures, even the lower official measure, overstate poverty among the elderly because they ignore the wealth that senior citizens are likely to hold and that can be used to augment income.

Table 8.1

Percent in Poverty Under Different Poverty Measures

Official Poverty Measure

NRC Measure: Deducts medical

expenses, includes geographic adjustment

Supplemental Poverty Measure

2008 13.2 12.8 N/A 2009 14.3 12.9 15.3 2010 15.1 13.4 16

233

Sources: “Tables of NAS-Based Poverty Estimates: 2010,” U.S. Census Bureau, Official and National Academy of Sciences-based Poverty Rates, 1999–2010, http://www.census.gov/hhes/povmeas/data/nas/tables/2010/index.html; and Kathleen Short, “The Research Supplemental Poverty Measure: 2010,” Current Population Reports P60-241, November 2011.

How Much Did Poverty Increase During the Recession?

The importance of using multiple measures of poverty is exemplified in discussions of how the most recent recession affected poverty. Using the official measure, poverty increased from 13.2 percent in 2008 to 14.3 percent in 2009. But because the official measure excludes in-kind benefits, it does not capture the impact of many government programs specifically employed as part of the 2009 federal stimulus package, such as unemployment benefits, tax credits, and food benefits. Alternative poverty measures, such as those suggested by the NRC panel that include these additional income sources, show virtually no change in poverty between 2008 and 2009, causing some to celebrate the benefits of the Recovery Act programs. The picture is further clouded by the inclusion, or exclusion, of medical expenses. For individuals who lose their jobs and associated health insurance, out-of-pocket medical expenses might go up (to compensate for lost insurance) or down (if families choose to simply go without health care); in the latter case, households would appear to have higher disposable income (and thus less likely to be in poverty) but might actually be worse off from an overall perspective. Finally, some analysts point out that home foreclosures hit record levels in 2009, and no poverty measure captures the loss in wealth associated with losing one’s home, although such a loss is certainly an indication of economic distress.

The lesson for researchers is that no measure of poverty tells the whole story by itself. Seemingly minor assumptions can lead to large differences in how poverty is measured, and different measures may be more or less appropriate in different situations. Analysts need to take care when using the official poverty rate and consider how policy conclusions may change if alternative measures are used instead.

Summary

The controversies here about wealth, income, and poverty highlight many data problems common to social science statistics: poor data, multiple ways to define similar concepts, and differences between relative and absolute measures. Studies of wealth are particularly hampered by poor

234

1.

data but at present, the barriers to better data appear insurmountable. The uneven distribution of wealth means that even extremely large surveys will tell us little about the very wealthy; it is unlikely that any of the Forbes 400 will be interviewed in the extremely large sample Current Population Survey, and even more improbable that they will be included in the smaller Federal Reserve survey.

To resolve questions about the income distribution and how it has changed over time, the problem is not so much a lack of data but an abundance of different data sources. As we have seen, there are different ways to measure “income”; what is, or is not, counted as income has implications for conclusions about when and how income changes over time, and how many Americans are living in poverty. Controversies about the distribution of income and the rate of poverty both demonstrate the need to test different assumptions in research projects. This kind of sensitivity analysis can preempt criticism about what may appear to be arbitrary choices, such as years chosen for study, inflation indexes, or poverty lines. Such care can help researchers reach an audience beyond the already convinced.

Finally, the controversies about income mobility and the plight of the middle class highlight the importance of clarifying when comparisons are being made on a relative versus an absolute scale. For example, although it may appear to be more difficult in the United States than in other countries for the poor to move into higher income classes, that does not necessarily mean the poor in America are not also doing better over time. To make sense of apparently contradictory statistics, researchers need to be thorough and up front in identifying what exactly is being compared.

Case Study Questions

Even though the Census Bureau spends considerable effort training its survey takers to gather accurate data, respondents may nonetheless misrepresent their income. one study comparing Current Population Survey (CPS) data with tax returns showed that actual self-employment income was 24 percent higher than reported in the survey; actual government transfer programs other than Social Security were 42 percent higher; and actual property income such as interest, rent, and dividends was 135 percent higher. How might these possible errors affect research on wealth and income distribution?

235

2.

3.

4.

5.

a.

b.

c.

Current Population Survey data on income exclude capital gains — that is, income derived from the profitable sale of homes, stocks and bonds, or other investments. What effect does this exclusion have on studies of the distribution of income?

When asked to guess the distribution of wealth in the United States, respondents in one survey estimated that the wealthiest 20 percent of Americans own 59 percent of the nation’s wealth and the poorest 20 percent own 3.7 percent. How does this compare to the actual distribution? Do you think the results would have been different if the survey had asked people to guess the distribution of income?

Many studies of income focus exclusively on family incomes. What biases are introduced by leaving out nonfamily households when we study the well-being of the “typical” household?

The gap between men’s pay and women’s pay has received much attention in recent years. The gap typically was measured by the ratio of women’s pay to men’s pay, which was approximately 77 percent in 2009. But there are several potential problems with this measure.

Most researchers use weekly earnings measured in the Current Population Survey. However, annual earnings that include second jobs measure a slightly larger pay gap. Why?

Most researchers compare pay for only full-time, year- round workers. Why would the pay gap appear larger if all workers were studied? What are the reasons for studying only full-time, year-round workers? What would be the reasons for studying all workers?

In 1996 women’s median weekly wages fell compared to men, contrary to an increase in the ratio of women’s to men’s full-time, year-round earnings. Experts believe that welfare reform may have been responsible. Why?

References

Data Sources (167)

Wealth (167)

236

Federal Reserve Survey of Consumer Finances (SCF) in Brian Bucks, et al., “Changes in U.S. Family Finances from 2004 to 2007: Evidence from the Survey of Consumer Finances,” Federal Reserve Bulletin, vol. 95 (February 2009), pp. A1-A55. SCF data sample in ibid., p. A27. Wealth surveys summarized in Edward N. Wolff, “Recent Trends in the Size Distribution of Household Wealth,” Journal of Economic Perspectives 12 (Summer 1998): 131–50; nineteenth-century census described in Joseph W. Duncan and William C. Shelton, Revolution in United States Government Statistics (Washington, DC: U.S. Government Printing Office, 1978), p. 5.

Survey of Income and Program Participation (167)

SIPP data sample in Rakesh Kochhar et al., “Wealth Gaps Rise to Record Highs Between Whites, Blacks, and Hispanics,” Pew Research Center, July 26, 2011, http://www.pewsocialtrends.org/2011/07/26/wealth-gaps-rise-to-record-highs- between-whites-blacks-hispanics/.

Indirect Estimates from Tax Records (167)

Indirect estimates in James D. Smith and Stephen D. Franklin, “The Concentration of Personal Wealth,” American Economic Review 64, no. 4 (May 1974): 162–67; James D. Smith, “Recent Trends in the Distribution of Wealth in the U.S.: Data, Research Problems, and Prospects,” in International Comparisons of the Distribution of Household Wealth, ed. Edward N. Wolff (Oxford: Clarendon Press, 1987), pp. 72–89. Methods summarized in Lars Osberg, Economic Inequality in the United States (Armonk, NY: M.E. Sharpe, 1984), pp. 38–45; Barry W. Johnson, “Updating Techniques for Estimating Wealth from Federal Estate Tax Returns,” Statistics of Income Division, IRS, http://www.irs.gov/pub/irs- soi/perwltes.pdf. Lampman data sample in Edward W. Wolff and Marcia Marley, “Introduction and Overview,” in Wolff, International Comparisons, p. 1.

Direct Counts (168)

“The Forbes 400: The Richest People in America,” Forbes, September 21, 2011, http://www.forbes.com/forbes-400/.

Income (168)

U.S. Census data described in Chapter 2. ACS data sample in Amanda Noss, “Household Income for States: 2009 and 2010,” American Community Survey Briefs, ACSBR/10–02 (U.S. Census Bureau, September 2011), p. 2. CPS described in Chapter 9. CPS changes in Charles R. Nelson, “Revising the CPS March Income Supplement to Accommodate State-Level Social Assistance Programs,” Poverty Research News 2 (1): 14–15. CPS data sample in Carmen DeNavas-Walt, et al.,

237

“Income, Poverty, and Health Insurance: Coverage in the United States: 2010,” Current Population Reports, P60–239 (Washington, DC: U.S. Bureau of the Census, September 2011), pp. 41–44. Establishment Survey data sample in Current Employment Statistics Online, U.S. Labor Department, Bureau of Labor Statistics, http://www.bls.gov/ces/. PSID summarized in “An Overview of the Panel Study of Income Dynamics,” http://www.isr. umich.edu/src/psi/overview/html. PSID compared with other sources in Christopher Jencks, Who Gets Ahead: The Determinants of Economic Success in America (New York: Basic Books, 1979), pp. 274–75.

Controversies (169)

Wealth (169)

What Is Wealth? (169): SCF data in Brian Bucks, et al., “Changes in U.S. Family Finances from 2004 to 2007,” Federal Reserve Bulletin. Distribution in Edward N. Wolff, Top Heavy: A Study of the Increasing Inequality of Wealth in America (Washington, DC: Brookings Institution Press, 1994); Edward N. Wolff, “Recent Trends in Household Wealth in the United States: Rising Debt and the Middle- Class Squeeze— an Update to 2007,” Levy Economics Institute of Bard College, Working Paper No. 589, March 2010, http://www.levyinstitute.org/pubs/wp_589.pdf. Retirement wealth in Martin Feldstein, “Social Security, Induced Retirement, and Aggregate Capital Accumulation,” Journal of Political Economy 82 (September 1974): 905–26; and “Perceived Wealth in Bonds and Social Security: A Comment,” Journal of Political Economy 84 (April 1976): 331–36. See also Sylvia A. Allegretto, “The State of Working America’s Wealth, 2011,” Economic Policy Institute, EPI Briefing Paper #292, March 23, 2011, http://epi.3cdn.net/2a7ccb3e9e618f0bbc_3nm6idnax.pdf.

How Wealthy Is Wealthy? (171): Wealth and income difference in Catherine Rampell, “Inequality Is Most Extreme in Wealth, Not Income,” Economix (blog), March 30, 2011, http://economix.blogs.nytimes.com/2011/03/30/inequality-is- most-extreme-in-wealth-not-income/; Robert Gebeloff and Shaila Dewan, “Measuring the Top 1% by Wealth, Not Income,” Economix (blog), January 17, 2012, http://economix. blogs.nytimes.com/2012/01/17/measuring-the-top-1-by- wealth-not-income/; see also Sylvia A. Allegretto, “The State of Working America’s Wealth, 2011,” EPI Briefing Paper #292. SCF v. tax data in Barry Johnson and Kevin Moore, “Consider the Source: Differences in Estimates of Income and Wealth from Survey and Tax Data,” Statistics of Income Division, IRS, http://www.irs.gov/pub/irs-soi/johnsmoore.pdf.

Income (172)

238

Does Income of $250,000 Mean You Are Rich? (172): öbama definitions of rich in Dave Eberhart, “McCain: öbama’s Slippery Definition of ‘Rich,’” Newsmax, öctober 30, 2008, http://www.newsmax.com/InsideCover/obama-definition- rich/2009/12/12/ id/341693; Andrew Ross Sorkin, “Rich and Sort of Rich: How Did $250,000 Become the Magic Number?” New York Times, May 14, 2011, http://www.nytimes. com/2011/05/15/weekinreview/15tax250copy.html. Gallup surveys in Lydia Saad, “Americans Still Split About Whether Their Taxes Are Too High,” Gallup Politics, April 18, 2011, http://www.gallup.com/poll/147152/Americans-Split-Whether-Taxes-High.aspx; Catherine Rampell, “Rich People Still Don’t Realize They’re Rich,” Economix (blog), April 29, 2011, http://economix.blogs.nytimes.com/2011/04/19/ rich- people-still-dont-realize-theyre-rich/; Andrew Gelman, “Upper-Income People Still Don’t Realize They’re Upper-Income,” Statistical Modeling, Causal Inference, and Social Science (blog), April 20, 2011, http://www.stat.columbia.edu/~cook/ movabletype/archives/2011/04/upper-income_pe.html. See also Catherine Rampell, “Who Counts as ‘Rich’?” Economix (blog), December 9, 2011, http://economix.blogs. nytimes.com/2011/12/09/who-counts-as-rich/, and Bruce Bartlett, “Who Counts as ‘Rich’? Continued,” Economix (blog), December 13, 2011, http://economix.blogs.nytimes.com/2011/12/13/who-counts-as-rich- continued/.

Are the Rich Getting Richer? (172): Divergence in wealth and income in Robert Frank, “The Wealth Data Puzzle,” The Wealth Report (blog), April 10, 2007, http://blogs.wsj.com/wealth/2007/04/10/the-wealth-data-puzzle/. Wolff, “Recent Trends in Household Wealth,” Levy Economics Institute of Bard College, Working Paper No. 589; Sam Pizzigati, “Wall Street’s Meltdown Increased Wealth Concentration,” Campaign for America’s Future, April 27, 2010, http://www.ourfuture.org/blog-entry/2010041727/wall-streets-meltdown-and- wealths-maldistribution.

Is Rags-to-Riches Just an American Myth? (174): James Q. Wilson, “Angry About Inequality? Don’t Blame the Rich,” Washington Post, January 26, 2012, http://www.washingtonpost.com/opinions/angry-about-inequality-dont-blame-the- rich/2012/01/03/gIQA9S2fTQ_story.html; Thomas A. Garrett, “U.S. Income Inequality: It’s Not So Bad,” Inside the Vault 14, no. 1 (Spring 2010): pp. 1–3, http://www.stlouisfed.org/publications/itv/articles/?id=1920; Kevin D. Williamson, “The Rich Aren’t Getting Richer,” National Review Online, April 11, 2011, http://www. nationalreview.com/articles/264296/rich-aren-t-getting-richer-kevin-d- williamson#. Intergenerational mobility in Gary Solon, “Intergenerational Income Mobility in the United States,” American Economic Review 82 (June 1992): 393– 408; David J. Zimmerman, “Regression Toward Mediocrity in Economic Stature,” American Economic Review 82 (June 1992): 409–29; “The Born Wealthy or Poor Usually Stay So, Studies Say,” New York Times, May 18, 1992, p. A1; Chul-In Lee and Gary Solon, “Trends in Intergenerational Income Mobility,” Review of Economics and Statistics 91, no. 4 (November 2009): 766–72; Bhashkar

239

Mazumder, “Fortunate Sons: New Estimates of Intergenerational Mobility in the United States Using Social Security Earnings Data,” Review of Economics and Statistics 87, no. 2 (May 2005): 235–55. International comparisons in Jason DeParle, “Harder for Americans to Rise from Lower Rungs,” New York Times, January 4, 2012, http://www.nytimes.com/2012/01/05/us/harder-for-americans-to- rise-from-lower-rungs. html?pagewanted=all.

Are We Better Off? (177)

Per capita income and earnings data in Council of Economic Advisers, Economic Report of the President, 2011 (Washington, DC: U.S. Government Printing Office, 2011), http://www.gpoaccess.gov/eop/. Income per person data advocated in John E. Schwarz and Thomas J. Volgy, “The Myth of America’s Economic Decline,” Harvard Business Review, September-October 1985, pp. 101–2; Jerry Flint, “How Are We Doing?” Forbes, July 13, 1987, p. 94; and Charles Murray, Losing Ground (New York: Basic Books, 1984). David M. Gordon, Fat and Mean: The Corporate Squeeze of Working Americans and the Myth of Managerial “Downsizing” (New York: Free Press, 1996), pp. 24–25. Congressional Budget Office, “Changes in the Distribution of Workers’ Hourly Wages Between 1979 and 2009,” February 16, 2011, http://www.cbo.gov/publication/22010. Family income data in U.S. Bureau of the Census, Historical Income Tables, table F-6, http://www.census.gov/hhes/www/income/data/historical/ families/. Labor force participation in Women in America: Indicators of Social and Economic Well- Being, report prepared by the U.S. Department of Commerce, Economics and Statistics Administration and the Executive Office of the President, Office of Management and Budget for the White House Council on Women and Girls, March 2011, p.http://www.whitehouse.gov/administration/eop/cwg/data-on- women. Limitations of data in Christopher Jencks, “The Politics of Income Measurement,” in The Politics of Numbers, ed. William Alonso and Paul Starr (New York: Russell Sage, 1987), pp. 83–131. Size of families in Urie Bronfenbrenner et al., The State of Americans (New York: Free Press, 1996), p. 59. Terry Fitzgerald, “Where Has All the Income Gone?” The Region, September 2008, pp. 24–57; U.S. Department of Commerce, Middle Class in America, September 2010, http://www.esa.doc.gov/Reports/ middle-class-america; summarized in “Middle Class in America,” Focus 27, no. 1 (Summer 2010): 1–8. See also Uwe E. Reinhardt, “What Does ‘Economic Growth’ Mean for Americans?” Economix (blog), September 2, 2011, http://economix.blogs.nytimes.com/2011/09/02/what-does-economic-growth- mean-for-americans/.

Poverty (181)

orshansky method described in Leonard Beeghley, “The Measurement of Poverty,” Social Problems 31, no. 3 (February 1984): 322–33. NRC/NAS (National Research

240

Council/National Academy of Sciences) report in Measuring Poverty: A New Approach, ed. C.F. Citro and R.T. Michael (Washington, DC: National Academy Press, 1995); summarized in Focus 19, no. 2 (Spring 1998). Poverty line is too low argued in Patricia Ruggles, Drawing the Line: Alternative Poverty Measures and Their Implications for Public Policy (Washington, DC: Urban Institute Press, 1990); Ruggles, “The Poverty Line—Too Low for the 90s,” New York Times, April 26, 1990, p. A23. Poverty line is too high argued in Rose Friedman, Poverty: Definitions and Perspectives (Washington, DC: American Enterprise Institute, 1965); June o’Neill, “Poverty: Programs and Policies,” in Thinking About America, ed. Annelise Anderson and Dennis Bark (Stanford, CA: Hoover Institute, 1988); NRC recommendations in “Definitional Issues in Establishing a New Poverty Measure,” Focus 19 (Spring 1998): 7–9. Supplemental measure in Kathleen Short, “The Research Supplemental Poverty Measure: 2010,” Current Population Reports, P60–241(Washington, DC: U.S. Bureau of the Census, November 2011); see also “Poverty—Experimental Measures,” U.S. Bureau of the Census, http://www.census.gov/hhes/povmeas/ and “Experimental Poverty Measures,” Bureau of Labor Statistics, http://www.bls.gov/pir/spmhome.htm. Comparison of measures in Nancy Folbre, “What Percentage Lives in Poverty?” Economix (blog), November 14, 2011, http://economix.blogs.nytimes.com/2011/11/14/what- percentage-lives-in-poverty/; Sabrina Tavernise and Robert Gebeloff, “New Way to Tally Poor Recasts View of Poverty,” New York Times, November 7, 2011, http://www.nytimes.com/2011/11/08/us/poverty-gets-new-measure-at-census- bureau.html?_r=2&scp=1&sq=poverty&st=cse. Nancy Folbre, “Is Poverty Up or Not?” Economix (blog), February 7, 2011, http://economix.blogs.nytimes.com/2011/02/07/is-poverty-up-or-not/; Arloc Sherman, “Despite Deep Recession and High Unemployment, Government Efforts —Including the Recovery Act—Prevented Poverty from Rising in 2009, New Census Data Show,” Center on Budget and Policy Priorities, January 5, 2011, http://www.cbpp.org/files/1–5-11pov.pdf.

Case Study Questions (185)

1. Lars Osberg, Economic Inequality in the United States (Armonk, NY: M.E. Sharpe, 1984), p. 25.

2. “Income Distribution,” Dollars and Sense, February 1983, pp. 6–7.

3. Drake Bennett, “Commentary: The Inequality Delusion,” BusinessWeek, October 21, 2010, http://www.businessweek.com/magazine/content/10_44/b4201- 008238184.htm.

4. Christopher Jencks, “The Politics of Income Measurement,” in The Politics of Numbers, ed. William Alonso and Paul Starr (New York: Russell Sage, 1987), pp. 92–105.

5. “But What of the Wage Gap?” BusinessWeek, November 3, 1997, p. 30.

241

9

Labor Statistics □□□□

This chapter examines labor statistics, a broad range of data on unemployment, the number of jobs, occupations, union membership, and workplace safety. Because these statistics directly affect so many lives, they are mainstays of public policy. U.S. macroeconomic policies, government training and education programs, equal rights efforts, and regulation of labor relations and the workplace all depend on labor statistics. In each case, however, there are significant measurement problems that complicate policy choices.

Where the Numbers Come From

Organizations Data sources URL Bureau of the Census, U.S. Department of Commerce

U.S. Census of Population; American Community Survey; Economic Censuses

www.census. gov

Bureau of Labor Statistics, U.S. Department of Labor

Current Population Survey; Establishment Survey; Occupational Safety and Health statistics; international comparisons

www.bls.gov

Data Sources

U.S. Bureau of Labor Statistics

A re searcher’s first stop often will be the U.S. Department of Labor’ s Bureau of Labor Statistics (BLS), which publishes a wide range of

242

information about the U.S. workforce. The best-known statistics on unemployment are computed by BLS based on survey data from the U.S. Census Bureau. In other cases, BLS simply publishes statistics collected by other U.S. government agencies. These include the American Time Use Survey and union membership estimated by the U.S. Census Bureau. Also, BLS compiles the Current Employment Statistics, based on a sample of over 400,000 individual worksites.

Data Sample: For June 2011 the Bureau of Labor Statistics estimated the unemployment rate for married men at 5.8 percent; for widowed, divorced, or separated men at 10.5 percent; and for single, never- married men at 15.7 percent.

U.S. Census Bureau

The Bureau of Labor Statistics publishes the most current labor data; the data published by the U.S. Census Bureau are the most comprehensive, although they are less timely than those of the BLS. Supplementing the decennial census is the annual American Community Survey (see Chapter 2).

Data Sample: The 2008 American Community Survey counted 639,620 male electricians. only 1.8 percent of individuals in this occupation were women.

At five-year intervals the U.S. Census Bureau surveys businesses in the Economic Censuses. These surveys are also quite detailed, providing labor force data by economic sector.

Data Sample: In the 2010 U.S. Census of Government, Kentucky local libraries employed 1,303 full-time workers and 1,061 part-time workers.

Controversies

Unemployment

on the first or second Friday of every month, the U.S. Bureau of Labor Statistics announces the unemployment rate for the previous month. It is a touchstone for U.S. economic policy; changes of a fraction of 1 percent are headline news—enough to bolster or shake confidence in a national

243

administration. When national employment statistics were delayed because of the three-week 1995–1996 budget impasse furlough, the stock market had jitters until the economic uncertainty cleared. But practically unnoticed by the media and many users of BLS data is fundamental uncertainty about the interpretation and accuracy of the measured unemployment rate.

How It Came to Be

The Bureau of Labor Statistics came into prominence because of an embarrassing gap in U.S. social statistics. At the height of the Great Depression of the 1930s, no one knew how many people were out of work. Limited funding prevented the BLS from conducting the kind of nationwide survey necessary to measure unemployment accurately. Instead, the BLS relied on partial surveys of the employed by individual states where strong labor lobbies had pressured for comprehensive workplace surveys. These data were misleading as a measure of unemployment; for example, in January 1930, President Hoover declared, “the tide of employment has changed,” based on an apparent upswing in the job count, although within one year more than 3 million people lost their jobs.

After a 1937 postcard survey yielded worthless returns by leaving out everyone without a permanent residence, the Roosevelt administration finally agreed to a large-scale, house-to-house survey. Ironically, the task was assigned to the Works Progress Administration, itself a government effort to provide jobs. The resulting 15 percent unemployment rate estimate proved that the country still suffered from economic depression. Once it had been demonstrated that a direct sampling of individuals could measure unemployment, the U.S. Census Bureau took over, initiating the monthly Current Population Survey (CPS) in 1947. Today it is the largest regularly conducted poll in the world, covering about 60,000 households, and is used to estimate many important social science statistics (see chapters 2, 4, 5 and 8), but a key focus remains the unemployment rate. In the CPS, unemployment has a specific, technical definition: those respondents who are not working (but are not on sick leave) and who actively searched for work during the last four weeks. The unemployment rate is the number of unemployed divided by the labor force, consisting of the employed plus the unemployed. Researchers disagree whether this method accurately measures the level of economic troubles.

244

Undercount

Two well-documented shortcomings in the official definition of unemployment cause it to underestimate the actual number of people in need of work. First, the official statistic leaves out part-timers who would like full-time work; anyone who works, even as little as one hour per week, is counted as employed. Second, discouraged workers who are not actively seeking a job are left out of the labor force entirely; that is, they are neither employed nor unemployed. only those who made a specific effort to find work, such as writing letters, canvassing, or reporting to an agency, are counted as unemployed.

The BLS recognizes these potential limitations to the official unemployment rate. Indeed, the BLS introduced new survey questions in January 1994 that partially corrected the undercount of unemployed homemakers. Nonetheless, according to 2011 BLS statistics, there were almost 9 million part-timers who desired full-time employment and about 2.5 million workers who had looked for work in the last year, but not in the last four weeks, and thus were not counted as unemployed. These and other adjustments to the unemployment rate are published by BLS under the labels U-1 through U-6. The news media typically report only the official one, called U-3, but the other rates provide researchers with important—although seldom used—statistics.

Box 9.1 How to Survey the Unemployed

Old survey: “What were you doing most of last week? … ‘working,’ ‘keeping house,’ ‘going to school.’”

New survey: “Last week did you do any work for pay?”

In January 1994 the Bureau of Labor Statistics introduced new survey questions intended to remove biases in the older format. Most worrisome was the manual instruction about a respondent who “appears to be a homemaker,” who was asked, “What were you doing most of last week—keeping house or something else?” whereas others were asked, “What were you doing most of last week— working or something else?” As a result, too many women were counted outside the labor force and thus technically not unemployed,

245

even though they were actively looking for work at the same time that they maintained a household. To the surprise of BLS analysts, the new survey did not lead to a significant overall increase in unemployment. The reason the new survey failed to change the measured unemployment among homemakers may have been because survey takers had been alerted in advance to the problem. Even before the new questions were introduced, survey takers tried to make their questions less gender biased, thereby raising the measured unemployment rate during the last few months of 1993 so that the new question wording had little effect.

There were other minor but unexpected changes in the measured unemployment rate. For example, the recently retired showed an increase in measured unemployment, probably because in the old format some respondents became weary of questions irrelevant to their retired status. As result, they failed to report attempts to find work when the interviewer finally reached questions about unemployment. These events show the critical nature of survey questions, a topic revisited in Chapter 12. Minor changes in question wording can affect the outcome, as can the attitude of survey takers— often, as in this case—with quite unpredictable outcomes.

Sources: Anne E. Polivka and Stephen M. Miller, “The CPS After the Redesign: Refocusing the Economic Lens” (U.S. Department of Labor, Bureau of Labor Statistics), March 1995; “Jobless Rate Misstated: U.S. Cites Survey Bias,” New York Times, November 17, 1993, p. A1; “Labor Department Reports U.S. Unemployment Rate of 6.7%,” New York Times, February 5, 1994, p. A1. Undercount data in Bureau of Labor Statistics Economic News Release, Employment Situation Summary, table A. Household data, seasonally adjusted, December 2011.

Table 9.1

Unemployment (November 2010)

Official unemployed (actively seeking a job) 15,041,000 Not counted as unemployed … But searched during past 12 months

246

and not currently looking for a job 2,531,000 Because discouraged over job prospects 1,282,000 At work part time, want a full-time job 8,960,000 Because of slack work or business conditions 6.025,000 Because could only find part-time work 2,557,000

Source: U.S. Department of Labor, Bureau of Labor Statistics, The Employment Situation, Summary, table A, “Household data, seasonally adjusted,” December 2011.

Box 9.2 Is Unemployment Structural or Cyclical?

Everyone agreed that unemployment during the Great Recession of 2007–2009 was extraordinarily high, reaching 10.1 percent at its peak in October 2009 and remaining near 9 percent through 2011. However, some economists maintained that most of this unemployment was structural—that is, a mismatch of worker skills and job requirements so that unemployment likely would remain high even when prosperity returned. As evidence that unemployment increasingly was a matter of skills mismatch, researchers pointed to the rising number of job openings, higher than one might expect during a recession. Based on such data, Minneapolis Federal Reserve Bank president Narayana Kocherlakota and Philadelphia Federal Reserve Bank president Charles Plosser argued that there was little government policy could do to counter high unemployment. In their view, high unemployment persisted because workers with high mortgages could not move to locations with more jobs, and workers in the now-decimated home construction industry could not transfer their skills to occupations with more employment.

Other economists strongly disagreed, arguing that changes in structural unemployment were quite small and that most unemployment was still cyclical in nature, meaning it was caused by insufficient demand in the economy and therefore amenable to appropriate government policy. As evidence, they pointed out that even though job openings had increased slightly from about 1.8 percent (of those working plus unfilled positions) in 2009 to 2.3

247

percent in 2010, job openings were still several times lower than the number of unemployed people in every major industrial sector.

As in other recessions, the debate was about the efficacy of government intervention to improve economic conditions. Those who claim a rise in structural unemployment are pessimistic about either government spending or Federal Reserve policies as a way to create better employment conditions. Those who see minimal structural unemployment maintain that the insufficient demand is responsible for the predominantly cyclical unemployment that can be helped with new government efforts.

Sources: Narayana Kocherlakota, “President’s Speeches: Inside the Federal Open Market Committee (FOMC),” Marquette, MI, August 17, 2010, http://www.minneapolisfed.org/news_events/pres/kocherlakota_speech_08172010.pdf; Fed policy in “Jobs Are There, Qualified Applicants Aren’t,” Businessweek, September 26, 2010, p. 17; Plosser in “Fed’s Plosser Says U.S. Deflation Risk Gone: Report,” Reuters.com, February 12, 2011, http://www.reuters.com/article/2011/02/12/us-usa-fed-plosser- idUSTRE71B0GK20110212; critics in John Schmitt and Kris Warner, “Deconstructing Structural Unemployment,” Washington, DC: Center for Economic and Policy Research, May 24, 2011 (corrected version).

Employment

Paralleling the unemployment report is a monthly employment estimate, also by the Bureau of Labor Statistics but based on the BLS Establishment Survey. What often is confusing in media reports is seeming divergence between the unemployment and employment reports. For instance, in April 2011, the unemployment rate increased by 0.2 percentage points, even though 244,000 new jobs were created—the most of the year to that date and above the level needed just to employ the growing population. The apparent contradiction can be explained by movement in and out of the labor force: If more potential workers are encouraged by economic conditions to look for jobs, then the labor force will expand, raising the level of unemployment at the same time that new jobs are created.

As a general rule of thumb, for month-to-month changes economists look to the employment levels rather than the unemployment rate. One reason is that the standard sampling error in the unemployment rate is

248

0.12, variation caused simply by which households are chosen at random for the Current Population Survey. As a result, a 0.1 percent change in the unemployment rate, often headlined as a sign of improving or deteriorating job opportunities, can be expected to occur purely at random about one- third of the time.

Box 9.3 Too Few U.S. Engineers?

Amid fears about U.S. competitiveness and outsourcing of research and manufacturing, the popular media and politicians pointed to statistics showing that China graduated 600,000 new engineers every year compared to only 70,000 in the United States. India also appeared to be ahead with 350,000 annual engineering graduates. At issue was government support, such as a 2011 proposal by President Obama to train an additional 10,000 new engineers a year. However, the engineering gap may have been exaggerated. A Duke University team led by Vivek Wadhwa found that Chinese and Indian data counted two-and three-year degrees, not the four-year bachelor’s degree required in the United States, and that the definition of engineer was vague enough to include Chinese motor mechanics. Correcting for such errors, Wadhwa concluded that the United States actually graduates more engineers than India, and while Chinese graduates were about three times the U.S. level, the quality level in China was far lower than in the United States. At the master’s and doctoral levels, American universities graduated similar numbers to Indian and Chinese universities.

Sources: Carl Bialik, “Counting Engineers,” Wall Street Journal, March 16, 2007, http://blogs.wsj.com/numbersguy/counting-engineers-62/; Vivek Wadhwa, “Seeing Through Preconceptions: A Deeper Look at China and India,” Issues in Science and Technology, http://www.issues.org/23.3/wadhwa.html. Vivek Wadhwa, “About That Engineering Gap …,” Businessweek, December 13, 2005, http://www.businessweek.com/smallbiz/content/dec2005/sb20051212_623922.htm. President Obama’s proposal in Patrick Thibodeau, “Obama: ‘We Don’t Have Enough Engineers,’” Computerworld, June 14, 2011, http://www.computerworld.com/s/article/9217624/Obama_We_don_t_have_enough_engineers_.

249

Debates about labor market statistics teach several lessons about the use of data. The Bureau of Labor Statistics employment survey provides the best guide to short-term, monthly economic changes, albeit subject to later corrections. Although less accurate on a month-to-month basis, the Current Population Survey is useful for broader questions, including extensive demographic information about characteristics of the employed and unemployed. For studies of unemployment, researchers should note that BLS provides a number of unemployment rates, of which U-3, the official and most frequently used rate, is only one (see Table 9.1, p. 196).

The Minimum Wage and Jobs

In the social sciences, the results of a single study seldom change the discipline’s viewpoint on an important issue. Why? Primarily because controlled experiments are rare: Unlike the natural sciences, it is difficult to find a situation in the social sciences in which only the variable under study changes while everything else stays the same. One exception was research on the minimum wage conducted by economists David Card and Alan B. Krueger and published in their 1995 book Myth and Measurement: New Economics of the Minimum Wage. Card and Krueger altered the widely accepted consensus view—by 90 percent of economists in one poll—that increasing the minimum wage reduced employment. By 2006 five Nobel Prize winners and six past presidents of the American Economic Association signed a statement that a higher minimum wage “can significantly improve the lives of low-income workers and their families, without the adverse effects that critics have claimed.”

Card and Krueger’s results prompted a bitter debate, including competing data that seemed to counsel precisely the opposite policy advice. At issue was widely accepted economic theory that wages above those set by the market would reduce an employer’s ability to hire as many workers. In this view, the minimum wage is a misguided attempt to help the poor; it results in a few gaining somewhat higher pay at the expense of others losing their jobs. Card and Krueger’s book introduced contrary evidence based on a 1992 “natural experiment,” when the minimum wage was raised in New Jersey but not in neighboring Pennsylvania. Using telephone surveys of more than 400 fast food restaurants, Card and Krueger found that employment expanded with the increase in the minimum wage.

250

Many economists responded to this finding with disbelief. The Employment Policies Institute, a group partially funded by the fast food industry, pointed out that the Card and Krueger telephone data included odd results such as sudden shifts from zero to 35 full-time employees at one Wendy’s restaurant and an increase from 6.5 to 30 employees at one Burger King. At the forefront of the academic response were economists David Neumark and William Wascher, who claimed to have data showing job loss after the increased minimum wage. They based their claim on payroll records that were allegedly more accurate than the telephone survey data used by Card and Krueger. However, critics charged that part of Neumark and Wascher’s data was suspect because it had been collected by the partisan Employment Policies Institute. Card and Krueger were more diplomatic, pointing out simply that “there are reasons to doubt the representative nature of their data.” In particular, the deleterious effect of the minimum wage disappeared if one left out results from a single Pennsylvania Burger King franchise owner.

Box 9.4 How Many Jobs in a Lifetime?

Career advisers recommend flexibility for workers who want to succeed in an economy that may require new skills and new careers over a typical lifetime. The Current Population Survey asks how long workers have been at the same job, with 2008 data indicating a trend toward job stability likely reflecting the aging workforce. There was a slight increase in jobs held more than 20 years and a small decrease in workers holding their current job one year or less. The number of jobs held by a typical worker over time requires longitudinal data that tracks the same respondent over many years. One of the few such studies is the National Longitudinal Survey of Youth, following over 12,000 individuals since 1979. It found an average of 11 jobs per person for those between the ages of 18 and 44.

Sources: U.S. Department of Labor, Bureau of Labor Statistics, National Longitudinal Surveys, “Frequently Asked Questions: Number of Jobs Held in a Lifetime,” http://www.bls.gov/nls/nlsfaqs.htm#anch41; “Seven Careers in a Lifetime: Think Twice, Researchers Say,” Wall Street Journal, September 4, 2010,

251

http://online.wsj.com/article/SB10001424052748704206804575468162805877990.html.

Card and Krueger further confirmed their results when the minimum wage was raised from $4.25 to $5.15 in 1996, creating yet another natural experiment because the minimum wage changed in Pennsylvania but not in New Jersey, where it already was higher. This time Card and Krueger used payroll records as recommended by Neumark and Wascher and once again they found a small positive employment effect. Neumark and Wascher conceded that their research never found very strong negative effects at fast food establishments, but that bigger effects are evident in what they term “more reliable tests,” primarily analyses in other countries.

In a debate about the accuracy of underlying raw data, it is difficult to ascertain who is correct. The lesson for researchers is that careful work as conducted by Card and Krueger affected the profession’s consensus view, although Neumark, Wascher and other more conservative economists remain unconvinced—as do many business representatives. Card and Krueger earned respect from many economists by making their data public and honestly assessing its shortcomings. In addition, Card and Krueger were clear that their agenda was not simply to gain acceptance of minimum wage policies. They argued from the outset that the minimum wage was an ineffective way to bring about greater equality because it transferred income indiscriminately to many individuals who were not poor. Card and Krueger maintain that attention paid to the minimum wage is out of proportion to its social importance. In their view, economists should focus their research efforts on studying programs more relevant to the poor such as Medicaid and Workers’ Compensation Insurance.

Unions

The U.S. Bureau of Labor Statistics is the main source of national data on labor unions. Recent problems in measuring union membership and strike activity illustrate how researchers must pay careful attention to changes in data-collection methods.

Membership

U.S. union membership has declined dramatically in recent years. The percentage of wages and salary workers who were union members fell to

252

11.9 percent in 2010 from 20.1 percent in 1983. Prior to 1973 the only data came directly from union reports of their own membership. After 1973 the Current Population Survey added a question on union membership, and beginning in 1981 this survey data replaced the direct count in government publications. Some social scientists charge that anti-union sentiment in the Reagan administration prompted the abandonment of the direct count. Yet even the critics admit that there were problems with the older data source, in particular an upward bias caused by exaggerated union membership reports from unions wanting to present a strong image to employers, as well as a downward bias when local unions failed to report all members to the national union in an effort to avoid dues payment. No one knows the extent of either error.

By avoiding self-serving membership reports, Current Population Survey data may be more accurate. Moreover, CPS unionization data can be correlated with CPS questions on industry of employment, occupation, race, sex, and other variables. But such survey data are only as good as the informants’ knowledge. There is evidence that some respondents may not know whether they or other household members actually belong to a union. The primary problem is that many unions, most notably teacher and nursing “associations,” do not use the title “union.” The original 1973 CPS question about union membership was changed in 1976 to include “associations.” But it was found that respondents answered “yes” if they simply belonged to a workplace club or association. So the question was changed once again in 1979, this time to read “union or employee association similar to a union.”

Strikes

In the event of a labor strike, the CPS sample is too small to estimate with any accuracy the number of striking workers. Consequently, the BLS relies on newspaper, magazine, and government reports to assemble data on labor strikes, including the number of strikes held, the number of workers involved in each strike, and the resulting days idle. These data are comprehensive for strikes involving more than 1,000 workers; however, budget cutbacks in 1981 caused termination of data on smaller strikes (involving at least six workers for one eight-hour shift). According to BLS officials, these data were deleted in order to preserve the quality of such statistics as the Consumer Price Index, which is more important for policy decisions. Thus, a full picture of labor relations requires information that is no longer available.

253

Data on strikes are important for labor statistics such as the unemployment rate. For example, in September 2011, the apparent good news of an employment increase of 103,000 was offset by the return of 45,000 Verizon workers who had been on strike and thus counted as unemployed in the previous month. The Verizon workers’ reentry into the labor force skewed employment data for that month. Excluding those 45,000 returning workers’ positions, only about half as many “new” jobs were created as the employment increase might indicate.

Is the Workplace Safe?

Concern about occupational safety and health is a long-standing BLS tradition. In 1909 a BLS study documented phosphorus poisoning in the match industry, causing U.S. Congress to impose a tax on the dangerous product. In 1910 the BLS began publishing accident statistics based on reports from individual states and insurance companies. Congress refused authorization for a separate division on safety, so in subsequent decades safety data covered only 25 percent of the workforce. Under the 1970 Occupational Safety and Health Administration (OSHA) legislation, BLS coverage was expanded, but it still excluded most small farms, the self- employed, and government workers. In 1992 BLS redesigned the survey in two ways: (1) Deaths are now counted in a Census of Fatal Occupational Injuries covering almost all workers, and (2) injuries and illnesses are now measured in a sample survey of 250,000 private businesses that includes previously uncollected demographic data on workplace victims, as well as better information on the severity and circumstances of the injury or illness.

Even with its improvements, BLS data still undercount workplace injury and illness. These data, unlike the fatality count, continue to exclude the self-employed and workers in households and on small farms. Beginning in 2008 BLS expanded its sample to include state and local government employees and explored ways to collect data on Federal workers. Prior to this expanded data base, the excluded groups caused an estimated 25 percent undercount in workplace injuries and illness. More difficult to correct is a count of illnesses with long latency periods such as exposure to carcinogens. According to BLS staff, these cases simply cannot be estimated with an employer-based recording mechanism.

In 2009 and 2010 workplace fatalities dropped to the lowest level since the Census of Fatal Occupational Injuries was initiated in 1992. The

254

overall trend had been downward since 1992, but BLS researchers ascribed some of the recent decline to low employment levels, particularly in some historically high-risk industries. The biggest decline in deaths occurred for vehicular highway accidents, an indication of fewer goods being delivered. Fishing-related jobs remained the most dangerous, with two fatalities per 1,000 workers, more than triple the death rate for the next most dangerous jobs in logging and aircraft piloting.

Squirrel Cage or Easier Times?

Do we work more than in the past so that our rising standard of living comes at the cost of reduced family and leisure time? Is our “better life” the result of an imbalance in which men or women share an unequal burden? Answering these and other questions requires data on how people spend their time. Aggregate data from business firms, accurate for work hours in the total economy, is not useful for analyzing individual work time because firm payroll accounts omit the self-employed and do not identify the increasing number of workers with two jobs. Most studies rely on the CPS, a large sample with additional demographic data on individuals and households. However, the CPS asks respondents to recall episodes of work across a specified time period—and to do so in a few seconds. Not surprisingly, research experts have identified a number of errors in this data source, notably the tendency for respondents to exaggerate their number of work hours, particularly for overtime.

The primary alternative to the CPS survey and a check on its accuracy is the American Time Use Survey (ATUS), conducted by the Census Bureau for the Bureau of Labor Statistics. First collected in 2003, the ATUS uses a time diary approach in which respondents recall activities for a specific period. In comparison with CPS data, the ATUS measures fewer hours at paid and unpaid work and is thought to be more accurate. One reason is that time diaries eliminate errors from respondents who may mistakenly include as “work time” hours spent commuting, on breaks, or taking care of unexpected personal crises. Also, diaries take into account multitasking, for example, doing laundry, providing child care, and preparing food, all at the same time. Correct measurement of multitasking is important for research on housework, causing the potentially inaccurate assessment of contributions by men and women if one group is more likely to multitask. Both men and women report less household work in time diaries than in surveys: Men record 56 percent less time; women, 47 percent.

255

Box 9.5 Are Men Overworked?

Time diary studies show that the hours women spend on household work fell in recent decades, while contributions by men increased. In The Myth of Male Power, Warren Farrell uses this data to conclude that men work more total hours in an average week than women, including both paid work and housework. He ascribes many of men’s mental and physical health problems to their high work burden. However, using the same data source, Carin Rubenstein, author of The Sacrificial Mother, points out that women’s average contribution to household work is still more than double that of men and four times as much when it comes to child care. In her view, men need a “reality check” if they believe that family life is fair and equitable.

Sources: Warren Farrell, The Myth of Male Power (New York: Simon and Schuster, 1993); Carin Rubenstein, “Superdad Needs a Reality Check,” New York Times, April 16, 1998, p. A17.

Also at issue in time-use studies is the trend over time: Are Americans working more than in the past? Using CPS data, Harvard University economist Juliet Schor’s book The Overworked American: The Unexpected Decline of Leisure found Americans working an average of 163 more hours per year in 1990 than in 1970, or what Schor calls an “extra month of work” every year. However, using time diary data, University of Maryland sociologist, John P. Robinson maintains that Schor’s “overworked” American is primarily due to inaccurately reported work hours so that we “have more time than ever before.”

Ideally, we need additional data sources to determine the truth about workers and labor statistics in the twenty-first century. Researchers have experimented with alternatives such as the “experience sampling method,” in which respondents record their activity every time an electronic beeper goes off, or an even more costly approach in which an investigator shadows a subject, recording activity every ten seconds. In the absence of such expensive-to-collect data, researchers will need to rely on a combination of CPS-type surveys (useful because of their large sample

256

sizes), as well as time diaries such as the ATUS (potentially more accurate in terms of actual work hours).

International Labor Statistics

Comparisons of labor statistics between countries are tricky because similar-sounding statistics, such as the labor force, unemployment, and employment, may differ in meaning. Researchers would be at a loss if it were not for the help available from statistical agencies. Regularly published BLS bulletins adjust data on the labor force, employment, and unemployment so that it is comparable between countries. Additional comparison data are assembled by the International Labor Organization and the Organisation for Economic Cooperation and Development.

Unemployment

Several problems confound direct comparison of unemployment rates in different countries. The labor force is defined variously to exclude those people under 16 years old (in the United States), those under 15 years old (in many other countries), full-time students looking for work (Sweden), and those waiting for a job to begin (Japan). Also, the definition of those available for work varies from those who actively searched for a job during a four-week period in the United States, to a 60-day period in Sweden, to an unspecified amount of time in Japan.

A second problem is variation in the level of distress caused by unemployment in different countries. Relatively comprehensive social welfare systems in Western Europe cause unemployment rates to mean quite different economic circumstances than they would in the United States. In the Netherlands, for example, it is estimated that reported unemployment would rise by 3 percent if it were not for relatively generous benefits that enable otherwise unemployed workers to retire and thus leave the labor force.

Employment

Which country creates the most jobs? During the 1980s, supporters of the Reagan and Bush administrations pointed proudly to Establishment Survey statistics measuring more new jobs in the United States than similar surveys counted in Western Europe and Japan combined. But data incompatibility clouds interpretation of these intercountry comparisons.

257

The European workforce is older than the U.S. workforce, so a smaller proportion of the population is in the traditional working-age group for whom jobs must be created. And in both Europe and Japan, many fewer teenagers work, again reducing the need for new jobs. (Japanese teenagers have never worked in large numbers.) In Germany, however, it appears that many teenagers have left the labor force precisely because work was unavailable.

Work Hours

Do Americans work more than Europeans, possibly trading off a slightly higher GDP per capita for much less leisure time? In Were You Born on the Wrong Continent? How the European Model Can Help You Get a Life, labor lawyer Thomas Geoghegan argues that higher U.S. average incomes come at the cost of substantially longer U.S. work hours: 1,778 per year compared to 1,554 hours in France, 1,419 hours in Germany, and 1,377 hours in the Netherlands. U.S. Department of Labor statistician Susan E. Fleck concludes that the differences are real, but that inconsistencies in the way data are collected may exaggerate the U.S. number and underestimate European estimates. The reason is that establishment and survey data used in the United States lead to slightly higher work-hour estimates than national account data used in most European countries.

Unionization

Comparison of union membership between countries requires additional caution. The U.S. Labor Department publishes periodic overviews of union membership. One such study from 2006 compared 24 countries. Foremost, data in some countries, including the United States, is based on household surveys, whereas counts in other countries must rely on union’s administrative reports. Household survey questions differ in the way they deal with subtleties such as multiple job holders and membership in “staff associations” rather than official unions. Even greater error occurs in administrative totals that are pieced together from annual reports, websites, and tax declarations. These records may include nonpaying members, supporters outside the labor force, or simply inflated roles in order to improve the union’s apparent status. Adjusting for these factors, the Labor Department study estimates that actual union membership should be reduced by an average of 24.2 percent for 14 European countries. Based on these adjusted figures, unionization ranged from 8.3 percent in France to 78 percent in Sweden. The U.S. rate was 12.4 percent.

258

1.

2.

3.

In summary, international comparisons require careful attention to how data are collected. Again, this problem is recognized by the U.S. Bureau of Labor Statistics as well as by the international statistical organizations that provide standardized statistics for most major countries.

Summary

Seemingly small changes in labor statistics can have dramatic human consequences. Consider the U.S. unemployment rate, where every 1 percent change translates into at least 1 million more jobs. An increase of 1.0 per 100,000 in U.S. job-related fatalities means more than 1,000 additional deaths. But as this chapter demonstrates, there are many problems in measuring labor statistics, ranging from definitional issues for the unemployment rate to changing survey techniques for union membership. The challenge for researchers is to understand the limitations of these numbers without losing sight of their effects on human lives. For instance, the official U.S. unemployment rate is an extremely limited and easily criticized single statistic. But it is possible to obtain a fascinating, and relatively complete, picture of the U.S. labor markets— if researchers look at all the statistics available, including alternative unemployment statistics, employment counts, and the causes and duration of unemployment. Similarly, there are flaws in the most commonly cited statistics about job quality, union membership, strikes, and workplace safety. International data compound each of these issues, requiring yet more skepticism about the comparisons of simple statistics. But, properly interpreted, the vast quantity of data collected has the potential to illuminate our understanding of labor issues.

Case Study Questions

The official unemployment rates for teenagers are quite high, 25 percent in 2011. But this statistic does not mean 25 percent of all teenagers are out of work. What does it in fact measure?

Using the CPS employment questions, explain why a high percentage of farm women with children in the North Central farm states are reported as “working.” How do these responses affect the measured unemployment rate for women with children?

In Japan between 1955 and 1975, there was a twofold increase in

259

4.

the number of women attending college. Also, many women moved from farms, where they had combined farm work with child care, to cities, where they had no paid employment. How might these factors explain declining labor force participation for Japanese women, the opposite of the trend in most other countries? (Hint: Few Japanese college students have outside jobs.)

The number of fatalities for cell-tower workers jumped to 18 in 2006 from 7 in 2005, leading the Occupational Health and Safety Administration head to call cell-tower climbing “the most dangerous job in America.” What additional data is needed to support this claim?

References

Data Sources (192)

U.S. Bureau of Labor Statistics (192)

Detailed official descriptions in U.S. Department of Labor, Bureau of Labor Statistics, BLS Handbook of Methods, http://www.bls.gov/opub/hom/. Data sample in Bureau of Labor Statistics, Labor Force Statistics from Current Population Survey, Household Data Not Seasonally Adjusted, Table A-29: “Unemployed persons by marital status, race, Hispanic or Latino ethnicity, age and sex,” http://www.bls.gov/web/empsit/cpseea29.pdf.

U.S. Census Bureau (193)

On U.S. Census, see notes to Chapter 2. Data sample in Jennifer Cheeseman Day and Jeffrey Rosenthal, “Detailed Occupations and Median Earnings: 2008” (highlights from the American Community Survey), U.S. Bureau of the Census, www.census.gov/hhes/www/ioindex/acs08_detailedoccupations.pdf. On Economic Censuses, see “Overview of the 2012 Economic Census,” U.S. Bureau of the Census, http://www.census.gov/econ/census12/. Data sample in 2010 Annual Survey of Public Employment and Payroll, U.S. Bureau of the Census, http://www2.census.gov/govs/apes/101ocky.txt.

Controversies (193)

Unemployment (193)

Measuring unemployment during the Great Depression, see William T. Moye and

260

Joseph Goldberg, The First Hundred Years of the Bureau of Labor Statistics (Washington, DC: U.S. Government Printing Office, 1985), pp. 125–77; Margo J. Anderson, The American Census: A Social History (New Haven, CT: Yale University Press, 1988), pp. 162–89. “The Tide of Employment” in Moye and Goldberg, The First Hundred Years, p. 130. History of Current Population Survey in John E. Bregger, “The Current Population Survey: A Historical Perspective and BLS’s Role,” Monthly Labor Review 107, no. 6 (June 1984): 8–14. April 2011 data on employment in Heidi Shierholz, “Jobs Data Send Mixed Messages,” Economic Policy Institute, May 6, 2011, http://www.epi.org/publication/jobs_data_send_mixed_messages/; sampling error in U.S. Department of Labor, Bureau of Labor Statistics, Employment and Earnings, February 2006, p. 195.

The Minimum Wage and Jobs (199)

David Card and Alan B. Krueger, Myth and Measurement: The New Economics of the Minimum Wage (Princeton, NJ: Princeton University Press, 1997); Employment Policies Institute, The Minimum Wage Debate: Questions and Answers, 3rd ed. (Washington, DC: Employment Policies Institute, 1997); Richard Berman, “The Crippling Flaws in the New Jersey Fast Food Study,” 2nd ed. (Washington, DC: Employment Policies Institute, 1996), http://epionline.org/studies/epi_njfastfood_04–1996.pdf; Card and Krueger, “A Reanalysis of the Effect of the New Jersey Minimum Wage Increase on the Fast- Food Industry with Representative Payroll Data,” National Bureau of Economic Research (NBER) Working Paper 6386, January 1998, p. 18. David Neumark and William Wascher, “Minimum Wages and Employment,” American Economic Review 90, no. 5 (December 2000): 1362–420; Liana Fox, “Minimum Wage Trends: Understanding Past and Contemporary Research,” Economic Policy Institute Briefing Paper #178, October 25, 2006.

Unions (201)

Membership(201): Union membership data in Edward C. Kokkelenberg and Donna R. Sockell, “Union Membership in the United States, 1973–1981,” Industrial and Labor Relations Review 38 (July 4, 1985): 497–533; Henry S. Farber, “The Recent Decline of Unionization in the United States,” Science, November 13, 1987, pp. 915–20.

Strikes(202): Strike data described in U.S. Department of Labor, Bureau of Labor Statistics, Monthly Labor Review 110, no. 1 (January 1987): 78. Impact of Verizon strike in U.S. Department of Labor, Bureau of Labor Statistics, “Employment Situation News Release,” October 7, 2011.

Is the Workplace Safe? (202)

261

History of BLS safety and health investigations in Moye and Goldberg, The First Hundred Years, pp. 58–61, 99–101, 132–33, and 251–53. Debate about effectiveness of OSHA in Kenneth B. Noble, “For OSHA, Balance Is Hard to Find,” New York Times, January 10, 1988, p. E5. New OSHA surveys and data in U.S. Department of Labor, Workplace Injuries and Illnesses in 1992, USDL-93–55 (Washington, DC: U.S. Government Printing Office, 1993); BLS, “Census of Fatal Occupational Injuries,” www.stats.bls.gov/oshfatl.htm, and “Workplace Injury and Illness Summary,” http:/stats.bls.gov/news.release/osh.nws.htm. Justin Lahart, “Workplace Fatalities Fall to Historic Low,” Wall Street Journal, August 20, 2010, p. A3. BLS changes in John W. Ruser, “Examining Evidence on Whether BLS Undercounts Workplace Injuries and Illnesses,” Monthly Labor Review, 131, no. 8 (August 2008): 20–32. Most dangerous jobs in Carl Bialik, “Do Cell-Tower Climbers Have the Nation’s Deadliest Job?” Wall Street Journal, July 21, 2008, http://blogs.wsj.com/numbersguy/do-cell-tower-climbers-have-the-nations- deadliest-job-381/.

Squirrel Cage or Easier Times? (203)

Time diary in John P. Robinson and Ann Bostrom, “The Overestimated Workweek? What Time Diary Measures Suggest,” Monthly Labor Review, August 1994, pp. 11–20, and John P. Robinson et al., “The Overestimated Workweek Revisited,” Monthly Labor Review, June 2011, pp. 43–53. Juliet B. Schor, The Overworked American: The Unexpected Decline of Leisure (New York: HarperCollins, 1991). Multitasking in Robert W. Drago and Jay C. Stewart, “Time- Use Surveys: Issues in Data Collection on Multitasking,” Monthly Labor Review, August 2010, pp. 17–31, and Nancy Folbre et al., “By What Measure? Family Time Devoted to Children in the United States,” Demography 42, no. 2 (May 2005): 373–91.

International Labor Statistics (204)

International labor statistics described in BLS Handbook of Methods, http://www.bls.gov/opub/hom/. Bush on job creation in the New York Times, August 19, 1988, p. A1; European workforce in McMahon, “An International Comparison of Labor Force Participation,” Monthly Labor Review, May 1986, pp. 3–12. Data on annual work hours in Organisation for Economic Co-operation and Development StatExtracts, “Average Annual Hours Actually Worked per Worker, 2000–2010,” http://stats.oecd.org/Index.aspx?DatasetCode=ANHRS. Thomas Geoghegan, Were You Born on the Wrong Continent? How the European Model Can Help You Get a Life (New York: New Press, 2010). Susan E. Fleck, “International Comparisons of Hours Worked: An Assessment of the Statistics,” Monthly Labor Review, May 2009, pp. 3–28. Union membership in Jelle Visser, “Union Membership Statistics in 24 Countries,” Monthly Labor Review, January 2006.

262

Case Study Questions (207)

1. See Eleni Theodossiou, “U.S. labor makert shows gradual improvement in 2011,” Monthly Labor Review, March 2012, pp. 3–23.

2. “Clippings,” Ms., October 1993, p. 89.

3. Women in Japanese labor force in International Labor Office, World Labor Report 1 (Geneva, Switzerland: International Labor Office, 1984), p. 54. Council of Economic Advisers, Economic Report of the President, 1990 (Washington, DC: U.S. Government Printing Office, 1990), p. 340.

4. Carl Bialik, “Do Cell-Tower Climbers Have the Nation’s Deadliest Job?” Wall Street Journal, July 21, 2008, http://blogs.wsj.com/numbersguy/do-cell- tower-climbers-have-the-nations-deadliest-job-381/.

263

10

Business Statistics □□□□

Information about business is itself a very profitable business, serving a large market of information-hungry investors. Thus, for research purposes, primary data directly from businesses are plentiful, as are secondary sources interpreting these data. The Data Sources section that follows reviews this vast field. The Controversies section focuses first on the many ways to measure business size, including sales, assets, stock market price, and profits. These statistics are important as indicators of the success of individual firms and for public policy to regulate private business, including antitrust law. other controversies covered in this chapter are the number of jobs created by small versus large businesses, the effectiveness of government bailout programs after the 2008 financial crisis, the actual level of corporate profits, and measuring success in the stock market.

Where the Numbers Come From Organizations Data sources URL U.S. Securities and Exchange Commission

Corporate reports

www.sec.gov

U.S. Federal Trade Commission

Corporate reports

www.ftc.gov

Fortune, Forbes magazines

Fortune 500 Forbes Global 2000

www.money.cnn.com/magazines/fortune/ www.forbes.com

U.S. Small Business Administration

Statistics of U.S. Businesses, Business

www.sba.gov

264

Dynamics Statistics (BDS), and Business Employment Dynamics (BED)

Data Sources

The availability of data depends critically on the type of business being studied. The approximately 32 million U.S. businesses filing tax returns are divided into two basic forms of legal organization. Corporations are least in number, about one-fifth of all businesses, but predominant in economic importance, accounting for over 80 percent of receipts. Corporate data are voluminous and usually publicly available. In contrast, more than 26 million sole proprietorships and partnerships account for less than 20 percent of receipts. Data on these mostly smaller businesses are much less accessible.

Public Corporations

SEC Disclosure

The financial distress of the Great Depression, coupled with complaints about corporate misrepresentation, led to the creation of the Securities and Exchange Commission (SEC) in 1933. In addition to monitoring the stock exchange and other capital markets, the SEC sets strict guidelines for corporate disclosure—that is, what corporations must tell the public about the type of goods or services produced, the names of corporate officers, the names of major stockholders, and most important, the state of corporate finances. Lists of companies reporting to the SEC—and the data as well— are available directly from the SEC. A few large corporations are privately held, meaning they do not sell stock to the public. Data on these private firms are considered separately in this chapter.

Annual Reports

The SEC’s financial disclosure requirements for stockholders are fulfilled by the annual report, a publication expanded by most corporations to a multipage glossy magazine or online report that advertises the company’s investment potential. The annual report is an easily accessible source of data, available in many public libraries or from the corporation itself.

265

Nevertheless, experts warn users not to rely entirely on annual reports. Additional, possibly more revealing, information is available to the public in SEC reports. Moreover, the text that fills most space in annual reports can be misleading; corporations may downplay unsavory prospects in this written material.

Data Sample: From the 2007 Lehman Brothers Annual Report: “In 2007, Lehman Brothers produced another year of record net revenues, net income, and earnings per share and successfully managed through the difficult market environment…. We effectively managed our risk, balance sheet, and expenses. Ultimately, our performance in 2007 was about our “One Firm” sense of shared responsibility and careful management of our liquidity, capital commitments, and balance sheet positions. We benefited from our senior level focus on risk management and, more importantly, from a culture of risk management at every level of the Firm.” In 2008, Lehman Brothers went bankrupt.

On the other hand, Marie Callender, a restaurant and pie-making business facing impending bankruptcy, warned in its 2007 annual report: “We have experienced a decline in same store sales and an increase in net losses in recent quarterly periods…. Management expects these initiatives together with the Company’s cash provided by operations and borrowing capacity to provide sufficient liquidity for at least the next twelve months. However, there can be no assurance as to whether these or other actions will enable us to generate sufficient cash flow to fund our operations and service our debt…. If the current economic conditions persist or worsen, our revenues are likely to continue to suffer and our losses could increase.”

Handbooks

For research on a number of corporations, it is most convenient to consult handbooks containing data collected by private companies such as Moody’s, Standard & Poor’s, Value Line, and Dun & Bradstreet. Each of these handy reference books is updated frequently to include all key data from SEC reports, as well as recent stock prices, debt, capitalization, and often an impartial assessment of each corporation’s current financial status. Supplementary handbooks help researchers find corporate addresses, subsidiary ownership, corporate leadership, and product lines.

266

Business Magazines

Based on the popularity of the Fortune 500 largest industrial corporations, first introduced in 1955, there is now also a Forbes 2000 global ranking with data on corporate sales, profits, assets and market value.

Privately Held Corporations

Some major U.S. corporations are nearly exempt from public scrutiny because they do not have publicly traded stock. Most are companies owned by only a few shareholders such as the Mars Candy Company, which is owned by members of the Mars family. Forbes magazine uses outside sources, and in some cases voluntary disclosure, to estimate the size of the 500 largest privately held firms, topped in recent years by Cargill, Inc., a Minneapolis grain-trading company with about $110 billion revenues in 2010.

For data on privately held companies, researchers should first check out business magazine and newspaper indexes to see if a reporter has already conducted the proposed research. Next, they should look in directories such as Dun & Bradstreet and Thomas Register, which include privately held firms. Finally, additional information can be gained from records of corporate formations filed with state offices, and real estate, litigation, and court judgments in local jurisdictions.

Data Sample: In D&B’s Million Dollar Directory, we learn that Kut Rate Fashions, Inc., of Columbia, South Carolina, had 85 employees and annual sales of $11 million.

Small Businesses

The great majority of U.S. businesses are small, meaning they employ fewer than 500 workers. Small businesses constitute over 99 percent of all businesses, although these small firms comprise less than one-half of total U.S. payrolls. For data on individual small businesses, a researcher may need to do some data digging. Local libraries are a good place to start; librarians often are familiar with sources on local businesses, and reference works on local economies are frequently in their collections. Private data firms, most notably Dun & Bradstreet, collect data on small firms, but getting access to the data is expensive.

Aggregate Statistics

267

Many research projects require aggregate statistics—that is, combined data for groups of businesses, perhaps in a single location or in a single line of business. For these numbers, U.S. government statistics are almost certainly a researcher’s starting point. The U.S. Census Bureau’s Statistics of U.S. Businesses and Business Dynamic Statistics provide a core set of data, as does the U.S. Bureau of Labor Statistics’ Business Employment Dynamics program. For information on businesses without employees, see the Census Bureau’s Nonemployer Statistics program.

Data Sample: The Census Bureau’s 2009 Nonemployer Statistics counted 548 single-person fishing firms, generating receipts of $36,649,000.

Controversies

Who Is the Biggest of Them All?

The size of a business can be measured in many ways: by the value of a firm’s sales, by the value of a firm’s assets, by the value of a firm’s profits, or by the number of a firm’s employees. As shown in Table 10.1, these measures of size cause different rankings of U.S. companies; Walmart, Exxon/Mobil, and Fannie Mae can all legitimately claim the number one spot. There are advantages and disadvantages to each measure of size, so researchers must consider carefully which to use.

Table 10.1

U.S. Corporate Size, 2011 (in billions)

Sales Assets Stock market value

Profits

1.Walmart $422 1.Fannie Mae $3,222

1.Exxon/Mobil $415

1.Exxon/Mobil $30

2.Exxon/Mobil $355

2.Bank of America $2,265

2.Apple $324 2.AT&T $20

3.Chevron $196 3.Freddie Mac $2,263

3.Microsoft $215 3.Chevron $19

268

4.ConocoPhillips $185

4.JP Morgan/Chase $2,118

4.Chevron $214 4.JP Morgan/Chase $17

5.Fannie Mae $154 5.Citigroup $1,914

5.Berkshire Hathaway $211

5.IBM $15

Source: Fortune 500 (2011).

Sales

The oldest and most famous ranking of U.S. corporations, the Fortune 500, is based on annual revenue. When the list was first published in 1955, economist M.A. Adelman wrote to the magazine to complain that sales data are “almost meaningless” because they include production by the company’s suppliers that is passed on in the company’s sales. As an example, Adelman pointed to the Du Pont Corporation with $1.7 billion sales, which appeared smaller than meat processor Swift & Co. with $2.5 billion sales. But almost all of Du Pont’s sales were based on its own production, whereas 90 percent of Swift’s sales came from supplier costs passed on to customers. Fortune admitted the data were imperfect but defended them as the best available. Today, Walmart is similarly placed as one of the largest business based on its revenues, whereas it is not in the top five if size is measured by stock market value or profits.

Assets

In theory, a second measure of size, the value of assets, should give a good comparative-size statistic valid between different sectors of the economy. The problem with assets (as every accounting student quickly learns) is that, unlike the items measured in sales data, few assets are actually bought or sold during a given year. As a result, accountants must estimate asset value based on changes in value since they were first purchased. The procedure is called depreciation, a complex formula oriented toward taxation, not actual demise of the asset. Smart investors know to investigate closely before trusting the measured assets as the actual value of a company. Despite this drawback, asset value is commonly used in research on diverse sectors of the economy where sales data would be misleading. For instance, the debate about the concentration of corporate power discussed later is based on estimates of assets owned by the largest corporations.

269

Market Value

A third measure of size is based on the combined value of all stock owned in the corporation. This number reflects both the value of funds invested in the corporation when stock was first issued and subsequent changes in the corporation’s value based on the rise and fall of the stock price. In theory, the combined knowledge of all stock investors may provide a better measure of corporate values than any single number in the corporate accounts. The drawbacks to stock values as a measure of size are twofold. First, changing stock prices mean that the “value” of the firm changes literally from moment to moment as the stock price varies. Researchers typically choose an arbitrary day for measuring stock prices; even a short time earlier or later, stock prices may be significantly different.

Second, aside from day-to-day fluctuations, there are uncertainties about the meaning of long-run stock prices, in particular whether this is actually a measure of underlying corporate values. There is ongoing debate among economists and business analysts about the efficiency of the stock market, both as an indicator of differences between corporations in terms of their productive capacity and the ability of the stock measure to assess the economy’s overall strength.

In summary, stock prices are a convenient statistic in the sense that they are widely reported and therefore are readily available in print or electronic format. The problem is that no one knows for certain what these voluminous data are telling us.

Employment

Finally, the number of employees is an indicator of size. For organizing data on varied types of firms, employment is a convenient size criterion for several reasons. First, counting employees is less ambiguous than using accounting principles to estimate sales or assets. Second, employment data are applicable to all firms, whether incorporated or not. And third, employment size is a statistic likely to be known by employees, and thus can be measured in household surveys as well as business surveys.

Two precautions in using employment data are required. First, employment gives only a general indication of size, differentiating between “large” and “small” businesses, but it cannot be used for comparing two companies of similar scale. For example, in 2007 the labor- intensive retailer Target generated one-third the revenue of the manufacturing and finance conglomerate General Electric, even though

270

both had about 300,000 employees. Second, researchers must distinguish between business establishments (locations where the business operates) and enterprises (legally constituted business that may have many establishments). The distinction is important—especially in household surveys where respondents may not know about enterprise employment beyond their workplace establishment.

Implications

one obvious solution to the problem of too many ways to measure size is to use more than one statistic; for instance, the Forbes 2000 is a composite in which the magazine’s staffers add the ranks for sales, profits, assets, and market value, so that JP Morgan Chase ranked number one despite not leading in any single category. When it is necessary to choose one of the rankings, however, researchers should ascertain which indicator is most appropriate. For example, investigation of market control within an industry typically looks at sales data, whereas research on the source of profitability—is a big firm more likely to be successful than a small one? —uses asset data; as a third option, studies of corporate takeovers are based on market value. A well-justified research project will explain why one size measure was chosen for a given case instead of the others.

Are the Big Too Big?

Probably the most important public policy use for data on business size is antitrust law. At issue is government intervention to limit the power of one business entity if it is too big or to restrict the power of several corporations if together they wield too much economic power. Not surprisingly, the criteria for “too big” and “too much economic power” are a sticking point—and the source of much statistical controversy.

In some historic cases, there was consensus that corporations had become too big, as in the nearly total monopoly gained by the American Tobacco Company and the Standard Oil Company. In 1911 both these firms were forced to split into smaller, competing entities. Most antitrust cases, however, involve a middle ground in which no single firm dominates an industry; instead, a few large firms share the market. Antitrust laws are vague about what to do in such situations. The Sherman Antitrust Act forbids the “attempt to monopolize,” while the Clayton Antitrust Act forbids business behavior that will “substantially lessen competition.” Consequently, the U.S. courts and federal regulators have attempted to develop more specific guidelines for antitrust policy based on

271

measures of market shares of sales.

Box 10.1 Which Is Larger: GM or Switzerland?

A frequently cited criticism of the power of large corporations is their size, measured by annual sales, in comparison to the gross domestic product (GDP) of relatively large countries. In this manner, Walmart, the largest U.S. corporation in sales in 2007, ranked ahead of Austria, Greece, Denmark, and South Africa.

GDP or Annual Sales—2007 (billions)

1. U.S. $13,811   .   .   . 25. WALMART STORES $378 26. Austria $377 27. EXXON/MOBIL $373 28. Greece $360 29. ROYAL DUTCH/SHELL $356 30. Denmark $308 31. BRITISH PETROLEUM $291 32. South Africa $278

Technically this list is an invalid comparison of statistical apples and oranges. GDP measures a country’s total production, whereas corporate sales include that corporation’s production plus production by other firms used as raw materials, machinery, and other inputs. A correct comparison would measure corporate production by what economists call “value added,” the concept used in GDP accounts to avoid the double counting that arises in corporate data.

Source: Bradley R. Schiller, The Micro Economy Today (New York: McGraw-Hill, 2010), p. 225.

Starting in 1982 federal regulators began using a measure of market

272

shares called the Herfindahl-Hirschman Index (HHI). Essentially this index has the same impact as the former criteria based simply on market share of sales controlled by the top four or eight firms, a statistic published in the Commerce Department’s Economic Censuses (see Box 10.2 for examples). But these seemingly straightforward guidelines have been complicated in actual practice.

Box 10.2 Top of the World (2011)

Sales Assets Market Value Profits 1. Walmart 1. Fannie Mae 1. Exxon/Mobil 1. Nestle

(Switzerland) 2. Royal Dutch

Shell (The Netherlands)

2. BNP Paribus (France)

2. Apple 2. Exxon/Mobil

3. Exxon/Mobil 3. Deutsche Bank (Germany

3. PetroChina 3. Gazprom (Russia)

4. BP (United Kingdom)

4. HSBC (United Kingdom)

4. ICBC (China) 4. PetroChina

5. Singopee- China Petroleum

5. Barclays (United Kingdom)

5. Petrobras- Petroleo (Brazil)

5. Petrobras- Petroleo (Brazil)

Source: “The World’s Biggest Public Companies,” Forbes Global 2000 http://www.forbes.com/global2000/.

Note: United States-based corporation unless indicated.

What is a market? HHI guidelines were introduced to take the uncertainty out of antitrust cases. However, one of the major problems with the HHI criterion was determining the relevant market that the firms under investigation are alleged to control. For example, the relevant market figured in the proposed merger of the office supply firms Staples and Office Depot (a merger not allowed), and Whole Foods and Wild oats (a merger permitted). In both cases the firms wanting to merge argued that

273

the HHI was low because office supplies and food were available at many other retail outlets. Critics maintained that the businesses had a large market share if the market was defined more narrowly as only those stores selling bulk office supplies or high-end organic food.

What is competitive? Even when the relevant market is identified, antitrust enforcement must still determine how much market share will be accepted. Once again, U.S. Supreme Court rulings on the issue have vacillated, shifting from permissiveness during the 1920s, when J.P. Morgan’s U.S. Steel was permitted to survive with a 60 percent market share, to increased enforcement beginning in the late 1930s. Antitrust enforcement peaked during the 1960s, when, for example, the government blocked the merger of Brown and Kinney Shoes, and between Vons Grocery and Shopping Bag supermarkets, even though the combined market shares in each case were less than 10 percent. By the 1980s the pendulum had swung back in the other direction, epitomized by the 1982 dismissal of a 13-year suit against IBM’s domination of the computer industry. Merger challenges increased slightly during the George H.W. Bush and Clinton administration, only to fall back again during George W. Bush’s years in office.

In 2010, after a year of consultation with economists and lawyers, the U.S. Department of Justice and the Federal Trade Commission issued new antitrust guidelines. They retained the HHI measure as a way to weed out mergers “unlikely to have adverse competitive effects.” However, the threshold was raised to a 200-point increase to an HHI above 2500 from the 1992 guideline of a 100 point increase above 1800. In addition, the measurement of market shares was identified as “useful to the extent it illuminates the merger’s likely competitive effects” but not “an end in itself.” Consequently, the antitrust authorities were given far more discretion to consider other factors in their antitrust decisions.

Implications

The HHI is an important research tool, used not only for antitrust analysis, but also by the Federal Reserve when reviewing bank mergers. As the examples above demonstrate, the definition of the market greatly affects HHI measurement, and its policy implications are the subject of political debate.

Is Small Beautiful?

274

Support for small business spans the political spectrum, often backed by claims that small businesses are the economy’s main job producer. Every U.S. president since Ronald Reagan has cited small business job creation as crucial for the U.S. economy. However, careful analysis suggests that small business job creation often is greatly exaggerated and that several data issues need to be resolved in order to understand the role of small business in job creation. The small business job creation thesis was initially presented in the late 1970s by business professor David Birch, who exploited a previously unexplored data source from Dun & Bradstreet Corporation’s credit reports tracing the birth, growth, and death of business firms, including information on the number of employees. This research culminated in Birch’s 1987 book, Job Generation in America, and was the basis for assertions by the U.S. Small Business Administration and other groups that small businesses create the most jobs.

Box 10.3 North American Industry Classification System

Consistent classification of U.S. industry is so important that responsibility is entrusted to the President’s Office of Management and Budget (OMB). The Statistical Policy Office of the OMB oversees the North American Industry Classification System (NAICS), the system introduced in 1997 to replace the Standard Industrial Classification (SIC) in use through the government and in most private-sector research since the 1930s. As with the SIC, NAICS uses numerical coding in which the first two digits of a code refer to economy sectors and subsequent digits define the category. Thus, 51 is the information sector, and 514191 is online Information Services. In this case, NAICS replaced a much broader, outdated SIC category, 7375 for all Information Retrieval Services. Both government and private-sector users laud NAICS because it accommodates new electronic and information industries that had been introduced piecemeal into the SIC with overly broad groupings. In particular, a single services SIC sector was split into eight NAICS sectors, and the manufacturing sector was better defined to include production of prepackaged computer software (previously in the business services sector) as well as retail bakeries, dental labs, and tire retreading.

Replacing SIC with NAICS was a multiyear process, requiring

275

revisions by many government statistical agencies; it was completed by the Census Bureau, the Commerce Department’s Bureau of Economic Analysis, and the Labor Department’s Bureau of Labor Statistics in the early 2000s. As late as 2011, conversion was still under way in some agencies such as the Occupational Health and Safety Administration. Thus, researchers today usually can count on NAICS data, but they should be aware that historical data may be organized in the older SIC framework.

Sources: John R. Kort, “The North American Industry Classification System,” Survey of Current Business, May 2001, pp. 7–13; OSHA

use in “Occupational Injury and Illness Recording and Reporting Requirements,” The Federal Register, June 22, 2011, 76 FR 36414.

By the 1990s accumulating evidence began to cast doubt on the role of small businesses in job formation. Critics at the Brookings Institution, the National Bureau for Economic Research (NBER), and Dun & Bradstreet pointed out flaws in Birch’s method. Foremost, regression to the mean biased the results, identified by economist Milton Friedman as “the most common fallacy in the statistical analysis of economic data.” In this case the error occurred when a firm was labeled small because of a temporary event or measurement error, then appeared larger in the subsequent time period; at the same time, other firms were temporarily or mistakenly labeled as large, only to later decrease in size. In this way job growth was erroneously attributed to small firms while large firms took unwarranted blame for job losses. Correcting for the error by measuring firms at their average size before and after the time period studied, NBER researchers concluded that there was no relationship between firm size and job creation in the manufacturing sector. However, extending the analysis to services as well as manufacturing, other economists found that small firms did create more jobs, although not nearly to the extent found by Birch.

A second problem with measuring job creation occurs when jobs are lost as well as created. Overall between 1975 and 2005, business establishments created jobs at an annual rate of 17.6 percent, but shed jobs at a 15.4 percent annual rate for a net job gain of 2.2 percent per year. Birch and other researchers focused on this net job creation rate, a statistic that can be quite deceptive over short stretches of time. When large firms laid off workers during a recession, small businesses appeared to create more than 100 percent of the jobs gained during that period. Moreover,

276

Birch’s method overstated job creation by new start-ups that, by definition, couldn’t destroy any jobs but did have high job destruction rates in subsequent years. In other words, small businesses create new jobs, especially when they first start up, but then most of these jobs are lost as small businesses fail or trim their staff. Economists under contract for the Small Business Administration found that over four-year periods, most net new jobs were created by a relatively small number of what they termed “high impact firms.” About half the firms were small, with fewer than 500 employees, and half were large, above 500 employees.

The small business job creation controversy shows social scientists at work, reaching consensus on a tricky data issue. Understanding job creation requires (1) careful attention to statistical theory such as regression to the mean, and (2) the necessity to look at a variety of measures, in this case job destruction as well as job creation, and to do so over different time periods and for firms of different ages. The problem is that these subtle, yet overly complex, results have not affected policy debate in which the myth of the small business job engine remains a popular political refrain. Researchers should also be aware of related research on the quality of small business jobs that are likely to be less desirable because of poorer working conditions, less job training and less generous health, vacation and pension benefits.

Box 10.4 Do 90 Percent of New Businesses Fail?

A common tagline for services to help new businesses is a 90 percent failure rate. One of the few studies that measured such poor performance looked at Australian start-ups over a very long time period, finding that only 10 percent were still in business after 10 years. However, over shorter stretches of time, new businesses do much better. Based on the U.S. Census Bureau Business Dynamics Statistics, 70 percent of new firms last at least two years, while half survive five years. In the short run, even restaurants don’t fail 90 percent of the time; a study of restaurants in Columbus, Ohio, found that 25 percent went out of business in the first year and 60 percent did so after three years. Higher failure rates occur only for self- employed, one-person enterprises that start and stop frequently.

Sources: Survival rate in U.S. Small Business Administration,

277

http://www.sba.gov/advocacy/7495/8430. Restaurants in H.G. Parsa et al., “Why Restaurants Fail,” Cornell Hotel and Restaurant Administration Quarterly 46, no. 3 (August 2005): 304–22.

Were the Bailouts a Success?

During the 2008 financial crisis, the U.S. Treasury and Federal Reserve came to the assistance of large banks, insurance, and car companies, providing trillions of dollars in loans. Since then, nearly all loans have been paid back. However, repayment does not mean that there was, in the end, no cost to the public. Because the government provided the loans at extraordinarily low interest rates—far less than a rate at which any business could possibly borrow for ordinary purposes— banks and other businesses received access to funds that could have been made available to others, such as consumers or households facing foreclosure. This led to what economists call an opportunity cost: Even if loan repayment meant that there was no increase in the federal deficit, there was still the cost of what else could have been done with these resources.

Bloomberg Markets magazine estimated the benefit to banks of Federal Reserve loans between 2007 and 2009 at $13 billion, based on banks’ typical profit on loans adjusted for taxes. J.P. Morgan countered that the actual benefits were less than $13 billion because the emergency loans were for short-term investments that earned less-than-average returns. However, New York University economist Viral Acharya pointed out that simply having access to funds from the Federal Reserve is a financial boon for banks, making them appear to be less risky to investors and thus lowering banks’ borrowing costs. After all, banks themselves charge corporations for lines of credit, whereas the Federal Reserve was offering a similar service to banks at no cost. The bottom line is that when evaluating subsidies, researchers need to look beyond claims that “all loans were repaid.” While perhaps necessary in this case to prevent a repeat of the 1930s-style Great Depression, the bailouts came at a substantial cost.

How Much Profit?

How Profitable Is a Professional Sports Team?

When a professional sports franchise reports low profits, it may be doing so in an effort to improve its negotiating position with players, or to win public subsidies for a new stadium. To an even greater degree than for

278

other businesses, sports profitability is difficult to measure. One reason is that profits exclude “salaries” that owners allocate to themselves, such as the $25 million fee that the former New York Yankees owner George Steinbrenner paid himself for negotiating the team’s cable contract. In addition, reported profits can be reduced by inflated long-term expenses like stadium costs and, in an unusual tax deduction exclusive to professional sports, may even depreciate future payments owed to players.

As an alternative to a team’s unreliably reported profits, analysts often measure a team’s net worth—a statistic that is also subject to uncertainties about the future but useful for understanding the benefits that accrue to the owners. Forbes magazine estimates the value of the world’s sports teams, led in 2010 by the Manchester United soccer team at $1.83 billion, followed by the Dallas Cowboys football team at $1.65 billion and the New York Yankees baseball team at $1.6 billion.

Are Profits Too High?

American corporations have paid for newspaper and magazine space to respond to what they perceive as widespread misperception on the part of the public about the level of corporate profits. Opinion polls show that the public estimates typical profit rates at nearly 30 percent, whereas, according to the Chevron Corporation in one advertisement, profits really average only 5 percent.

The number reported by Chevron is profit as a percentage of sales, based on a statistic in corporate reports called “net income as a percentage of sales.” Indeed, for many corporations this figure is no more than 5 percent. This statistic has limited applicability for comparative purposes, however, because sales cannot be compared between different sectors of the economy (see earlier discussion of sales as a measure of corporate size). For example, the profit rate on sales can be misleadingly low if the business has a fast turnover that enables high profits to be earned despite a low profit margin. As an alternative, business analysts calculate profits as a percentage of how much has been invested in the firm, a profit rate that averages over 10 percent. For example, in 2010 Chevron’s return on shareholder equity was 19.3 percent. It is this second definition of profits that the public probably had in mind in the opinion polls because it is similar to the concept of the rate of return received on individual investments (the survey questions asked specifically about profit as a percentage of sales, but this qualification was likely ignored or not understood by many respondents). Moreover, the public probably also

279

used an entirely different concept of profits that included pay and other benefits accruing to high-level management. In standard corporate accounting, however, these are considered “employee costs” and are never counted as profits. Unlike small businesses, with which the public is more familiar and managers and owners are the same people, corporate accounting limits the concept of profit to earnings paid as dividends or retained for corporate use. For example, in 2009 Ray Elliott, CEO of medical supply company Boston Scientific, was the second highest paid of all U.S. corporate executives with $33.4 million in salary, bonuses, and stock options—even though in strict accounting terms, his company earned no profits that year.

In summary, Chevron used a standard definition of profits in its advertisement, although the corporation chose to measure profits as a percentage of sales, a procedure that minimized their apparent profit rate. The public had a much looser sense of what should be included in profits and misunderstood the accounting definition. Yet the public was correct in the sense that profits are best measured as a proportion of what has been invested. The contrast between these two uses of the same concept is instructive about the need for caution when strict accounting terms are compared with everyday usage.

How Now Dow?

The Dow Jones Industrial Average (DJIA) is probably the most commonly reported of all economics or business statistics, but few viewers of the nightly news know either the meaning of the “Dow” or its limitations. The Dow Jones Industrial Average is simply the total stock prices of 30 representative stocks corrected for all the accumulated stock splits and changes in the 30 representatives.

Stock market experts agree that the Dow is a poor gauge of the overall market. By simply adding each stock price, the Dow gives greater weight to higher-priced stocks. For example, stocks in both McDonald’s and the pharmaceutical firm, Pfizer, had similar growth of about 10 percent in 2011. However, when McDonald’s stock sold at $80 per share, it exerted a much stronger upward pull on the Dow because it was selling at about four times as much as Pfizer, which was $20 per share. Even greater bias occurred because IBM, with its highly priced shares, counted more than ten times as much as low-share priced General Electric, Cisco, Alcoa, or Bank of America.

280

An additional drawback to the Dow Jones Industrial Average is its coverage of only 30 companies. The choice of 30 stocks stems from the Dow’s origin in 1884, when it was viewed as an easily calculated measure of overall stock activity. Today, with computers, it is possible to include many stock prices in an index and to weight them by corporate size. Stock indexes such as the New York Stock Exchange Index and the Standard and Poor’s Index use these more sophisticated techniques.

Picking Stock Winners

Can anyone predict which stocks will do well? Or, as some researchers argue, do stock prices move so randomly that monkeys throwing darts at the stock page would do as well as stock advisers—or even better once we subtract the fees charged by human stock pickers? It seems common sense that some experts predict stocks well; after all there are some investors who become fabulously rich. It is also true that if we have a large number of monkeys, a few of them also will succeed. Since there are thousands of stock advisers, perhaps we should not be surprised that some do well.

To separate luck from skill, analysts measure the repeat success of investment advisers: do they accurately predict the market year after year? Business magazines often feature the record of the previous year’s top predictors with sobering results; most successful advisers fail to maintain their success from one year to another. For example, a Business Week study found that only four out of over 900 funds stayed in the top quarter of funds for every year between 2001 and 2006.

Determining whether or not these few, apparently “star” funds are successful because of luck or skill is a matter of academic debate that requires careful attention to the data. On the surface it appears that stock pickers such as Fidelity Magellan fund’s Peter Lynch must be skilled because of his year-after-year success. University of Maryland finance expert Russ Wermers points out that simply by chance we would expect about one in 200 stock pickers to earn the extraordinary returns of 10 percent per year after expenses for five consecutive years. His studies find that more than three times as many funds actually do this well, suggesting that “a sizeable minority” of managers pick stocks well enough to more than cover their costs. At the same time, he counsels that it is difficult to identify such a fund in advance.

Data Sample: Market recommendations based on astrology, as advised in Crawford Perspectives, earned a ranking of 7 out of 65 in

281

1.

2.

3.

Hulbert Financial Service’s 1992 study—an embarrassment to newsletters that claim to use more scientific principles.

Summary

Many of the controversies reviewed in this chapter involve data reported directly by businesses. Overall, researchers using business statistics are more fortunate than most data users in the sense that few other data sources are subject to such intensive review by outside experts or to such well-established reporting criteria. It is true that investors sometimes complain about corporate attempts to hide unfavorable information, as in the case of misleading language in annual reports. Nonetheless, the underlying data are probably the most completely documented field covered in this book, in terms of both the volume of data and their interpretation by experts. As with many other social science data, a telephone call to the appropriate business reporter, corporate or government official often will clarify problems encountered with these numbers. Researchers also should not overlook public libraries, which often have extensive holdings of business publications, including many that are too expensive for an individual to buy.

Disagreements about business statistics occur primarily because of varying definitions rather than uncertainty about measurement. Thus, size comparisons depend on which statistic is used: assets, market value, profits or employment. The profit rate depends on the denominator: profits as a percentage of sales or how much was invested. And stock market success requires a choice of time period in order to separate random variation from stock-picking skills.

Case Study Questions

Several cell phone providers claim to provide the most “coverage.” How is it possible for different firms to claim they have service available in more of the United States than their competitors? Which measure of “coverage” is most meaningful?

Design an advertising campaign for a business that would make use of North American Industry Classification System (NAICS) data.

On January 31, 2012, the Dow Jones Industrial Average was down by 0.16 percent, while the New York Stock Exchange

282

4.

5.

Composite Index was up by 0.05 percent. Explain the difference.

In an antitrust suit, the U.S. Supreme Court ruled against the merger of the Aluminum Company of America (Alcoa) with a much smaller aluminum company called Rome. The following market shares for Alcoa were cited in the case:

Aluminum bare conductor wire and cable: 27.8 percent

Aluminum and copper bare conductor wire and cable: 10.3 percent

Aluminum and copper bare and insulated wire and cable: 1.8 percent

(Rome’s market share was 2 percent or less in all markets.)

Using these data, construct cases both for and against the merger. (The Supreme Court ruled against it.)

In his 2010 State of the Union address, President Obama maintained most new jobs are created by “small businesses.” In the 2012 address, he altered the statement to most new jobs are created by “start-ups and small businesses.” Which statement was most accurate? Why?

References

Data Sources (211)

On number and types of businesses, see U.S. Census Bureau, Statistical Abstract of the United States: 2011, table 743, p. 491.

Public Corporations (212)

Lehman Brothers in Annual Report, 2007, http://lehman.rclclients.com/annual/2007/. Marie Callender, Annual Report, 2009, http://sec.edgar-online.com/perkins-marie-callenders-inc/10- k-annual-report/2009/03/30/section22.aspx.

Privately Held Corporations (213)

Cargill in “America’s Largest Private Companies,” Forbes, November 3, 2010, http://www.forbes.com/2010/11/01/largest-private-companies- business-private-companies-10-intro.html.

Small Businesses (214)

283

For current data, see U.S. Small Business Administration, Data on Small Business, http://www.sba.gov/advocacy/. Data sample in Dun’s Marketing Services, Million Dollar Directory, 1989 (Parsippany, NJ: Dun’s Marketing Services, 1989), p. 2886.

Aggregate Statistics (214)

Data sample in U.S. Census Bureau, 2009 Nonemployer Statistics, http://censtats. census.gov/cgi-bin/nonemployer/nondetl.pl.

Controversies (214)

Who Is the Biggest of Them All? (214)

Debate on use of Fortune sales data in letter to the editor by M.A. Adelman, Fortune, September 1955, p. 20; and in F.M. Scherer, Industrial Market Structure and Economic Performance, 2d rev. ed. (Boston, MA: Houghton Mifflin, 1980), p. 47. Target and General Electric in “America’s Largest Private Companies,” Forbes, 2007 http://money.cnn.com/magazines/fortune/global500/2007/full_list/index.html.

Are the Big Too Big? (217)

On history of antitrust, see William E. Kovacic and Carl Shapiro, “Antitrust Policy: A Century of Economic and Legal Thinking,” Journal of Economic Perspectives 14, no. 1 (Winter 2000): 43–60. On HHI see U.S. Department of Justice, “Herfindahl-Hirshman Index” at http://www.justice.gov/atr/public/guidelines/hhi.html. Market-share use in F.M. Scherer and David Ross, Industrial Market Structure and Economic Performance, 3d rev. ed. (Boston, MA: Houghton Mifflin, 1990), pp. 72– 96 and 184–85. Policy of the 1980s debated in Journal of Economic Perspectives 1, no. 2 (Fall 1987): 3–54. Office Depot and Staples in U.S. Federal Trade Commission, “FTC Rejects Proposed Settlement in Staples/Office Depot Merger,” April 4, 1997, http://www.ftc.gov/opa/1997/04/stapdep.shtm. Whole Foods and Wild Oats in U.S. Federal Trade Commission, “FTC Consent Order Settles Charges That Whole Foods’ Acquisition of Rival Wild Oats Was Anticompetitive,” March 6, 2009, http://www.ftc.gov/opa/2009/03/wholefoods.shtm. On new guidelines, Larry Fullerton, “Introduction: 2010 Horizontal Merger Guidelines,” Antitrust, 25, no. 1 (Fall 2010): 8–9; see also Carl Shapiro, “The 2010

284

Horizontal Merger Guidelines: From Hedgehog to Fox in Forty Years,” Antitrust Law Journal 77, no. 1 (2010): 49–107.

Is Small Beautiful? (220)

“Every U.S. president since Ronald Reagan” in Ruth Marcus, “Mythbusting on Small Business and Job Creation,” Pasadena Star News, September 15, 2010, p. A8; David Birch, Job Generation in America (New York: Free Press, 1987). The “most common fallacy” in Milton Friedman, “Do Old Fallacies Ever Die?” Journal of Economic Literature 30, no. 4 (December 1992): 2131. Steven J. Davis, John C. Haltiwanger, and Scott Schuh, Job Creation and Destruction in U.S. Manufacturing (Cambridge: MIT Press, 1998); summary in Davis, Haltiwanger, and Schuh, “Small Business and Job Creation: Dissecting the Myth and Reassessing the Facts,” Small Business Economics 8 (August 1996): 297–315. Response in Martin Carree and Luuk Klomp, “Small Business and Job Creation: A Comment,” Small Business Economics, 8 (August 1996): 317–22; David Neumark, Brandon Wall, and Junfu Zhang, “Do Small Businesses Create More Jobs?” NBER Working Paper No. 13818, February 2008, http://www.nber.org/papers/w13818; Zoltan J. Acs et al., “High-Impact Firms: Gazelles Revisited,” U.S. Small Business Administration, Office of Advocacy, June 2008, http://archive.sba.gov/advo/research/rs328tot.pdf. On net and gross, see Cordelia Okolie, “Why Size Class Methodology Matters in Analyses of Net and Gross Job Flows,” Monthly Labor Review 127, no. 7 (July 2004): 3–12.

Were the Bailouts a Success? (223)

Bob Ivry, Bradley Keoun, and Phil Kuntz, “Secret Fed Loans Gave Banks $13 Billion Undisclosed to Congress,” Bloomberg Markets, November 27, 2011, http://www.bloomberg.com/news/2011-11-28/secret-fed-loans- undisclosed-to-congress-gave-banks-13-billion-in-income.html. Acharaya in Ivry,

How Much Profit? (224)

New York Yankees in Sam Pizzigati, Greed and Good: Understanding and Overcoming the Inequality That Limits Our Lives (Lanham, MD: Apex Press, 2004), p. 296. Subsidized stadiums in D. Stanley Eitzen, “Public Teams, Private Profits,” Dollars and Sense, March/April 2000, pp. 21–23. Chevron in “Chevron Energy Report,” New York Times, November

285

18, 1980, p. A20. Public estimate in James D. Gwartney and Richard Stroup, Economics (San Diego, CA: Harcourt Brace Jovanovich, 1987), p. 493. Chevron 2010 profits in Chevron Corporation, “Chevron Financial Highlights,” 2010 Annual Report, p. 4, http://www.chevron.com/documents/pdf/Chevron2010An-nualReport.pdf. Ray Elliott in “20 Highest Paid CEOs,” CNNMoney, http://money.cnn.com/galleries/2010/news/1004/gallery.top_ceo_pay/2.html.

How Now Dow? (225)

Dow described at http://averages.dowjones.com/mdsidx/downloads/meth_info/Dow_Jones_Industrial_Average_Methodology.pdf. Weights in “Index Component Weights of Stocks in the Dow Jones Industrial Average,” updated May 16, 2012, http://indexarb.com/indexComponentWtsDJ.html.

Picking Stock Winners (226)

Overview of research in Burton G. Malkiel, A Random Walk Down Wall Street (New York: Norton, 1992), and Peter Bernstein, Capital Ideas: The Improbable Origins of Modern Wall Street (New York: Free Press, 1991). “Mutual Funds: Few Consistent Winners,” Business Week, August 16, 2006, http://www.businessweek.com/investor/content/aug2006/pi20060816_494365.htm. Mark Hulbert, “The Index Funds Win Again,” New York Times, February 22, 2009; Robert Kosowski, Allan Timmermann, Russ Wermers, and Hal White, “Can Mutual Fund ‘Stars’ Really Pick Stocks?” Journal of Finance 61, no. 6 (December 2006): 2551–95. Data sample in Mark Hulbert, The Hulbert Guide to Financial Newsletters (Chicago, IL: Dearborn Financial Publishing, 1993), pp. 34–40.

Case Study Questions

1. Several independent websites compare cell phone coverage.

2. On NAICS see John R. Kort, “The North American Industry Classification System,” Survey of Current Business, May 2001, pp. 7–13.

3. On Dow Jones Industrial Average see http://www.dowjones.com; on NYSE Composite Index see http://www.nyse.com/about/listed/nya.shtml.

4. See

286

http://bulk.resource.org/courts.gov/c/US/377/377.US.271.204.html.

5. See http://www.whitehouse.gov/the-press-office/remarks-president- state-union-address; and http://www.whitehouse.gov/the-press- office/2012/01/24/remarks-president-state-union-address.

287

11

Government □□□□

By this point in the book, readers should be convinced that few social statistics can be taken at face value. Thus, it will come as no surprise that there is controversy about statistics concerning the operations of the U.S. government itself. And again, there is little debate about the sincerity of those who gather the data. In fact, many of the problems were first identified by government statisticians and have been discussed in government publications.

This chapter reviews disputes about military spending, welfare spending, the U.S. deficit, the money supply, and voter participation. Important policy decisions depend on how we measure each of these variables. In addition, a final section reviews problems raised in earlier chapters about inflation and currency-exchange rate adjustments. As emphasized in earlier chapters, the difficulty for researchers is choosing the correct statistic from the many numbers published by the government.

Where the Numbers Come From Organizations Data sources URL Office of Management and Budget, Executive Office of the President

Budget of the United States

www.whitehouse.gov/omb

Social Security Administration, U.S. Department of Health and Human Services

Social security and other program expenditures

ssa.gov

Bureau of Labor Statistics, U.S.

Consumer Expenditure Survey and other BLS

www.bls.gov

288

Department of Labor surveys Bureau of Economic Analysis, U.S. Department of Commerce

Deflators from national income and product accounts

www.bea.gov

Stockholm International Peace Research Institute

Summaries of government reports

www.sipri.org

Board of Governors, U.S. Federal Reserve

Reports from banks; flow of funds accounts

www.federalreserve.gov

U.S. Bureau of the Census

Current Population Survey

www.census.gov

American National Election Studies

Surveys www.electionstudies.org

United States Elections Project

Election results elections.gmu.edu

Data Sources

Data collected by the U.S. government about the government are readily available online and are distributed to Federal Depository libraries. Moreover, federal agencies and individual government statisticians generally are receptive to direct inquiries.

U.S. Budget

Detailed analysis of spending and taxation is contained in the annual Budget of the U.S. Government submitted by the president to Congress in January and then revised substantially during the year. It is a massive book with additional volumes of appendices, special analyses, and historical tables. Summary data are published in a more manageable “United States Budget in Brief,” the annual Economic Report of the President, and the March issue of the U.S. Commerce Department’s Survey of Current Business. These data are published in fiscal years. From 1844 until 1976, the fiscal year ran from July 1 to June 30; in 1976, it was changed to run from October 1 to September 30. Thus, fiscal year 2012 included October 1, 2011, through September 30, 2012. The slight differences from the calendar year are overlooked for most research purposes.

Military Spending

Data on military spending for 171 countries for the period since 1988 is

289

available from the Stockholm International Peace Research Institute (SIPRI).

Data Sample: The SIPRI Military Expenditure Database reports Denmark’s 2011 military spending at just under 4.6 billion US$, or 1.4 percent of Danish GDP.

Money

The money supply is measured by the Board of Governors of the U.S. Federal Reserve. Both its short-term estimates and longer-term official figures are widely available in official government publications, with analysis in many business periodicals.

Data Sample: The U.S. Federal Reserve reported $4.5 billion in travelers checks in May 2011, down from $4.9 billion a year earlier.

Prices

Most price data are collected by the U.S. Labor Department’s Bureau of Labor Statistics (BLS) for its own price indexes, numbers that also form a large part of the U.S. Commerce Department’s inflation adjustments for GDP accounting. The BLS surveys more than a million price quotations to compute consumer and producer price indexes for 211 item categories and 38 geographic areas for a total of about 8,000 basic indexes. These data are widely publicized: both up-to-date estimates and historical data are available online.

The relative value of one currency, such as the U.S. dollar, with another, such as the euro, is determined by market exchanges around the world. These exchange rates are widely publicized in newspapers and business periodicals. Composite exchange rates measuring values relative to several currencies at one time are calculated by central banks, including the U.S. Federal Reserve and the European Central Bank.

Data sample: The Consumer Price Index for all urban consumers for fresh sweet rolls, coffee cakes, and doughnuts increased by 2.7 percent from May 2010 to May 2011.

Voting Data

The Current Population Survey, conducted by the U.S. Bureau of the

290

Census for the Bureau of Labor Statistics, includes a November supplement on voting and registration by various demographic and socioeconomic characteristics. Voter interview data is also available in the American National Election Studies (ANES). Actual election-day data are collected in CQ Vital Statistics on American Politics Online Edition and by the United States Elections Project.

Data sample: As reported in the United States Elections Project, 3,148,613 U.S. citizens were ineligible to vote in the 2010 general election because of a felony conviction.

Controversies

How Much for the Military?

Two fundamental questions arise in assessing the appropriate size of military spending. First, what is the total level of U.S. military spending? And second, what was the cost of the wars in Afghanistan and Iraq?

How Big Is the U.S. Military Budget?

The 2012 U.S. government budget listed defense spending at $676 billion, which included a reversal of the Bush administration practice of putting the cost of the wars in Afghanistan and Iraq in a separate “emergency appropriation” category on the grounds that they were temporary, unexpected expenditures. However, the $676 billion leaves out spending technically outside the U.S. Department of Defense, but arguably still military in purpose. Some of these items clearly belong in the defense- spending category, such as roughly $20 billion for U.S. Energy Department weapons development and about $40 billion for the Central Intelligence Agency, which after 50 years of secrecy now has a disclosed budget. Added together, these items, plus military expenditures for counterterrorism and homeland security, raise U.S. 2012 military spending to about $800 billion, a more useful figure for comparing military spending with other items in the U.S. budget and with spending by other nations.

Some analysts maintain that the military portion of the budget also should include $130 billion for veterans’ benefits, as well as $185 billion for the interest payments on borrowing to pay for past war spending, bringing total 2011 U.S. military spending to more than $1,100 billion. By this total, military spending is far in front of any other federal budget

291

category—Social Security is second at around $700 billion. Even at the officially reported $676 billion level, U.S. military spending is more than military spending by all other countries in the world combined.

A counter-argument by Mallory Factor, Forbes magazine commentator, maintains that the official military budget overstates “legitimate defense activities” because it includes nation-building expenses such as policing, humanitarian missions as well as transportation and protection for U.S. officials traveling in war zones. Such expenditures are not counted separately, but according to Factor are “substantial” and should be taken into account.

How Much Did the Wars in Afghanistan and Iraq Cost?

In 2003 the Bush administration maintained the newly declared war in Iraq would cost $50 to $60 billion. By 2011 even supporters of the war effort agree that the war added up to at least ten times as much. However, economists Joseph Stiglitz and Linda Bilmes put the cost at well over $3 trillion, a many-fold difference because they included items not in the standard defense budget, such as future costs of caring for wounded soldiers, interest payments on debt incurred to pay for the war, the effect of the war on the price of oil, and the cost of the war for the Iraqi people.

Economists Stephen Davis, Kevin Murphy and Robert Topel, agreeing that the wars might have even cost as much as $600 billion, compared this expense with the alternative of trying to contain the Iraqi regime without war. In their view the war was cost-effective because of the huge expenses they estimated for alternative policies other than war, as well as the damage caused by Iraqi dictator Saddam Hussein had he stayed in power. An even more optimistic assessment came from political scientist Charles Hill, who argued that the war had multitrillion-dollar benefits because it spurred Mideast governmental transformation.

Analysts on all sides agree that original Bush administration estimates were far too low, below even the eventual direct war costs. But questions remain about how many of the indirect costs to include and how to measure them. For example, it is difficult to determine the total human cost of the war on soldiers or on Iraqi civilians (see Box 11.1). Even more problematic are efforts to cost out counterfactual history. We simply don’t know what would have happened to oil prices, or with Hussein’s or other Mideast government policies had the war not occurred.

How Much for Welfare?

292

In common usage, welfare means assistance to the poor or, in more technical terms, “means-tested transfer payments”—that is, programs for which recipients must prove they are poor. Using this definition, conservatives argue that the United States spends too much on welfare, with little effect. In its 2009 report “Obama to Spend $10.3 Trillion on Welfare,” the Heritage Foundation maintained that President Obama’s proposed budgets allocated over $714 billion to welfare programs in 2008, with future growth putting the ten-year total over $10.3 trillion. Even at the current level, spending amounted to $28,000 for each lower-income family, far more than needed to pull them out of poverty. However, the Heritage Foundation total includes noncash benefits at their face value, most importantly health care, over one-half of the welfare budget. Thus, most welfare was in fact payment to others, primarily in the health care industry, not direct income to low-income families.

By an alternative definition, welfare can be used to describe government benefits for all citizens, a definition commonly used to describe European public support for housing, education, health and child care, regardless of the recipient’s income. When this broader definition is applied to the United States, welfare includes programs that benefit the well-to-do, not the poor. For example, the federal government tax deduction for mortgage interest amounts to over $100 billion per year, primarily compensating high-income households who have larger mortgages and fall in higher marginal tax brackets (see Chapter 3). Even greater benefits for the well- off occur as a result of the exclusion of employer-provided health insurance from taxes, valued at more than $200 billion annually. Thus, researchers need to be aware that welfare as defined in everyday usage as “aid to the poor” leaves out welfare for the nonpoor, a significant portion of government support.

Box 11.1 How Many Iraqi Civilian Deaths?

The estimates of civilian deaths caused by the Iraq war range from 4,000 or 600,000. Such a huge span highlights the difficulty of gathering data during a war as well as the ambiguity about what constitutes a death caused by war. At one extreme, the very low estimate came from the Iraqi Health Ministry and appeared to be a politically motivated undercount, lower than the one month United

293

Nations estimate for August 2006. More credible is the Iraq body count by a nongovernmental British research group, estimating around 50,000 deaths based on media accounts. Similar estimates came from morgue reports, but these were likely a minimum count as well: Data was sparse from outside the capital Baghdad, and not all families were willing to register deaths at a central morgue. Stiglitz and Bilmes (see p. 235) relied on a study published in the British medical journal Lancet that estimated over 100,000 civilian deaths caused by the war during its first two years. Researchers from Johns Hopkins University suggested up to 600,000 Iraqi civilians had died as a result of the war based on a survey of 1,849 families across the country. Extrapolating results from such a survey is a standard statistical technique, but one that requires careful attention to make certain that the sample is representative and that respondents answer correctly—assumptions that may not hold in a wartorn country.

Sources: Sabrina Tavernise and Donald G. McNeil, Jr., “Iraqi Dead May Total 600,000, Study Says”, New York Times, October 11, 2006, p. A16; Peter Steinfels, “In the Brutality of War, the Innocents Have Become Lost in the Crossfire,” New York Times, November 20, 2004, p. B6; Lawrence K. Altman and Richard A. Oppel, Jr., “WHO Says Iraq Civilian Death Toll Higher Than Cited,” New York Times, January 10, 2008, p. A14.

How Big Is the Deficit?

In national surveys, U.S. citizens identify the deficit as the major problem facing the country, and nearly all national political candidates present plans to bring the budget into balance. Indeed, the U.S. deficit, the difference between federal government spending and revenues, reached a record $1.3 trillion in 2011. At the same time, many state and local governments faced deficits—smaller in size, of course, but still prompting onerous cutbacks in public services. However, measuring the deficit, and the related statistic, the debt (all past deficits and surpluses added together), raises accounting issues about its actual size and the appropriate policies needed to manage it.

Off-Budget Items

One possible discrepancy in budgeting involves moving items in or out of the official budget, in what could be manipulation or could follow standard

294

accounting practice if both on–and off-budget items are reported and measured consistently from year to year. In the U.S. federal budget, the principle for excluding some items from the budget is to separate programs such as the Social Security trust fund. It alone has an accumulated surplus adding up to more than $2.5 trillion by 2011, specifically collected to pay for the expected Social Security deficit when baby boomers retire in the following years. There is also a smaller off-budget deficit for the U.S. Postal Service, a surplus for the U.S. Federal Reserve, and a several-year deficit for Fannie Mae and Freddie Mac.

State and local governments also have off-budget items, primarily funds supporting long-term investments. The accounting argument is that such projects cannot have a yearly balanced budget as required in most states since they require high initial expenditures; examples include building bridges or port facilities that will yield revenue only in the future. However, problems arise when, in order to achieve a quick budget fix, funds are moved on– or off-budget, as occurred in the New York and Connecticut budget battles of 2009. Or, in a reverse move, items such as federal reimbursement can be moved off-budget if political leaders want to be able to boast that they lowered overall spending, as occurred in New Jersey in 2009 when Governor Corzine put federal Medicaid and school spending off-budget.

Valuing Pensions

Potential accounting problems also are associated with state-funded pensions when regulators must determine if sufficient funds are being set aside for future pension payments. The main statistical issue is measuring how much today’s contributions will increase in value as they are invested over time. By valuing contributions at a risk-free U.S. bond rate rather than the traditionally used higher rate of return possible with stocks, economists Robert Novy-Marx and Joshua Rauh calculated a potential shortfall in state pensions of over $3 trillion. Economist Dean Baker, codirector of the Center for Economic and Policy Research, allows that some states were far too optimistic in assuming their pension funds would grow at stock market boom rates, but highly conservative assumptions of risk-free bonds portray a too-bleak picture of pensions. In his view, states face manageable pension payouts even if we assume only moderate returns on state-pension-fund stock market investments.

Selling Off Government Assets

295

A third accounting issue is how to deal with the sale of assets, something that all governments do from time to time, but can cause the budget temporarily to look too rosy. During the 1980s the Reagan administration improved the apparent budgetary balance by planning to sell U.S. government assets— including large blocks of public lands, the government-owned freight railroad system Conrail, and farm and school loans (to be sold to private investors for collection). More recently, conservative think tanks have proposed sales of federal lands as a way to decrease government involvement in the economy as well as reduce the federal debt. On the local level, Chicago sold its Skyway toll road, underground parking garages, and even city parking meters, receiving over $3 billion in total. And Europe’s most economically troubled countries such as Greece have contemplated sales of airports, national lotteries, post offices, roads, and utilities as a way to reach financial stability.

Whatever the advantages or disadvantages of public ownership for these assets—for example, environmentalists fought bitterly against the sale of public land—the sale of government assets was criticized by accountants as a one-shot injection of revenue. The deficit would fall only once and reappear the following year. In the private sector there are strict rules to prevent companies from using the sale of assets to cover up deficits.

Implications

In all these situations, the lesson for researchers is to look at several statistics in order to understand the total financial picture of governments. On– and offbudget deficits, often overlooked in popular accounts of the U.S. budget, are available in the annual Economic Report of the President. Unfortunately, for state and local governments, data often are reported in an inconsistent manner; one useful policy change would be for more transparency in these budgets. Measuring pension reserves requires researchers to consider a range of likely values for the return on today’s contributions. Finally, analysis of debt must take into account sale of assets, a common feature in corporate accounting but one that is often overlooked in government finance.

A Debt Monster?

The U.S. federal debt topped $14 trillion in 2011, a number guaranteed to shock those unfamiliar with aggregate macroeconomic data. It is the total of all past years’ deficits, subtracting the few years when the U.S. government ran a surplus, as in fiscal year 2000. As with the deficit, there

296

are a number of accounting issues in measuring the debt.

Foremost, economists point out that federal debt is quite different from any other debt. Thus, when President Obama stated, “Government has to start living within its means, just like families do,” economist Paul Krugman pointed out that “the government shouldn’t budget the way families do; on the contrary, trying to balance the budget in times of economic distress is a recipe for deepening the slump.” In this view, the federal government, with its power to borrow and to create money, has the unique responsibility to balance the economy and not be distracted by manageable debt numbers if correctly measured.

More than $6 trillion—about 40 percent of U.S. government borrowing as of 2011—was bought by the federal government itself. The purchases were made primarily by the Social Security Trust Fund, investing surpluses built up over past years, and by the Federal Reserve in its efforts to prop up the economy with an increased money supply. The remaining $8 trillion held by the public is a more reasonable estimate of debt that the U.S. government will pay back over time. Even that number, while impressively large, is not necessarily economically harmful. In the private sector, if no one borrowed for investment, our economy would collapse: there would be no home building, no telephone system, no railroads, and no other projects requiring long-term financing. Many government expenditures such as bridge and road building clearly qualify as investment, as do education expenditures, a major portion of state and local budgets.

The federal government might follow the corporate model and keep a separate capital budget. In fact, many local governments already do so. However, some economists oppose extension of the concept to the federal government because politicians would be too likely to manipulate the process to spend money without collecting taxes. Brookings Institution economist Henry J. Aaron points out that there is no independent accounting authority as exists for private corporations to ensure proper bookkeeping.

As in the case of all macroeconomic numbers, the debt needs to be put into perspective; as a percentage of U.S. GDP, the 2011 federal debt was higher than in recent years but actually lower than it was following expenditures during World War II. Based on historical experience in other countries, economists Carmen Reinhart and Kenneth Rogoff set a threshold of danger at a debt-to-GDP ratio of 90 percent, about the level of

297

total U.S. debt in 2011. However, other economists point out that Reinhart and Rogoff’s analysis was based on past experiences with debt owed to the public, only 69 percent in 2011.

In summary, like the federal deficit, the federal debt is a large number readily exploited for dramatic effect by ill-informed politicians. These multitrillion dollar figures need to be viewed with a full understanding of the size of the U.S. economy. In addition, care must be taken to separate debt owned by the government itself. And, most important, the negative images conjured up by the words deficit and debt must be balanced by their necessity and sometimes positive effects in a modern economy.

Taxes

It seems a simple matter to understand taxes. After all, tax rates are public knowledge, and there are complete government records about the amount and kind of taxes collected. But two factors complicate recent analysis of the U.S. tax system: first, how to measure who actually bears the burden of the tax system, and second, how to measure the “fairness” of taxes.

Who Bears the Tax Burden?

The overall distribution of taxes among income groups is measured by what economists call the progressivity or regressivity of the tax system. Taxes are progressive if tax rates increase for higher-income groups, while tax rates are regressive if they fall as one moves up the income ladder. The most frequently used categories for high and low incomes are quintiles, a value determined by dividing households into five groups of 20 percent each from poorest to richest. In order to make sense of the uneven distribution of income in the nation, the top 1 percent is analyzed separately.

The quintiles plus 1 percent categories are used by the Congressional Budget Office (CBO) in its annual estimate of the tax rates for different income groups, ranging from the bottom 20 percent to the top 1 percent. Particularly surprising, and often featured in conservative publications, is the near-zero federal income tax rate for nearly one-half of all households, or actually a rebate once earned income credits are taken into account. At the same time, the top 1 percent paid nearly 40 percent all federal income tax.

These numbers applied only to the federal income tax, leaving out other more regressive taxes like the federal payroll tax, state sales taxes, and

298

excise taxes (primarily taxes on gas, alcohol, and tobacco). The CBO report also measured the impact of all taxes but, in order to do so, needed to make assumptions about who actually bears the burden of these other taxes. For example, the payroll tax includes an employer’s share that technically is paid by the employer, but the CBO estimates count it as an employee burden, on the assumption that it is actually passed along in the form of lower wages. Similarly, excise taxes paid by businesses when they produce goods and services were assumed to fall on consumers in the form of higher prices. On the other hand, in CBO estimates, corporate income taxes were borne by owners even though economists agree that at least some of this burden is passed along to consumers through higher prices. Finally, the CBO also adjusted for the effect of different household sizes. A one-person household clearly is better off with the same income as a large family, so the CBO divided household income by the square root of the household size, a reasonable if arbitrary assumption, and then ranked households in order to determine which were in the top 1 percent, top 20 percent, and so on.

Each of these CBO assumptions could be made slightly differently, changing the measured effect of taxes on income distribution. Or alternative adjustments could be added. For example, the CBO makes no attempt to correct for income understatement at tax time, evasion that economists Slemrod and Johns note is stronger for higher-income households and therefore causes their tax rates to be slightly exaggerated.

In summary, CBO tax analysis is widely accepted because it is based on generally accepted economic principles; additionally, it is recognized as relatively balanced because of the CBO’s bipartisan makeup. Nonetheless, researchers need to be aware that published assessments of the tax burden by necessity require that the underlying data be manipulated in various ways.

What’s a Fair Share for the Rich?

The major debate about tax fairness focuses on the appropriate way to evaluate the effect of taxes on different income groups. According to calculations by the Citizens for Tax Justice, looking at all federal, state and local taxes (and thus different from Figure 11.1 for only federal taxes), the lower 20 percent of households pays about 17 percent of household income in taxes, compared to 29 percent for the top 1 percent. By this view, the overall tax system appears quite progressive. However, if one compares the share of taxes paid by each group with the share of that

299

group’s total income, then there is very little difference across the income spectrum. The poor pay 2 percent of taxes and receive 3.5 percent of income, while the top 1 percent pays 22 of all taxes and receives 20 percent of all income. The difference arises because of the underlying highly unequal income distribution in the United States (see Chapter 8). It allows for the best-paid Americans to pay a share of taxes collected, not only because they pay higher tax rates, but also because this sliver of U.S. households receives an equally disproportionate share of all income.

Figure 11.1 U.S. Federal Taxes 2009

Source: Congressional Budget Office, Data on the Distribution of Federal Taxes and Household Income, April 6, 2009, table 1.

These different ways of viewing taxes are the basis of recent debates about the income tax. During the 2008 campaign, Barack obama proposed to return the top personal marginal income tax bracket back to the 39.6 percent level of the Clinton administration. opponents emphasized that the most well-to-do 1 percent of households, the only group affected by the new bracket, already paid nearly 40 percent of the federal income tax. Supporters of the higher tax pointed out that the rich paid more primarily because they received a disproportionate share of income and that this share had increased by 275 percent between 1979 and 2007, far more than any group lower on the income ladder.

Measuring Money

300

In the United States, measurement of financial statistics rests primarily with the Federal Reserve System, a quasi-independent branch of government (simi-lar central banking organizations exist in most other countries). The Federal Reserve, or the Fed, as it is more commonly known, runs its own statistical-gathering operation, assembling data on national production (see Chapter 7), financial affairs, and most important, measurement of the money supply.

Box 11.2 Do the Rich Pay 90 Percent in Taxes?

In a highly personal account, Gregory Mankiw, the former chief economic adviser to President George W. Bush, described how he turns down extra work because of the potential 90 percent tax on this added or marginal income. Readily admitting that he can afford the tax, he nonetheless used this case study to argue against President Barack Obama’s proposed increase in the highest income tax bracket, returning it to 39.6 percent, the rate before it was lowered during the Bush administration.

Critics took Mankiw to task, pointing to the controversial accounting behind the alleged 90 percent marginal tax rate. To reach that number, Mankiw included not only the federal income tax, but also the state tax, payroll tax, and the reduced inheritance his children would receive 30 years from now (the last item accounting for about one-half the total tax). Smart investing would target the much lower capital gains tax rate, only 15 percent in 2011, and even then the inheritance tax applied only to the amount of bequeaths above $10,000,000 for a married couple. Finally, as economist Uwe Reinhart points out, higher tax rates should be judged on their capacity to fund desirable government programs, not simply for their impact on the inheritance received by Mankiw’s children.

Sources: N. Gregory Mankiw, “I Can Afford Higher Taxes. But They’ll Make Me Work Less,” New York Times, October 9, 2010, p. BU-3; John Schmitt, “Three-Card Mankiw,” http://www.cepr.net/index.php/blogs/cepr- blog/three-card-mankiw; Michael Kinsey, “Rich Professors Doth Protest Too Much,” Politico, October 19, 2010, http://dyn.politico.com/printstory.cfm? uuid=C0EF3853-AFCF-6810-DA278243FD3915F4; Uwe Reinhart in Schmitt.

301

Box 11.3 Taxation at the Margin: Working More Does Earn You More

Because the U.S. income tax is progressive, additional income may be taxed at a rate higher than the overall, or average, rate faced by the typical taxpayer. The additional income is taxed at the marginal tax rate with brackets ranging from 10 percent to 35 percent in 2011. When the marginal rate is confused with the average rate, taxpayers may believe that the additional income will cause their overall earnings to fall—an error unfortunately perpetuated in the national media. USA Today and ABC News all made this mistake, implying that readers should pass up a raise because it could “bump you into the next tax bracket.” While it is true that some additional income may be taxed at the higher rate, overall the amount earned after taxes always increases because the remaining income is still taxed in lower rate brackets. When such articles are tied to political commentary, as for example ABC News reporting on President Obama’s proposed (but not passed) plan to raise the marginal tax on annual incomes over $250,000, their incorrect math obscured the underlying policy debate.

Sources: Dean Baker, “USA Today Is Too Dumb for Words When It Comes to Taxes” in Beat the Press, Center for Economic and Policy Research, September 4, 2011, http://www.cepr.net/index.php/blogs/beat-the- press/usa-today-is-too-dumbfor-words-when-ot-comes-to-taxes; Jamison Foser, “Does ABC News Understand How Income Tax Works,” Mediamatters for America, March 3, 2009, http://mediamatters. org/blog/200903030013.

What Is Money?

Economists use the word money in a far more complex manner than everyday usage would suggest. As defined by economists, money includes not just the coins and currency commonly called money but also bank deposits that exist only on bank accounting sheets. The measurement challenge is determining which bank accounts should count as money. At stake are important economic policies, as well as many research projects

302

that require a measure of the money supply.

Federal Reserve publication of money data began during the 1940s, based on surveys of banking institutions. Today, daily reporting from approximately 10,000 large banks is supplemented by less regular reporting from smaller banks. As of 2011, the two major definitions were labeled simply M1 and M2, based on the ease with which funds can be spent, a criterion called liquidity. M1, about $2 trillion in 2011, includes only coins, currency, and checking accounts, funds readily available for use; M2, about $9 trillion in 2011, includes all money in M1 plus savings accounts and money market accounts—funds less likely to be spent in the near future on goods and services. A third, broader definition of money, M3, was discontinued in 2006, but is still available from private sources.

At issue in these statistics is the Federal Reserve’s monetary policy, the use of various tools to affect the U.S. economy. Economists debate which statistics should be used to guide Fed policies, including the relevance of M1 and M2 as targets. During the early 1980s, the Federal Reserve set quite restrictive goals for M1 and M2 in an attempt to lower the inflation rate. However, at almost precisely the same time that the Federal Reserve started to target the money supply, the relationship between the money- supply measures and the overall economy became especially erratic. Because of bank deregulation initiated during the late 1970s, there was an unprecedented movement of money into new types of bank and money market accounts. As a result, M1 and M2 fluctuated when funds moved between bank accounts, confounding attempts to use them as indicators of the inflationary pressures in the underlying economy. By the 1990s, the Fed had abandoned money supply targets and, after 2000, was no longer required to report target ranges to Congress. Fed response to the 2007 financial crisis showed up not in M1 or M2, but in the Federal Reserve’s balance sheet that rose from a steady approximately $800 billion to well over $2 trillion by 2008 and close to $3 trillion in 2011.

Missing Currency?

Do you have $3,000 in your pocket? Probably not, although more than this amount of U.S. coins and currency is in circulation for every U.S. adult. Not all of this cash belongs to individuals; some cash lies in vending machines, business cash registers, and banks. But economists believe these locations account for little of the currency excess. Instead, some of the missing money is likely held by drug dealers and other members of the underground economy who hold on to large amounts of cash because of

303

disclosure rules when these sums are deposited in banks. No one knows how much money can be accounted for in this way. More of the missing money is currency held by foreigners, including about two-thirds of the nation’s $100 bills, comprising most of the U.S. currency supply. A number of important issues are at stake in tracking this money. As described in Chapter 7, the appropriate response to the underground economy depends on its size, often estimated by the value of unaccounted- for currency. Also, dollars that flow overseas create a temporary windfall, over $20 billion for the U.S. Treasury in 2010.

Inflation

One of the most common problems in social science research and policymak-ing is how to adjust variables for the effect of inflation. Although inflation indices are readily available from U.S. government agencies, the correct use of these adjustments is critical for accurate research as well as for appropriate inflation indexing in many labor union contracts, the Social Security system, and the income tax.

For measuring U.S. inflation, nearly all indexes are variations of the U.S. Bureau of Labor Statistics Consumer Price Index (CPI). In recent years, the CPI has come under attack for measuring too much inflation, then too little inflation, and for making seemingly arbitrary assumptions. Although there are crucial issues that researchers need to understand when using price indexes, the Bureau of Labor Statistics is well aware of these potential pitfalls and, in fact, publishes alternative statistics to deal with them.

Quality. over time, goods and services change: Candy bars become smaller; the typical new car has improved standard features such as automatic braking systems; computers get faster. To adjust for these quality changes, BLS uses a number of methods. When a good simply gets larger or smaller, as in the case of a candy bar, the price can be adjusted up or down based on the new size. Somewhat trickier are new items such as automatic braking systems for which the BLS adjusts prices based on the item’s production cost. When goods are improved, as in the case of faster computers, BLS uses a technique called hedonic pricing (see Chapter 3), comparing what consumers had been willing to pay for similar, already improved products.

Substitution. Consumers switching to less expensive items requires another adjustment by the BLS. It makes sense for the BLS to decrease the

304

weight given in the index to products when consumers stop buying them, as for example when inexpensive downloads of movies threatened to replace mail order DVDs. on the other hand, if inflation itself causes consumers to switch away from a desirable item, then it is less obvious that the now less-used product should be removed from the index. Doing so would reduce the apparent inflation rate and not take into account decreased consumer well-being. Media accounts focused on such an alleged bias, calling it the “steak/hamburger” trade-off when budget- conscious consumers substituted hamburger for higher-priced steak. In fact, the BLS never made this adjustment. Instead the CPI is adjusted only for item weights within narrow groups, such as flank steak versus filet mignon. Moreover, the total effect of these substitutions amounted to less than 0.3 percent per year. Thus, although substitution effects make it difficult to measure inflation precisely, the impact of substitution on the CPI is minimal and to a degree is already taken into account.

Rental equivalence. In an attempt to more accurately measure changes in the prices of consumer goods, the BLS changed its measure of homeownership costs in 1983 to count how much it would cost the homeowner to rent a similar home rather than the actual cost of buying that house. The BLS wanted to omit the investment aspect of homeownership so that the CPI measured only the effect of prices on consumer goods and services. For the same reason, other investments such as stock and bond prices are not part of the CPI. Researchers need to know that, as a result of the CPI’s focus on consumption to the exclusion of investment, the CPI understated the full cost of living during the housing price bubble that occurred before 2007 and will overstate costs when home prices fall as they did after 2007 (see Chapter 3).

Which CPI?

BLS offers a number of inflation indexes, each valid in different contexts. The two main indexes are the CPI-U for all urban consumers, covering 87 percent of the population, and the CPI-W for urban wage earners and clerical workers, covering 32 percent of the population. These two indexes track one another closely: the CPI-U rose 3.6 percent in the year before June 2011, while the CPI-W went up by 4.1 percent, reflecting purchases by wage earners with slightly faster rising prices than the items bought by all urban consumers. The CPI-U is used for adjusting federal tax brackets, while the CPI-W is used to adjust Social Security and federal retirement benefits.

305

A third index, the chained CPI-U, first published in 2002, fully reflects consumer substitution of less expensive alternatives, not only the switching between closely related items described above. A 2011 proposal to use this index for government programs generated a strident debate. At that time the chained CPI was running about half a percentage point lower than the unchained CPI. Consequently, some political leaders urged a switch to the chained index in order to save more than $200 billion in the federal budget over the next 10 years. Critics pointed out that the change meant a $600 reduction in average annual Social Security retiree benefits. Moreover, if the purpose was to adjust Social Security payments to what seniors actually purchased (potentially inaccurate with the regular CPI-U because of the substitution effect), then it made more sense to use a CPI index already constructed for the elderly—one that showed a 0.2 percent higher annual rate of inflation because of older persons’ purchases of medical care with quickly rising prices.

Even if the switch to the chained CPI may have been politically motivated in the 2011 debate, the statistic is nonetheless of potential use for researchers. Chain-style indexes were incorporated into GDP calculations in 1996 (see Chapter 7) and the chained CPI is appropriate for research requiring a cost-of-living measure.

Currency Rates

Converting currency values between countries is a common social science problem, providing a challenge in the interpretation of data from different countries. Many comparisons depend critically on how prices are converted to a common currency.

Box 11.4 Annual Inflation Rate

It is common in economic research to use annual inflation rates. In addition to choosing the appropriate index, researchers also must select between various starting points for the year. The Bureau of Labor Statistics (BLS) offers a series called the “annual average” that measures the percent change between (1) the average of the consumer price index (CPI) in one year’s 12 months and (2) the average of the CPI in the previous year’s 12 months. By this measurement, the percent change in the CPI-U (for all urban consumers) from 2009 to

306

2010 was 1.6 percent. Alternatively, in the monthly reports, inflation is measured by the percent change in the CPI compared with 12 months earlier. In December 2010, the CPI-U was 1.5 percent higher than a year earlier. These differences usually are minor but can lead to differing inflation adjustments. Measurements used in research projects should be consistent.

Source: U.S. Department of Labor, Bureau of Labor Statistics, “Math Calculations to Better Utilize CPI Data,” http://www.bls.gov/cpi/cpimathfs.pdf.

Box 11.5 Effect of Rounding

In reporting data, it is necessary to round off so that numbers are readable and to avoid the false impression that extra decimal places convey meaningful information. For example, the U.S. Bureau of Labor Statistics rounds off the reported CPI to one decimal place. For nearly all purposes, this is all the accuracy needed—or warranted— by the underlying measurement. However, problems can occur because of rounding, especially for small monthly movement in prices. In 2005, financial markets raised interest rates on the mistaken impression that inflation unexpectedly had jumped to 0.3 percent in February from its 0.2 percent level during previous months. BLS economist Elliot Williams points out that by using the unrounded monthly index numbers, the February inflation rate actually was only 0.2 percent; rounding down of the index in the previous month and rounding up in the current month caused the misleading 0.3 percent change and an unwarranted financial market response. In addition, BLS warns that even if a 0.1 percent change had occurred, it could have been the result of random error that typically occurs at least once every 20 months simply because of sampling error in their estimation procedures.

Sources: Elliot Williams, “The Effects of Rounding on the Consumer Price Index,” Monthly Labor Review, October 2006, pp. 80–89. Sampling error in U.S. Department of Labor, Bureau of Labor Statistics, “Note on Sampling Error in the Consumer Price Index,” CPI Detailed Report Data for May 2007, http://www.bls. gov/cpi/cpid0705.pdf.

307

The Problem of Many Currencies

Newspaper and magazine financial pages show the daily ups and downs of world currencies relative to one another. For example, on December 21, 2011, the U.S. dollar was up slightly against the Canadian dollar while down relative to the Japanese yen. But many research projects require a statistic for the overall movement in the dollar. Common sense suggests that in measuring the value of the U.S. dollar, we should count currency changes more for top trading partners, such as China and Canada, than changes relative to currencies in Sweden or Argentina, with whom the United States does less trade. In fact, just such trade-weighted indexes are calculated by the U.S. Federal Reserve Bank, the European Central Bank, and several commercial enterprises, but differences in the methods for calculating these indexes cause research and policymaking confusion.

Prior to 1998 the most commonly cited index was the Federal Reserve Board weighted value based on trade with ten countries: Germany, France, the United Kingdom, Belgium, Italy, Sweden, Spain, Switzerland, Canada, and Japan. This short list was sensible in 1972, when these countries accounted for most U.S. trade. However, by 1998 the index had become obsolete, first because five of the ten countries were about to consolidate their currencies in the new euro, and second, because of increased U.S. trade with Korea, Mexico, Brazil, and especially China, countries not on the original list.

The Federal Reserve now publishes a Broad index with 26 currencies, each exceeding one-half percent of total U.S. imports or exports, comprising well over 90 percent of total U.S. trade. A second index for Major Currencies includes only the euro, Canadian dollar, Japanese yen, British point, Swiss franc, Australian dollar, and Swedish krona, selected because these currencies are traded widely in open exchange markets. The appropriate index depends on the research purpose. For analyzing worldwide trade, the Broad index likely will be best, whereas studies of currency speculation, limited to key currencies, might use the Major Currency index. Both indexes are available in nominal and real (adjusted for inflation) terms, helpful when greatly different inflation rates between countries may distort currency values. One potential drawback to these indexes is that the trade weights are updated annually, an important adjustment to take into account new trade patterns. For example, China’s weight, the highest of all countries at 19.9 percent in 2011, was only fifth

308

in the world at 8.8 percent in 2001. However, if one is studying events over time, for example future commodity trades, then it may be appropriate to use alternatives such as the privately produced Intercontinental Exchange, which keeps a fixed weight for each currency so that dollar values don’t change when the Federal Reserve makes its annual adjustment for new trade weights.

Implications

Correcting prices for the effects of inflation and different currency values is a problem in research on a wide range of issues. Cost data on health, housing, and crime, as well as economic data on incomes, businesses, and government, usually must be adjusted for inflation (if different time periods are compared) or for currency values (if different countries are compared). The controversies described here serve as an introduction to these adjustments and as a caution to their use. The overall lesson is that researchers should investigate alternative measures of inflation and currency value that would demonstrate how results change when different adjustments are used. Alternatively, if one measure of inflation or currency value is considered to be more accurate, then its use should be carefully justified.

Fewer Voters?

Is U.S. electoral participation on the decline? Newspaper headlines notwithstanding, it isn’t clear that citizens are becoming more apathetic. The voter participation rate is surprisingly difficult to measure, prompting a debate among political scientists about its trend.

Even the numerator, the number of voters, isn’t a simple matter to count. Ideally, we would want to know the total number of citizens who voted in a given election, but this number is not reported in 13 states. Instead, political science researchers usually count the vote for the highest office on the ballot. The problem is that not everyone votes for every office, so the count of votes for the highest office will miss some voters who didn’t make a choice among those highest office candidates. This factor alone leads to a more than 2 percent undercount of voters.

However, the main problem is what to put in the denominator of the voting rate. Should the number of voters be divided by the population of voting-age citizens, or should it be divided by the number eligible to vote? Traditionally, political scientists measured voter participation as a

309

percentage of the voting-age population, an easy-to-find and agreed-upon number showing voter turnout at lower levels in recent decades. However, some analysts have shifted to a count of eligible voters compiled by political scientist Michael P. McDonald. on this basis, the apparent decline in voting participation disappears. The reason the percent of eligible voters differs from the percentage of the voting-age population is that noncitizens rose to 8 percent of the population in 2000 from 2 percent in 1966, and ineligible felons increased to 1.4 percent from 0.5 percent. Thus, while overall voter participation as a percentage of everyone over 18 has fallen, there has not been a similar decrease in voting by those who can vote.

McDonald maintains that if we are interested in measuring civic interest (or apathy), it makes sense to look only at the percentage of eligible voters. After all, noncitizens or otherwise ineligible voters may have great concern about political matters that simply cannot be expressed by voting. With this correction, voter turnout has changed somewhat over the years, declining when 18 year olds were enfranchised in 1970, but then mostly offset by an increase in participation, particularly in presidential elections. As a result, the 2008 presidential turnout, as a percentage of eligible voters was above the average for the previous century and the 2010 midterm turnout was about average compared to the previous 100 years.

Political scientist Martin Wattenberg responds that the voting decline is real no matter which denominator is used. In his analysis, nationwide statistics obscure the difference between regions in which a huge gain in Southern states’ voting, up by 16 percent between 1990 and 2004, was offset by a 10 percent decline outside the South. Thus, in his view, higher Southern voting caused by the civil rights movement masks an overall actual decline in civic participation.

Wattenberg prefers the voting-age population denominator also on the grounds that voting statistics should show us what proportion of the affected population is making our political choices. California, for example, appears quite average in its 57 percent participation rate by eligible voters, but only 47 percent of its voting-age population took part in the 2004 presidential election because a high percentage of its population were noncitizens.

Both McDonald and Wattenberg agree on one point: Far greater than any change in overall voting participation over time are the differences in voting between states. Minnesota tops the country with 78 percent of eligible voters in the 2008 presidential election, attributed to easy voter

310

registration and an educated population. On the other hand, in the same election, only 55 percent voted in Texas, perhaps because of state statutes that make it more burdensome to vote as well as a less educated population. Also, voting participation varies by age, especially for midterm elections when voters over 65 turn out at more than double the rate for those 18 to 25 years of age.

The debate about U.S. voter participation underscores the need for researchers to look at multiple statistical measures. In particular, analysis of the historical trend is ambiguous, dependent on whether voter participation is measured as a percentage of the population or as a percentage of those eligible to vote. Researchers may need to look at both rates in order to fully understand when and why Americans vote.

Summary

Data covered in this chapter have the advantage of accuracy in the sense that the numbers come from complete counts rather than surveys. Even inflation measures that do rely on survey data are derived from such massive samples that sampling error is not an issue. Instead, the problem for researchers is choosing which numbers to use. In the case of government budgets, inflation rates, and currency values, there is a choice between conflicting official statistics. For military spending and the deficit, official data are challenged by alternative data based on very different accounting techniques. Finally, welfare spending, tax rates, and voter participation rates have been shown to generate different statistics and opposite conclusions depending on how the terms are defined.

As this chapter demonstrates, official government data are not the only data available; nor are they always the most appropriate. But when an alternative statistic is used, explicit discussion of the choice is necessary. The best solution is to test research hypotheses based on varying assumptions. In the case of inflation adjustment where several indexes are available, research results are far more convincing if they do not change significantly when alternative statistics are used.

Case Study Questions

The proportion of the U.S. currency in $100 bills increased to 75 percent in 2010, up from 52 percent in 1990. What might explain this change?

311

Studies of taxes often focus on federal taxes. But state and local taxes often take a bigger bite out of income, especially for low-income groups. Which taxes in your state are progressive? Which are regressive? Why?

Most state and local governments separate their current budgets from capital budgets that are reserved for spending on long-lived projects such as roads and buildings. If the U.S. government were to issue a capital budget, list the types of items that would be included. How would this new budget technique affect the measurement of the U.S. budget deficit?

The proportion of the U.S. budget designated as investment depends critically on whether military spending is counted as investment or as consumption. Which components of military spending might be counted correctly as productive investment?

References

Data Sources (233)

Military spending data sample in SIPRI Military Expenditure Database http://milex-data.sipri.org/result.php4/. Travelers checks data sample in U.S. Federal Reserve, “Money Stock Measures H6,” http://www.federalreserve.gov/releases/h6/current/. Consumer and producer price indexes described in U.S. Department of Labor, B ureau of Labor Statistics, http://www.bls.gov/bls/inflation.htm. Data sample in U.S. Department of Labor, Bureau of Labor Statistics, “CPI Detailed Report— May 2011,” table 3, p. 8, http://www.bls.gov/cpi/cpid1105.pdf. Ineligible to vote data sample in United States Election Project, “2010 General Election Turnout Rates,” http://elections.gmu.edu/Turnout_2010G.html.

Controversies (234)

How Much for the Military? (234)

How Big Is the U.S. Military Budget? (235): Summary in Chuck Spinney, “Madison’s Nightmare: How Much Should We Spend for National Insecurity?” The Atlantic, http://www.theatlantic.com/politics/archive/2011/02/madisons-nightmare- how-much-should-we-spend-for-national-insecurity/70687/. On possible additions to military budget, see Chris Hellman, “$1.2 Trillion for National

312

Security,” TomDispatch.com, http://www.tomdispatch.com/archive/175361/.

How Much Did the Wars in Afghanistan and Iraq Cost? (235): Summary of debate in Gregory D. Hess, ed., Guns and Butter: The Economic Causes and Consequences of Conflict (Cambridge: MIT Press, 2009); “United States: Blood and Treasure—Paying for Iraq,” The Economist 379, April 8, 2006, p. 53, and Paul Sol-man, “Economics of War,” Journal of Economic Education 39, no. 4 (Fall 2008): 391–400. Original Bush administration estimate in Joseph E. Stiglitz and Linda J. Bilmes, “The True Cost of the Iraq War: $3 Trillion and Beyond,” Washington Post, September 5, 2010, p. B04. Anna Bernasek, “An Early Calculation of Iraq’s Cost of War,” the New York Times, October 22, 2006. Steven J. Davis, Kevin M. Murphy, and Robert H. Topel, “War in Iraq versus Containment,” National Bureau of Economic Research (NBER) Working Paper No. 12092, March 2006. Charles Hill in Solman, “Economics of War.”

How Much for Welfare? (236)

Robert Rector, Katherine Bradley, and Rachel Sheffield, “Obama to Spend $10.3 Trillion on Welfare,” Washington, DC: The Heritage Foundation, 2009, http://www.heritage.org/Research/Welfare/sr0067.cfm. Cost of mortgage interest deduction, Paul Sullivan, “Despite Critics, Mortgage Deduction Resists Change,” the New York Times, November 8, 2011, p. F4. Health insurance exclusion in Jonathan Gruber, “The Tax Exclusion for Employer-Sponsored Health Insurance,” NBER Working Paper No. 15766, February 2010.

How Big Is the Deficit? (237)

U.S. deficit in U.S. Department of the Treasury, Financial Report of the United States Government, http://www.fms.treas.gov/finrep11/citizenguide/fr_citizen_guide.html. off- budget items in New York in Tom Precious, “Democrats Agree on Plan to Fix Deficit,” Buffalo News, February 4, 2009, p. 1; in Connecticut, “State Budget Vote Expected Today,” Connecticut Post, February 25, 2009; in New Jersey, John Reitmyer, “State Budget Math Isn’t a Simple Equation,” The Record, April 12, 2009, p. 1. Valuing pensions in Robert Novy-Marx and Joshua D. Rauh, “The Liabilities and Risks of State-Sponsored Pension Plans,” Journal of Economic Perspectives 23, no. 4 (Fall 2009):

313

191–210; Dean Baker, “The origins and Severity of the Public Pension Crisis,” Washington, DC: Center for Economic and Policy Research, February 2011, www.cepr.net/documents/publications/pensions-2011– 02.pdf. Selling off government assets in Mark Guarino, “The Great Sell- off: Chicago Auctions City Assets,” Christian Science Monitor, June 24, 2009; John Kendrick, “Federal Government Could Reduce Debt by $1.5 Trillion with a Sale of Unneeded Assets,” The Foundry: Conservative Policy News Blog from The Heritage Foundation, http://blog.heritage.org.

A Debt Monster? (240)

$14 trillion in U.S. Department of the Treasury, Bureau of Public Debt, “Debt to the Penny and Who Holds It,” http://www.treasurydirect.gov/NP/BPDLogin?application=np; obama and Krugman in Paul Krugman, “What obama Wants,” the New York Times, July 7, 2011, p. A23. Debt held by public, http://www.treasurydirect.gov/NP/BPDLogin?application=np. Aaron in Letters, the New York Times, March 12, 1980, IV, p. E24. Carmen M. Reinhart and Kenneth S. Rogoff, “Growth in a Time of Debt,” American Economic Review 100, no. 2 (2010): 573–78; criticism in Josh Bivens and John Irons, “Government Debt and Economic Growth,” Economic Policy Institute Briefing Paper #271, July 26, 2010, http://www.epi.org/publication/bp271/.

Taxes (241)

CBo methodology in U.S. Congressional Budget office, “Historical Effective Federal Tax Rates: 1979–2004,” http://www.cbo.gov/publication/18278; CBo, “Average Federal Tax Rates in 2007,” http://www.cbo.gov/publication/21521; Citizens for Tax Justice, “Who Pays Taxes in America?” http://ctj.org/ctjreports/2012/04/who_pays_taxes_in_america.php. Top 1 percent share of income increase in CBo, “Trends in the Distribution of Household Income Between 1979 and 2007,” http://cbo.gov/publication/42729. Slemrod in “The Distribution of Income Tax Noncompliance,” (with Andrew Johns), National Tax Journal, 63, no. 3 (September 2010): 397–418.

Measuring Money (243)

Size M1, M2, in U.S. Federal Reserve Statistical Release, “H.6: Money

314

Stock Measures,” http://www.federalreserve.gov/releases/h6/current/default.htm; Federal Reserve balance sheet at “Credit and Liquidity Programs and the Balance Sheet,” http://www.federalreserve.gov/monetarypolicy/bst_recenttrends.htm. “Currency in Circulation: Value,” http://www.federalreserve.gov/paymentsystems/coin_cur-rcircvale.html; Binyamin Appelbaum, “As Plastic Reigns, Printing of Money Slows,” New York Times, July 7, 2011, p. B1.

Inflation (246)

John S. Greenlees and Robert B. McClelland, “Addressing Misconceptions About the Consumer Price Index,” Monthly Labor Review, August 2008, pp. 3–19. Chained CPI in J. Steven Landefeld, Brent R. Moulton, and Cindy M. Vojtech. “Chained-Dollar Indexes,” Survey of Current Business, November 2003, pp. 8–16. CPI and Social Security in Dean Baker, “On Using the Chained CPI for Social Security Cost of Living Adjustments,” http://mrzine.monthlyreview.org/2011/baker080711.html. CPI for the elderly in Kenneth J. Stewart, “The Experimental Consumer Price Index for Elderly Americans (CPI-E): 1982–2007,” Monthly Labor Review, April 2008, pp. 19–24.

Currency Rates (248)

Change in TWI in Mico Loretan, “Indexes of the Foreign Exchange Value of the Dollar,” Federal Reserve Bulletin, 91, no. 1 (Winter 2005): 1–8. China’s weight in Board of Governors, U.S. Federal Reserve, “Foreign Exchange Rates: H.10,” http://www.federalreserve.gov/releases/h10/weights/previousweights.htm; Intercontinental Exchange at http://www.theice.com/about.jhtml.

Fewer Voters? (251)

Number of voters in Michael P. McDonald, “Voter Turnout in the 2010 Midterm Election,” The Forum, 8, no 4, article 8, http://www.bepress.com/forum/vol8/iss4/art8; and Michael P. McDonald and Samuel L. Popkin, “The Myth of the Vanishing Voter,” American Political Science Review 95, no. 4, (December 2001): 963–74. Number eligible to vote in ibid.; Martin P. Wattenberg, “Elections: Turnout in the 2004 Presidential Election,” Presidential Studies Quarterly 35, no. 1

315

(March 2005): 138–36. Minnesota and Texas turnout in Nonprofit Voter Engagement Network, America Goes to the Polls: A Report on Voter Turnout in the 2008Election, March 2009, p. 5. Over-65 voting rate in McDonald, p. 7.

Case Study Questions (253)

1. Board of Governors of the U.S. Federal Reserve, “Currency in Circulation: Value,” February 22, 2011, www.federalreserve.gov/paymentsystems/coin_cur-rcircvalue.htm.

2. See Citizens for Tax Justice, “Who Pays Taxes in America?” http://ctj.org/ctjreports/2012/04/who_pays_taxes_in_america.php.

3. See http://clinton3.nara.gov/pcscb/report.html.

4. See Robert Eisner, The Misunderstood Economy: What Counts and How to Count It (Boston: Harvard University Press), 1994).

316

12

Public Opinion Polling □□□□

Public polling is a twentieth-century science first developed as a method to predict elections and later expanded to include opinion polling on a wide variety of issues. By the early twenty-first century, many businesses were hiring pollsters to take the public pulse about their products and the advisability of introducing new items. A number of political and business polls are proprietary, released to the public only when it serves the purpose of the poll’s subscriber. Nonetheless, enormous numbers of election and public opinion polls are available to researchers. These include several decades of data predicting elections and then analyzing voter behavior afterward. In addition, private and university-affiliated opinion surveys are designed to trace public opinion over long periods of time. They cover most current public policy issues, ranging from support for the U.S. president to the public’s general happiness. All are well indexed and readily accessible for research use.

Where the Numbers Come From

Organizations Data sources URL The Gallup Organization Gallup poll www.gallup.com Roper Center for Public Opinion Research

Most major U.S. polls and polls from several other countries

www.ropercenter.uco- nn.edu

American National Election Studies

American National Election Studies

www.electionstudies.o- rg

Center for the Study of Politics and Society, National Opinion Research

General Social Survey www.norc.org

317

Center

Data Sources

Private Polling Organizations

The U.S. polling industry comprises more than 200 private companies including the household names Gallup, Nielsen, and Roper. Beginning with the 1960 presidential election, pollsters have assisted every major U.S. political campaign. Even on the local level, political candidates typically spend 5 to 15 percent of campaign funds on polls.

Data Sample: A 2010 Gallup poll found that 40 percent of U.S. respondents chose a “creationist” explanation for the descent of humankind: “God created human beings pretty much in their present form at one time within the last 10,000 years.” 16 percent chose a secular version: “Human beings have developed over millions of years from less advanced forms of life. God had no part in this process”; while 38 percent preferred: “Human beings developed over millions of years from less advanced forms of life, but God guided this process.”

Media Polls

CBS entered the market first in 1967 with its own election-polling unit and created a partnership with the New York Times in 1975 that soon became the major force in media polling. Similar collaboration has been established between ABC and the Washington Post and NBC and the Wall Street Journal. The increase in polls was particularly large in the 1990s: a study of media polling by the Roper Center measured an increase to more than 4,000 questions per year in 1990 from about 1,000 poll questions asked by news organizations in the mid-1970s.

Research Centers

The Center for Political Studies at the University of Michigan’s Survey Research Center and Stanford University’s Institute for Research in the Social Sciences collect economic data from consumers and voting and attitude data in the American National Election Studies, which are conducted every other year.

318

Data Sample: In the 2008 American National Election Studies, 60 percent of respondents agreed with the statement, “Public officials don’t care much what people like me think,” more than double the percent agreeing with that statement in 1956.

The National Opinion Research Center (NORC) at the University of Chicago collects data for the U.S. government, including the National Longitudinal Survey (see Chapter 5). The NORC’s General Social Survey (GSS), conducted since 1972, is a major data source for public behavior, attitudes, and happiness with various aspects of life. It contains a core of replicated questions that enable researchers to track changing attitudes over time.

Data Sample: In the GSS, willingness to vote for a woman for president rose from 70.5 percent in 1972 to 93 percent in 2008.

Controversies

In 1936 a poll taken by the Literary Digest was sensationally wrong, predicting a 14 percent margin for Alf Landon over incumbent President Franklin D. Roosevelt, who won with an even larger margin than the Digest predicted for Landon. This tremendous error cast a shadow of doubt on political polling that has yet to disappear and that caused the Digest to go out of business the following year.

Strangely, the straw poll methodology that killed the Digest was successful in four previous presidential elections, even predicting Roosevelt’s 1932 margin with only a 0.7 percent error. What went wrong in 1936? The Literary Digest poll was exceptionally large, with ballots sent to more than 10 million households. But large samples are not necessarily better: The Literary Digest poll suffered from sample bias because names were drawn from telephone books and auto registration lists, a group decidedly better off than the typical Depression-era voter. Even with its biased sample, the Digest poll would have correctly projected Roosevelt the winner if everyone had returned the survey. Of the 10 million ballots distributed, however, only 2 million were returned, most in support of Landon, whereas those who did not return the Digest ballots were overwhelmingly Roosevelt voters. From 1920 to 1932, the sample bias toward wealthier Republicans was offset by a higher response rate from Democrats who wanted to voice their opposition to incumbent Republican presidents. In 1936, though, the two biases reinforced each

319

other when disgruntled Republican subscribers responded in higher numbers than the Democratic majority.

The Literary Digest fiasco shows how political polling—and polling in general—can go wrong. Lesser known is the stunning success of George Gallup, who offered money back to his poll subscribers unless he was more accurate than the 1936 Digest poll. Gallup won his audacious gamble by predicting a Roosevelt victory, albeit by a margin that was 7 percent too small. Even though Gallup’s sample was a fraction of the size of that used by the Digest, it was less biased because of a quota system that guaranteed representative proportions based on economic class, age, gender, and political preference. In the 1948 presidential election, however, Gallup received his comeuppance when he, along with most other pollsters, predicted that Harry Truman would lose to Thomas E. Dewey. Compared to 1936, Gallup came closer to the actual vote in 1948 with just a 5 percent error, but this was small consolation for missing the winner.

In this chapter we highlight many of the issues that plague surveys and polls, both within and outside the political arena. Sample bias is only one of many factors that can influence whether a poll’s results accurately represent the true opinions of the population.

Sampling

Are polls today superior to those conducted in 1936 or 1948? The answer is, “usually”: Because of superior sampling methods, modern polls rely on a sample of about 1,000 individuals. Statistical theory suggests that the error caused by random chance in these polls should be plus or minus three percentage points in 19 out of 20 samples. However, there are many reasons why a sample may not be truly random: when a sample systematically differs from a random draw of the population, we say the results are biased. Bias most commonly arises when a sample excludes particular segments of the population, when members of the intended sample choose not to respond to survey questions, and when survey respondents volunteer to participate.

Cell Phones v. Landlines

Most opinion polling in the early 2000s is conducted by telephone, a shift from the early Gallup and Roper polls, when pollsters traveled around the country interviewing people face to face. The reason for the change is cost; as one pollster explains, “only the federal government can come up with

320

the bucks for in-person interviewing.” By the 1970s more than 90 percent of the population could be reached by telephone, and the bias toward the well-to-do of the 1930s Literary Digest poll could be avoided. Even those who have unlisted phone numbers are included in most surveys by randomly choosing telephone numbers. For example, the Gallup poll chooses the first digits of phone numbers to assure geographic representation and then randomly selects the last digits to include both listed and unlisted numbers.

However, the increasing use of cell phones, particularly a growing “cell phone only” population, raises new concerns about the representativeness of landline phone surveys. Recent estimates suggest that around 13 percent of U.S. households are cell-only, but, not surprisingly, there are large differences across age groups, with over 25 percent of those under age 30 abandoning landlines. There are other demographic differences as well, with cell-phone only households more likely to be urban and low-income. In several dual frame surveys conducted in 2006, both cell phone and landline numbers were included. The inclusion of cell-only respondents did not change the overall estimates very much; however, important differences were found between cell-only and landline respondents, suggesting that bias may increase as the cell-only share of the population grows. There is also evidence that there is bias in estimates for subgroups with higher cell-phone usage, such as young adults and lower-income households. Many pollsters apply some form of weights to account for such things as demographic variation in nonresponse for landline surveys, but it is unclear that these weights are sufficient to capture what is going on with rapidly changing cell-phone only populations. For example, a 2009 study found that voters over 30 are abandoning landlines at a faster rate than younger voters, and differences between landline and cell-only users are larger for the older voters. This runs counter to what many analysts expect and suggests that age weights in landline samples may not account for the appropriate level of bias.

The rise of cell-phone only respondents poses a particular problem for political pollsters. At the national level, many polling organizations now routinely call both landlines and cell phones but given the higher cost associated with calling cell phones, landline-only surveys are still common at the state level. One study found that failing to include cell phones in voter polls might lead to a bias against Democrats, even after demographic weighting had been applied. However, Nate Silver, a statistician with a popular blog that regularly analyzes political polling, points out that

321

weighting techniques can vary a lot from pollster to pollster, so not all landline-only polls will suffer from the same bias. It is also difficult to compare landline and cell-phone voter surveys because of differences in methodology; for example, federal law prohibits the use of automated dialers when calling cell phones, so cell-phone samples all use live operators, which can affect response rates.

Nonresponse

Another problem for modern pollsters is an increase in the refusal rate, driven in part because respondents confuse legitimate polls with the large number of telephone sales pitches. This problem has particularly grown as caller ID technology has made it easier for potential respondents to avoid picking up the phone in the first place. For example, the response rate for the University of Michigan’s Survey of Consumer Attitudes declined sharply after 1997 from just above 60 percent to just below 48 percent in 2003. When more than half of the sample does not respond, there is much higher risk of misinterpretation when the results are extrapolated to the population as a whole.

Many polls correct for nonresponse by weighting the final results to compensate for those who refused to participate. For example, those respondents who say they were not home on the previous evening may be counted again as proxies for those who were not home on the day of the survey. The assumption is that individuals who were not home previously share similar characteristics with those who cannot be reached for the survey. Similarly, pollsters can adjust survey data for misrep-resentative samples, adjusting the total to include the known proportion of men and women, Republicans and Democrats, or those likely to vote. This weighting is standard practice, used in most polling by Gallup, ABC/Harris, and CBS/New York Times. Clearly weighting is better than an original sample that may be biased because it includes only those who have telephones, answered their phones, and were willing to talk. But it is not often obvious how large an impact weighting might have. Researchers also need access to the original polling methodology to assess the significance of weighting, and few polls publish such information.

Self-Selected Internet Polls

The rise of the World Wide Web has been a large factor in the explosion of polls in the last decade. Web-based surveys allow traditional polling

322

organizations to reach larger samples with lower cost and faster data collection. But Internet polls are likely to suffer from the same type of sampling bias that affects cell phones and landlines. Although there has been rapid growth in Internet access across the United States and the world, there are still large shares of the population that remain without access, and the characteristics of those who do and do not have access to the technology are systematically different. Some polling firms attempt to deal with this by providing Internet access for those without it; others use weighting schemes to compensate for differences between web-based samples and the general population.

Sampling bias is likely to be exacerbated further when self-selected volunteers are the ones answering questions. Many professional polling organizations start with a representative intended sample and recruit respondents in various ways, in some cases providing Internet access when necessary, but there are also many Internet surveys that rely entirely on samples of volunteers. The resulting responses are even less likely to represent the general population. Web-based surveys also make it possible for every media outlet, big or small, to poll their audience on a wide range of issues, but many of these surveys are not scientific. Researchers using information from any web-based survey should look carefully at how a sample was selected, who responded and who did not, and whether/ how the reported results have been weighted by the pollster.

Survey Design

Even when a survey sample is truly random and representative of the population of interest, poll responses still may not always accurately reflect public opinion. Several issues related to how surveys are designed can affect the validity of the results.

Question Wording

Survey experts identify question wording as the most critical factor in poll results—and the one most easily manipulated to gain desired results. Sometimes the effect is dramatic, as in the following examples. The General Social Survey asks respondents if “we spend too little on”:

Assistance to the poor

65%

Welfare 20%

323

Assistance to big cities

18%

Solving problems of big cities

48%

Assistance to blacks 27% Improving condition of blacks

37%

Box 12.1 Bias from Self-Reporting

In 2000, data from the Nielson ratings service showed that 30 to 35 million people watched the nightly news on an average weekday; however, according to the National Annenberg Election Survey, between 85 and 110 million people said they watched the news daily, with higher levels of overreporting among higher-income viewers and viewers in households with children. Interestingly, the amount of overreporting seems to be stable over time; that is, surveys taken at different points in time exhibit the same average level of overreporting. There is also evidence that people overreport listening to National Public Radio and how often they attend church, and either men overreport the number of their sex partners or women underreport.

It may be that people honestly misestimate their own behavior (for example, believing they watch more news than they actually do), or they may respond in ways that they believe reflect a more socially acceptable answer. In support of the latter explanation, there is evidence that people answer surveys differently when they are administered by a live caller rather than an automated poll. For example, in 2010, there was a stark contrast among surveys about California’s Proposition 19, which would have legalized marijuana. Traditional live interviews found that only 41 percent of respondents supported the measure and 46 percent did not; however, automated polls found 56 percent in support and only 41 percent against. Thus, predictions about whether the poll would pass varied greatly, depending on which poll was used. On the other hand, Prop 19 ultimately failed with 53 percent of the voters voting against it, so it

324

is less clear why the automated polls found such high levels of support.

Sources: Markus Prior, “The Immensely Inflated News Audience: Assessing Bias in Self-Reported News Exposure,” Public Opinion Quarterly 73, no. 1 (2009): 130–43. Prop 19 polls in Kevin Drum, “Will California Legalize Pot?” Mother Jones, October 24, 2010, http://motherjones.com/kev- in-drum/2010/10/will-california-legalize-pot.

Box 12.2 Randomizing Response Techniques

As suggested by the Prop 19 polls (see Box 12.1), self-reporting can be particularly problematic for questions that are highly personal or morally sensitive. The Randomized Response Technique (RRT) provides a way for respondents to reveal potentially embarrassing information in a way that is anonymous to the person administering the survey. The RRT presents the respondent with two questions, one that is the true question of interest and one that has the same answer choices but is completely benign. For example, in one study, the question of interest was “Have you ever had premarital sex?” and the benign question was “Were you born in an even year?” The respondent is then asked to flip a coin and, without revealing the outcome of the coin toss to the interviewer, answer Question 1 if the coin turns up heads and Question 2 if the coin turns up tails. Thus, the interviewer only hears “yes” or “no” but does not know which question is being answered. Simple probabilities can then be used to calculate the percentage of people who have actually had premarital sex.

In many studies, respondents are much more likely to report “undesirable” behaviors when RRT is used than when asked direct questions. However, a handful of studies have found no significant differences between RRT and direct questions. For example, a 2010 study of voter turnout found that the RRT led to higher estimated rates of voter turnout than direct self-reports, the opposite of what analysts usually expect. The researchers note that the surveys were conducted via telephone and the Internet, and respondents were simply more likely to ignore the RTT questions. Some of the

325

researchers posit that there might have been implementation problems; at the very least, the inconsistent results suggest the RTT may not provide more accurate results in all circumstances.

Sources: M.B. Soudarssanane et al., “CME: Randomized Response Technique— An Innovative Method to Measure Culturally Sensitive Variables: Result from a Pilot Study,” Indian Journal of Community Medicine 28, no. 3 (2003): 138–40; Allyson L Holbrook and Jon A. Krosnick, “Measuring Voter Turnout by Using the Randomized Response Technique: Evidence Calling into Question the Method’s Validity,” Public Opinion Quarterly 74, no. 2 (2010): 328–43.

These are predictable wording effects, well known to pollsters. Welfare has negative connotations, as do terms such as big business or big labor; positive verbs, such as solving or improving, increase the popularity of an answer. Most recently, in polls that ask if people support “creating jobs” versus “cutting spending,” or “reducing unemployment” versus “reducing the deficit,” creating jobs and reducing unemployment have sizable majorities; however, when the same choices are framed as “spending on recovery” versus “reducing the deficit,” support flips to deficit reduction.

In some cases, the effects of question wording are quite subtle. In response to the question, “Do you usually think of yourself as a Republican, a Democrat, an Independent, or what?” political scientist Barry Burden found that responses pointed to a gender gap, with women more likely to identify themselves as Democrats. However, replacing the word think with the word feel (that is, “Do you feel that you are a Republican … ?”) shrinks that gap considerably, with women much more likely to move toward a Republican identity.

Stanford professor Jon Krosnick suggests that when carefully interpreted, polls can actually tell us what the public thinks—even when changes in question wording give apparently contradictory answers. The problem tends to lie with media headlines that simplify poll results and infer opinions that are not actually expressed (or even asked about). For example, some headlines have suggested a growing share of the population does not consider global warming as a real problem and are skeptical about climate change. These poll results have led some policymakers to back off from environmental regulation. But in surveys conducted by Krosnick and colleagues, large majorities (well over 70 percent) agree that the earth’s temperature has been heating up over the last century, that human behavior

326

is substantially responsible for that warming, and that government should limit business’s emissions and should provide tax breaks to encourage alternative fuels and production of environmentally friendly goods.

The apparent discrepancy can perhaps be resolved when one takes a closer look at the actual questions asked in the polls. Among the polls that were interpreted as showing growing public skepticism about climate change, two of the questions asked were: “From what you’ve read and heard, is there solid evidence that the average temperature on earth has been getting warmer over the past few decades, or not?” and “Thinking about what is said in the news, in your view, is the seriousness of global warming generally exaggerated, generally correct, or is it generally underestimated?” The first question asks for the respondent’s perception of evidence, not their personal opinion about whether the earth is getting warmer; similarly, the second question asks for the respondent’s perception of media coverage, not their opinion about global warming itself. It is certainly possible that some people may not know whether there is solid evidence, and they may believe that the idea of global warming is overhyped in the media, yet still believe climate change is a real problem. When more direct questions are asked—“Do you believe global warming is happening?” “Do you support government action to address global warming?”—there is widespread agreement that the problem is real and needs to be addressed.

Careful attention to question wording also explains apparently conflicting attitudes on other contentious issues such as abortion rights. A CBS poll found more than 70 percent supporting the “right of a woman to have an abortion,” but a Harris poll found the public evenly split on “legalized abortion.” Such results provide simplistic propaganda for both sides in the dispute, but not the sophisticated views of thoughtful respondents. Michael Kagay, New York Times director of polling, concludes: “When people want to make such distinctions, pollsters should too—if they are to give an accurate picture of public opinion.” A 2009 Gallup poll found only 23 percent of U.S. respondents opposing abortion in all circumstances and only 22 percent permitting abortion in all circumstances. Given the opportunity, most respondents reported a shaded position in between these two extremes.

These examples demonstrate the critical nature of question wording. The American Association for Public Opinion Research (AAPOR) standards recommends disclosure of full question wording, a requirement that is not always followed.

327

Question Order

Even when survey results report exact question wording, they rarely tell us the order in which questions were asked, a difference that can also affect the outcome. On October 13, 1982, the New York Times reported Toby Moffett in the lead over incumbent Senator Lowell Weicker of Connecticut, while the Hartford Courant found Weicker—the eventual winner—ahead by 16 percentage points.

To pollsters Irving Crespi and Dwight Morris, the results were extraordinary because both polls used large samples and similar questions. Even the keypunching of the data was double-checked and verified. Careful detective work revealed a subtle difference—and a warning to poll users. The Courant poll asked respondents for their vote in the Senate race first, then in the governor’s contest; the Times reversed the order, thereby inadvertently skewing the results between Moffett and Weicker. Apparently voters who had first indicated a preference for Democratic Governor O’Neill were reluctant to admit that they planned to vote for the Republican Weicker. A large number of middle-of-the-road voters supported the moderate Democrat O’Neill and moderate Republican Weicker—but did not want to appear inconsistent to the poll takers.

A second problem involving question order occurs when polls first ask respondents what they know about an issue or a candidate before asking further questions. In 1992 surveys about Ross Perot, respondents who first admitted lack of knowledge about Perot were less likely to indicate support for him than in a poll that simply asked for whom they would vote. Similar results occurred in 2000, when voters were asked if they supported Al Gore or George W. Bush; respondents were significantly less likely to report being “undecided” if they were first asked their opinions of each candidate.

A third problem occurs when prior questions alter a respondent’s opinion on subsequent questions. The National Crime Victimization Survey (see Chapter 6) measures more crime victimization if preceded by attitude questions about crime that apparently stimulate memory and willingness to report more personal experiences. Similarly, in a study of perceptions of interracial prejudice, University of Delaware political scientist David C. Wilson found that perception of prejudice depended on whether a respondent was asked about their own group’s bias first or second. When white survey respondents were asked, “How many Whites dislike Blacks?” and then asked, “How many Blacks dislike Whites?” the

328

perception of out-group dislike (Blacks dislike Whites) was higher than when the questions are reversed; on the other hand, when black respondents were asked these two questions, perception of outgroup dislike (Whites dislike Blacks) was higher when they were asked about their own group (“How many Blacks dislike Whites?”) first.

Finally, when a poll asks for a choice of answers, the order of the answers can lead to bias. Researchers have documented a substantial recency bias in which respondents tend to choose the answer listed last. Thus a 1988 Roper poll found a swing of eight percentage points if Vice President George H.W. Bush was listed before, rather than after, his opponent in that year’s presidential contest, Governor Michael Dukakis. To complicate matters, there is also a primacy effect in which respondents choose the response listed first, as during the 2000 presidential election, when Al Gore did significantly better when his name came before George W. Bush in poll responses.

These examples show the tremendous complexity of question-order effects. Because these effects are so unpredictable, an accurate interpretation of surveys requires that the format be varied to see if there are important order effects. Once again, such tests rarely are reported in survey results.

Stem Cell Research: Support or Strongly Support?

Because few opinions fit comfortably in a “yes or no” framework, pollsters usually provide a range of answers ranging from “strongly agree” to “strongly disagree.” The problem arises when researchers attempt to summarize this continuum by grouping different opinions into what pollsters call “cutting points.” For example, support for a position may include just those who “strongly support” or those who “strongly support” and “support.” The problem here is that the change in cutting points may lead to entirely opposite conclusions. In 2006 President Bush vetoed a bill that would have expanded federal funding of stem cell research, a position that some argued was counter to the opinion of the majority of Americans. But whether a majority of Americans supported stem cell research depended on how cut points were defined. A 2005 survey asked:

“On the whole, how much do you favor or oppose medical research that uses stem cells from human embryos—do you strongly favor, somewhat favor, somewhat oppose, or strongly oppose this?”

329

Strongly favor or somewhat favor, 58%

Strongly favor, 27%

Using the cutting point “strongly favor or somewhat favor” would allow supporters of the stem cell bill to claim majority support; focusing on just “strongly favor” would allow opponents to claim that support is far lower. The lesson for researchers is that the choice of cutting points can critically affect the interpretation of results.

Don’t Knows: Ignorance or Honesty?

When confronted by a pollster, no one wants to appear ignorant and thus may attempt to answer questions despite a lack of knowledge. The extent of the problem is revealed in polling experiments that deliberately ask questions to which no one is likely to have an answer. For example, in 1978 the U of M Survey Research Center asked about the Agricultural Trade Act, a bill so little known that the Survey Research Center presumed “virtually no respondents were familiar with its nature or contents.” Even so, more than 30 percent reported an opinion—19 percent in favor, 11 percent opposed—whereas 70 percent were honest enough to admit that they did not know about the bill. Survey takers sometimes worsen the problem by asking questions that no one could reasonably be expected to answer. For example, pollsters asked the public to predict whether the Clintons would stay together during the Monica Lewinsky scandal, or how long the war in Afghanistan would last. A more common example of unanswerable questions occurs in polls taken a long time before elections that ask citizens to make a choice before they know anything about most of the candidates.

One solution is to provide “don’t know or no opinion” as an explicit choice. Typically this option increases the number of individuals professing ignorance or uncertainty by about 20 percent over the number who spontaneously provide this response. For questions with a strong moral content, such as abortion, the “don’t know” answer may be used by well-informed respondents who have a highly ambivalent response to the issue that cannot be pigeonholed as “yes” or “no.”

A second option is to ask respondents if they are familiar with a topic or are interested enough to have an opinion. The use of these “filter” questions raises additional problems of interpretation: Will respondents vote for a candidate even though the filter suggests they do not have

330

enough knowledge about the individual to make an informed decision? Such an effect may have occurred in 1992, when voters knew little about candidate Ross Perot’s positions on issues but nonetheless voted for him in unexpected numbers. Should we count only opinions of informed respondents, unfortunately often a small fraction of the population? Repeatedly, polls find that respondents will give an opinion on issues about which they know little, including such major national security issues as whether Congress should declare war: in September 2001, a CNN/Time poll found that 62 percent of respondents were willing to say “yes, we should go to war,” but 61 percent did not know against whom war should be declared.

Predicting Elections

All of the issues that plague public opinion polls, in general, apply equally to political polls, with the added concern that when polls attempt to predict the outcome of an election, they can potentially have an impact on the election itself. In this section, we review controversies that are particularly relevant to election polls.

Exit Polling

Television and radio news predict results on election day through exit polls in which respondents are surveyed outside voting sites. The individual interviews required in exit polling are so expensive that major television networks pool their resources to pay for the surveys, a dangerous precedent according to some critics because there is only one data source for all news programs. The major statistical problem with exit polling is nonresponse bias: 30 to 50 percent of those approached refuse to participate, sometimes in protest against the ability of the media to predict elections before all citizens have voted. Because refusal rates are higher for some groups, such as older people, exit polls may provide a biased picture of the electorate unless the results are weighted to take into account those who do not participate.

Aside from voters who refuse to respond to polls as they leave the voting booth, there is another group of voters who may not be sampled at all because they are physically absent. Absentee voters, who submit their ballots by mail, traditionally have been a small percentage of voters and not included in exit polls. However, they are a growing share of the voting population, and their exclusion from exit polls actually led to one of the

331

more infamous miscalls in recent election history. In the presidential election of 2000, many news outlets used exit polls prematurely to project a win in Florida for Al Gore over George W. Bush. The state (and the national election) was later determined to go to Bush, after a controversial vote recount and Supreme Court battle. A significant problem with the exit poll projections stemmed from the unusually large number of absentee ballots in that election (almost double what was expected), and the polling service had no way to reliably integrate those votes. In response to that fiasco, the networks reorganized their exit polling, hiring two private firms who now include telephone polls in key states to reach absentee voters. In subsequent elections, the networks have been much more cautious about projecting winners in each state and have adopted procedures that are more transparent.

Box 12.3 Black Candidates/White Voters

Some of the largest election polling errors occur when black candidates face a white opponent. For example, in 1982 popular Los Angeles mayor Tom Bradley, an African American, held a solid lead over his white opponent in all polls leading up to California’s gubernatorial election, which he then lost by roughly 100,000 votes. One theory was that white voters, sensitive to charges of racism, misrepresented their voting intentions by telling survey takers they would vote for the black candidate. This phenomenon is now commonly referred to as “the Bradley effect.” Henry E. Brady and Gary R. Orren argue that the effect is more subtle than that—more likely a result of previously undecided white voters who disproportionately made a last-minute decision to vote for the white candidate. Evidence for such switching also occurred in a 1989 Chicago mayoral election, when undecided black voters supported black independent Timothy Evans at the last moment, making a closer race than had been expected against white Democrat Richard Daley.

Not surprisingly, the Bradley effect was often discussed during Barack Obama’s historic campaign for the U.S. presidency. During the Democratic primary, two researchers at the University of Washington did find evidence of a Bradley effect in several states,

332

but they also found a “reverse Bradley effect” in 12 states, where pre- election support for Obama was significantly lower than his actual support on Election Day 2008. This reverse Bradley effect was particularly strong in southern states, where exit polls showed both blacks and whites voted for Obama more often than previous polls had predicted. The researchers speculated that social pressures were a large driver of these errors. What is clear is that when an election includes sensitive issues such as race, we need to be prepared for the possibility of increased survey error.

Sources: Elections in Henry E. Brady and Gary R. Orren, “Polling Pitfalls: Sources of Error in Public Opinion Surveys,” in Media Polls, pp. 82–85; Obama in Ben Smith, “A Reverse Bradley Effect?” Politico, October 9, 2008, http://www.politico.com/blogs/bensmith/1008/A_reverse_Bradley_Ef- fect.html?showall.

Bandwagon and Underdog Effects

Another complaint about television coverage of the 2000 presidential election was that networks were predicting the outcome in some states before the polls officially closed. Even the most unbiased poll may influence election results if voters jump on the bandwagon, voting for the apparent election winner. However, there is also the possibility that the reverse may be true: voters empathize with the underdog, voting for the candidate who is behind in the polls as a protest against those with political power. Considerable research evidence indicates that both effects occur.

Surveys taken after an election typically count more respondents claiming to have voted for the winner than actually occurred, a potentially serious problem for research based on the National Election Studies and other postelection surveys. The error is especially high for House of Representative elections, for which political scientist Gerald C. Wright finds about 20 percent of those on the losing side say they voted for the winner. Also, 25 to 30 percent of nonvoters claim to have voted, introducing another bias if these respondents also claim to have voted for the winner more often than those who actually went to the polls. Misreporting is slightly less serious for Senate races and quite low for U.S. presidential elections—with the exception of 1964, when an additional 6 percent failed to report their vote for Barry Goldwater in his overwhelming

333

loss to incumbent Lyndon Johnson.

The reasons for misreporting include the eagerness of respondents to be on the winning side and, for those who do not remember much about the election, better name recall for the winner. If a bandwagon effect exists after elections, then likely it exists in pre-election polls as well, increasing support for the candidate who is reported to be most popular. Several studies have verified the existence of bandwagon effects in experiments; for example, one study found that when university students were given polling information that Bill Clinton was leading George Bush, they became more likely to vote for Clinton, even if originally intending to vote for Bush.

The opposite, or underdog, effect—support for the apparent losing candidate—shows up in the common tendency of political races to tighten during the last days of a campaign. In particular, nonincumbents are able to gain ground by picking up the undecided votes. Thus the Mason-Dixon Political/Media Research polling firm maintains that “incumbents who fail to get more than 50 percent in late polls usually lose.” For example, one week before the 1993 New Jersey gubernatorial election, incumbent Jim Florio’s 46 percent support was insufficient, despite a 13 percent lead, because so many voters were undecided. Indeed, he lost to challenger Christine Todd Whitman.

Because the bandwagon and underdog effects may counteract one another, measuring the net result would require an enormous study, tracking voter reaction over time and separating these effects from other swings in voter preference. One study, proposed in 1950 but abandoned because of cost, would cost about $500,000 in today’s buying power. As a result, political scientists believe that both bandwagon and underdog effects occur, but they are uncertain about their total effect.

Horserace Journalism and Polls as News

For some political critics, the problem with media reporting of poll results runs deeper than bandwagon or underdog effects. In this view, polling has become so prevalent that instead of studying the pros and cons of candidates’ positions, the media become fixated on who is ahead, treating elections as “horse races.” One study found that at least half of election news stories are horse-race reporting, taking up as much news time as all election issues combined. Moreover, the reported changes in support for candidates may not be statistically significant, but while some journalists

334

may remember to mention the margin of error, many do not.

The rise in media attention paid to political polls is largely an American phenomenon. Although journalists elsewhere regularly report the results of election polls, such coverage is not nearly as extensive or as uncritical as in the United States. In the interest of so-called neutrality and balanced coverage, American journalists are less likely than their European counterparts to analyze issues and more likely to simply report the numbers.

Some also see problems with the rise of polls themselves as news. Technological advances have expanded the number of news outlets and made it increasingly easy for those outlets to conduct their own surveys, even if the samples are self-selected and not “real polls.” Put that together with the need to fill a 24-hour cable news cycle and we see a huge increase in the reporting of polls that, back in the 1990s, would not have been considered publishable. Some media outlets try to distinguish between pseudo and scientific polls, referring to the former as “votes” or “questions” and adding disclaimers that the results of such surveys are not representative. But as Mark A. Schulman, past president of the American Association for Public Opinion Research, has put it, “Our fear is that in the public’s mind a poll is a poll is a poll.” When a news organization conducts even an informal poll, viewers are unlikely to differentiate between the credibility of the poll and the credibility of the organization itself, and most people do not bother to read the fine print.

Political Bias

A related problem is that an increasing number of news organizations cater to viewers and/or readers with a specific political viewpoint; that bias is then reflected in responses to the self-selected surveys conducted by those organizations. But even with larger scientific polls, we know it is possible for polls to be politically biased. Pollsters are able to promise candidates, “Tell me the results you want, and I can get them.” Political candidates, ill served by inaccurate polls, may be offered two poll results, one for public relations and one to report how things really are. Even independent polls are subject to manipulation, as occurred in 1968 when Richard Nixon vied with Nelson Rockefeller for the Republication presidential nomination. Embarrassed by repeated polls showing Rockefeller as the candidate more likely to beat Democratic nominee Hubert Humphrey, it is alleged that Nixon arranged to be endorsed by former President Eisenhower just before

335

the next Gallup poll would be conducted. Indeed Nixon’s support jumped, and according to some observers, this knocked Rockefeller out of contention for the nomination.

Pollsters often work exclusively for one political perspective, providing yet another source of potential bias. For example, Louis Harris worked extensively for the Kennedy brothers, including campaign surveys for John F. Kennedy’s 1960 presidential nomination campaign. During the 1980 primaries, when Harris published three syndicated columns a week, critics charged that he featured polls that showed Edward (Teddy) Kennedy leading, even if insignificantly, and did not report his polls that showed incumbent President Carter with as much as a 30 percentage point lead. Even more overt manipulation of pre-election polls was reported in Mexico’s 1997 presidential elections in which clients paid for phony polls, sometimes never even conducted, that showed their candidate ahead.

Most pollsters do not proclaim their political bias as brazenly as the Harris poll, but researchers need to be aware that pollsters have political or even economic ties to candidates. They may not use obviously inaccurate survey techniques—as did the Mexican consultants or the pollster who promised “any result you want”—but when pollsters write their own columns, their selection of polls introduces a bias to their reporting. Researchers need to consult the full set of polls that usually are available from the pollsters but are not all featured in the media (see sources in “Where the Numbers Come From” at the beginning of this chapter).

Polling Standards

The fine print at the bottom of media polls often reads something like as follows:

The Gallup Organization interviewed a representative national sample of 606 adults by telephone March 26 and March 27. The margin of error is plus or minus 5 percentage points. Some “Don’t know” responses were omitted.

Box 12.4 What Do Women Want?

Many of the problems described in this chapter are encapsulated in a

336

dispute about female sexuality that followed the publications of Shere Hite’s books, two Hite Reports and Women and Love. Hite found that most women were unfulfilled in their sexuality and had sexual interests far more diverse than generally believed. An ABC/Washington Post survey conducted in 1987 produced entirely contradictory results:

Hite ABC/Washington Post Percentage of women satisfied with their relationships

16 93

Percentage of women married 5 years having sex outside marriage

70 6

Who was right? Critics attacked Hite’s methodology as unscientific because she distributed her surveys on a nonrandom basis to women’s organizations and women’s magazines. Hite’s network was large, including more than 100,000 surveys in one study, but the low 3 percent response rate led to an unrepresentative sample. Finally, contrary to standard polling practice, Hite changed her questions midway through one of her studies and encouraged respondents to answer whatever questions they liked. Whenever Hite equated her book with scientific studies, the critics were unrelenting, calling her results “garbage.”

Unfortunately, this dispute about Hite’s sampling method distracted attention from scholarly praise. In the tradition of qualitative research, Hite uncovered depths of feelings and insights from women that would never be measured in the traditional ABC/Washington Post poll. Unquestionably the ABC/Post poll met standard polling criteria, with a 1,505-person sample selected carefully as a cross-section of the country’s age, education, and race. But did scientific method make the poll more accurate? Hite parodied its format: “Oh, so you call people on the telephone, you don’t know who’s home with them, they don’t know who you are, and then you say, ‘Are you having extramarital sex?’ and you expect them to tell you.” It is remarkable that anyone answered yes under such circumstances. Independent evidence suggests that the ABC/Post poll significantly overstated marital bliss. For example, the Kinsey Institute survey found the rate of extramarital affairs to be double the

337

ABC/Post count, and a National Opinion Research Center Survey, published in 1994, found that 85 percent of married women 18 to 59 years old reported being faithful to their husbands.

The intensity of ABC’s response was out of proportion to the importance of the scientific error in Hite’s survey methodology. ABC and the Washington Post may have won the technical battle (even adjusting for the problems of telephone polling on private matters, their data were closer to the numerical truth about marital infidelity), but if our goal is to understand social behavior, then Hite’s research is more helpful. By limiting the debate to numerical estimates on which Hite probably erred, the ABC/Post survey failed to look at women’s perceptions, which were Hite’s equally important research goals.

Sources: On Hite, see Shere Hite, The Hite Report (New York: Dell, 1977); on Hite/ABC controversy, see Moore, Superpollsters, pp. 6–22. See also Susan Faludi, Backlash: The Undeclared War Against American Women (New York: Crown, 1991), pp. 4–9; on National Opinion Research Center, see “Sex in America: Faithfulness in Marriage Thrives After All,” New York Times, October 7, 1994, p. A1.

The National Council on Public Polls (NCPP) and the American Association for Public Opinion Research (AAPOR) have adopted principles of disclosure, encouraging pollsters to include reference in their polls to the survey’s sponsorship, date of interviewing, method of interviewing, population sampled, sample size, complete question wording, and percentages on which conclusions are based. In 2010 AAPOR revised their Code of Professional Ethics and Practices to include reporting the response rate, data weighting methods, and the source of samples, including the method to recruit participants for opt-in surveys.

The AAPOR’s revisions were prompted in part by a controversy over 2008 pre-election polls, in which several polling firms were asked by AAPOR to provide information about their methodology; at least one firm refused to provide the information and was censured by the organization. Then in 2010 polling analyst Nate Silver developed a model to rate the track record of various polling organizations and found that pollsters who are members of NCPP or AAPOR’s transparency initiative have better records.

338

Unfortunately, even when pollsters provide press releases that comply with AAPOR requirements, media reports usually edit out some (if not most) of this information. Some media outlets such as the New York Times and ABC News have their standards posted on their websites, but television news often provide the scantiest background help, and newspapers often omit important details; for example, sampling error is left out about 84 percent of the time. Researchers need to be prepared to contact the survey source itself for a poll’s complete background information.

Summary

In summary, researchers face a daunting task in using public opinion data. Should we throw up our hands and dismiss poll results? As this book describes, data problems plague all social science statistics, and they are particularly bothersome when first learned by users unaccustomed to data- collection techniques. Just as we accept population counts, GDP, and crime rates despite their limitations, we should do the same with poll and survey data. Kathleen Frankovic, director of CBS News polling, attributes part of the problem to survey researchers, because “even when we know our methods cannot produce precision, we allow those who read and use our results to think they do.” We should join the careful experts who use these statistics to great advantage, albeit with appropriate caution.

Opinion polling is a well-developed science that has put in place methodological standards to reduce the problems summarized in this chapter. Sampling error is the best-understood source of error—one that pollsters can control with reasonable sample sizes as long as nonresponse rates are not too high. Unfortunately, refusal rates appear to be increasing, and, unlike old-fashioned face-to-face interviewers, we often do not even know the basic demographic data about those who decline to participate.

Question wording is a critical feature of opinion polling, one that is susceptible to purposeful or inadvertent manipulation. Poll research provides some guidelines about which words produce which response, but the examples in the text and the added complication of the order effect show that we must continuously watch for subtle complications. Pollsters are well aware of these limitations in their surveys, and national standards require that question wording be available when survey results are released. Nevertheless, it is often necessary for researchers to ask for such details directly from pollsters rather than rely on media reports.

339

Also troublesome is the problem of bias in polling. In addition to outright favoritism toward political candidates, opinion polling selects which questions will be asked and how they will be interpreted. Researchers need to be active social critics, thinking about why particular questions are used in surveys and why other questions are omitted. We need to speculate about what alternative surveys should be conducted. At times it may be possible to conduct our own local surveys; step-by-step procedures are outlined in a several guidebooks. But national surveys are too expensive for most independent researchers, so we remain dependent on the resources of private pollsters and academic and government research centers. Fortunately, their data are largely open to scrutiny. Even when the media report only scant information about these surveys, far more detail is available from the poll takers themselves or the resources listed in each chapter of this text in the “Where the Numbers Come From” box.

Political polling and horserace journalism also present unique issues. Some would argue that horserace journalism has gone too far; perhaps we would be better off following the lead of other countries, banning polls just before election time. But election polling does provide data that can shed light on the political process. Rather than avoid polls, it would be better for journalists to use them more judiciously and with a more critical eye. The National Council on Public Polls provides a primer of the questions that a journalist should ask about poll results. Researchers would be well served to ask many of those same questions (reproduced below) before using polling data for analysis:

Who did the poll?

Who paid for the poll and why was it done?

How many people were interviewed for the survey?

How were those people chosen?

What area (nation, state, or region) or what group (teachers, lawyers, Democratic voters, and so on) were these people chosen from?

Are the results based on the answers of all the people interviewed?

Who should have been interviewed and was not? Or do response rates matter?

When was the poll done?

340

1.

2.

3.

4.

How were the interviews conducted?

What is the sampling error for the poll results?

What other kinds of factors can skew poll results?

What questions were asked?

In what order were the questions asked?

What other polls have been done on this topic? Do they say the same thing? If they are different, why are they different?

What else needs to be included in the report of the poll?

Case Study Questions

A study by psychologist Elizabeth Loftus asked respondents: “Do you get headaches frequently, and if so, how often?” Respondents reported on average 2.2 headaches per week. To a slightly different question—“Do you get headaches occasionally, and if so, how often?”—respondents reported only 0.7 headaches per week. Explain this difference in response.

Pay calls to express political and other views are common on television. After the 1980 Carter-Reagan debate, 727,000 calls at 50 cents per call gave candidate Ronald Reagan a two-to-one margin over President Jimmy Carter. Why might these results be biased? Today, viewers are more likely to be asked to send text messages to register their opinion. Why might those results be biased?

Both Pete Rose and “Shoeless Joe” Jackson were deemed ineligible for baseball’s Hall of Fame because of charges related to gambling on games. When asked whether Pete Rose should be eligible, 64 percent of survey respondents said yes. However, when asked whether Shoeless Joe should be eligible and then asked about Pete Rose, only 52 percent said Rose should be eligible. Discuss the discrepancy between the two results.

One question in the National Race and Crime Survey asked: “Some people want to increase spending for new prisons to lock up violent criminals. Other people would rather spend this money for antipoverty programs to prevent crime. What about you? If you had to choose, would you rather see this money

341

5.

6.

spent on building new prisons, or on antipoverty programs?” In some versions of the survey, the first sentence was altered to, “Some people want to increase spending for new prisons to lock up violent inner-city criminals.” Why might the responses to the two versions of this question differ?

The Kaiser Family Foundation’s Health Tracking Poll asks a number of questions about the federal health care overhaul known as the Affordable Care Act. One question is: “Given what you know about the health reform law, do you have a generally favorable or generally unfavorable opinion of it?” Another question asks: “What would you like to see Congress do when it comes to the health care law?” with answer options of “Expand law,” “Keep as is,” “Repeal and Replace with Republican- sponsored alternative,” and “Repeal and not replace.” Which question more accurately gauges public opinion?

William the Conqueror’s 1086 c.E. “Doomsday Survey” is one of the most valuable sources of data on life in medieval England. Design three survey questions for today’s world that would be useful to social scientists in the year 3086.

References

Data Sources (257)

Gallup data sample in Frank Newport, “Four in 10 Americans Believe in Strict Creationism,” Gallup, December 17, 2010, http://www.gallup.com/poll/145286/fo- ur-americans-believe-strict-creationism.aspx. Media polls in Thomas E. Mann and Gary R. Orren, Media Polls in American Politics (Washington, DC: Brookings Institution Press, 1992). National Election Studies described in Warren E. Miller, American National Election Studies Data Sourcebook (Cambridge, MA: Harvard University Press, 1989); American National Election Studies, “The ANES Mission,” http://www.electionstudies.org/overview/overview.htm. ANES data sample in American National Election Studies, “The ANES Guide to Public Opinion and Electoral Behavior,” table 5B.3, http://electionstudies.org/nesguide/to- ptable/tab5b_3.htm. General Social Survey (GSS) data summarized in Richard G. Niemi, Trends in Public Opinion: A Compendium of Survey Data (Westport, CT: Greenwood Press, 1989). GSS data sample in Tom W. Smith, “Trends in Willingness to Vote for a Black and Woman for President, 1972–2008,” GSS Social Change Report No. 55, August 2009, http://publicdata.norc.org:41000/gss/- documents//TOPL/SC55%20Trends%20in%20Willingness%20to%20Vote%20for- %20a%20Black%20and%20Woman%20for%20President,%201972–2008.pdf.

342

Controversies (259)

Literary Digest in David W. Moore, Superpollsters (New York: Four Walls Eight Windows, 1992), pp. 32–71; Peverill Squire, “Why the 1936 Literary Digest Poll Failed,” Public Opinion Quarterly 52 (1988): 125–33; Don Cahalan, “The Digest Poll Rides Again!” Public Opinion Quarterly 53 (1989): 129–33. On Gallup in 1936, see Moore, Superpollsters, pp. 2–32; Gallup’s own account in George Gallup and Saul Rae, The Pulse of Democracy (New York: Simon and Schuster, 1940), pp. 38–48. Truman/Dewey in Moore, Superpollsters, pp. 68–71.

Sampling (260)

Cell Phones v. Landlines (260): Random dialing in Moore, Superpollsters, pp. 270–73. Summary of major media poll techniques in Herbert B. Asher, Polling and the Public (Washington, DC: CA Press, 1988), pp. 140–41; see also “Design of the Sample” in The Gallup Poll Monthly. Cell phone bias in Paul J. Lavrakas et al., “The State of Surveying Cell Phone Numbers in the United States, 2007 and Beyond,” Public Opinion Quarterly 71, no. 5 (2007): 840–54; Scott Keeter et al., “What’s Missing from National Landline RDD Surveys? The Impact of the Growing Cell-Only Population,” Public Opinion Quarterly 71, no. 5 (2007): 772– 92; Stephen J. Blumberg and Julian V. Luke, “Coverage Bias in Traditional Telephone Surveys of Low-Income and Young Adults,” Public Opinion Quarterly 71, no. 5 (2007): 734–49; Stephen Ansolabehere and Brian F. Schaffner, “Residential Mobility, Family Structure, and the Cell-Only Population,” Public Opinion Quarterly 74, no. 2 (2010): 244–59. See also Carl Bialik, “Pollsters Go Mobile,” The Numbers Guy (blog), December 2, 2011, http://blogs.wsj.com/numb- ersguy/pollsters-go-mobile-1103, and Bialik, “Survey Says: Cellphones Annoy Pollsters,” Wall Street Journal, December 3, 2011, http://online.wsj.com/article/S- B10001424052970204012004577072913322217338.html. Political issues in Michael Mokrzycki, Scott Keeter, and Courtney Kennedy, “Cell-Phone-Only Voters in the 2008 Exit Poll and Implications for Future Noncoverage Bias,” Public Opinion Quarterly 73, no. 5 (2009): 845–65; Nate Silver, “Bypassed Cellphones: Biased Polls?” FiveThirtyEight (blog), October 14, 2010, http://fiveth- irtyeight.blogs.nytimes.com/2010/10/14/bypassed-cellphones-biased-polls/.

Nonresponse (261): Richard Curtin, Stanley Presser, and Eleanor Singer, “Change in Telephone Survey Nonresponse over the Past Quarter Century,” Public Opinion Quarterly 69, no. 1 (2005): 87–99; Eleanor Singer, “Nonresponse Bias in Household Surveys,” Public Opinion Quarterly 70, no. 5 (2006): 637–45; Asher, Polling and the Public, pp. 140–41.

Self-Selected Internet Polls (262): Mick P. Couper and Peter V. Miller, “Web Survey Methods: Introduction,” Public Opinion Quarterly 72, no. 5 (2009): 831– 35; LinChiat Chang and Jon A. Krosnick, “National Surveys via RDD Telephone Interviewing Versus the Internet: Comparing Sample Representativeness and

343

Response Quality,” Public Opinion Quarterly 73, no. 4 (2009): 641–78; Nate Silver, “Before Citing a Poll, Read the Fine Print,” FiveThirtyEight (blog), January 15, 2012, http://fivethirtyeight.blogs.nytimes.com/2012/01/15/before-citing-a-pol- l-read-the-fine-print/.

Survey Design (262)

Question Wording (262): Barry C. Burden, “The Social Roots of the Partisan Gender Gap,” Public Opinion Quarterly 72, no. 1 (2008): 55–75. Jon A. Krosnick, “The Climate Majority,” New York Times, June 8, 2010, http://www.nytimes.com/- 2010/06/09/opinion/09krosnick.html; Dan Vergano, “Some Scientists Misread Poll Data on Global Warming Controversy,” USA Today, March 9, 2010, http://www.u- satoday.com/tech/science/columnist/vergano/2010-03-05-global-warming-doubt_- N.htm. Abortion in Asher, Polling and the Public, p. 124; Lydia Saad, “More Americans ‘Pro-Life’ Than ‘Pro-Choice’ for First Time,” Gallup, May 15, 2009, h- ttp://www.gallup.com/poll/118399/more-americans-pro-life-than-pro-choice-first-- time.aspx. Michael R. Kagay, “Variability Without Fault: Why Even Well- Designed Polls Can Disagree,” in Media Polls, ed. Mann and Orren, p. 118.

Question Order (266): Connecticut race discussed in Irving Crespi and Dwight Morris, “Question Order Effect and the Measurement of Candidate Preference in the 1982 Connecticut Elections,” Public Opinion Quarterly 48 (1984): 578–91; Perot in Henry E. Brady and Gary R. Orren, “Polling Pitfalls: Sources of Error in Public Opinion Surveys,” in Media Polls, ed. Mann and Orren, p. 77; Gore/Bush in Monika L. McDermott and Kathleen A. Frankovic, “The Polls—Review: Horserace Polling and Survey Method Effects: An Analysis of the 2000 Campaign,” Public Opinion Quarterly 67, no. 2 (2003): 244–64. National Crime Survey in Howard Schuman and Stanley Presser, Questions and Answers in Attitude Surveys (New York: Academic Press, 1981), p. 45; Wilson study in David C. Wilson, “Perceptions About the Amount of Interracial Prejudice Depend on Racial Group Membership and Question Order,” Public Opinion Quarterly 74, no. 2 (2010): 344–56; Bush-Dukakis in Irving Crespi, Public Opinion, Polls, and Democracy (Boulder, CO: Westview Press, 1989), p. 69.

Stem Cell Research: Support or Strongly Support? (267): Lydia Saad, “Stem Cell Veto Contrary to Public Opinion,” Gallup, July 20, 2006, http://www.gallu- p.com/poll/23827/stem-cell-veto-contrary-public-opinion.aspx.

Don’t Knows: Ignorance or Honesty? (268): Agricultural Trade Act in Schuman and Presser, Questions, pp. 148–50; Don’t know and filters in Schuman and Presser, Questions and Answers, pp. 114–46. Lori Robertson, “Poll Crazy,” American Journalism Review, January/February 2003, http://www.ajr.org/article.a- sp?id=2748.

Predicting Elections (269)

344

Exit Polling (269): Kathleen A. Frankovic, “Technology and the Changing Landscape of Media Polls,” in Media Polls, ed. Mann and Orren, pp. 34–40. Refusal rate in ibid., pp. 37–39. Bush-Gore in Howard Kurtz, “Networks Vow Caution in Calling Election,” Washington Post, October 12, 2004, p. A07, http://w- ww.washingtonpost.com/wp-dyn/articles/A25309–2004öct11.html.

Bandwagon and Underdog Effects (270): Misreporting in Gerald C. Wright, “Errors in Measuring Vote Choice in the National Election Studies, 1952–1988,” American Journal of Political Science 37 (February 1993): 291–316; and Wright, “Misreports of Vote Choice in the 1988 NES Senate Election Study,” Legislative Studies Quarterly 55 (November 1990): 543–62. Clinton-Bush in Vicki G. Morwitz and Carol Pluzinski, “Do Polls Reflect Opinions or Do Opinions Reflect Polls? The Impact of Political Polling on Voters’ Expectations, Preferences, and Behavior,” Journal of Consumer Research 23, no. 1 (1996): 53–67. Mason-Dixon in “When Voters Tell Polls They’re Undecided,” letter from J. Bradford Coker and Robert L. Joffee, New York Times, November 11, 1993, p. A18. Net effect in Larry J. Sabato, The Rise of Political Consultants (New York: Basic Books, 1981), p. 107. The $500,000 study in Michael W. Traugott, “The Impact of Media Polls on the Public,” in Media Polls, ed. Mann and Orren, p. 138.

Horserace Journalism and Polls as News (272): Thomas E. Patterson, “Of Polls, Mountains: U.S. Journalists and Their Use of Election Surveys,” Public Opinion Quarterly 69, no. 5 (2005): 716–24; McDermott and Frankovic, “Review: Horserace Polling,” 244–64; Tom Rosenstiel, “Political Polling and the New Media Culture: A Case of More Being Less,” Public Opinion Quarterly 69, no. 5: 698–715; Mark A. Schulman in Lori Robertson, “Poll Crazy,” http://www.ajr.org/- article.asp?id=2748.

Political Bias (272): On Nixon, see Moore, Superpollsters, p. 95; on Harris and Kennedys, see ibid., pp. 110–21; Mexico in “The Pollsters’ Greatest Enemy: Themselves,” Washington Post Weekly Edition, February 23, 1998, p. 35.

Polling Standards (273)

Gallup explanation in Asher, Polling and the Public, p. 63. Standards in ibid., pp. 78–81; David Hill, “AAPOR updates poll standards,” The Hill, May 18, 2010, htt- p://thehill.com/opinion/columnists/david-hill/98489-aapor-updates-poll-standards. The 2008 controversy in Clint Hendler, “How Do You Know What a Poll Number Is Worth?” Columbia Journalism Review, July 21, 2010, http://www.cjr.org/cam-- paign_desk/how_do_you_know_what_a_poll_nu.php; Nate Silver, “Pollster Ratings v4.0: Results,” FiveThirtyEight (blog), June 6, 2010, http://www.fivethirt- yeight.com/2010/06/pollster-ratings-v40-results.html; Mark Blumenthal, “Does Transparency Increase Accuracy?” National Journal, June 21, 2010, http://conven- tions.national-journal.com/njonline/mysterypollster.php.

Summary (276)

345

Frankovic in “The Pollsters’ Greatest Enemy,” p. 35. For summary of research on polling, see Graham R. Walden, Public Opinion Polls and Survey Research: A Selective Annotated Bibliography of U.S. Guides and Studies from the 1980s (New York: Garland, 1990). How to do your own survey in Celinda Lake, Public Opinion Polling (Washington, DC: Island Park Press, 1987); and Thomas I. Miller, Citizen Surveys (Washington, DC: International City Managers Association, 1991). Survey kits, including “How to Ask Survey Questions,” “How to Conduct Self- Administered and Mail Surveys,” and “How to Sample in Surveys,” are available from Sage Publications. National Council on Public Polls questions in Sheldon R. Gawiser and G. Evans Witt, “20 Questions a Journalist Should Ask About Poll Results,” National Council on Public Polls, http://www.ncpp.org/?q=node/4.

Case Study Questions (277)

1. Herbert H. Clark, “Asking Questions and Influencing Answers,” in Questions About Questions: Inquiries into the Cognitive Bases of Surveys, ed. Judith M. Tanur (New York: Russell Sage Foundation, 1992), p. 21.

2. David W. Moore, Superpollsters (New York: Four Walls Eight Windows, 1992), pp. 287–89; Kathleen A. Frankovic, “Technology and the Changing Landscape of Media Polls,” in ibid., pp. 52–53.

3. David W Moore, “Measuring New Types of Question-Order Effects, Additive and Subtractive,” Public Opinion Quarterly 66 (2002): 80–91.

4. Jon Hurwitz and Mark Peffley, “Playing the Race Card in the Post-Willie Horton Era: The Impact of Racialized Code Words on Support for Punitive Crime Policy,” Public Opinion Quarterly 69, no. 1 (2005): 99–112.

5. “Kaiser Health Tracking Poll,” Kaiser Family Foundation, September 2011, http://www.kff.org/kaiserpolls/8230.cfm.

6. The Doomsday Book Online, http://www.domesdaybook.co.uk.

346

13

Conclusions □□□□

Students of statistics soon learn that there is a dazzling array of mathematical techniques for analyzing data, testing hypotheses, and estimating the probability of error. Nevertheless, the controversies reviewed in this book suggest that students and practitioners alike need to look more closely at the limitations of the data to which the sophisticated techniques are applied. Repeatedly, the origin of policy disputes can be traced to questions about the underlying data: Why are only certain data collected? Why are the data organized into particular categories? Why are some data reported as fact to the exclusion of other equally reputable data? And why do the data so often generate apparently conflicting statistics? In order to answer these questions, it is convenient to divide the debates discussed in the previous chapters into categories based on the core source of controversy in the underlying data.

Poor/Missing Data

Novice and expert researchers often are surprised at the absence of data needed to answer important policy questions. For example, researchers face a near dead end in trying to study the extent of white-collar crime. For the study of individual wealth, there are some sources of data; however, as illustrated by the controversy about the interpretation of Federal Reserve Board survey data, these statistics are quite incomplete. Basic data on divorce, abortions, and HIV/AIDS are unavailable for some states, causing unknown error in national statistics and making it difficult to analyze these important matters for the data-poor states. Education statistics are especially noteworthy for their relative inadequacies in coverage, timeliness, and standardization. Despite increased attention to education in the political arena, financial support for good data has not been forthcoming; instead, funding was reduced drastically during the 1980s

347

and has been only partially replaced.

In some cases, data is missing in a way that can lead to a consistent bias in one direction. For example, in the U.S. Census and in political polling, there is concern that certain demographic groups are less likely to participate in the surveys. Both omissions are widely recognized, and there are attempts to decrease the number of individuals who are missed or to adjust the statistics to take into account the responses expected from those not interviewed. Nonetheless, Chapters 2 and 12 reviewed concerns that such efforts potentially could still leave biased statistics. There is not much an individual researcher can do except to be aware when the errors are large enough to affect conclusions.

Several of these cases involve data about those with power; it does not require a conspiracy theory to argue that in the cases of measuring wealth or corporate finances, these individuals and groups are unwilling to participate in data-collection efforts in part out of a desire to protect their own status. From their perspective, little would be gained by sharing data on their position; on the contrary, social statistics might be used to argue that limits should be placed on their status. overall, the poor and the powerless are subject to considerably more data-collection efforts than the well-to-do and powerful. There are numerous studies of poverty, government income-maintenance programs, the homeless, and the unemployed, especially in comparison to the handful of studies of wealth and high incomes. The ability of the powerful to avoid statistical scrutiny raises the question of what interest the poor or middle-income groups have in complying with data collection. The controversies described in this book suggest that the numbers are beneficial for these groups. An accurate measure of the poverty rate can be used to argue for more poverty programs; a complete picture of housing inadequacy (not just homelessness) can be used to argue for new housing programs; a complete count of unemployment can be used to argue in favor of new jobs programs. of course, alternative statistics have been used to argue against each of these programs. But, on balance, better numbers weigh in favor of those without other resources to present their case.

Improved Data

During the last two decades there have been major improvements in some areas of U.S. social statistics. Most notably, in the study of crime, the new National Incident-Based Reporting System (NIBRS) greatly expands data

348

for research purposes, covering more crimes with additional information on perpetrators and victims although it is not yet available for all jurisdictions. Similarly, there are now improved data sets for HIV/AIDS that track individuals (confidentially), although these reports also do not cover the entire nation. occupational Health and Safety data now include some public-sector workers previously omitted. The American Time Use Survey (ATUS), first collected in 2003, provides more accurate data about how Americans spend their time than previous studies that relied on respondent recall of past events. And polling methods will be more transparent to researchers because of the 2010 revised Code of Professional Ethics and Practices requiring pollsters to better report their methodology.

It is interesting to note that improved data collection may introduce new problems for researchers if current numbers cannot be accurately compared with earlier series. Sometimes the problem is inherent in the data. For example, price indices will measure different goods and services as some items are no longer purchased or are changed in quality. The issue is well recognized by government economists and partially accounted for in alternative series such as the chain-weighted index.

In other cases, new data series are introduced to improve on past deficiencies. Government statisticians sometimes offer the revised and original series side by side, or provide adjustments to make them comparable. For example, when new survey questions were introduced to measure unemployment in the United States, the Bureau of Labor Statistics reported differences in the unemployment rate with the old and new questions. Similarly, new poverty measures implemented by the U.S. Census Bureau prompted side-by-side lists of the poverty rate. New boundaries for geographic areas require careful attention for which government statistical agencies provide considerable support. In some cases, new categories introduce major challenges for researchers when the data are discontinuous and cannot be synchronized backward or forward. For instance, changes in race and ethnicity questions in the decennial Census mean that studies by race or ethnicity over time may encompass groups with slightly different definitions—in addition to differences caused by respondents’ changing self-identification. In summary, researchers need to be alert to such data discontinuities and may need to consult government officials for appropriate ways to combine data that was collected in different formats.

349

Conflicting Definitions

A common frustration in news reports is apparently even-handed presentation of competing statistics. When experts on each side offer numbers in support of opposing sides in a debate, it may be tempting to ignore the evidence altogether. The controversies in the previous chapters suggest a more productive approach. Investigation can reveal why the experts disagree and thus enable us to evaluate what otherwise appears to be contradictory evidence.

In a few instances, one side is simply incorrect, or at least using data that has been updated or otherwise proved to be no longer appropriate. For example, estimates for human trafficking, the largest problems in schools, and the size of U.S. military spending (omitting the Iraq and Afghanistan wars) were proved wrong.

More commonly, competing numbers are both accurate but involve prior decisions about how a concept should be defined and measured. In these cases, data users need to evaluate these underlying judgments, which are sometimes (though not always) political or ethical in nature. For example, should business size be measured by sales, assets, market value, or number of employees? Should student achievement be measured using state standardized tests or national assessments? Is per capita income, earnings, or family income the best way to measure economic well-being? In each of these examples, a case can be made for using each of the potential variables, but the results may change dramatically depending on which is selected. Researchers should clearly state their reasons for using one definition over the others and, ideally, discuss the implications of using alternative definitions.

Comparing Apples and Oranges

Even when researchers are in general agreement about which variable to use, many social statistics are available in more than one variation such as the unemployment rate, infant mortality rate, voter participation rate, and inflation rate. The choice between them is critical for the conclusions reached but first requires decisions about what matters. Are part-time workers who desire full-time work actually unemployed (and if so, should they be counted in the same way as those with no work at all?). Should infant mortality count births of fetuses that are less than full-term? Is it more important that more of those individuals eligible to vote are voting, or that fewer citizens are eligible to vote? Should the inflation rate be

350

adjusted when consumers switch to lower priced goods and services? These important assumptions are rarely described in media accounts of dueling statistics, but they are necessary if readers are to understand why the experts use different numbers.

Which parts make up the whole? For some controversies, the problem lies in classifying the underlying data. In studying crime, for example, researchers must define which crimes will be counted as crime before there can be any consideration of the crime rate, the causes of crime, or the effectiveness of responses to crime. Most important, the manner in which white-collar crime is treated—or ignored—in official statistics causes entirely different analyses of who commits crimes and the efficacy of current allocation of resources to deal with crime. Specifically, if white- collar crime is counted, then much more crime is associated with high- income and high-status groups, and police resources aimed at white-collar crime are lacking compared with the amount spent to deal with traditionally defined crimes. Crime rates are also affected by the definition of rape, which changed in recent government statistics so that it is no longer limited to specific sexual acts by men against women. The revision came about in part because of wider social acceptance that violence against women could take other forms and that men, especially in prison, could suffer rapes.

For some statistical controversies, disagreements about how to measure variables are based on conceptual problems in economic theory. The clearest examples came from GDP statistics, for which there are debates about how to account for pollution, resource depletion, the underground economy, and nonmarket production. Welfare spending measurement required assumptions about what constitutes welfare: is it aid to the poor, as in everyday U.S. usage, or is it public support available to anyone, as in many European countries? For military spending, choices need to be made about whether to include research with potential military applications, expenditures on veterans’ benefits, and the cost of debt incurred to pay for past wars. The questionable accuracy of economic data may be humbling to economists, but it is often suggested within the profession as a necessary corrective to the advanced state of economic mathematical techniques. Without certainty about the underlying data, the most sophisticated economic model cannot be put to practical use.

Percentage of what? Many statistics require a choice for the denominator. Are taxes fair when they are equal as a percentage of the

351

payers’ income, or as a percentage of all taxes paid? Should homeownership be measured as percentage of households or the population? Is the high school graduation rate a percentage of those who begin ninth grade or those still enrolled in the same school four years later? Is voter participation a ratio of those eligible to vote or a ratio of the adult population affected by the vote? Are profits a percentage of a firm’s revenue, assets, or market equity? What is the relevant market in determining a firm’s dominance? In each case the simple decision about which value goes in the denominator underlies the statistical controversy.

Fluid categories. For some controversies, apparently conflicting statistics occur because of the different ways in which the underlying data were organized into categories. This critical step of conceptualization is often overlooked, even though it is fundamental to all subsequent use of data. Racial classification is an obvious instance in which there has been little recognition in research projects of the social construction of the categories. In retrospect, we can easily see the arbitrary and racist character of the 1790 U.S. Census three-fifths apportionment allocation for each black slave and the subsequent pseudo-scientific classification based on “quadroons” and “octoroons.” But the current use of self-classification still reflects the older racist standard in the sense that partial black background usually means the individual is considered black. As a rebuttal to scientific racism, it is important to remember the social, not genetic, basis of the black and white racial categories.

The classification of Hispanics and Asians further demonstrates the arbitrary choice of which groups will be counted as separate racial identities and the problems in defining the boundaries between these groups. As in the black-white dichotomy, the classification began from a white perspective in which all Hispanics were grouped together and all Asians were grouped together, even though many individuals in the groups did not recognize such classification, typically preferring a country-of- origin label. Because of successful political efforts by Hispanic and Asian groups, these designations have changed with almost every census, which has been an inconvenience for research covering different time periods but is nonetheless an important reminder of the social origins of racial and ethnic designations.

Comparing Apples and Orangutans

Many data controversies arise when analysts appear to be answering the

352

same empirical question but, in reality, are making very different comparisons by using different statistical measures, subgroups, or time periods. Because similar language is used by both sides, situations that appear to conflict may actually be perfectly consistent but yield different results because of differences in selecting what is compared to what.

Absolute versus relative. An important distinction analysts must make is the choice between absolute and relative statistics. In order to measure the extent of poverty, researchers must choose between various poverty lines: Some of these measurements are constant, or nearly so, in terms of the standard of living they represent, while others measure poverty as relative to current living standards, even if those standards have increased in terms of buying power. The debate about the trend in infant mortality also hinged on the distinction between absolute and relative change. Compared to high levels of the past, U.S. infant mortality has declined, an absolute improvement. But on a relative basis compared to other countries, U.S. infant mortality is a singular measure of health care failure. Similarly, the trend in male college enrollment depends on whether we look at the number of men enrolled, which has been rising, or the enrollment rate relative to women, by which standard male college attendance is falling.

Means versus subgroups. In some cases controversies arise because certain analysts focus on the average and others focus on subgroups or different parts of the distribution. In the debate about residential segregation, those who focused on the average level of segregation were willing to claim the “end” of segregation had arrived, while others emphasized that many African Americans still live in neighborhoods that are almost entirely black. In discussions of health outcomes, statistics for the mean or median often mask large variations in the outcomes across the population, such as the report that 1 in 11 women will develop breast cancer, versus much lower probabilities when looking at specific age groups.

Time period. When comparing any statistic over time, decisions must be made about when to begin and end the analysis. Short-run statistics often differ from long-run statistics; for instance, the rate of growth for small businesses or the rate of failure for new businesses depends critically on whether one measures those rates over one year or multiple years. The choice of base year for comparison can also have a huge impact on how a variable appears to trend over time. For example, the government can only

353

claim to be winning “the war on drugs” by comparing recent drug use to the exceptionally high levels of use recorded in the late 1970s. Choice of time period is also pivotal for discussing trends in labor productivity; one could show productivity to be either increasing or decreasing over the last decade by highlighting different years.

The Mathematics of Social Science Statistics

For some statistics, the choices that analysts must make depend on mathematical distinctions such as how to aggregate multiple variables into an index, whether and how to adjust for confounding variables, and whether to use the mean or the median of a sample. These mathematical complexities still involve the problem of conceptualization, as they require the researcher to make choices about theoretical categories. Because of the complexities involved in learning the mathematical techniques, it is easy to overlook underlying assumptions about how to organize the data; still, these assumptions are critical for understanding precisely what is being measured and why there often are conflicting statistics for the same social issue.

Index numbers. Many statistical measures aggregate a series of data, usually counting some of those data more than others—a technique called an index number by statisticians. For example, the inflation rate must take into account the effect of many products with changing prices. However, the way in which these numbers are added and weighted can be quite controversial, as in the debate about the size and growth of the Chinese economy. In this case the relative weights for different areas of the country and hard-to-measure services caused quite different estimates for how fast China was growing and when its economy would surpass the U.S. economy in total size.

The most extreme index number problem occurs in the Places Rated example, in which hundreds of cities could be ranked number one—or dead last—depending on the weights chosen for each factor. Estimates for the measurement of currency rates in different countries, as well as the measurement of productivity for an economy with a changing product mix, also involved similar indexing problems for which there was no single correct solution.

Adjustments. When comparing social statistics for two geographic areas, it may be necessary to adjust the data to take into account different

354

attributes that the researcher wants to exclude from consideration. Thus in comparing death rates, the data are usually age-adjusted because an elderly population will have a higher death rate, obscuring environmental or other health risks that the researcher wants to understand. Similarly, and often more controversially, some research projects comparing crime rates adjust for poverty and other demographic characteristics in order to investigate, for example, the impact of different policing.

Means and medians. The issue of means and medians was critical for several statistical disputes discussed in this book, such as longevity for cancer patients, the distribution of income, or the number of crimes per prison inmate. These dissimilar social issues share the common problem of an asymmetric distribution in which there is a large cluster in the low end (short survival, low incomes, few crimes), and a few individuals at the high end (long survival, high incomes, many crimes). The resulting difference between the mean and median requires researchers to pay careful attention to which measure is appropriate for a particular project.

What is the baseline? The statistical problem of whether to define categories at the start or the end of a time period affected several disputes including income mobility, small business job creation, drug use, voting rates, and labor productivity. In all cases, different conclusions were reached depending on whether individuals (or companies) were defined as rich (or large) at the beginning or at the end of the time period under analysis. Standard statistical practice requires that the outcome be reported using a variety of techniques to prove the robustness of the results. The failure of the initial investigators to do so led to suspicion about their conclusions.

Undue accuracy. The abundance and frequency of publicly provided statistics can cause undue response based on small numerical changes that do not in fact have meaningful social impact. In a particularly illustrative case, rounding in the reported inflation rate—smoothing out changes that could occur through random measurement error—caused investors to change their behavior unnecessarily. Similarly, small changes in the ranking of countries based on test scores and monthly movement in the unemployment rate are reported with more unwarranted importance. In the case of human trafficking, the CIA data included “confidence intervals,” implying that the reported numbers were far more accurate than actually proved to be the case.

355

The Reporting of Controversies

As the discussion thus far suggests, controversies over social science statistics most often arise because different researchers have made different choices in the course of their analysis. Careful researchers will be aware of and upfront about those choices and, ideally, test the robustness of their results under different choices. Unfortunately, the full story of what researchers do is not always what gets reported in the media. Instead, popular discussions of social science statistics often exacerbate rather than clarify the controversies.

Complicated cases oversimplified. Many of the controversies recounted in this book involved complex data issues. Notably, case studies of racial discrimination in bank mortgage lending, the effect of the minimum wage on jobs, small business job creation, and cancer incidence rates all prompted a sophisticated debate between researchers. However, coverage in the popular press often misrepresented the controversies, merely highlighting the conflicting numbers rather than explaining the underlying differences in the data upon which the conflicting numbers rested. Thus, for example, even though there was disagreement among medical researchers about the trend in cancer incidence, all experts know that these rates depend greatly on the types of patients tested.

Similarly, media coverage of statistical debates has often contributed to oversimplified analysis of policy issues. In the study of homelessness, the number of hungry Americans, the extent of hate crimes, and data on sexual orientation, there was an attempt to reduce a social problem to a single number. Unfortunately, there was more public debate about a rate that could not be measured precisely than about the more fundamental social concern. Thus, there was a bitter dispute about the number of homeless people without much understanding that homelessness was a transitory phase for a much larger number of individuals facing the problem of housing displacement. Similarly, the media reported the debate about the number of hungry Americans rather than the impact of food insecurity on families’ mental and physical health. All these cases require that we recognize the need for agreed-upon single numbers in order to track changes over time, but policymaking also requires more complex analysis recognizing the limits of oversimplified single numbers. The lesson for researchers is to be skeptical about apparently simple resolution to controversies about which reputable experts disagree. Although important policy questions may be at stake, it may be the case that social scientists

356

are unable to offer unambiguous answers.

Headline makers. A few of the controversies involve egregious examples of misused statistics, the extreme of Disraeli’s “lies, damn lies, and statistics.” In this category were the number of missing children, the extent of human trafficking, the shortage of U.S. engineers, and the possibility of lower earnings at a higher marginal tax rate. In each case the controversies were started by eye-catching popular articles that carried alarming messages, although not ones supported by careful examination of the actual data.

Moreover, each example exploited fear: our children will be abducted, the United States’ technological edge will be lost, and taxes will undermine our earnings. Such reinforcement of stereotypical views is a tempting trap for newspapers and magazine writers who are looking for provocative headlines. Because articles that told the opposite side were published in academic journals, but not in the popular media, most readers learned only the original misleading statistics.

Truth by repetition. Another group of statistical controversies were not actually incorrect, only misleading. The examples are numerous and include the personal savings rate, the number of homeless people, the cost of immigrants, the size of China’s economy, average family income, international test scores, productivity rates, small business job creation, and the cancer survival rate. Also, there were debates in which short-term changes were mistaken for long-term trends. Such examples included quarterly estimates of GDP, monthly trade statistics, annual SAT scores, yearly crime rates, monthly unemployment rates, and yearly educational attainment statistics.

These statistics were cited in popular and academic articles even though experts warned about their misleading characteristics. Readers may consult discussion in previous chapters for the reasons why these statistics are wrong-headed. The question remains: Why were the discredited statistics still used? The answer appears to be a situation similar to the problem with sensationalist statistics: The numbers tell a good story. It was far easier to lament the decline in personal savings or the poor performance of U.S. school children on international tests than to investigate the more complicated but truthful interpretation of these statistics. In an ironic twist, misleading statistics such as the Dow Jones average and Fortune 500 list have become important simply because they are perceived as important.

357

Thus, the stock market responds to the Dow Jones, even though it is a poor measure of the overall stock market. And corporations take note of their status in the Fortune 500, even though accountants know it can be a misleading measure of corporate success.

Opinion polling. Polling sits at a unique intersection of social science analysis and media reporting. There is ample evidence that polls affect public opinion through the bandwagon and underdog effects. Equally troubling is the way in which public opinion polling sometimes undercuts thoughtful debate by asking for overly simplified answers from an underinformed public. For example, after the Persian Gulf War, some media polls focused more on what the public thought about the events than what actually had occurred. Researchers need to identify when surveys create opinion, rather than measure it. Finally, in election contests, coverage of poll results may replace substantive election campaigning so that “who is ahead” becomes more important than “who stands for what.” Researchers may want to consider whether their projects foster “horserace journalism” instead of substantive analysis of the candidates and issues.

Biased Analysis: Slanting the Numbers

The popular media cannot bear the entire blame for misreporting and sensationalizing these data controversies. In the United States, social scientists are fortunate to be served by public servant statisticians dedicated to getting the numbers right and taking on public questions or criticisms about their conclusions. But it is also not unusual for the data in official reports by data-collecting agencies to be slanted for political gain by government officials in public announcements. For example, drug use data have been used to promote a sometimes unwarranted victory in the war on drugs. By selectively choosing the years studied, the type of drug used, or age group of the drug user, government officials have trumpeted declining drug abuse even though the data show relatively stable use over time.

In the field of cancer research, experts understand the limitation of survival-rate statistics. Yet there is evidence that those who benefited from existing research budgets were willing to use oversimplified survival-rate statistics to lobby for additional research funds on ways to extend cancer survival at the expense of funding research on cancer prevention. In criminology, it is widely recognized that the predominant measures of crime overemphasize some crimes while overlooking others. But official

358

reports on crime, usually put out by the same agencies that operate the criminal justice system, typically ignore these caveats that might question their funding priorities. Similarly, key economic variables such as GDP and unemployment are reported by government officials in a manner that supports the success of their economic policies, even though professional economists are well aware of the limitations of these statistics.

It is important to distinguish between official government reports that use data responsibly and those with a political slant. One obvious difference is the author: U.S. government statisticians not only tend to offer balanced accounts but also details about the manner in which the data has been collected and warnings about its limitations. Thus, the FBI’s Uniform Crime Reports has a “pop-up” caution about misleading perceptions. The Bureau of Labor Statistics warns that the Consumer Price Index is “called a cost-of-living index, but it differs in important ways from a complete cost-of-living measure.” A careful reading of websites or even a call to the government agency can clarify the appropriate use of statistical reports.

What’s a Researcher to Do?

The major problems with the social statistics listed above all point to one remedy: data literacy. By better understanding the data used to create social statistics, we will be better equipped to understand complex social issues.

Ironically, the starting point for a critical view of the data may be an appreciation of the extensive data available to us. Many readers likely will share the amazement we felt in preparing this book at the sheer volume of U.S. social statistics. In addition to the well-known U.S. Census and Current Population Survey, there are large-scale U.S. surveys of housing, education, crime, crime victims, health, small business, large business, and a variety of demographic characteristics. Although some areas, such as education, suffer from less-complete data than, say, health, nonetheless researchers in every field are confronted with ever-growing quantities of data, many of which are published by the government without further expert analysis. For social science students, there is little prospect that they will run out of numbers to examine; the challenge is to use them correctly.

A second step toward data literacy is an appreciation of why we have so many data. Without political pressure, few of the vast U.S. data resources would have been collected. For example, the U.S. Census, which we now

359

take for granted, was not undertaken to provide a data set for researchers. Instead, the men who drafted the U.S. Constitution realized that the new democratic features of the government required statistics to determine the proportional representation from each state. Several methods were considered, including property values, before the concept of representation based on population was adopted (that is, except for slaves, who were counted as three-fifths of a person). Thus, although women, children, alien residents, and slaves could not vote, representational apportionment still required a complete population count.

Political pressure during the Progressive Era on the dangers of U.S. workplaces prompted the first U.S. Bureau of Labor Statistics investigation and publication of workplace health and safety data. In 1930 political leaders wanted statistical evidence of the severity of what would come to be known as the Great Depression and campaigned for better employment statistics in the Census. Continued political lobbying led to the post-World War II Current Population Survey and its emphasis on employment statistics. Similarly, political pressure led to a major overhaul in the Department of Labor’s occupational Safety and Health statistics in the 1990s. In education, the Department of Education stepped up efforts to improve and standardize education statistics after the federal No Child Left Behind Act of 2001 elevated the role of testing and data for tracking student achievement. In business statistics as well, it took pressure from those concerned about the power of large corporations to generate data used for antitrust enforcement and general research on the structure of markets in the U.S. economy.

These examples underscore the social nature of social science statistics. A society chooses what to measure—or better stated, groups within society struggle about what will be measured. on the one hand, the decisions to count all residents in the census, to document workplace hazards, to survey unemployment in great detail, and to measure how corporations dominate certain industries are all evidence that groups traditionally without power can use statistics as a resource to their advantage. on the other hand, the cutbacks in statistical efforts described earlier show how the numbers can be taken away. overall, good statistics bring us closer to the truth, even when that truth undermines the authority of those in power.

A third step toward data literacy is an appreciation of the people behind the numbers. Among those who have traditionally fought for more and better social statistics are the men and women who collect and analyze the numbers for federal, state, and local governments. one might be tempted to

360

respond that these individuals benefit personally from increased statistical collection; that is, more numbers means more jobs. But anyone who has consulted government statisticians is unlikely to take such a cynical view. The jobs are not well paid, the work is often tedious, and the rewards in terms of recognition are slim—unless a statistician makes an error.

Most of all, many of these individuals are eager to talk about the numbers to which they have dedicated their work lives. Certainly no one is better prepared to discuss problems in the data, and quite often it is government statisticians who warn outside users about the limits of the statistic s for research and policy purposes. In other words, for the most part, government statisticians are not blind to errors in the numbers they carefully generate but tend to be advocates of careful scholarship and the appropriate use of statistics. Public opinion polling data are maintained by several university-affiliated groups. They, too, provide open access to the data, as well as freely give advice about its use.

Researchers can learn from these vital sources of knowledge. Many helpful reference works are published by statistical agencies, and it is often possible to consult directly with the government and university officials responsible for a particular data series. The Internet gives unprecedented access to data by nearly any student or researcher. Government and some private websites are described at the beginning of each chapter through which it is possible to download not only raw data, but sometimes the tools with which to analyze it. In addition, many scholars provide copies of their original data via the Internet, often a requirement for public grants or publication in professional journals. For example, the data on both sides in the minimum wage jobs debate is available for public use, limited only by coded names so that it is impossible to identify the ownership of the individual fast food establishments.

Data literacy is critical for playing the Data Game. By understanding what data are available, how they came to be collected, and who is responsible for their dissemination, we can begin to use social statistics to understand the society in which we live. Statistics alone will not provide a road map to a better world; they can only set a framework for our powers of analysis. But the more data literate we are, the more power we have to create a world of our own choosing.

361

Index □□□□

Italic page references indicate figures, tables, and boxes.

AAPOR, 266, 275, 285

ABC polling, 258, 275

Abortion data, 63, 265–266

Absolute statistics, 288

Accuracy, undue, 290–291

Acharya, Viral, 223

ACS, 8–9, 13, 38, 168, 193

Adelman, M.A., 215

Adjustments to data, 119–120, 145, 290

Afghan war cost, 235–236

African Americans. See Blacks

Aggregate statistics, 214

Agricultural Trade Act, 268

Ahlburg, Dennis A., 15

AHS, 37, 38

Alan Guttmacher Institute, 63

Allegretto, Sylvia, 103

Allen, Walter R., 10–11

American Association for Public Opinion Research (AAPOR), 266, 275, 285

American Community Survey (ACS), 8–9, 13, 38, 168, 193

American Housing Survey (AHS), 37, 38

American National Election Studies (ANES), 234, 258

American Time Use Survey (ATUS), 192, 203, 284

362

American Tobacco Company, 217–218

Anderson, Margo, 13

Anderson, Marian, 182

ANES, 234, 258

“Angry About Inequality? Don’t Blame the Rich” (Wilson), 175

Annual reports, 212–213

Annual Survey of Manufacturers (ASM), 141

Antitrust regulations, 218, 220

Apple (technology firm), 158–159

Arias, Elizabeth, 65

Asians, 20–21, 43

ASM, 141

Assets data

corporate, 215–216

government, 239

Association of Multiethnic Americans, 18

ATUS, 203, 284

Bailar, John, 70

Bailout (banks) data, 223–226

Baker, Dean, 41–42, 155, 238–239

Bandwagon effect, 270–272, 293

Banks

bailout data, 223–226

racial discrimination by, 44–46, 45

Baselines, 290

Bassett, John, 93

BEA, 141–142, 144–145

Benefit-cost analysis, 77–80

Berger, Mark C., 52

Berliner, David C., 97, 99

363

Bernanke, Ben, 42–43

“Best city” indices, 51–52, 53

Bias

in analysis, 293–294

political, 272–273

results and, 260–262

Biddle, Bruce J., 97, 99

Bilmes, Linda, 235, 237

Birch, David, 220–222

Blacks

achievement gap and, 99

crime and, 123–124

educational attainment and, 90

HIV/Aids and, 72

interracial prejudice and, 267

life expectancy and, 66

as political candidates, 270

race classification, 19–20

racial discrimination by banks and, 45

segregation and, 49–50

Simpson’s Paradox and, 101

undercount of population and, 10

in U.S. Census (1840), 20

Block, Gladys, 74

Blomquist, Glenn C., 52

Bloomberg Markets magazine, 223

BLS, 18–19, 36, 142, 151, 153, 169, 192–193, 195, 198, 201–203, 214, 234, 247, 248, 249, 250, 294–295

Blue, Laura, 65

Blumstein, Alfred, 116

Boom in population, predicting, 14–15

364

Boston Federal Reserve Bank, 45–46

Boston Scientific, 225

Bosworth, Barry, 158

Bradley effect, 270

Bradley, Tom, 270

Brady, Henry E., 270

Breast cancer data, 71

Broad index, 250

Burden, Barry, 265

Bureau of Economic Analysis (BEA), 141–142, 144–145

Bureau of Justice, 113, 124

Bureau of Labor Statistics (BLS), 18–19, 36, 142, 151, 153, 169, 192–193, 195, 198, 201–203, 214, 234, 247, 248, 249, 250, 294–295

Bureau of the Budget. See Office of Management and Budget

Burt, Martha, 48

Bush, George H.W., 220, 271

Bush, George W., 220, 235–236, 244, 266–267, 269

Business Employment Dynamics program (BLS), 214

Business savings, 156

Business statistics. See also specific corporation

Case Study Questions, 227–228

controversies

bailout’s success, 223–226

corporate size, 214–222, 215, 218, 219, 221

stock picks, 226–227

new business, 223

overview, 211, 227

sources, 211–214, 211

Business Week study, 226

California’s Proposition 19 survey, 263

Callender, Marie, 213

365

Calment, Madame Jeanne, 68

Cancer data, 69–72, 71

Capital punishment data, 4, 126–128

Car accidents data, 75–77

Card, David, 199–201

Career data, number of jobs in, 200

Cargill, Inc., 213

Carry concealed weapons (CCW), 129

Carter, Jimmy, 273

Case, Karl, 39

Case-Shiller Home Price Index, 39–40, 40, 42

Case Study Questions, 6. See also specific subject

Categories, fluid, 287–288

Causes of Cancer, The (Doll and Peto), 70

CBO, 17, 178, 241–242

CBS poll on abortion, 265–266

CBSA, 48–49

CCW, 129

CDC, 19–20, 60, 62, 72–73, 130

Cell phone data, 75–77, 260–261

Census of Fatal Occupational Injuries, 202–203

Central Intelligence Agency (CIA), 118, 120, 143, 235

Cervical cancer data, 72

Chain–type inflation adjustment, 145

Challenger explosion, 79

Charter schools data, 101–102

Chevron Corporation, 224–225

Chicago heat wave deaths data, 74–75

Chinese currency data, 149–150, 250

Chinese economic data, 146, 149–150

Chodorow–Reich, Gabriel, 158

366

CIA, 118, 120, 143, 235

Citizens for Tax Justice, 242

City ranking data, 51–52, 53

Clayton Antitrust Act, 218

Clinton, Bill, 220, 268, 271

CMSAs, 48–49

Code of Professional Ethics and Practices (AAPOR), 275, 285

Cohabitation data, 27

Cohen, Patricia Cline, 20

College Board, 100

Collins, Susan M., 158

“Color of Money, The” series, 45

Common Core of Data, 89–90

Congressional Budget Office (CBO), 17, 178, 241–242

ConocoPhillips, 215

Consolidated metropolitan statistical areas (CMSAs), 48–49

Consumer confidence data, 156–157

Consumer Price Index (CPI), 39, 180, 247–248, 249, 294

Consumer Sentiment Index (University of Michigan), 156

Controversies about data, 6, 291–293. See also specific subject

Corak, Miles, 176

Corcoran, Sean, 103

Core based statistical area (CBSA), 48–49

Corporation size data

assets, 215–216

classification system, 221

employment, 216–217

implications, 217

market value, 216

overview, 214–215, 215

sales, 215, 218

367

small businesses, 220–222

“too big” issue, 217–220, 218, 219

Cost effectiveness, 79

Country of origin concept, 159

County Business Patterns, 141–142

CPI, 39, 180, 247–248, 249, 294

CPI-U, 248

CPI-W, 248

CPS, 13, 25, 27, 67, 111, 168–169, 194, 198, 200, 201, 203, 234, 294–295

CQ Press, 119

CQ Vital Statistics on American Politics Online Edition, 234

Crespi, Irving, 266

Crime data

Case Study Questions, 133–134

controversies

blacks and crime, 123–124

capital punishment and crime, 4, 126–128

decline in crime, 116–117

female criminals, 117–118

guns and crime, 128–129, 130, 131

hate crimes, 125–126

human trafficking, 118–119, 120

poverty and crime, 122–123

prison’s effectiveness, 124–125

rankings of crime, 119–121

rape, 121–122

UCR, NCVS, or NIBRS, 115–116

white-collar crime, 131–132

gun research, 130

Improving Crime Data project and,119–120

murder, by stranger versus acquaintance, 130

368

in New York City, 118

overview, 111, 132–133

recidivism, 127

sex offenders, 127

sources, 111–115, 111

“victimless crimes,” 125

wrongful incarceration and, 128

Crime Index, 112

Crime rankings, 119–121

Culhane, Dennis, 47

Currency rates, 248, 250–251

Current Employment Statistics, 192

Current Population Survey (CPS), 13, 25, 27, 67, 111, 168–169, 194, 198, 200, 201, 203, 234, 294–295

Currie, Elliott, 122, 124

Curtin, Richard T., 157

Cyclical unemployment, 196–197

Dallas Cowboys, 224

Daly, Richard, 270

“Dark Side of Numbers, The” (Seltzer and Anderson), 13

Data literacy, 294–296

Davis, Devra Lee, 70

Davis, Stephen, 236

Davis, T. Cullen, 92

Death penalty data, 4, 126–128

Deaton, Angus, 149

Decline in crime data, 116–117

Defensive gun use (DGU), 129

Deficit data, U.S., 237–241

Definitions of concepts, conflicting, 285–286

Demographic data

369

Case Study Questions, 30–31

controversies

households and families, 25–28, 26

population counts, 9–25

head of household designation, 26

historical perspective, 11

migration, interstate, 14

overview, 7, 29–30

sources, 7–9, 7

undercounts, 10, 11, 12, 120

Depenalization and drug use data, 78, 263, 264

Depreciation, 216

Development, Relief, and Education for Alien Minors (DREAM) Act, 15

Dewey, Thomas E., 259

DGU, 129

Digest of Education Statistics, 90

DiIulio, John, 116, 125

DIJA, 225–226

Direct counts, 165, 168

Disraeli, Benjamin, 3

Divorce data, 27–28

DNA evidence, 128

DOD, 4, 142

Doll, Richard, 70

Donohue, John, 116

Dow Jones Industrial Average (DIJA), 225–226

DREAM Act, 15

Drug use data, 77, 78, 263, 264

Du Pont Corporation, 215

Duggan, Mark, 129

Dukakis, Michael, 267

370

Dun & Bradstreet, 213–214, 220–221

Duneier, Mitchell, 75

Easton, Joseph, 132

Eaton, David, 132

Economic data. See National economy data; specific type

Economic mobility, 174–176

Economic Policy Institute (EPI), 90, 103

Economic Report of the President, 239

Education data

Case Study Questions, 106

controversies

charter schools, 101–102

educational attainment, 90–93, 94, 98, 99

gender differences, 94, 98

poor data, 89–90

teacher compensation, 102–105

testing, 93, 95–97, 95, 99–100

overview, 87, 105

school ratings, 92

sources, 87–89, 87

value-added tests, 104

Educational attainment data, 90–93, 94, 98, 99

Educational Testing Service, 97

Ehrlich, Isaac, 126–127

Elderly data, 66, 68

Election predictions, 259, 269–273

Elliott, Ray, 225

Employment data, 197–198, 198, 205, 216–217

Engineer jobs data, 198

EPI, 90, 103

371

Establishment Survey (BLS), 169, 197, 205

Estate-multiplier method, 167–168

Estate tax data (ETD), 168

ETD, 168

Ethnicity data. See Race and ethnicity data

European Central Bank, 234

Evans, Timothy, 270

Exit polling, 269–270

Exxon/Mobil, 215, 215, 218, 219

Factor, Mallory, 235

FAIR, 17

Family data

cohabitation, 27

defining family, 25–26

divorce, 27–28

head of household designation, 26

households, 25–28, 26

income, 178–179

marriage, 28

Fannie Mae, 215, 215

Farley, Reynolds, 10–11

Farrell, Warren, 204

FBI, 111–113, 121

Fed, The. See U.S. Federal Reserve Bank

Federal Bureau of Investigation (FBI), 111–113, 121

Federal Children’s Bureau survey, 62

Federal Housing Finance Agency (FHFA), 39–41, 40, 42

Federal Trade Commission, 220

Federation for American Immigration Reform (FAIR), 17

Feige, Edgar, 147

372

Female criminals data, 117–118

Fenelon, Andrew, 65

FFQ data, 74

FHFA, 39–41, 40, 42

Financial wealth, 170. See also Wealth data

Fitzgerald, Terry, 179–180

Fixed weight method, 144–145, 251

Fleck, Susan E., 206

Florio, Jim, 271

Fluid categories, 287–288

Food consumption surveys, 74

Food Frequency Questionnaire (FFQ) data, 74

“Food insecure” data, 67

Forbes 2000, 217

Fortune 500 rankings, 215

Frankovic, Kathleen, 276

Freeman, Harry L., 158

Freudenberg, William R., 79

Friedman, Jeffrey, 73

Friedman, Milton, 222

Friedman, Rose, 182

Gallup, George, 259

Gallup polls, 259, 266

Gates, Gary, 24

Gay population data, 24–25

GDP data. See Gross domestic product data

Gender differences data

educational, 94, 98

math versus science skills, 98

workload, 204

373

General Electric, 217

General Social Survey (GSS), 258

Geoghegan, Thomas, 205–206

Geographic units, 5, 48–49

Gertz, Marc, 129

Gilbert, Neil, 121

Glaeser, Edward, 50

Glenn, John, 92

Globalization, 158

GNP data, 141, 144

Goldwater, Barry, 271

Gonzales, Alberto, 118

Gore, Al, 266–267, 269

Gould, Stephen Jay, 66–67

Government data. See also specific agency

assets, selling, 239

Case Study Questions, 253

controversies

currency rates, 248, 250–251

deficits, U.S., 237–241

inflation, 246–248, 249

Iraqi civilian deaths, 236, 237

military spending, 234–236

money measurements, 243, 245–246

taxes, 241–243, 243, 244–245

voter participation, 251–253

welfare spending, 236–237

overview, 232, 253

sources, 232–234, 232

Great Recession (2007–2009), 196–197

Green, Mark, 79

374

Greenspan, Alan, 41

Gross domestic product (GDP) data

advance, 144

celebrations/holidays and, 148

Chinese, 146, 150

Chinese versus United States, 149–150

controversies, 143–150, 145, 148

data sources, 141

gross national product versus, 141

inflation and, 144–145

intercountry comparisons, 147–149

problems with, 145–147

products and, changing, 153

real, 144, 145

sales of corporations and, annual, 218

underground economy and, 146–147

Gross national product (GNP) data, 141, 144

GSS, 258

Guns and crime data, 128–129, 130, 131

Handbooks, corporate, 213

Harris, Louis, 273

Harris poll on abortion, 266

Hate crimes data, 125–126

Hawkins, Gordon, 124–125

Hayford, Sarah R., 27

Head of household designation, 26

Headline makers, 292

Health data

benefit-cost analysis and, 77–80

Case Study Questions, 81–82

375

controversies

cancer, 69–72, 71

car accidents, 75–77

drug use, 77, 78

heat-related deaths, 74–75

HIV/AIDS, 72–73, 283–284

infant mortality, 62–63, 64

life expectancy, 3, 63–69

obesity, 73–74, 74

food consumption surveys, 74

overview, 60, 81

sources, 60–62, 60

tamper–proof closures, cost of, 80

worldwide, 61–62

Healy, Bernadine, 62

Heat-related deaths data, 74–75

Heat Wave (Klinenberg), 74–75

Heckler, Margaret M., 71

Hedonic Index, 52

Hedonic pricing, 52, 247

Heimer, Karen, 117

Herfindahl-Hirschman Index (HHI), 218–219

Heritage Foundation, 236

Heston, Alan, 149

HHI, 218–219

Hill, Charles, 236

Himmelberg, Charles, 41–42

Hispanics

educational attainment and, 90

immigrants and, unauthorized, 16

life expectancy and, 65

376

race classification, 21–23

racial discrimination by banks and, 46

Simpson’s Paradox and, 101

Hite Reports (Hite), 274

Hite, Shere, 274–275

HIV/AIDS data, 72–73, 283–284

HMDA (1975), 45

Hodgkinson, Harold, 105

Hoehn, John P., 52

Holmes, James, 8

Holz, Carsten, 150

Home Mortgage Disclosure Act (HMDA) (1975), 45

Homeless population data, 46–48

Homeownership data, 43–44

Homosexual population, 24–25

Horserace journalism, 272

Household data, 25–28, 26

Housing bubble data, 39–43, 40

Housing data

Case Study Questions, 54–55

controversies

city rankings, 51–52, 53

geographic units, 5, 48–49

homeless population, 46–48

homeownership, 43–44

housing crisis, 39–43, 40

racial discrimination by banks, 44–46, 45

segregation, 49–50

overview, 36, 53

price indices, 42

quality of housing, 42

377

sources, 36, 37–39, 38

Housing starts measures, 37

HUD, 37, 47

Human cost of war, 236, 237

Human Development Index, 143

Human life, value of, 77–79

Human trafficking data, 118–119, 120

Humphrey, Hubert, 273

Hunger data, 67

Hurricane Katrina, 132

ICP, 149

IEA, 95–96

Immigrant population data, undocumented, 15–17

Improved data, 284–285

Improving Crime Data project, 119–120

Incidence rate (cancer), 71

Income capitalization, 168

Income data

Case Study Questions, 185–186

controversies, 172–176

family, 178–179

individual earnings, 177–178

mean, 171

median, 171

mobility and, economic, 174–176

overview, 166, 185

per capita, 177

rich classification and, 172

rich getting richer issue and, 172–174

sources, 168–169

378

standards of living and, 177–181, 178

Index numbers, 289–290. See also specific index

Index of Leading Economic Indicators, 37

Individual earnings, 177–178

Industry statistics, 141–142

Infant mortality data, 62–63, 64

Inflation data, 5, 144–145, 246–248, 249

Innocence Project, 128

Inside the box effect, 153

Institute for Research in the Social Sciences (Stanford University), 258

Integrated Postsecondary Education Data System (IPEDS), 92–93

Inter-University Consortium for Political and Social Research website, 112

Interagency Technical Working Group on Developing a Supplemental Poverty Measure, 183

Intercountry comparisons, 147–149

International Association for the Evaluation of Educational Achievement (IEA), 95–96

International Comparison Program (ICP), 149

International Labor Organization, 205

International Monetary Fund, 149, 158

International statistics, 143, 157–159, 159

Internet polls, self–selected, 262

Interstate migration, 14

IPEDS, 92–93

Iraqi civilian deaths data, 236, 237

Iraqi war cost, 235–236

Jaeger, Richard, 96

Jefferson, Thomas, 11

Jesilow, Paul, 132

Job Generation in America (Birch), 220–221

Jobs data, 198, 199

379

John M. Olin Foundation, 130

Johns Hopkins University, 237

Johnson, Lyndon, 271

Joint Economic Committee of Congress, 173

JP Morgan Chase, 217, 223

Kagay, Michael, 266

Kaplan, Greg, 14

Kennedy, Edward (Teddy), 273

Kennedy, John F., 273

Keyfitz, Nathan, 15

Kinsey, Alfred, 24

Kinsey Institute, 274–275

Kleck, Gary, 129

Klinenberg, Eric, 74–75

Kocherlakota, Narayana, 196

Krosnick, Jon, 265

Krueger, Alan B., 199–201

Krugman, Paul, 240

Kuznets, Simon, 141, 143, 145

Labor statistics

career, number of jobs in, 200

Case Study Questions, 206–207

controversies

employment, 197–198, 198, 205

minimum wage and jobs, 199–201

safety in the workplace, 202–203

unemployment, 193–194, 195, 196–197, 197, 205, 285

unions, 201–202

engineer jobs, 198

international, 204–206

380

men’s workload, 204

overview, 192, 206–207

sources, 192–193, 192

standard of living, 203–206

Ladd, Helen F., 46

Lampman, Robert J., 168

Land phone data, 260–261

Landon, Alf, 259

Langer, Gary, 157

Lauritsen, Janet, 117

Lave, Charles A., 76

Lehman Brothers, 212

Levitt, Steven, 116

Lewinsky (Monica) scandal, 268

Life expectancy data, 3, 63–69

Lindsey, Lawrence B., 46

Liquidity, 245

Literary Digest poll (1936), 259–260

Lott, John R., Jr., 129, 130, 131

Love, Richard, 71

Ludvigson, Sydney, 157

Lynch, James P., 117

Lynch, Peter, 226

MA, 48

MacAvoy, Paul, 80

MacCoun, Robert J., 78

MacDonald, Heather, 123

Maddison, Angus, 150

Magazines, business, 213

Major Currency index, 250

381

Manhattan Institute study, 50

Mankiw, Gregory, 244

Marginal tax rate, 244–245

Marijuana legalization, 78, 263, 264

Market value data, corporate, 216

Marriage data, 28

Mars Candy Company, 175, 213

Mason-Dixon Political/Media Research polling firm, 271

Massey, Douglas, 50

Mathematics of social statistics, 289–291

Maxfield, Michael, 114–115

Mayer, Christopher, 41–42

McDonald, Michael P., 251–252

McDonald’s stock, 226

McGahey, Richard, 128

McGraw-Hill Construction’s Dodge data, 38

Mean income, 171

Mean life expectancy, 66–68

Means, 66–68, 171, 288–290

Measure of Economic Welfare, 146

Media polls, 258

Median income, 171

Median life expectancy, 66–68

Medians, 66–68, 171, 290

Men’s workload data, 204

Mergers, 219–220

Metropolitan area (MA), 48

Metropolitan Statistical Areas (MSAs), 5, 48–50

Mexicans. See Hispanics

Microsoft (technology firm), 158

Middle class, disappearing, 179–180

382

Migration, 14, 65

Military spending data, 234–236

Military spending data sources, 233

Minimum wage data, 199–201

Mishel, Lawrence, 103

Mismeasuring Our Lifes (2010 French study), 146

Missing data, 283–284

Mobil Oil Corporation, 79. See also Exxon/Mobil

Mobility, economic, 174–176

Mode, 171

Moffett, Toby, 266

Money

Chinese, 149–150, 250

data sources, 233

measuring, 243, 245–246

rates, 248, 250–251

Money Magazine, 51

Monitoring the Future (MTF) survey, 61, 77

Moore, Michael, 65

More Guns, Less Crime (Lott), 129, 130, 131

Morgan, S. Philip, 27

Morris, Dwight, 266

Mortality rates

cancer, 69–71

infant, 62–63, 64

Mortgage racial discrimination, 44–46, 45

Mortgage tax deduction, 43–44

Motorcycles and car accidents data, 75–77

MSAs, 5, 48–50

MTF survey, 61, 77

Multiracial classification, 18–19, 23

383

Munnell, Alicia, 45

Murphy, Kevin, 236

Myers, Dowell, 43

Myth and Measurement (Card and Krueger), 199

Myth of Male Power, The (Farrell), 204

NAACP, 19

NAEP, 97, 99–100, 101

NAICS, 141, 221

National Advisory Commission on Civil Disorders, 49–50

National Annenberg Election Survey, 263

National Assessment of Educational Program (NAEP), 97, 99–100, 101

National Bureau for Economic Research (NBER), 221–222

National Cancer Act, 69

National Cancer Institute (NCI), 70–71

National Center for Education Statistics (NCES), 87–89

National Center for Health Statistics, 24, 61

National Council on Public Polls (NCPP), 275

National Crime Victimization Survey (NCVS), 111, 113–116, 125–126, 133, 267

National economy data

Case Study Questions, 161

controversies

consumer confidence, 156–157

gross domestic product, 143–150, 145, 148

international statistics, 157–159, 159

productivity measurements, 149, 150–154, 151, 152

savings rate, 154–156, 155

overview, 140, 159–161

sources, 140, 141–143

National Election Studies, 271

National Household Survey on Drug Use data, 77

384

National Incident–Based Report System (NIBRS), 111–112, 114–116, 126, 131–132, 284

National Income and Product Accounts, 141, 177

National Institute on Drug Abuse (NIDA), 61

National Law Center on Homelessness and Poverty, 47

National Longitudinal Survey, 88

National Longitudinal Survey of Youth, 200

National Opinion Research Center (NORC), 25, 258

National Research Council, 76, 183

National Rifle Association (NRA), 130

National Survey of Families and Households, 27

National Survey of Family Growth, 27

National Survey on Drug Use and Health (NSDUH), 61

National Violence Against Women Survey, 121–122

National White Collar Crime Center survey, 131

“Nation’s Report Card, The.” See SAT

NBER, 221–222

NCES, 87–89

NCI, 70–71

NCLB legislation, 90–91, 96–97, 295

NCPP, 275

NCVS, 111, 113–116, 125–126, 133, 267

Net worth, 169–170. See also Wealth data

Neumark, David, 199

New business data, 223

New York City crime data, 112, 116–117, 118

New York City census, 10

New York Times polls, 258, 275

New York Yankees, 224

NIBRS, 111–112, 114–116, 126, 131–132, 284

NIDA, 61

385

Niederle, Muriel, 98

Nixon, Richard, 50, 69, 273

No Child Left Behind (NCLB) legislation, 90–91, 96–97, 295

Nonemployer Statistics program (U.S. Census Bureau), 214

Nonresponse to polling, 261–262

NORC, 25, 258

Nordhaus, William D., 146, 156

North America Industry Classification System (NAICS), 141, 221

Notes section, 6

Novy-Marx, Robert, 238

NRA, 130

NSDUH, 61

Nurses’ Study, 74

Obama, Barack, 172, 236, 240, 243, 244–245, 270

Obesity data, 73–74, 74

Occupational Safety and Health Administration (OSHA), 79, 202, 284

OECD, 92, 149, 158, 205

Off–budget deficits, 238

Office of Civil Rights, 88

Office of Management and Budget (OMB), 18–19, 22, 221

Office of National Drug Control Policy (ONDCP), 77

OMB, 18–19, 22, 48, 221

ONDCP, 77

O’Neill, Barry, 92

O’Neill, Tip, 266

Opportunity costs, 223

Organisation for Economic Co-Operation and Development (OECD), 92, 149, 158, 205

Orren, Gary R., 270

Orshansky, Mollie, 181

Orshansky poverty line, 181–182

386

OSHA, 79, 192, 202, 284

Over the Edge (Burt), 48

Oversimplification, 291–292

Overworked American, The (Schor), 204

Palloni, Alberto, 65

Panel Survey of Income Dynamics, 167, 169

Paradox Lost (Palloni and Arias), 65

PCE, 180

Penn World Tables, 143, 150

Pensions, valuing, 238–239

Pepinsky, Harold, 132

Per capita income, 177

Percentages, 287

Perception of evidence, 265

Perception of media coverage, 265

Perot, Ross, 266, 268–269

Personal consumption expenditures (PCE), 180

Personal savings, 155–156

Peto, Richard, 70

Pew Hispanic Center, 16

Pfizer’s stock, 226

Piehl, Anne Morrison, 125

PIRLS, 95

PISA, 95

Places Rated Almanac, 51–52

Plosser, Charles, 196

PMSAs, 48–49

Podgursky, Michael, 103

Political bias, 272–273

Pollan, Michael, 74

387

Polls as news, 272

Poor data, 89–90, 157–158, 283–284

Pope, Devlin, 98

Population count controversies

boom, predicting, 14–15

gaps between counts, 12

gay population, 24–25

immigrants, undocumented, 15–17

number of people, 9–11, 11

prison population, 11–12

privacy in census, 13–14

race and ethnicity, 17–24

Poveda, Tony G., 128

Poverty data

Case Study Questions, 185–186

controversies, 181–184, 184

crime and, 122–123

homelessness and, 47

line of poverty, 181–184, 184

measuring, 181–184, 184, 285

overview, 166, 185

PPP, 149, 159

Price data, 5, 38–39, 234

Price indices, 42

Primary metropolitan statistical areas (PMSAs), 48–49

Prison and crime data, 124–125

Prison Policy Initiative, 11

Prison population, counting, 11–12

Prison sentences, wrongful, 128

Privacy in U.S. Census, 13–14

Private corporation data sources, 213–214

388

Private polling organizations, 257–258

Private surveys, 61

Productivity data and measurements, 142, 150–154, 151, 152

Profits data, 224–225

Progessivity of tax system, 241, 244

Programme for International Student Assessment (PISA), 95

Progress in International Reading Literacy Study (PIRLS), 95

Project RACE (Reclassify All Children Equally), 18

Prostate cancer data, 71–72

PSA blood test, 72

PSID (University of Michigan), 169, 176

Public corporation data sources, 212–213

Public Health Service, 124

Public opinion polling

black candidates/white voters, 270

Case Study Questions, 277–278

controversies

election predictions, 259, 269–273

overview, 259

sampling, 260–262

standards, polling, 273, 275–276

survey design, 262–269

data sources, 257–258, 257

ignorance versus honesty, 268–269

impact of, 293

nonresponse and, 261–262

overview, 257, 276–277

Randomizing Response Technique, 264

self-reporting, 263

self-selected Internet polls, 262

women’s wants and satisfaction, 273–274

389

Purchasing power parity (PPP), 149, 159

Quality Counts reports (Education Week), 88

Quality of goods, 247

Quindlen, Anna, 92

Race and ethnicity data

Asian classification, 20–21

black classification, 19–20

Hispanic classification, 21–23

implications, 23–24

multiracial backgrounds, 18 19

self-identifying, 18, 21

trends in, 24

Racial discrimination by banks, 44–46, 45

Rand Corporation, 128

Randomizing Response Technique (RRT), 264

Rape data, 121–122

Rauh, Joshua, 238

Reagan, Ronald, 151, 239

Recidivism data, criminal, 127

Redefining Progress in Its Genuine Progress Indicator, 146

Redlining, 45

Regressivity of tax system, 241

Reiman, David, 132

Reinhart, Carmen, 240–241

Relative statistics, 288

Rental equivalence, 247–248

Reporting controversies, 6, 291–293. See also specific subject

Research centers, 258

Return migration, 65

Reuter, Peter, 78

390

Reverse Bradley effect, 270

Risk assessment, 79–80

Roberts, Paul Craig, 45

Robinson, John P., 204

Robinson, Matthew B., 77

Rockefeller, Nelson, 273

Rogoff, Kenneth, 240–241

Roosevelt, Franklin D., 194, 259

Roper Center polls, 258, 267

Rosenfeld, Richard, 123

Rothstein, Richard, 50

Rounding, 249

Rowan, Carl, 92

RRT, 264

Rubenstein, Carin, 204

Sacrificial Mother, The (Rubenstein), 204

Safety data, workplace, 202–203

Sales rankings, corporate, 215, 218

Samplings, 260–262

Sarkozy, Nicholas, 146

SAT, 21, 97, 99–100, 100

Savings rate measurements, 154–156, 155

SCF, 167, 169, 171–172, 174

Scherlen, Renee, 77

Schneider, Friedrich, 147

Scholastic Aptitude Test (SAT), 21, 97, 99–100, 100

School ratings, 92

Schor, Juliet, 204

Schulhofer-Wohl, Sam, 14

Schulman, Mark A., 272

391

Securities and Exchange Commission (SEC), 212

Segregation in housing data, 49–50

Self-reporting, 263

Self-selected Internet polls, 262

Seltzer, William, 13

Sen, Amartya, 146

Services productivity, 153–154

Sex offenders data, 127

Sex trafficking data, 118–119, 120

Shelter Index, 39

Shepherd, Joanna, 127

Sherman Antitrust Act, 218

Shiller, Robert, 39

SIC, 221

Sicko (film), 65

Silver, Nate, 261

Simpson’s Paradox, 101

Sinai, Todd, 41–42

SIPP, 167, 183

SIPRI, 253

Size of corporations. See Corporation size

Skin cancers, 72

Slanting the numbers, 293–294

Small Business Administration, 143

Small business data, 220–222

Small business data sources, 214

Smith, Charlie, 68

Social Security Trust Fund, 240

Social statistics. See also specific subject

absolute, 288

aggregate, 214

392

biased analysis of, 293–294

classifying, 286–288

conflicting definitions and, 285–286

correct use of, 4–5

data literacy and, 294–296

improved data and, 284–285

interpreting, 3, 283

mathematics of, 289–291

means and, 288–289

missing data and, 283–284

poor data and, 283–284

relative, 288

reporting controversies and, 291–293

slanting, 293–294

subgroups and, 288–289

time period and, 289

variables and, 5, 286

views of, 3

SOI, 168

Solon, Gary, 176

Sources of data, 5. See also specific subject and organization name

Speed and car accidents data, 75–77

Sports team profits, 224

Standard and Poor’s Case-Shiller Home Price Index, 39

Standard Industrial Classification (SIC), 221

Standard of living, 177–181, 178, 203–206

Standard Oil Company, 218

Standards, polling, 273, 275–276

Stanford University’s Institute for Research in the Social Sciences, 258

State education data sources, 89

Statistics. See specific type

393

Statistics of Income (SOI), 168

Steinbrenner, George, 224

Stem cell research survey design, 267–268

Stevenson, Betsey, 13, 27

Stiglitz, Joseph, 146, 235, 237

Stock market, 225–227

Stock picks, 226–227

Stockholm International Peace Research Institute (SIPRI), 253

Straw poll method, 259

Structural unemployment, 196–197

Subgroups, 288–289

Substitution of goods, 247

Summers, Larry, 98

Supplemental Poverty Measure, 183

Survey design

order of questions, 266–267

overview, 262

stem cell research, 267–268

wording of questions, 262–266

Survey of Consumer Attitudes (University of Michigan), 261

Survey of Consumer Finances (SCF), 167, 169, 171–172, 174

Survey of Current Business, 233

Survey of Income and Program Participation (SIPP), 167, 183

Survey questions

order of, 266–267

wording of, 262–266

Survey Research Center (University of Michigan), 61, 258, 268

Surveys. See specific type

Sutherland, Edwin H., 131

Swift & Co., 215

Sydnor, Justin, 98

394

Tamper-proof closures, cost of, 80

Tax data

bearers of tax burden, 241–242, 243

defining wealth and, 171–172

fair share for rich, 242–243, 244

marginal tax rate, 244–245

mortgage deduction and, 43–44

progressivity of tax system and, 241, 244

regressivity of tax system and, 241

Tax Policy Center, 44

Tax records, 167–168

Teacher compensation data, 102–105

Teens and car accidents data, 75–77

Testing data, 93, 95–97, 95, 99–100

“Tests and Gender” symposium, 98

Texting and car accidents data, 75–77

Thomas Register, 213–214

Time period, 289

TIMSS, 95

Tittle, C.S., 122

Tobin, James, 146

Topel, Robert, 236

Trade data errors, 159

Trade statistics, 142, 157–159, 159

Trade-weighted indexes, 250

Trends in International Mathematics and Science Study (TIMSS), 95

Truth by repetition, 292–293

Twain, Mark, 3

UCR, 111–113, 115–116, 121, 123, 126, 130, 131, 294

UN, 14, 143

395

Undercounts, 10, 11, 12, 120, 194, 197

Underdog effect, 270–272, 293

Underground economy data, 146–147

Undue accuracy, 290–291

Unemployment data, 193–194, 195, 196–197, 197, 205, 285

Uniform Crime Reports (UCR), 111–113, 115–116, 121, 123, 126, 130, 131, 294

Union data, labor, 201–202

United Nations (UN), 14, 143

“United States Budget in Brief,” 233

United States Election Project, 234

University of Michigan

Consumer Sentiment Index, 156

Panel Study of Income Dynamics (PSID), 169, 176

Survey of Consumer Attitudes, 261

Survey of Consumer Finances, 167, 169, 171–172, 174

Survey Research Center, 61, 258, 268

University of Pennsylvania School of Social Work, 118–119

Urban Institute, 95

U.S. Budget, 253

U.S. Bureau of Labor Statistics. See Bureau of Labor Statistics (BLS)

U.S. Census, 6, 8–10, 13–14, 17–19, 20, 21, 25, 37, 48, 66, 168, 179, 294

U.S. Census Bureau, 5, 9, 13, 14, 15, 25, 28, 88, 111, 113, 193, 285

U.S. Centers for Disease Control and Prevention (CDC), 19–20, 60, 62, 72–73, 130

U.S. Commerce Department, 141–142

U.S. Customs data, 142

U.S. deficit data, 237–241

U.S. Department of Agriculture, 67

U.S. Department of Commerce, 141, 144, 153–155, 157, 177, 233

U.S. Department of Defense (DOD), 4, 142

U.S. Department of Education, 88–90

U.S. Department of Energy, 235

396

U.S. Department of Health and Human Services, 61, 69

U.S. Department of Homeland Security, 16

U.S. Department of Housing and Urban Development (HUD), 37, 47

U.S. Department of Justice, 111, 113, 118, 121, 124, 126, 220

U.S. Department of Labor, 4, 142, 152–154, 206

U.S. Department of Transportation, 76

U.S. Economic Census, 141

U.S. Federal Reserve Bank, 154, 157, 167, 169, 173, 223, 234, 240, 243, 245–246, 250

U.S. Federal Reserve Board, 142–143, 167, 233, 250–251, 283

U.S. government surveys, 61

U.S. Health Surveys, 60–61

U.S. Steel, 219–220

U.S. Treasury, 223

U.S. Vital Statistics, 7, 9, 27, 60

“User cost,” 41

Value added tests, 104

Variables, 5, 286

Vaupel, James W., 15

Vesterlund, Lise, 98

“Victimless crimes,” 125

Vigdor, Jacob, 50

Village Voice study, 119

Vital and Health Statistics Series, 61

Vital statistics, 9

Voter participation data, 251–253

Voting data, 234

Wadhwa, Vivek, 198

Wall Street Journal polls, 258

Walmart, 215, 215, 218, 219

397

War on Poverty, 181–182

Wascher, William, 199

Washington Post polls, 258, 274–275

Wattenberg, Martin, 252

Wealth data

Case Study Questions, 185–186

controversies

defining wealth, 169–172, 171

measuring wealth, 173, 174, 175

middle class, disappearing, 179–180

mobility, economic, 174–176

standard of living, 177–181, 178

overview, 166, 185

sources, 166, 167–169

Weicker, Lowell, 266

Welfare spending data, 236–237

White-collar crime data, 131–132

Whitman, Christine Todd, 271

WHO, 61–62, 73

“Why Don’t We Have the Prisons We Need?” (Reader’s Digest article), 124

Williams, Elliot, 249

Wilson, David C., 267

Wilson, James Q., 175

Winner, Langdon, 80

Wolff, Edward N., 170

Women and Love (Hite), 274

Women’s data

crime and, 117–118

educational attainment and, 94, 98

rape and, 121–122

wants and satisfaction surveys, 274–275

398

workload, 204

Women’s Health Initiative, 74

Workplace safety data, 202–203

Works Progress Administration, 194

World Bank, 148

World Bank data sets, 143

World Facebook (CIA), 143

World Health Organization (WHO), 61–62, 73

Worldwide health data, 61–62

Wright, Bruce, 123

Wright, Gerald C., 271

Wrongful incarceration data, 128

Yankelovich Partners consumers survey, 24

Yu, Zhou, 43

Zedlewski, Edwin, 124–125

Zimmerman, David J., 176

Zimring, Franklin E., 118, 124–125

399

About the Authors □□□□

Mark H. Maier is Professor of Economics at Glendale Community College, Glendale, California. He is the co-editor of Just-in-Time Teaching: Across the Disciplines, Across the Academy with Scott Simkins (2009), co-author of Introducing Economics: A Critical Guide for Teaching with Julie A. Nelson (2007), and author of City Unions: Managing Discontent in New York City (1987). He received his Ph.D. in economics from the Graduate Faculty, New School for Social Research in 1980.

Jennifer Imazeki is Professor of Economics at San Diego State University where she teaches courses in applied microeconomics. She has a B.A. from Pomona College and an M.A. and Ph.D. from the University of Madison-Wisconsin, all in economics. Her research focuses on the economics of K–12 education, including school finance reform, adequacy, and teacher labor markets. She has worked on several projects to encourage active learning in economics and writes about teaching economics at economicsforteachers. blogspot.com.

400

目录

Cover 2 Half Title 3 Title Page 4 Copyright Page 6 Table of Contents 8 List of Figures, Tables, and Boxes 16 Preface to the Fourth Edition 19 Acknowledgments 21 1. Introduction 23

The Purpose of This Book 23 How to Use This Book 25

2. Demography 28 Data Sources 29

U.S. Census 29 American Community Survey 29 Vital Statistics 30

Controversies 30 Population Counts 30 Privacy in the Census: Double-Edged Sword? 34 Will There Be a Population Boom? 37 Undocumented Immigrants 37 Race and Ethnicity 40 How Big Is the Gay Population? 47 Households and Families 48

Summary 53 Case Study Questions 54

3. Housing 62 Data Sources 63

U.S. Census 63 American Housing Survey 64

401

Other Industry Data 64 Price Data 65

Controversies 66 Housing Crisis 66 Homeownership 70 Racial Discrimination by Banks 72 Geographic Units 76 Segregation 78 Is Your City the Best Place to Live? 79

Summary 82 Case Study Questions 83

4. Health 90 Data Sources 91

U.S. Health Surveys 91 Other Government Surveys 91 Private Surveys 91 Worldwide Data 92

Controversies 92 Infant Mortality 92 Abortion 94 Are We Living Longer? 94 How to Measure Longevity 97 Cancer 101 HIV/AIDS 104 Are Americans Getting Fatter? 105 Chicago Heat Wave: What Caused the Tragedy? 107 What's Unsafe on the Road? Speed, Texting, Teens, Motorcycles, or Alcohol?

108

Drug Use 109 Benefit-Cost Analysis 110

Summary 114 Case Study Questions 114

5. Education 122 Data Sources 123

402

National Center for Education Statistics 123 U.S. Census Bureau 123 Other Surveys 124 State Data 124

Controversies 124 Poor Data 124 Educational Attainment 126 Testing 129 Charter Schools: Are They More Effective Than Regular Public Schools?

139

Teacher Compensation 140 Summary 143 Case Study Questions 143

6. Crime 150 Data Sources 150

Uniform Crime Reports 151 National Crime Victimization Survey 152 National Incident-Based Reporting System 153

Controversies 154 UCR, NCVS, or NIBRS? 154 Crime Is Down—And We Don't Know Why 155 Are There More Female Criminals? 157 Human Trafficking: How Often Does It Occur? 158 Where Is Crime the Worst? 159 Rape 161 Does Poverty Cause Crime? 162 Why Is the Black Crime Rate So High? 163 Does Prison Pay? 165 Hate Crimes 166 Does Capital Punishment Deter Murder? 167 More Guns/More Crime or Less Crime? 170 What About White-Collar Crime? 172

Summary 174 Case Study Questions 176

403

7. The National Economy 183 Data Sources 184

U.S. Commerce Department 184 U.S. Labor Department 185 U.S. Federal Reserve Board 186 U.S. Small Business Administration 186

International Statistics 186 Controversies 187

Which GDP? 187 Adjusting GDP Growth for Inflation 188 Problems with GDP 189 Underground Economy 190 Intercountry Comparisons 191 Has China Caught Up with the United States? 193 Measuring Productivity 195 The Savings Rate 199 Consumer Confidence 201 International Statistics 202

Summary 205 Case Study Questions 206

8. Wealth, Income, and Poverty 213 Data Sources 214

Wealth 214 Survey of Income and Program Participation 214 Indirect Estimates from Tax Records 215 Direct Counts 215 Income 215

Controversies 217 Wealth 217 Income 220 Are We Better Off? 225 Poverty 230

Summary 234 Case Study Questions 235

404

9. Labor Statistics 242 Data Sources 242

U.S. Bureau of Labor Statistics 242 U.S. Census Bureau 243

Controversies 243 Unemployment 243 The Minimum Wage and Jobs 250 Unions 252 Is the Workplace Safe? 254

Squirrel Cage or Easier Times? 255 International Labor Statistics 257

Summary 259 Case Study Questions 259

10. Business Statistics 264 Data Sources 265

Public Corporations 265 Privately Held Corporations 267 Small Businesses 267 Aggregate Statistics 267

Controversies 268 Who Is the Biggest of Them All? 268 Are the Big Too Big? 271 Is Small Beautiful? 274 Were the Bailouts a Success? 278 How Much Profit? 278 How Now Dow? 280

Picking Stock Winners 281 Summary 282 Case Study Questions 282

11. Government 288 Data Sources 289

U.S. Budget 289 Military Spending 289 Money 290

405

Prices 290 Voting Data 290

Controversies 291 How Much for the Military? 291 How Much for Welfare? 292 How Big Is the Deficit? 294 A Debt Monster? 296 Taxes 298 Measuring Money 300 Inflation 304 Currency Rates 306 Fewer Voters? 309

Summary 311 Case Study Questions 311

12. Public Opinion Polling 317 Data Sources 318

Private Polling Organizations 318 Media Polls 318 Research Centers 318

Controversies 319 Sampling 320 Survey Design 323 Predicting Elections 331 Polling Standards 336

Summary 339 Case Study Questions 341

13. Conclusions 347 Poor/Missing Data 347 Improved Data 348 Conflicting Definitions 350 Comparing Apples and Oranges 350 Comparing Apples and Orangutans 352 The Mathematics of Social Science Statistics 354 The Reporting of Controversies 356

406

Biased Analysis: Slanting the Numbers 358 What's a Researcher to Do? 359

Index 362 About the Authors 400

407

  • Cover
  • Half Title
  • Title Page
  • Copyright Page
  • Table of Contents
  • List of Figures, Tables, and Boxes
  • Preface to the Fourth Edition
  • Acknowledgments
  • 1. Introduction
    • The Purpose of This Book
    • How to Use This Book
  • 2. Demography
    • Data Sources
      • U.S. Census
      • American Community Survey
      • Vital Statistics
    • Controversies
      • Population Counts
      • Privacy in the Census: Double-Edged Sword?
      • Will There Be a Population Boom?
      • Undocumented Immigrants
      • Race and Ethnicity
      • How Big Is the Gay Population?
      • Households and Families
    • Summary
    • Case Study Questions
  • 3. Housing
    • Data Sources
      • U.S. Census
      • American Housing Survey
      • Other Industry Data
      • Price Data
    • Controversies
      • Housing Crisis
      • Homeownership
      • Racial Discrimination by Banks
      • Geographic Units
      • Segregation
      • Is Your City the Best Place to Live?
    • Summary
    • Case Study Questions
  • 4. Health
    • Data Sources
      • U.S. Health Surveys
      • Other Government Surveys
      • Private Surveys
      • Worldwide Data
    • Controversies
      • Infant Mortality
      • Abortion
      • Are We Living Longer?
      • How to Measure Longevity
      • Cancer
      • HIV/AIDS
      • Are Americans Getting Fatter?
      • Chicago Heat Wave: What Caused the Tragedy?
      • What's Unsafe on the Road? Speed, Texting, Teens, Motorcycles, or Alcohol?
      • Drug Use
      • Benefit-Cost Analysis
    • Summary
    • Case Study Questions
  • 5. Education
    • Data Sources
      • National Center for Education Statistics
      • U.S. Census Bureau
      • Other Surveys
      • State Data
    • Controversies
      • Poor Data
      • Educational Attainment
      • Testing
      • Charter Schools: Are They More Effective Than Regular Public Schools?
      • Teacher Compensation
    • Summary
    • Case Study Questions
  • 6. Crime
    • Data Sources
      • Uniform Crime Reports
      • National Crime Victimization Survey
      • National Incident-Based Reporting System
    • Controversies
      • UCR, NCVS, or NIBRS?
      • Crime Is Down—And We Don't Know Why
      • Are There More Female Criminals?
      • Human Trafficking: How Often Does It Occur?
      • Where Is Crime the Worst?
      • Rape
      • Does Poverty Cause Crime?
      • Why Is the Black Crime Rate So High?
      • Does Prison Pay?
      • Hate Crimes
      • Does Capital Punishment Deter Murder?
      • More Guns/More Crime or Less Crime?
      • What About White-Collar Crime?
    • Summary
    • Case Study Questions
  • 7. The National Economy
    • Data Sources
      • U.S. Commerce Department
      • U.S. Labor Department
      • U.S. Federal Reserve Board
      • U.S. Small Business Administration
    • International Statistics
    • Controversies
      • Which GDP?
      • Adjusting GDP Growth for Inflation
      • Problems with GDP
      • Underground Economy
      • Intercountry Comparisons
      • Has China Caught Up with the United States?
      • Measuring Productivity
      • The Savings Rate
      • Consumer Confidence
      • International Statistics
    • Summary
    • Case Study Questions
  • 8. Wealth, Income, and Poverty
    • Data Sources
      • Wealth
      • Survey of Income and Program Participation
      • Indirect Estimates from Tax Records
      • Direct Counts
      • Income
    • Controversies
      • Wealth
      • Income
      • Are We Better Off?
      • Poverty
    • Summary
    • Case Study Questions
  • 9. Labor Statistics
    • Data Sources
      • U.S. Bureau of Labor Statistics
      • U.S. Census Bureau
    • Controversies
      • Unemployment
      • The Minimum Wage and Jobs
      • Unions
      • Is the Workplace Safe?
    • Squirrel Cage or Easier Times?
      • International Labor Statistics
    • Summary
    • Case Study Questions
  • 10. Business Statistics
    • Data Sources
      • Public Corporations
      • Privately Held Corporations
      • Small Businesses
      • Aggregate Statistics
    • Controversies
      • Who Is the Biggest of Them All?
      • Are the Big Too Big?
      • Is Small Beautiful?
      • Were the Bailouts a Success?
      • How Much Profit?
      • How Now Dow?
    • Picking Stock Winners
    • Summary
    • Case Study Questions
  • 11. Government
    • Data Sources
      • U.S. Budget
      • Military Spending
      • Money
      • Prices
      • Voting Data
    • Controversies
      • How Much for the Military?
      • How Much for Welfare?
      • How Big Is the Deficit?
      • A Debt Monster?
      • Taxes
      • Measuring Money
      • Inflation
      • Currency Rates
      • Fewer Voters?
    • Summary
    • Case Study Questions
  • 12. Public Opinion Polling
    • Data Sources
      • Private Polling Organizations
      • Media Polls
      • Research Centers
    • Controversies
      • Sampling
      • Survey Design
      • Predicting Elections
      • Polling Standards
    • Summary
    • Case Study Questions
  • 13. Conclusions
    • Poor/Missing Data
    • Improved Data
    • Conflicting Definitions
    • Comparing Apples and Oranges
    • Comparing Apples and Orangutans
    • The Mathematics of Social Science Statistics
    • The Reporting of Controversies
    • Biased Analysis: Slanting the Numbers
    • What's a Researcher to Do?
  • Index
  • About the Authors