4 discussions due in 48 hours

profilecombs
new85743_04_c04_169-212_LowRes.pdf

Kim Steele/Photodisc/Getty Images

chapter 4

Survey Research—Describing and Predicting Behavior

Chapter Contents

• Introduction to Survey Research • Designing Questionnaires • Sampling From the Population • Analyzing Survey Data • Ethical Issues in Survey Research

CO_

CO_

new85743_04_c04_169-212.indd 169 6/18/13 12:02 PM

170

CHAPTER 4Section 4.1 Introduction to Survey Research

In a highly influential book published in the 1960s, the sociologist Erving Goffman (1963) defined stigma as an unusual characteristic that triggers a negative evaluation. In his words, “The stigmatized person is one who is reduced in our minds from a whole and usual person to a tainted, discounted one” (1963, p. 3). People’s beliefs about stigmatized characteristics exist largely in the eye of the beholder but have substantial influence on social interactions with the stigmatized (see Snyder, Tanke, & Berscheid, 1977). A large research tradition in psychology has been devoted to understanding both the origins of stigma and the consequences of being stigmatized. According to Goffman and others, the characteristics associated with the greatest degree of stigma have three features in common: They are highly visible, they are perceived as controllable, and they are misunderstood by the public.

Recently, researchers have taken considerable interest in people’s attitudes toward mem- bers of the gay and lesbian community. Although these attitudes have become more posi- tive over time, this group still encounters harassment and other forms of discrimination on a regular basis (see National Gay Task Force, 1984). One of the top recognized experts on this subject is Gregory Herek, professor of psychology at the University of Califor- nia at Davis (http://psychology.ucdavis.edu/herek/). In a 1988 article, Herek conducted a survey of heterosexuals’ attitudes toward both lesbians and gay men, with the goal of understanding the predictors of negative attitudes. Herek approached this research question by constructing a scale to measure attitudes toward these groups. In three stud- ies, participants were asked to complete this attitude measure, along with other existing scales assessing attitudes about gender roles, religion, and traditional ideologies.

Herek’s (1988) research revealed that, as hypothesized, heterosexual males tended to hold more negative attitudes about gay men and lesbians than heterosexual females. However, the same psychological mechanisms seemed to explain the prejudice in both genders. That is, negative attitudes were associated with increased religiosity, more traditional beliefs about family and gender, and fewer experiences actually interacting with gay men and lesbians. These associations meant that Herek could predict people’s attitudes toward gay men and lesbians based on knowing their views about family, gender, and religion, as well as their past interactions with the stigmatized group. Herek’s primary contribution to the literature in this paper was the insight that reducing stigma toward gay men and lesbians “may require confronting deeply held, socially reinforced values” (1988, p. 473). And this insight was possible only because people were asked to report these values directly.

4.1 Introduction to Survey Research

Whether you are aware of it or not, you have been encountering survey research for most of your life. Every time your telephone rings during dinnertime, and the person on the other end of the line insists on knowing your household income and favorite brand of laundry detergent, he or she is helping to conduct survey research. When news programs try to predict the winner of an election two weeks early, these reports are based on survey research of eligible voters. In both cases, the researcher is trying to make predictions about the products people buy or the candidates they will elect based on people’s reports of their own attitudes, feelings, and behaviors.

TX_

TX

new85743_04_c04_169-212.indd 170 6/18/13 12:02 PM

171

CHAPTER 4Section 4.1 Introduction to Survey Research

Surveys can be used in a variety of contexts and are most appropriate for questions that involve people describing their attitudes, their behaviors, or a combination of the two. For example, if you wanted to examine the predictors of attitudes toward the death penalty, you could ask people their opinions on this topic and also ask them about their political party affiliation. Based on these responses, you could test whether political affil- iation predicted attitudes toward the death penalty. Or, imagine you wanted to know whether students who spent more time studying were more likely to do well on their exams. This question could be answered using a survey that asked students about their study habits and then tracked their exam grades. We will return to this example near the end of the chapter as we discuss the process of analyzing survey data to test our hypoth- eses about predictions.

The common thread running through these two examples is that they require people to report either their thoughts (e.g., opinions about the death penalty) or their behaviors (e.g., the hours they spend studying). Thus, in deciding whether a survey is the best fit for your research question, the key is to consider whether people will be both able and willing to report these things accurately. We will expand on both of these issues in the next section.

In this chapter, we continue our journey along the continuum of control, moving on to survey research, in which the primary goal is either describing or predicting attitudes and behavior. For our purposes, survey research refers to any method that relies on people’s reports of their own attitudes, feelings, and behaviors. So, for example, in Herek’s (1988) study, the participants reported their attitudes toward lesbians and gay men, rather than these attitudes being somehow directly observed by the researchers. Compared with the qualitative and descriptive designs for observing behavior we discussed in Chapter 3, survey research tends to yield more control over both data collection and question con- tent. Thus, survey research falls somewhere between quantitative descriptive research (Chapter 3) and the explanatory research involved in experimental designs (Chapter 5). This chapter provides an overview of survey research from conceptualization through analysis. We will cover the types of research questions that are best suited to survey research and provide an overview of the decisions to consider in designing and conduct- ing a survey study. We will then cover the process of data collection, with a focus on selecting the people who will complete your survey. Finally, we will cover the three most common approaches for analyzing survey data, bringing us back full circle to addressing our research questions.

Distinguishing Features of Surveys

Survey research designs have three distinguishing features that set them apart from other designs. First, all survey research relies on either written or verbal self-reports of peo- ple’s attitudes, feelings, and behaviors. This means that researchers will ask participants a series of questions and record their responses. This approach has several advantages, including being relatively straightforward and allowing access to psychological processes (e.g., “Why do you support candidate X?”). However, researchers are also cautious in their interpretation of self-reported data because participants’ responses can reflect a combina- tion of their true attitude and their concern over how this attitude will be perceived. Scien- tists refer to this as social desirability, which means that people may be reluctant to report

new85743_04_c04_169-212.indd 171 6/18/13 12:02 PM

172

CHAPTER 4Section 4.1 Introduction to Survey Research

unpopular attitudes. So if you were to ask people their attitudes about different racial groups, their answers might reflect both their true attitude and their desire not to appear racist. We return to the issue of social desirability and discuss some tricks for designing questions that can help to sidestep these concerns and capture respondents’ true attitudes.

The second distinguishing feature of survey research is that it has the ability to access internal states that cannot be measured through direct observation. In our discussion of observational designs in Chapter 3, we learned that one of the limitations of these designs was a lack of insight into why people do what they do. Survey research is able to address this limitation directly: By asking people what they think, how they feel, and why they behave in certain ways, researchers come closer to capturing the underlying psychologi- cal processes.

However, people’s reports of their internal states should be taken with a grain of salt, for three reasons. First, as mentioned, these reports may be biased by social desirability concerns, particularly when unpopular attitudes are involved. Second, there is a large literature in social psychology suggesting that people may not be very accurate at under- standing the true reasons for their behavior. In a highly cited review paper, psychologists Richard Nisbett and Tim Wilson (1977) argued that we make poor guesses about why we do things, and those guesses are based more on our assumptions than on any real intro- spection. Thus, survey questions can provide access to internal states, but these should always be interpreted with caution. Third, on a more practical note, survey research allows us to collect large amounts of data with relatively little effort and few resources. However, their actual efficiency depends on the decisions made during the design process. In reality, efficiency is often in a delicate balance with the accuracy and completeness of the data.

Broadly speaking, survey research can be conducted using either verbal or written self- reports (or a combination of the two). Before we dive into the details of writing and for- matting a survey, it is important to understand the pros and cons of administering your survey as an interview (i.e., an oral survey) or a questionnaire (i.e., a written survey).

Interviews

An interview involves an oral question-and-answer exchange between the researcher and the participant. This exchange can take place either face-to-face or over the phone. So our telemarketer example from earlier in the chapter represents an interview because the questions are asked orally, via phone. Likewise, if you are approached in a shopping mall and asked to answer questions about your favorite products, you are experiencing a sur- vey in interview form because the questions are administered out loud. And, if you have ever taken part in a focus group, in which a group of people gives their reactions to a new product, the researchers are essentially conducting an interview with the group. (For a more in-depth discussion of focus groups and other interview techniques, see Chapter 3, Section 3.2, Qualitative Research Interviews.)

Interview Schedules Regardless of how the interview is administered, the interviewer (i.e., the researcher) has a predetermined plan, or script, for how the interview should go. This plan, or script, for the progress of the interview is known as an interview schedule. When conducting an inter- view—including those telemarketing calls—the researcher/interviewer has a detailed

new85743_04_c04_169-212.indd 172 6/18/13 12:02 PM

173

CHAPTER 4Section 4.1 Introduction to Survey Research

plan for the order of questions to be asked, along with follow-up questions depending on the participant’s responses.

Broadly speaking, there are two types of interview schedules. A linear (also called “struc- tured”) schedule will ask the same questions in the same order for all participants. In contrast, a branching schedule unfolds more like a flowchart, with the next question dependent on participants’ answers. A branching schedule is typically used in cases with follow-up questions that make sense only for some of the participants. For example, you might first ask people whether they have children; if they answer “yes,” you could then follow up by asking how many.

One danger in using a branching schedule is that it is based partly on your assumptions about the relationships between variables. Granted, it is fairly uncontroversial to ask only people with children to indicate how many children they have. But imagine the follow- ing scenario in which you first ask participants for their household income, and then ask about their political donations:

• “How much money do you make? $18,000? Okay, how likely are you to donate money to the Democratic Party?”

• “How much money do you make? $250,000? Okay, how likely are you to donate money to the Republican Party?”

The assumption implicit in the way these questions branch is that wealthier people are more likely to be Republicans and less wealthy people to be Democrats. This might be supported by the data or it might not. But by planning the follow-up questions in this way, you are unable to capture cases that do not fit your stereotypes (i.e., the wealthy Democrats and the poor Republicans). The lesson here is to be careful about letting your biases shape the data collection process, as this can create invalid or inaccurate findings.

Advantages and Disadvantages of Interviews Spoken interviews have a number of advantages over written surveys. For one, people are often more motivated to talk than they are to write. Let’s say that an undergraduate research assistant is dispatched to a local shopping mall to interview people about their experiences in romantic relationships. The researcher may have no trouble at all recruiting par- ticipants, many of whom will be eager to divulge the personal details about their recent relationships. But for better or for worse, these experi- ences will be more difficult for the researcher to capture in writing. Related to this observation, people’s oral responses are typically richer and more detailed than their written

Alina Solovyova-Vincent/E1/Getty Images

Conducting interviews may allow a researcher to gather more detailed and richer responses.

new85743_04_c04_169-212.indd 173 6/18/13 12:02 PM

174

CHAPTER 4Section 4.1 Introduction to Survey Research

responses. Think of the difference between asking someone to “Describe your views on gun control” versus “Indicate on a scale of 1 to 7 the degree to which you support gun con- trol.” The former is more likely to capture the richness and subtlety involved in people’s attitudes about guns.

On a practical note, using an interview format also allows you to ensure that respondents understand the questions. If written questionnaire items are poorly worded, people are forced to guess at your meaning, and these guesses introduce a big source of error vari- ance (variance from random sources that are irrelevant to the trait or ability the question- naire is purporting to measure). But if an interview question is poorly asked, people find it much easier to ask the interviewer to clarify. Finally, using an interview format allows you to reach a broader cross-section of people and to include those who are unable to read and write—or, perhaps, unable to read and write the language of your survey.

Interviews also have three clear disadvantages compared with written surveys. First, interviews are more costly in terms of both time and money. It certainly used more of my time to go to a shopping mall than it would have taken to mail out packets of surveys (but no more money—these research assistant gigs tend to be unpaid!). Second, the inter- view format allows many opportunities to glean personal bias from the interview. These biases are unlikely to be deliberate, but participants can often pick up on body language and subtle facial expressions when the interviewer disagrees with their answers. These cues may lead them to shape their responses in order to make the interviewer happier (the influence of social desirability again). Third, interviews can be difficult to score and interpret, especially with open-ended questions. Although administering them may be easy, scoring them is relatively more complicated, often involving subjectivity or bias in the interpretation. Because the researcher often has to make judgments based on personal beliefs about the quality of the response, multiple raters are generally used to score the responses in order to minimize bias.

The best way to understand the pros and cons of interviewing is that both are a con- sequence of personal interactions. The interaction between interviewer and interviewee allows for richer responses but also the potential for these responses to be biased. As a researcher, you have to weigh these pros and cons and decide which method is the best fit for your survey. In the next section, we turn our attention to the process of administering surveys in writing.

Questionnaires

A questionnaire is a survey that involves a written question-and-answer exchange between the researcher and the participant. The questionnaire can be in open-ended for- mat (e.g., the participant writes in his or her answer) or forced-choice response format (e.g., the participant selects from a set of responses, such as with multiple choice ques- tions, rating scales, or true/false questions), which will be discussed later in this chapter. The exchange is a bit different from what we saw with interview formats. In written sur- veys, the questions are designed ahead of time and then given to participants, who write their responses and return the questionnaire to the researcher. In the next section, we will discuss details for designing these questions. But before we get there, let’s take a quick look at the process of administering written surveys.

new85743_04_c04_169-212.indd 174 6/18/13 12:02 PM

175

CHAPTER 4Section 4.1 Introduction to Survey Research

Distribution Methods Questionnaires can be distributed in three primary ways, each with its own pattern of advantages and disadvantages.

Distributing by Mail

Until recently, one common way to distribute surveys was to send paper copies through the mail to a group of participants (see Section 4.3, Sampling From the Population, for more discussion on how this group is selected). Mailing surveys is relatively inexpensive and relatively easy to do, but is unfortunately one of the worst methods when it comes to response rates. People tend to ignore questionnaires they receive in the mail, dismissing them as one more piece of junk. There are a few tricks available to researchers to increase response rates, including providing incentives, making the survey interesting, and mak- ing it as easy as possible to return the results (e.g., with a postage-paid envelope). How- ever, even using all of these tricks, researchers consider themselves extremely lucky to get a 30% response rate from a mail survey. That is, if you mail 1,000 surveys, you will be doing well to receive 300 back. Because of this low return on investment, researchers have begun using other methods for their written surveys.

Distributing in Person

Another option is to distribute a written survey in person, simply handing out copies and asking participants to fill them out on the spot. This method is certainly more time- consuming, as a researcher has to be stationed for long periods of time in order to collect data. In addition, people are less likely to answer the questions honestly because the presence of a researcher makes them worry about social desirability. Last, the sample for this method is limited to people who are in the physical area at the time that ques- tionnaires are being handed out. As we will discuss later, this might lead to problems in the composition of the sample. On the plus side, however, this method tends to result in higher compliance rates because it is harder to say no to someone face-to-face than it is to ignore a piece of mail.

Distributing Online

Over the past 20 years, Internet, or Web–based, surveys have become increasingly common. In Web-based survey research, the questionnaire is designed and posted on a Web page, to which participants are directed in order to complete the questionnaire. The advantages of online distribution are clear: This method is easiest for both researchers and participants and may give people a greater sense of anonymity, thereby encouraging more honest responses. In addition, response times are faster and the data are easier to analyze because they are already in digital format. The disadvantages include the fol- lowing: Specific groups being underrepresented because they do not have access to the Internet, the researcher has little to no control over sample selection, and the researcher receives responses only from those who are interested in the topic—so-called self- selection bias. All these limitations could raise questions about the validity and reliability of the data collected. In addition, several ethical issues might arise regarding informed con- sent and the privacy of participants. So when considering conducting Web-based surveys, researchers should evaluate all the advantages and disadvantages, as well as any ethical or legal implications.

new85743_04_c04_169-212.indd 175 6/18/13 12:02 PM

176

CHAPTER 4Section 4.2 Designing Questionnaires

For readers interested in more information on designing and conducting Internet research, Sam Gosling and John Johnson’s 2010 book Advanced Methods for Conducting Online Behavioral Research is an excellent resource. In addition, several groups of psychological researchers have been attempting to understand the psychology of Internet users. (You can read about recent studies on this website: http://www.spring.org.uk/2010/10/internet -psychology.php.)

Advantages and Disadvantages of Questionnaires Just as with interview methods, written questionnaires have their own set of advantages and disadvantages. Written surveys allow researchers to collect large amounts of data with little cost or effort, and they can offer a greater degree of anonymity than interviews. Anonymity can be a particular advantage in dealing with sensitive or potentially embar- rassing topics. That is, people may be more willing to answer a questionnaire about their alcohol use or their sexual history than they would be to discuss these things face-to-face with an interviewer. On the downside, written surveys miss out on the advantages of interviews because no one is available to clarify confusing questions or to gather more information as needed. Fortunately, there is one relatively easy way to minimize this prob- lem: Write questions (and response choices, if using multiple choice or forced choice for- mats) that are as clear as possible. In the next section, we turn our attention to the process of designing questionnaires.

4.2 Designing Questionnaires

One of the most important elements when conducting survey research is deciding how to construct and assemble the questionnaire items. In some cases, you will be able to use questionnaires that other researchers have developed in order to answer your research questions. For example, many psychology researchers use standard scales that measure behavior or personality traits, such as self-esteem, prejudice, depres- sion, or stress levels. The advantage of these ready-made measures is that other people have already gone to the trouble of making sure they are valid and reliable. So if you are interested in the relationship between stress and depression, you could distribute the Perceived Stress Scale (Cohen, Kamarck, & Mermelstein, 1983) and the Beck Depression Inventory (Beck, Steer, Ball, & Ranieri, 1996) to a group of participants and move on to the fun part of data analyses. For further discussion, see Chapter 2, Section 2.2, Reliability and Validity.

However, in many cases there is no perfect measure for your research question— either because no one has studied the topic before or because the current measures do not accurately assess the construct of interest you are investigating. When this happens, you will need to go through the process of designing your own questionnaire. In this section, we discuss strategies for writing questions and choosing the most appropriate response format.

new85743_04_c04_169-212.indd 176 6/18/13 12:02 PM

177

CHAPTER 4Section 4.2 Designing Questionnaires

Five Rules for Designing Better Questionnaires

Each of the following rules is designed to make your questions as clear and easy to under- stand as possible in order to minimize the potential for error variance. We discuss each one and illustrate them with contrasting pairs of items, consisting of “bad” items that do not follow the rule, and then “better” items that do.

1. Use simple language. One of the simplest and most important rules to keep in mind is that people have to be able to understand your questions. This means you should avoid jargon and specialized language whenever possible.

BAD: “Have you ever had an STD?”

BETTER: “Have you ever had a sexually transmitted disease?”

BAD: “What is your opinion of the S-CHIP?”

BETTER: “What is your opinion of the State Children’s Health Insurance Program?”

It is also a good idea to simplify the language as much as possible so that people spend time answering the ques- tion rather than trying to decode your meaning. For example, words like assist and consider can be replaced with words like help and think. This may seem odd—or perhaps even conde- scending to your participants—but it is always better to err on the side of sim- plicity. Remember, if people are forced to guess at the meaning of your ques- tions, these guesses add error variance to their answers.

In addition, when developing and administering surveys to various pop- ulations, it is important to remember to use language and examples that are

not culturally or linguistically biased. For example, participants from various cultures may not be familiar with questions involving the historical development of the United States or the nature of gender roles in the United States. Thus, with the changing U.S. population and the different languages being spoken in the home, it is important to use language that does not discriminate against particular populations. Also, it is advisable to avoid slang at all costs.

iStockphoto/Thinkstock

Simple language is one characteristic of an effective questionnaire.

new85743_04_c04_169-212.indd 177 6/18/13 12:02 PM

178

CHAPTER 4Section 4.2 Designing Questionnaires

2. Be precise. Another way to ensure that people understand the question is to be as precise as possible in your wording. Questions that are ambiguous in their wording will introduce an extra source of error variance into your data because people may interpret these questions in varying ways.

BAD: “What kind of drugs do you take?” (Legal drugs? Illegal drugs? Now? In college?)

BETTER: “What kind of prescription drugs are you currently taking?”

BAD: “Do you like sports?” (Playing? Watching? Which sports?)

BETTER: “How much do you like watching basketball on television?”

3. Use neutral language. It is important that your questions be designed to mea- sure your participants’ attitudes, feelings, or behaviors rather than to manipu- late their responses. That is, you should avoid leading questions, which are written in such a way that they suggest an answer.

BAD: “Do you beat your children?” (Clearly, beating isn’t good; who would say yes?)

BETTER: “Is it acceptable to use physical forms of discipline?”

BAD: “Do you agree that the president is an idiot?” (Hmmm . . . I wonder what the researcher thinks. . . .)

BETTER: “How would you rate the president’s job performance?”

This guideline can also be used to sidestep social desirability concerns. If you suspect that people may be reluctant to report holding an attitude such as using corporal punish- ment with their children, it helps to phrase the question in a nonthreatening way—“using physical forms of discipline” versus “beating your children.” Many current measures of prejudice adopt this technique. For example, McConahay’s Modern Racism Scale contains items such as “Discrimination against Blacks is no longer a problem in the United States” (McConahay, 1986). People who hold prejudicial attitudes are more likely to agree with statements like this one than with more blunt ones like “I hate people from Group X.”

4. Ask one question at a time. One remarkably common error that people make in designing questions is to include a double-barreled question, which fires off more than one question at a time. When you fill out a new patient question- naire at your doctor’s office, these forms often ask whether you suffer from “headaches and nausea.” What if you suffer from only one of these? Or what if you have a lot of nausea and only an occasional headache? The better approach is to ask about each of these symptoms separately.

BAD: “Do you suffer from pain and numbness?”

BETTER: “How often do you suffer from pain?” “How often do you suffer from numbness?”

BAD: “Do you like watching football and boxing?”

BETTER: “How much do you enjoy watching professional football on TV?” “How much do you enjoy watching boxing on TV?”

new85743_04_c04_169-212.indd 178 6/18/13 12:02 PM

179

CHAPTER 4Section 4.2 Designing Questionnaires

5. Avoid negations. One final and simple way to clarify your questions is to avoid questions with negative statements because these can often be difficult to understand. In the following examples, the first is admittedly a little silly, but the second comes from a real survey of voter opinions.

BAD: “Do you never not cheat on your exams?” (Wait—what? Do I cheat? Do I not cheat? What is this asking?)

BETTER: “Have you ever cheated on an exam?”

BAD: “Are you against rejecting the ban on pesticides?” (Wait—so am I for the ban? Against the ban? What is this asking?)

BETTER: “Do you support the current ban on pesticides?”

Participant Response Options

In this section, we turn our attention to the issue of deciding how participants should respond to your questions. The decisions you make at this stage affect the type of data you ultimately collect, so it is important to choose carefully. We also review the primary decisions you will need to make about response options, as well as the pros and cons of each one.

One of the first choices you have to make is whether to collect open-ended or fixed- format responses. As the names imply, fixed-format responses require participants to choose from a list of options (e.g., “Pick your favorite color”), whereas open-ended responses ask participants to provide unstructured responses to a question or statement (e.g., “How do you feel about legalizing marijuana?”). Open-ended responses tend to be richer and more flexible but harder to translate into quantifiable data—analogous to the trade-off we discussed in comparing written versus oral survey methods. To put it another way, some concepts are difficult to reduce to a seven-point fixed-format scale, but these scales are easier to analyze than a paragraph of free-flowing text.

Another reason to think carefully about this decision is that fixed-format responses will, by definition, restrict people’s options in answering the question. In some cases, these restrictions can even act as leading questions. In a study of people’s perceptions of his- tory, James Pennebaker and his colleagues (2006) asked respondents to indicate the “most significant event over the last 50 years.” When this was asked in an open-ended way (i.e., “list the most significant event”), 2% of participants listed the invention of computers. In another version of the survey, this question was asked in a fixed-format way (i.e., “choose the most significant event”). When asked to select from a list of four options (World War II, Invention of Computers, Tiananmen Square, or Man on the Moon), 30% chose the inven- tion of computers! In exchange for having easily coded data, the researchers accidentally forced participants into a smaller number of options and ended up with a skewed sense of the importance of computers in people’s perceptions of history.

Fixed-Format Options Although fixed-format responses can sometimes constrain or skew participants’ answers, the reality is that researchers tend to use them more often than not. This decision is largely practical; fixed-format responses allow for more efficient data collection from a

new85743_04_c04_169-212.indd 179 6/18/13 12:02 PM

180

CHAPTER 4Section 4.2 Designing Questionnaires

much larger sample. (Imagine the chore of having to hand-code 2,000 essays!) But once you have decided on this option for your questionnaire, the decision process is far from over. In this section, we discuss three possibilities you can use to construct a fixed-format response scale.

True/False

One option is to ask questions using a true/false format, which asks participants to indicate whether they endorse a statement. For example:

“I attended church last Sunday” True False

“I am a U.S. citizen” True False

“I am in favor of abortion” True False

This last example may strike you as odd, and in fact illustrates an important point about the use of true/false formats: They are best used for statements of fact rather than attitudes. It is relatively straightforward to indicate whether you attended church or whether you are a U.S. citizen. However, people’s attitudes toward abortion are often complicated— one can be “pro-choice” but still support some common-sense restrictions, or “pro-life” but support exceptions (e.g., in cases of rape). The point is that a true/false question can- not even come close to capturing this complexity. However, for survey items that involve simple statements of fact, the true/false format can be a good option.

Multiple Choice

A second option is to use a multiple-choice format, which asks participants to select from a set of predetermined responses.

“Which of the following is your favorite fast-food restaurant?” a. McDonald’s b. Burger King c. Wendy’s d. Taco Bell

“Who did you vote for in the 2008 presidential election?” a. John McCain b. Barack Obama

“How do you travel to work on most days? (Select all that apply.)” a. drive alone b. carpool c. take public transportation

As you can see in these examples, multiple-choice questions give you quite a bit of free- dom in both the content and the response scaling of your questions. You can ask partici- pants either to select one answer or, as in the last example, to select all applicable answers. You can cover everything from preferences (e.g., favorite fast-food restaurant) to behav- iors (e.g., how you travel to work).

new85743_04_c04_169-212.indd 180 6/18/13 12:02 PM

181

CHAPTER 4Section 4.2 Designing Questionnaires

In addition to assessing preferences or behaviors, as in the preceding examples, multiple- choice questions are often used to examine knowledge and abilities. For example, intel- ligence and achievement tests utilize multiple-choice questions to assess the cognitive and academic abilities of individuals. In these cases, there is only one correct response choice and three to four incorrect or “distracter” response choices. Most research indicates that there should be at least three but no more than four incorrect response choices. Providing more than four incorrect response choices can make the question too difficult. It is also important that the correct response choice not obviously differ from the incorrect choices. Thus, all response choices, both correct and incorrect, should be approximately the same length and be plausible choices. If the incorrect response choices are obviously wrong, this will make the item too easy, which will affect its validity. In addition, including obviously wrong answers allows those individuals who do not know the answer to deduce the cor- rect response.

Regardless of whether you are developing multiple-choice questions to assess behaviors, knowledge, or abilities, all questions should be clear, unambiguous, and brief so that they can be easily understood. In addition, as with all types of question formats, it is best to avoid negative questions such as, “Which of the following is not correct?” because they can be confusing to test-takers and cause the question to be invalid (i.e., not measure what it is supposed to be measuring). Finally, it is best practice to avoid response choices that include “All of the above” or “None of the above.” Most individuals know that when such response choices are provided, they are usually the correct answer.

You may have already spotted a downside to multiple-choice formats. Whenever you pro- vide a set of responses, you are restricting participants to those choices. This is the problem that Pennebaker and colleagues encountered when asking people about the most signifi- cant events of the last century. In each of the preceding examples, the categories fail to cap- ture all possible responses. What if your favorite restaurant is In-and-Out Burger? What if you voted for Ralph Nader? What if you telecommute or ride your bicycle to work? There are two relatively easy ways to avoid (or at least minimize) this problem. First, plan care- fully when choosing the response options. During the design process, it helps to brain- storm with other people to ensure you are capturing the most likely responses. However, in many cases, it is almost impossible to provide every option that people might think of. The second solution is to provide an “other” response to your multiple-choice question. This allows people to write in an option that you neglected to include. For example, our last question about traveling to work could be rewritten as follows:

“How do you travel to work on most days? (Select all that apply.)” a. drive alone b. carpool c. take public transportation d. other (please specify):__________________

This way, people who telecommute, bicycle, or even ride their trained pony to work will have a way to respond rather than skipping the question. And, if you start to notice a pat- tern in these write-in responses (e.g., 20% of people adding “bicycle”), then you will have gained valuable knowledge to improve the next iteration of the survey.

new85743_04_c04_169-212.indd 181 6/18/13 12:02 PM

182

CHAPTER 4Section 4.2 Designing Questionnaires

Rating Scales

Last, but certainly not least, another option is to use a rating scale format, which asks participants to respond on a scale that represents a continuum.

“Sometimes it is necessary to sacrifice liberty in the name of security.”

1 2 3 4 5

strongly agree agree neither agree nor disagree disagree strongly disagree

“I would vote for a candidate who supported the death penalty.”

1 2 3 4 5

Always often about half of the time seldom never

“The political party in power right now has messed things up.”

1 2 3 4 5

strongly agree agree neither agree nor disagree disagree strongly disagree

This format is well suited to capturing attitudes and opinions and, in fact, is one of the most common approaches to attitude research. Rating scales are easy to score, and they give participants some flexibility in indicating their agreement with or endorsement of the questions. As a researcher, you have two critical decisions to make about the construc- tion of rating scale items. Both have implications for how you will analyze and interpret your results.

First, you’ll need to decide on the anchors, or labels, for your response scale. Rating scales offer a good deal of flexibility in these anchors, as you can see in the preceding examples. You can frame questions in terms of “agreement” with a statement or “likelihood” of a behavior; alternatively, you can customize the anchors to match your question (e.g., “not at all necessary”). Scales that use anchors of strongly agree and strongly disagree are also referred to as Likert scales. At a fairly simple level, the choice of labels affects the interpre- tation of the results. For example, if you asked the “political party” question, you would have to be aware that the anchors were phrased in terms of agreement with the state- ment. In discussing these results, you would be able to discuss how much people agreed with the statement, on average, and whether agreement correlated with other factors. If this seems like an obvious point, you would be amazed at how often researchers (or the media) will take an item like this and spin the results to talk about the “likelihood of vot- ing” for the party in power—confusing an attitude with a behavior! So, in short, make sure you are being honest when presenting and interpreting research data.

new85743_04_c04_169-212.indd 182 6/18/13 12:02 PM

183

CHAPTER 4Section 4.2 Designing Questionnaires

At a more conceptual level, you need to decide whether the anchors for your rating scale make use of a bipolar scale, which has polar opposites at its endpoints and a neutral point in the middle, or a unipolar scale, which assesses the presence or absence of a single con- struct. The difference between these scales is best illustrated by an example:

Bipolar: How would you rate your current mood?

1 2 3 4 5 6 7

very happy happy slightly happy neither happy nor sad slightly sad sad very sad

Unipolar: How would you rate your current mood?

1 2 3 4 5

not at all sad slightly sad moderately sad very sad completely sad

1 2 3 4 5

not at all happy slightly happy moderately happy very happy completely happy

With the bipolar option, participants are asked to place themselves on a continuous scale somewhere between sad and happy, which are polar opposites. The assumption in using a bipolar scale is that the endpoints represent the only two options—participants can be sad, happy, or somewhere in between. In contrast, with the unipolar option, participants are asked to rate themselves on a continuous scale, indicating their level of either sadness or happiness. The assumption in using a pair of unipolar scales is that it is possible to experi- ence varying degrees of each item: For example, participants can be moderately happy but also a little bit sad. The decision to use a bipolar or a unipolar scale comes down to the context. What is the most logical way to think about these constructs? What have previous researchers done?

In the 1970s, Sandra Lipsitz Bem revolutionized the way researchers thought about gen- der roles by arguing against a bipolar approach. Previously, gender role identification had been measured on a bipolar scale from “masculine” to “feminine,” the assumption being that a person could be one or the other. Bem (1974) argued instead that people could eas- ily have varying degrees of masculine and feminine traits. Her scale, the Bem Sex Role Inventory, asks respondents to rate themselves on a set of 60 unipolar traits. Someone with mostly feminine and hardly any masculine traits would be described as “feminine.” Someone with high ratings on both masculine and feminine traits would be described as “androgynous.” And, someone with low ratings on both masculine and feminine traits would be described as “undifferentiated.” You can view and complete Bem’s scale online at this website: http://garote.bdmonkeys.net/bsri.html.

new85743_04_c04_169-212.indd 183 6/18/13 12:02 PM

184

CHAPTER 4Section 4.2 Designing Questionnaires

The second critical decision in constructing a rating scale item is to decide on the number of points in the response scale. You may have noticed that all the examples in this section have an odd number of points (e.g., five or seven). This is usually preferable for rating scale items because the middle of the scale (e.g., “3” or “4”) allows respondents to give a neutral, middle-of-the-road answer. That is, on a scale from strongly disagree to strongly agree, the midpoint can be used to indicate “neither” or “I’m not sure.” However, in some cases, you may not want to allow a neutral option in your scale. By using an even number of points (e.g., four or six), you can essentially force people to either agree or disagree with the statement; this type of scaling is referred to as forced choice.

So how many points should your scale have? As a general rule, more points will translate into more variability in responses—the more choice people have, the more likely they are to distribute their responses among those choices. From a researcher’s perspective, the big question is whether this variability is meaningful. For example, if you wanted to assess college students’ attitudes about a student fee increase, opinions will likely vary depend- ing on the size of the fee and the ways in which it will be used. Thus, a five- or seven- point scale would be preferable to a two-point (yes or no) scale. However, past a certain point, increasing the scale range ceases to be linked to meaningful variation in attitudes. In other words, the difference between a 5 and a 6 on a seven-point scale is fairly intuitive for your participants to grasp. But what is the real difference between an 80 and an 81 on a 100-point scale? When scales become too large, you risk introducing another source of error variance as participants impose their interpretations on the scaling. In sum, more points do not always translate into a better scale.

Back to our question: How many points should you have? The ideal compromise sup- ported by many statisticians is to use a seven-point scale for bipolar scales. The reason has to do with the differences between scales of measurement. As you’ll remember from our discussion in Chapter 2, the way variables are measured has implications for data analyses. For the most popular statistical tests to be legitimate, variables need to lie on either an interval scale (i.e., with equal intervals between points) or a ratio scale (i.e., with a true zero point). Based on mathematical modeling research, statisticians have concluded that the variability generated by a seven-point scale is most likely to mimic an interval scale (see, e.g., Nunnally, 1978). So a seven-point scale is often preferable because it allows us the most flexibility in data analyses. Note, however, that whereas seven-point scales produce reliable results with bipolar scales, unipolar scales tend to perform the best with five-point scales.

Finalizing the Questionnaire

Once you have finished constructing the questionnaire items, one last important step remains before beginning to collect data. This section discusses a few guidelines for assem- bling the items into a coherent questionnaire. The main issues at this stage are to think carefully about the order of the individual items; how many items to include; whether to include open-ended, fixed-format, or multiple choice questions; and writing the instruc- tions in clear and concise language.

new85743_04_c04_169-212.indd 184 6/18/13 12:02 PM

185

CHAPTER 4Section 4.2 Designing Questionnaires

First, keep in mind that the first few questions will set the tone for the rest of the question- naire. It is best to start with questions that are both interesting and nonthreatening to help ensure that respondents complete the questionnaire with open minds. For example:

BAD OPENING: “Do you agree that your child’s teacher is incompetent?” (threatening and also a leading question)

BETTER OPENING: “How would you rate the performance of your child’s teacher?”

BAD OPENING: “Would you support a 1% sales tax increase?” (boring)

BETTER OPENING: “How do you feel about raising taxes to help fund education?”

Second, strive whenever possible to have continuity in the different sections of your ques- tionnaire. Imagine you are constructing a survey to give to college freshmen—you might have questions on family background, stress levels, future plans, campus engagement, and so on. It is best to have the questions grouped by topic on the survey. So, for instance, students would fill out a set of questions about future plans on one page and then a set of questions about campus engagement on another page. This approach makes it eas- ier for participants to progress through the questions without having to mentally switch between topics.

Third, remember that individual questions are always read in context. This means that if you start your college student survey with a question about plans for the future and then ask about stress, respondents will likely have their future plans in mind when they think about their level of stress. Another example is a graduate school that would administer a gigantic survey packet to every student enrolled in its Introductory Psychology course. One year, a faculty member included a measure of identity, asking participants to com- plete the statements “I am ______” and “I am not ______.” As the students started to ana- lyze data from this survey, they found that an astonishing 60% of students had filled in the blank with “I am not a homosexual.” This response seemed pretty unusual until they real- ized that the questionnaire immediately preceding this one in the packet was a measure of prejudice toward gay and lesbian individuals. So, as these students completed the identity measure, they had homosexuality on their minds and felt compelled to point out that they were not homosexual. This proves once again that results can be skewed by the context.

Finally, once you have assembled a draft version of your questionnaire, do a test run. This test run, called pilot testing, involves giving the questionnaire to a small sample of people, getting their feedback, and making any necessary changes. One of the best ways to pilot test is to find a patient group of friends to complete your questionnaire because this group will presumably be willing to give more extensive feedback. Another effective way to pilot test a questionnaire is to administer it to individuals in the target group. In soliciting their feedback, you should ask questions like the following:

Was anything confusing or unclear?

Was anything offensive or threatening?

new85743_04_c04_169-212.indd 185 6/18/13 12:02 PM

186

CHAPTER 4Section 4.2 Designing Questionnaires

How long did the questionnaire take you to complete?

Did it get repetitive or boring? Did it seem too long?

Were there particular questions that you liked or disliked? Why?

The answers to these questions will give you valuable information to revise and clarify your questionnaire before devoting the resources for a full round of data collection. In the next section, we turn our attention to the question of how to find and select participants for this stage of the research.

Research: Thinking Critically

“Beautiful People Convey Personality Traits Better During First Impressions”

Medical News Today

A new University of British Columbia study has found that people identify the personality traits of people who are physically attractive more accurately than others during short encounters.

The study, published in the December [2010] edition of Psychological Science, suggests people pay closer attention to people they find attractive, and is the latest scientific evidence of the advantages of perceived beauty. Previous research has shown that individuals tend to find attractive people more intelligent, friendly, and competent than others.

The goal of the study was to determine whether a person’s attractiveness impacts others’ ability to discern their personality traits, says Prof. Jeremy Biesanz, UBC Dept. of Psychology, who coauthored the study with PhD student Lauren Human and undergraduate student Genevieve Lorenzo.

For the study, researchers placed more than 75 male and female participants into groups of five to 11 people for three-minute, one-on-one conversations. After each interaction, study participants rated partners on physical attractiveness and five major personality traits: openness, conscientious- ness, extraversion, agreeableness, and neuroticism. Each person also rated his or her own personality.

Researchers were able to determine the accuracy of people’s perceptions by comparing participants’ ratings of others’ personality traits with how individuals rated their own traits, says Biesanz, adding that steps were taken to control for the positive bias that can occur in self-reporting.

Despite an overall positive bias towards people they found attractive (as expected from previous research), study participants identified the “relative ordering” of personality traits of attractive par- ticipants more accurately than others, researchers found.

“If people think Jane is beautiful, and she is very organized and somewhat generous, people will see her as more organized and generous than she actually is,” says Biesanz. “Despite this bias, our study shows that people will also correctly discern the relative ordering of Jane’s personality traits—that she is more organized than generous—better than others they find less attractive.”

The researchers say this is because people are motivated to pay closer attention to beautiful people for many reasons, including curiosity, romantic interest, or a desire for friendship or social status. “Not only do we judge books by their covers, we read the ones with beautiful covers much closer than others,” says Biesanz, noting the study focused on first impressions of personality in social situ- ations, like cocktail parties.

Although participants largely agreed on group members’ attractiveness, the study reaffirms that beauty is in the eye of the beholder. Participants were best at identifying the personalities of people they found attractive, regardless of whether others found them attractive.

(continued)

new85743_04_c04_169-212.indd 186 6/18/13 12:02 PM

187

CHAPTER 4Section 4.3 Sampling From the Population

Research: Thinking Critically (continued) According to Biesanz, scientists spent considerable efforts a half-century ago seeking to determine what types of people perceive personality best, to largely mixed results. With this study, the team chose to investigate this long-standing question from another direction, he says, focusing not on who judges personality best, but rather whether some people’s personalities are better perceived.

Think about it:

1. Suppose the following questions were part of the questionnaire given after the 3-minute one-on-one conversations in this study. Based on the goals of the study and the rules dis- cussed in this chapter, identify the problem with each of the following questions and suggest a better item.

a. Jane is very neat.

1 2 3 4 5

strongly agree agree neither agree or disagree disagree strongly disagree

main problem:

better item:

b. Jane is generous and organized.

1 2 3 4 5

strongly agree agree neither agree or disagree disagree strongly disagree

main problem:

better item:

c. Jane is extremely attractive TRUE FALSE

main problem:

better item:

2. What are the strengths and weaknesses of using a fixed-format questionnaire in this study versus open-ended responses?

3. The researchers state that they took steps to control for the “positive bias that can occur in self- reporting.” How might social desirability influence the outcome of this particular study? What might the researchers have done to reduce the effect of social desirability?

George, R. (2010, December 31). Beautiful people convey personality traits better during first impressions. Medical News Today. Retrieved from http://www.medicalnewstoday.com/articles/212245.php

4.3 Sampling From the Population

By now, you should have a good feel for how to construct survey items. Once you have finalized your measures, the next step is to find a group of people to fill out the survey. But where do you find this group? And how many of them do you need? On the one hand, you want as many people as possible in order to capture the full range of attitudes and experiences. On the other hand, researchers have to conserve time and other

new85743_04_c04_169-212.indd 187 6/18/13 12:02 PM

188

CHAPTER 4Section 4.3 Sampling From the Population

resources, which often means choosing a smaller sample of people. In this section, we will examine the strategies researchers can use in selecting samples for their studies.

Researchers refer to the entire collection of people who could possibly be relevant for a study as the population. For example, if you were interested in the effects of prison overcrowding in this country, you would want to study the population of prisoners in the United States. If you wanted to study voting behavior in the next U.S. presidential elec- tion, your population would be United States residents eligible to vote. And if you wanted to know how well college students cope with the transition from high school, your popu- lation would include every college student who was graduated from high school and is now enrolled in any college in the country.

You may have spotted an obvious practical complication with these populations. How on earth are you going to get every college student, much less every prisoner, in the coun- try to fill out your questionnaire? You can’t; instead, researchers will collect data from a sample, a subset of the population. Instead of trying to reach all prisoners, you might sample inmates from a handful of state prisons. Rather than attempt to survey all college students in the country, researchers might restrict their studies to a collection of students at one university.

The goal in choosing a sample for quantitative research is to make it as representative as possible of the larger population. This is the goal, though it is not always practical. That is, if you choose students at one university, they need to be reasonably similar to college students elsewhere in the country. If the phrase “reasonably similar” sounds vague, this is because the basis for evaluating a sample varies depending on the hypothesis and the key variables of your study. For example, if you wanted to study the relationship between family income and stress levels, you would need to make sure that your sample mirrored the population in the distribution of income levels. Thus, a sample of students from a state university might be a better choice than students from, say, Harvard (which costs about $50,000 per year). On the other hand, if your research question dealt with the pressures faced by students in selective private schools, then Harvard students could be a represen- tative sample for your study.

Figure 4.1 is a conceptual illustration of both a representative and nonrepresentative sam- ple, drawn from a larger population. The population in this case consists of 144 individu- als, split evenly between Xs and Os. Thus, we would want our sample to come as close as possible to capturing this 50/50 split. The sample of 20 individuals on the left is rep- resentative of the sample because it is split evenly between Xs and Os. But the sample of 20 individuals on the right is nonrepresentative because it contains 75% Xs. Because there are far fewer Os than we might expect in the right-hand population, this sample does not accurately represent the population. This failure of the sample to represent the population is also referred to as sampling bias.

new85743_04_c04_169-212.indd 188 6/18/13 12:02 PM

189

CHAPTER 4Section 4.3 Sampling From the Population

Figure 4.1: Representative and nonrepresentative samples of a population

So, where do these samples come from? As a researcher, you have two broad categories of sampling strategies at your disposal: probability sampling and nonprobability sampling.

Probability Sampling

Probability sampling is used when each person in the population has a known chance of being in the sample. This is possible only in cases where you know the exact size of the population. For instance, the 2010 population of the United States was 308,745,538 (United States Census Bureau, 2010). If you were to have selected a U.S. resident at ran- dom, each resident would have had a one in 308,745,538 chance of being selected. When- ever you have information about total population, probability-sampling strategies are the most powerful approach because they greatly increase the odds of getting a representative sample. Within this broad category of probability sampling are three specific strategies: simple random sampling, stratified random sampling, and cluster sampling.

Simple Random Sampling Simple random sampling, the most straightforward approach, involves randomly pick- ing study participants from a list of everyone in the population. The term for this list is a sampling frame (e.g., imagine a list of every resident of the United States). To have a truly representative random sample, several criteria need to be met: You must have a sam- pling frame, you must choose from it randomly, and you must have a 100% response rate from those you select. (As we discussed in Chapter 2, it can threaten the validity of your hypothesis test if people drop out of your study.)

POPULATION (50% X’s, 50% O’s)

X X XO O OXO OX XO OX XO OX

O O OX X XOX XO OX XO OX XO

X X XO O OXO OX XO OX XO OX

O O OX X XOX XO OX XO OX XO

X X XO O OXO OX XO OX XO OX

O O OX X XOX XO OX XO OX XO

X X XO O OXO OX XO OX XO OX

O O OX X XOX XO OX XO OX XO

X X X X X X X X X X

O O O O O O O O O O X X X X X X X X X X

X X X X XO O O O O

REPRESENTATIVE SAMPLE (50% X’s, 50% O’s)

NONREPRESENTATIVE SAMPLE

(75% X’s, ONLY 25% O’s)

new85743_04_c04_169-212.indd 189 6/18/13 12:02 PM

190

CHAPTER 4Section 4.3 Sampling From the Population

Stratified Random Sampling Stratified random sampling, a varia- tion of simple random sampling, is used when subgroups of the popu- lation might be left out of a purely random sampling process. Imagine a city with a population that is 80% Caucasian, 10% Hispanic, 5% Afri- can American, and 5% Asian. If you were to choose 100 residents at ran- dom, the chances are very good that your entire sample would consist of Caucasian residents. As a result, you would inadvertently ignore the perspective of all ethnic minority residents. To prevent this problem, researchers use stratified random sampling—breaking the sampling frame into subgroups and then sam- pling a random number from each subgroup. In the preceding city example, you could divide your list of residents into four ethnic groups and then pick a random 25 from each of these groups. The result would be a sample of 100 people who captured opinions from each ethnic group in the population.

Cluster Sampling Cluster sampling, another variation of random sampling, is used when you do not have access to a full sampling frame (i.e., a full list of everyone in the population). Imagine that you wanted to do a study of how cancer patients in the United States cope with their ill- ness. Because there is not a list of every cancer patient in the country, you have to get a little creative with your sampling. The best way to think about cluster sampling is as “samples within samples.” Just as with stratified sampling, you divide the overall population into groups; however, cluster sampling is different in that you are dividing into groups based on more than one level of analysis. In the cancer example, you could start by dividing the country into regions, then randomly selecting cities from within each region, and then randomly selecting hospitals from within each city, and finally randomly selecting cancer patients from each hospital. The result would be a random sample of cancer patients from, say, Phoenix, Miami, Dallas, Cleveland, Albany, and Seattle; taken together, these patients would constitute a fairly representative sample of cancer patients around the country.

Nonprobability Sampling

The other broad category of sampling strategies is known as nonprobability sampling. These strategies are used in the (remarkably common) case in which you do not know the odds of any given individual being in the sample. This is an obvious shortcoming—if you do not know the exact size of the population and do not have a list of everyone in it, there is no way to know that your sample is representative! But despite this limitation, research- ers use nonprobability sampling on a regular basis.

iStockphoto/Thinkstock

Stratified random sampling allows researchers to include all subgroups of a population in a study.

new85743_04_c04_169-212.indd 190 6/18/13 12:02 PM

191

CHAPTER 4Section 4.3 Sampling From the Population

Nonprobability sampling is used in qualitative, quantitative, and mixed methods research whose focus is on selecting relatively small samples in a purposeful manner rather than selecting samples that are representative of the entire population. Unlike probability sam- pling, nonprobability sampling does not include randomly selected, stratified samples that provide for generalization. Rather, the procedures used to select samples are purpose- ful; that is, the samples are selected to obtain information-rich cases that yield in-depth insights. The sampling techniques used for quantitative and qualitative research probably make up one of the biggest differences between the two methods. The following sections will discuss some of the most common nonprobability strategies. All of the nonprobability strategies described (except convenience sampling) are considered categories of purpo- sive sampling, or sampling with a purpose. Convenience sampling is considered a form of so-called accidental or haphazard (serendipitous) sampling, since the cases are selected based on ready availability. Even so, convenience sampling can be purposeful, in that researchers do their best to recruit in ways that are convenient and yet attract individuals that meet meaningful criteria.

In many cases, it is not possible to obtain a sampling frame. When researchers study rare or hard-to-reach populations, or investigate potentially stigmatizing conditions, they often recruit by word of mouth. The term for this is snowball sampling—imagine a snowball rolling down a hill, picking up more snow (or participants) as it goes. If you wanted to study how often homeless people took advantage of social services, you would be hard- pressed to find a sampling frame that listed the homeless population. Instead, you could recruit a small group of homeless people and ask each of them to pass the word along to others, and so on. The resulting sample is unlikely to be representative, but researchers often have to compromise for the sake of obtaining access to a population.

One of the most popular nonprobability strategies is known as convenience sampling, or simply enrolling people who show up for the study. Any time you see results of a viewer poll on your favorite 24-hour news station, the results are likely based on a convenience sample. CNN and Fox News do not randomly select from a list of their viewers; they post a question on-screen or online, and people who are motivated enough to respond will do so. For that matter, a large majority of psychology research studies are based on conve- nience samples of undergraduate college students. Often, experimenters in psychology departments advertise their studies on a bulletin board or website, and students sign up for studies for extra cash or to fulfill a research requirement for a course. Students often pick a particular study based on whether it fits their busy schedules or whether the adver- tisement sounds interesting. Another example of convenience sampling is a researcher seeking to collect information from human resources managers to study the issue of bul- lying in organizations. The researcher might use a database of potential participants who belong to one or more chapters of the Society of Human Resource Management (SHRM), an organization for human resources professionals. In this example, the convenience com- ponent of the sample is using a body of existing data; although these data are not repre- sentative of all human resources managers or all organizations, they are a useful proxy for a broad sample of human resource managers across industries.

Other popular nonprobability strategies include extreme or deviant case sampling, typ- ical case sampling, heterogeneous sampling, expert sampling, criterion sampling, and theory-based sampling. Whatever strategy you decide on, the goal here is to make you mindful that all the decisions that you make as a researcher have inherent strengths and weaknesses.

new85743_04_c04_169-212.indd 191 6/18/13 12:02 PM

192

CHAPTER 4Section 4.3 Sampling From the Population

Extreme or deviant case sampling is used when we want information-rich data on unusual or special cases. For example, if we are interested in studying university diversity plans, it would be important to examine plans that work exceptionally well, as well as plans that have high expectations but are not working for some reason or another. The focus is on examining outstanding successes and failures to learn lessons about those conditions.

On the other hand, sometimes researchers are not interested in learning about unusual cases but rather want to learn about typical cases. Typical case sampling involves sam- pling the most frequent or “normal” case. It is often used to describe typical cases to people unfamiliar with a setting, program, or process. For example, a typical case sample may be used to study the practicum experiences of psychology students from universities that are rated average. One of the biggest drawbacks of this method involves the difficulty of knowing how to identify or define a typical or normal case.

Heterogeneous sampling, which is also called maximum variation sampling, may include both extreme and typical cases. It is used to select a wide variety of cases in relation to the phenomenon being investigated. The idea here is to obtain a sample that includes diverse characteristics from multiple dimensions. For example, if the researchers are interested in examining the experiences of community health clinic patients, they may wish to conduct a focus group and then select 10 different groups that represent various demographics from several health clinics in the area. Heterogeneous sampling is beneficial when it is desirable to view patterns across a diverse set of individuals. It is also valuable in describ- ing experiences that may be central or core to most individuals. Heterogeneous sampling can be problematic in small samples, though, because it is difficult to obtain a wide variety of cases using this method.

Expert sampling involves sampling a panel of individuals who have known expertise (knowledge and training) in a particular area. While expert sampling can be used in both qualitative and quantitative research, it is used more often in qualitative research. Expert sampling involves a process of identifying individuals with known expertise in an area of interest, obtaining informed consent from each expert, and then collecting information from them either individually or as a group.

The goal of criterion sampling is to select cases “that meet some predetermined criterion of importance” (Patton, 2002, p. 238). A criterion could be a particular illness or experience that is being investigated, or even a program or a situation. For example, criterion sam- pling could be used to investigate a range of topics: the experiences of individuals who have attempted suicide, why students are exceedingly absent from online course rooms, or cases that have exceeded the standard waiting time at a doctor’s office. The idea with this method is that the participants meet a particular criterion of interest.

Theory-based sampling is a version of criterion sampling that focuses on obtaining cases that represent theoretical constructs. As Patton (2002) discussed, theory-based sampling involves the researcher obtaining “sampling incidents, slices of life, time periods, or peo- ple on the basis of their potential manifestation or representation of important theoretical constructs” (p. 238). This type of sampling arose from grounded theory research and fol- lows a more deductive or theory-testing approach. As Glaser (1978) described, theory- based sampling is “the process of data collection for generating theory whereby the analyst

new85743_04_c04_169-212.indd 192 6/18/13 12:02 PM

193

CHAPTER 4Section 4.3 Sampling From the Population

jointly collects, codes, and analyzes his data and decides which data to collect next and where to find them” (p. 36). Thus, data collection is driven by the emerging theory, and participants are selected based on their knowledge of the topic.

Choosing a Sampling Strategy

Although quantitative researchers strive for representative samples, there is no such thing as a perfectly representative one. There is always some degree of sampling error, defined as the degree to which the characteristics of the sample differ from the characteristics of the population. Instead of aiming for perfection, then, researchers aim for an estimate of how far from perfection their samples are. These estimates are known as the error of esti- mation, or the degree to which the data from the sample are expected to deviate from the population as a whole.

One of the main advantages of a probability sample is that we are able to calculate these errors of estimation (or margins of error). In fact, you have likely encountered errors of estimation every time you see the results of an opinion poll. For example, CNN may report that “Candidate A is leading the race with 60% of the vote, 6 3%.” This means Candidate A’s percentage in the sample is 60%, but based on statistical calculations, her real percentage is between 57% and 63%. The smaller the error (3% in this example), the more closely the results from the sample match the population. Naturally, researchers conducting these opinion polls want the error of estimation to be as small as possible; imagine how nonpersuasive it would be to learn that “Candidate A has a 10-point lead, 6 20 points.” In general, these errors are minimized when three conditions are met: The overall population is smaller; the sample itself is larger; and there is less variability in the sample data. When samples are created using a probability method, all of this information is available because these methods require knowing the population.

If probability sampling is so powerful, why are nonprobability strategies so popular? One reason is that convenience samples are more practical; they are cheaper, easier, and almost always possible to enroll with relatively few resources because you can avoid the costs of large-scale sampling. A second reason is that convenience is often a good starting point for a new line of research. For example, if you wanted to study the predictors of relationship satisfaction, you could start by testing hypotheses in a controlled setting using college student participants, and then you could extend your research to the study of adult mar- ried couples. Finally, and relatedly, in many types of qualitative research, it is acceptable to have a nonrepresentative sample because you do not need to generalize your results. If you want to study the prevalence of alcohol use in college students, it may be perfectly acceptable to use a convenience sample of college students. Although, even in this case, you would have to keep in mind that you were studying drinking behaviors among stu- dents who volunteered to complete a study on drinking behaviors.

There are also cases, however, where it is critical to use probability sampling despite the extra effort it requires. Specifically, researchers use probability samples any time it is important to generalize and any time it is important to predict behavior of a popula- tion. The best example for understanding these criteria is to think of political polls. In the lead-up to an election, each campaign is invested in knowing exactly what the voting public thinks of its candidate. In contrast to a CNN poll, which is based on a convenience

new85743_04_c04_169-212.indd 193 6/18/13 12:02 PM

194

CHAPTER 4Section 4.3 Sampling From the Population

sample of viewers, polls conducted by a campaign will be based on randomly selected households from a list of registered voters. The resulting sample is much more likely to be representative, much more likely to tell the campaign how the entire population views its candidate, and therefore, much more likely to be useful.

Determining a Sufficient Sample Size

Several factors come into play when determining what a sufficient sample size is for a research study. Probably the most important factor is whether the study is a quantitative or a qualitative one. In quantitative research, there is one basic rule for sample size: The larger, the better. This is because a larger sample will be more representative of the popu- lation being studied. Having said that, even small samples using quantitative data can yield meaningful correlations if the researcher makes efforts to achieve a representative sample or a meaningful convenience sample.

How does one make a practical decision, though, regarding the exact sample size a par- ticular research situation requires? Gay, Mills, and Airasian (2009) provided the following guidelines for determining sufficient sample size based on the size of the population:

• For populations that include 100 or fewer individuals, the entire population should be sampled.

• For populations that include 400–600 individuals, 50% of the population should be sampled.

• For populations that include 1,500 individuals, 20% of the population should be sampled.

• For populations larger than 5,000, about 8% of the population should be sampled.

Thus, as we can see, the larger the population size, the smaller the percentage required to obtain a representative sample. This does not mean that the sample shrinks as the popula- tion increases, but rather the opposite: The sample size required actually increases when studying larger populations.

Although larger sample sizes are generally preferred in quantitative studies, the size of an adequate sample also depends on the similarities and dissimilarities among members of the population. If the population is fairly diverse for the construct being measured, then a larger sample will likely be required in order to obtain a representative sample. Popula- tions whose members have similar characteristics will require smaller samples. Research- ers have developed fairly sophisticated methods for determining sufficient sample sizes through the use of statistical power analysis; however, such methods are beyond the scope of this book. Gay et al.’s (2009) guidelines are sufficient for most types of research.

As discussed in Chapter 3, recommended samples for qualitative research are typically much smaller than for quantitative research, which strives to be representative of larger populations. In qualitative research, sample sizes are not based on numbers but rather on how well the variables of interest are represented (Houser, 2009). Unlike quantitative approaches, which require large samples, qualitative techniques do not have any rules regarding sample size. Thus, sample size depends more on what the researcher wants to know, the purpose of the inquiry, what the findings will be useful for, how credible they

new85743_04_c04_169-212.indd 194 6/18/13 12:02 PM

195

CHAPTER 4Section 4.4 Analyzing Survey Data

will be, and what can be done with available time and resources. Qualitative research can be very costly and time-consuming, so choosing information-rich cases will yield the great- est return on investment. As noted by Patton (2002), “The validity, meaningfulness, and insights generated from qualitative inquiry have more to do with the information-richness of the cases selected and the observational/analytical capabilities of the researcher than with sample size” (p. 245).

Nonresponse Bias in Survey Research

Sometimes participants do not submit a survey or do not fill it out completely. When this occurs, it is considered nonresponse bias, which can affect the size and characteristics of the sample, as well as the external validity of the study. External validity refers to how well the results obtained from a sample can be extended to make predictions about the entire population, which will be further discussed in Chapter 5. Nonresponses are particularly problematic if a large number of participants or a specific group of participants fails to respond to particular questions. For example, perhaps all females skip a certain question, or several participants skip certain questions because they are too personal. This omission creates bias not only in the characteristics of the sample (e.g., the sample may no longer be representative) but also in the size of the sample that is required for the study.

Nonresponses can occur for many reasons, including the survey being too long, the ques- tions being worded awkwardly, the survey topic being uninteresting, or, in the case with Web-based surveys, the participants not knowing how to access or log into the website to complete the survey.

There are a few ways to minimize the threat of nonresponse bias. These include increas- ing the sample size to account for the possibility of nonresponses, making sure that survey directions and questions are worded clearly, making sure the survey is not too long, providing rewards or incentives for completing the survey, sending out reminders to complete the survey, and providing a cover letter that describes the exact reasons for conducting the survey.

4.4 Analyzing Survey Data

Once you have designed a survey, chosen an appropriate sample, and collected some data, now comes the fun part. As with the quantitative descriptive designs covered in Chapter 3, the goal of analyzing survey data is to subject your hypothe- ses to a statistical test. Surveys can be used both to describe and predict thoughts, feelings, and behaviors. However, since we have already covered the basics of descriptive analysis in Chapter 3, this section will focus on predictive analyses, which are designed to assess the associations between and among variables.

Researchers typically use three approaches to test predictive hypotheses: correlational analyses, chi-square analyses, and regression analyses. Each one has its advantages and disadvantages, and each is most appropriate for a different kind of data. Correlational analysis allows one to examine the strength, direction, and statistical significance of a

new85743_04_c04_169-212.indd 195 6/18/13 12:02 PM

196

CHAPTER 4Section 4.4 Analyzing Survey Data

relationship; chi-square analysis determines whether two nominal variables are indepen- dent from or related to one another; and simple linear regression is the method used to “predict” scores for one variable based on another. In this section, we will walk through the basics of each analysis.

Correlational Analysis

In the beginning of this chapter, we encountered an example of a survey research ques- tion: What is the relationship between the number of hours that students spend studying and their grades in the class? In this case, the hypothesis claims that we can predict some- thing about a student’s grades by knowing how many hours he or she spends studying.

Imagine we collected a small amount of data to test this hypothesis, shown in Table 4.1. (Of course, if we really wanted a good test of this hypothesis, we would need more than 10 people in the sample, but this will do as an illustration.)

Table 4.1: Data for quiz grade/hours studied example

Participant Hours Studied Quiz Grade

1 1 2

2 1 3

3 2 4

4 3 5

5 3 6

6 3 6

7 4 7

8 4 8

9 4 9

10 5 9

The Logic of Correlation The important question here is whether and to what extent we can predict grades based on study time. One common statistic for testing these kinds of hypotheses is a correla- tion, which assesses the linear relationship between two variables. A stronger correlation between two variables translates into a stronger association between them. Or, to put it differently, the stronger the correlation between study time and quiz grade, the more accu- rately you can predict grades based on knowing how long the student spends studying.

Before we calculate the correlation between these variables, it is always a good idea to visualize the data on a graph. The scatterplot in Figure 4.2 displays our sample data from the studying/quiz grade study.

new85743_04_c04_169-212.indd 196 6/18/13 12:02 PM

197

CHAPTER 4Section 4.4 Analyzing Survey Data

Figure 4.2: Scatterplot for quiz grade/hours studied example

Each point on the graph represents one participant. For example, the point in the top right corner represents a student who studied for 5 hours and earned a 9 on the quiz. The two points in the bottom right represent students who studied for only 1 hour and earned a 2 and a 3 on the quiz.

There are two reasons to graph data before conducting statistical tests. First, a graph con- veys a general sense of the pattern—in this case, students who study less appear to do worse on the quiz. As a result, we will be better informed going into our statistical calcula- tions. Second, the graph ensures that there is a linear relationship between the variables. This is a very important point about correlations: The math is based on how well the data points fit a straight line, which means nonlinear relationships might be overlooked. Figure 4.3 demonstrates a robust nonlinear finding in psychology regarding the relation- ship between task performance and physiological arousal. As this graph shows, people tend to perform their best on just about any task when they have a moderate level of arousal.

When arousal is too high, it is difficult to calm down and concentrate; when arousal is too low, it is difficult to care about the task at all. If we simply ran a correlation with data on performance and arousal, the correlation would be zero because the points do not fit a straight line. Thus, it is critical to visualize the data before jump- ing ahead to the statistics. Otherwise, you risk overlooking an important finding in the data.

Q u

iz G

ra d

e

Hours Studied

Quiz Grade and Hours Spent Studying

10

8

6

4

2

0

0 1 2 3 4 5 6

arousal

performance

Figure 4.3: Curvilinear relationship between arousal and performance

new85743_04_c04_169-212.indd 197 6/18/13 12:02 PM

198

CHAPTER 4Section 4.4 Analyzing Survey Data

Interpreting Coefficients Once we are satisfied that our data look linear, it is time to calculate our statistics. This is typically done using a computer software program, such as SPSS (IBM), SAS/STAT (SAS), or Microsoft Excel. The number used to quantify our correlation is called the correlation coefficient. This number ranges from 21 to 11 and contains two important pieces of information:

• The size of our relationship is based on the absolute value of our correlation coef- ficient. The farther our coefficient is from zero in either direction, the stronger the relationship between variables. For example, both a 1 .80 and a 2 .80 indicate strong relationships.

• The direction of the relationship is based on the sign of our correlation coefficient. A 1 .80 would indicate a positive correlation, meaning that as one variable increases, so does the other variable. A 2 .80 would indicate a negative correlation, meaning that as one variable increases, the other variable decreases. (Refer back to Section 2.1, Overview of Research Designs, for a review of these two terms.)

So, for example, a 1 .20 is a weak positive relationship and a 2 .70 is a strong negative relationship.

When we calculate the correlation for our quiz grade study, we get a coefficient of .96, indicating a strong positive relationship between studying and quiz grade. What does this mean in plain English? Students who spend more hours studying tend to score higher on the quiz.

How do we know whether to get excited about a correlation of .96? As with all of our sta- tistical analyses, we look up this value in a critical value table, or, more commonly, let the computer software do this for us. This gives us a p value representing the odds that our correlation is the result of random chance. In this case, the p value is less than .001. This means that the chance of our correlation being a random fluke is less than 1 in 1,000, so we can feel pretty confident in our results. Note, however, that because the p value associated with r is dependent on sample size, even a tiny, unimportant correlation can be statisti- cally significant when drawing conclusions from a large population.

We now have all the information we need to report this correlation in a research paper. The standard way of reporting a correlation coefficient includes information about the sample size and p value, as well as the coefficient itself. Our quiz grade study would be reported as shown in Figure 4.4.

So where does this leave our hypothesis? We started by pre- dicting that students who spend more time studying would per- form better on their quizzes than those who spend less time study- ing. We then designed a study to test this hypothesis by collecting data on study habits and quiz grades. Finally, we analyzed these

Statistical symbol for the correlation coefficient, always

italicized and lowercase

Degrees of freedom, always

n–2 for a correlation

Correlation Coefficient

p value

r (8) = .964, p < .001

Figure 4.4: Correlation coefficient diagram

new85743_04_c04_169-212.indd 198 6/18/13 12:02 PM

199

CHAPTER 4Section 4.4 Analyzing Survey Data

data and found a significant, strong, positive correlation between hours studied and quiz grade. Based on this study, our hypothesis has been supported—students who study more have higher quiz grades! Of course, because this is a correlational study, we are unable to make causal statements. It could be that studying more for an exam helps you to learn more. Or, it could be the case that previous low quiz grades make students give up and study less. Or, the third variable of motivation could cause students to both study more and perform better on the quizzes. To tease these explanations apart and determine causality, we will need an experimental type of research design, which we will cover in Chapter 5.

Regression Analysis

Correlations are the best tool for testing the linear relationship between pairs of quantita- tive variables. However, in many cases, we are interested in comparing the influence of several variables at once. Imagine that you wanted to expand the investigation about hours studying and quiz grade by looking at other variables that might predict students’ quiz grades. We have already learned that the hours students spend studying positively cor- relate with their grades. But what about SAT scores? We might predict that students with higher standardized test scores will do better in all of their college classes. Or what about the number of classes that students have previously taken in the subject area? We might predict that increased familiarity with the subject would be associated with higher scores. In order to compare the influence of all three variables, we will use a slightly different analytic approach. Multiple regression analysis is a variation on correlational analysis; in it, more than one predictor variable is used to foresee a single outcome variable. In this example, we would be attempting to predict the outcome variable of quiz scores based on three predictor variables: SAT scores, number of previous classes, and hours studied.

Multiple regression analysis requires an extensive set of calculations; consequently, it is always done using computer software. A detailed look at these calculations is beyond the scope of this book, but a conceptual overview will help you understand the unique advan- tages of this form of analysis. Essentially, the calculations for multiple regression are based on the correlation coefficients between each of our predictor variables, and between each of these variables and the outcome variable. These correlations for our revised quiz grade study are shown in Table 4.2. If we scan the top row, we can see the correlations between quiz grade and the three predictor variables: SAT score (r 5 .14), previous classes (r 5 .24), and hours studied (r 5 .25). The remainder of the table shows correlations between the various predictor variables; for example, hours studied and previous classes correlate at r 5 .24. When we conduct multiple regression analysis using computer software, the soft- ware will use all of these correlations in performing its calculations.

Table 4.2: Correlations for a multiple regression analysis

Quiz Grade SAT Score Previous Classes Hours Studied

Quiz Grade — .14 .24* .25*

SAT Score — .02 2.02

Previous Classes — .24*

Hours Studied —

new85743_04_c04_169-212.indd 199 6/18/13 12:02 PM

200

CHAPTER 4Section 4.4 Analyzing Survey Data

The advantage of multiple regression is that it considers both the individual and the combined influence of the predictor variables. Figure 4.5 is a visual diagram of the indi- vidual predictors of quiz grades. The numbers along each line are known as regression coefficients, or beta weights. These are standardized coefficients that allow comparison across predictors. Their values are very similar to correlation coefficients but differ in an important way: They represent the effects of each predictor variable while controlling for the effects of all the other predictors. That is, the value of b 5 .21 linking hours studied with quiz grades is the independent contribution of hours studied, controlling for SAT scores and previous classes. If we compared the size of these regression coefficients, we would see that, in fact, hours spent studying were still the largest predictor of quiz grades (b 5 .21) compared with both SAT scores (b 5 .14) and previous classes (b 5 .19).

Figure 4.5: Predictors of quiz grades

Even if individual variables have only a small influence, they can add up to a larger com- bined influence. So, if we were to analyze the predictors of quiz grades in this study, we would find a combined multiple correlation coefficient of r 5 .34. The multiple correla- tion coefficient represents the combined association between the outcome variable and the full set of predictor variables. Note that in this case, the combined r of .34 is larger than any of the individual correlations (r) in Table 4.2, which ranged from .14 to .25. These numbers mean that we are better able to predict quiz grades from examining all three variables than we are from examining any single variable. Or, as the saying goes, the whole is greater than the sum of its parts!

Multiple regression analysis is an incredibly useful and powerful analytic approach, but it can also be a tough concept to grasp. Before we move on, let’s revisit the concept in the form of an analogy. Imagine you’ve just eaten the most delicious hamburger of your life and are determined to understand what made it so good. Lots of factors will contribute to the taste of your hamburger: the quality of the meat, the type and amount of cheese, the freshness of the bun, perhaps the smoked chili peppers layered on top. If you were to approach this investigation using multiple regression analysis, you would be able to sepa- rate out the influence of each variable (e.g., How important is the cheese compared with the smoked peppers?) as well as take into account the full set of ingredients (e.g., Does the freshness of the bun really matter when the other elements taste so good?). Ultimately, you would be armed with the knowledge of which elements are most important in craft- ing the perfect hamburger. And you would understand more about the perfect hamburger than if you had examined each ingredient in isolation.

SAT Score

Previous Classes

Hours Studied

Quiz Grade

b = .14, p > .05

b = .19, p < .05

b = .21, p < .05

new85743_04_c04_169-212.indd 200 6/18/13 12:02 PM

201

CHAPTER 4Section 4.4 Analyzing Survey Data

Chi-Square Analyses

Both correlations and regressions are well suited to testing hypotheses about prediction, as long as it is possible to demonstrate a linear relationship between two variables. But linear relationships require that variables be measured on one of the quantitative scales— that is, ordinal, interval, or ratio scales (see Section 2.3, Scales and Types of Measurement, for a review). What if we wanted to test the association between nominal, or categori- cal, variables? In these cases, we would need an alternative statistic called the chi-square statistic, which determines whether two nominal variables are independent from or related to one another. Chi-square is often abbreviated with the symbol x2, which shows the Greek letter chi with the superscript 2 for squared. (This statistic is also referred to as the chi-square test for independence—a slightly longer but more descriptive synonym.)

The idea behind this test is similar to that of the correlation coefficient. If two variables are independent, then knowing the value of one variable does not tell you anything about the value of the other. As we will see in the following examples, a larger chi-square reflects a larger deviation from what we would expect by chance and is thus an index of statistical significance.

The Logic of Chi-Square To determine whether two variables are associated, the chi-square works by comparing the observed frequencies (collected data) with the expected frequencies if the variables were unrelated. If the results significantly deviate from these expected frequencies, then we conclude that our variables are related. And, consequently, we are able to predict one variable based on knowing the values of the other. Let’s look at a couple of examples to make this more concrete.

First, let’s say we wanted to know whether gender is related to political party affiliation. We might randomly select 100 men and 100 women and ask them whether they identified as Republican or Democrat. Because both of these variables are nominal—that is, they identify only categories, not quantitative measures—chi-square will be our best choice to test the association between them. The first step in conducting this analysis is to arrange our data in a contingency table, which displays the number of individuals in each of the combinations of our nominal variables. We encountered these tables before in our exam- ples of observational studies in Chapter 3 (Section 3.1) but stopped short of conducting the statistical analyses. So imagine we get the results shown in Table 4.3a from our survey of gender and party affiliation.

Table 4.3a: Gender and party affiliation

Male Female

Democrat 60 60

Republican 40 40

In this case, there is no association between sex and party affiliation. It does not matter that the sample consists of 60% Democrats and 40% Republicans. What matters for our

new85743_04_c04_169-212.indd 201 6/18/13 12:02 PM

202

CHAPTER 4Section 4.4 Analyzing Survey Data

hypothesis test is that the pattern for males is the same as the pattern for females: Our sample consists of 1.5 times the number of Democrats for both sexes. In other words, knowing a person’s sex does not tell us anything about their political affiliation.

For illustration purposes, imagine we tested the same hypothesis again but could recruit only 50 women, compared with 100 men. Now, if we found the same 60%/40% split among men again, and assuming that the variables were still not related, here’s the ques- tion: What would we expect the split to look like among women? If the ratio of Democrats to Republicans remains at 1.5 to 1, we would expect to see the women divided into 30 Democrats and 20 Republicans (i.e., the same ratio; shown in Table 4.3b). This concept is referred to as the expected frequency, or the frequency you would expect to see if the vari- ables were not related. In this example, we have a 60/40 split among men. If gender were unrelated to party affiliation, we would expect to see the same pattern among women.

Table 4.3b: Gender and party affiliation with unequal ns

Male Female

Democrat 60 30

Republican 40 20

The chi-square statistic is calculated by comparing our observed data with these expected frequencies. In our gender and party affiliation example, the observed data match the expected frequencies, meaning that the variables are not related. Let’s walk through another example and see how these calculations work.

Calculating Chi-Square In this second example, imagine we wanted to know whether people in rural or urban areas were more likely to support a sales tax increase. It would be easy to speculate why either group might be more likely to do so—perhaps people living in cities are more politi- cally liberal or perhaps people living in small towns are better able to see the benefits of higher local taxes. So, once again, imagine we surveyed a sample of 100 people, asking them to indicate both their party affiliation and their support for a sales tax proposal. We get the following contingency table of results (in Table 4.4a). Notice that we have more urban than rural residents, reflecting the higher population density in cities. But, as with our preceding gender and political affiliation example, the raw numbers are less impor- tant than the ratios within each group.

new85743_04_c04_169-212.indd 202 6/18/13 12:02 PM

203

CHAPTER 4Section 4.4 Analyzing Survey Data

Table 4.4: Chi-square example: Support for a sales tax increase

4.4a: Observed data

Rural Urban Total

Support 10 45 55

Don’t support 30 15 45

Total 40 60 100

4.4b: Expected frequencies

Rural Urban Total

Support 10 (24.75) 45 (33) 55

Don’t support 30 (18) 15 (27) 45

Total 40 60 100

4.4c: Calculating deviations between observed and expected values

Rural Urban Total

Support (10 2 24.75)2/24.75 (55 2 33)2/33 55

Don’t support (30 2 18)2/18 (15 2 27)2/27 45

Total 40 60 100

4.4d: Deviations between observed and expected values

Rural Urban Total

Support 8.79 14.67 65

Don’t support 8 5.33 45

Total 40 60 100

The first stage in calculating our chi-square is to determine the expected frequencies. We begin by calculating the sums across each row in column, as shown in Table 4.4a. This gives us a sense of the overall patterns in the data. Overall 40% of the sample consisted of rural residents, compared with 60% urban residents. And, overall, 55% supported the sales tax increase, while 45% did not. But, as in our previous example, these descriptive statistics do not tell us anything about the relationship between the two variables. If the variables are independent, then the 55/45 split in support for the sales tax will not differ based on where people live.

So we need to determine how much these observed data differ from what would be expected under independence. That is, what would these cells look like if there were no relationship? These expected frequencies are calculated using the following formula:

Expected Frequency 5 R 3 C Total N

new85743_04_c04_169-212.indd 203 6/18/13 12:02 PM

204

CHAPTER 4Section 4.4 Analyzing Survey Data

For each of the four cells, we multiply the row total (R) by the column total (C), and then divide by the total N in the sample. For example, in the rural resident/support cell, we would multiply the row total (55) by the column total (40), and then divide by the total sample size (100); (55 3 40) 4 100 5 22. Conceptually, this makes perfect sense: If the two variables are unrelated, we can guess the value of each cell using the overall totals. Table 4.4b shows expected frequencies for each cell in parentheses.

The second stage in calculating chi-square is to determine the extent to which our observed data deviate from these expected frequencies. We will need to calculate this deviation in each cell and then add them up for a total chi-square. So, for each cell we (1) subtract the expected from the observed value; (2) square the difference to remove any negative num- bers; and (3) divide by the expected frequency in order to standardize the deviation. The final chi-square value is obtained by adding up all four of these deviation scores, which translates into the following formula:

x2 5 a 1observed 2 expected2 2

expected

For example, in our rural resident/support cell, we calculated an expected frequency of 2, representing the number we would expect under independence. But in our sample, there were 10 people in this cell. To calculate how much this deviates from what is expected, we (1) subtract 22 from 10 (5 212), (2) square this difference to remove the negative num- ber (5144), and (3) divide this by the expected frequency to standardize the deviation (144 4 22 5 6.55). Tables 4.4c and 4.4d illustrate the steps for obtaining these deviation scores in each of the four cells.

Finally, we add up all four of our deviation scores (one for each cell) to get the total chi- square value:

x2 5 a 1observed 2 expected2 2

expected 5 6.55 1 14.67 1 8 1 5.33 5 34.55

Our final chi-square value, 34.55, represents the sum of our deviations from the expected value. The larger this number, the more our observed data differ from the expected fre- quencies. Remember that these expected frequencies represent our null hypothesis—we would expect these frequencies only if the variables were unrelated. So the greater our chi-square value, the more our variables are related to one another. In the present exam- ple, this means we can predict a person’s support for a sales tax increase based on where he or she lives, which is consistent with our initial hypothesis.

But how do we know whether our value of 34.55 is meaningful? As with the other statisti- cal tests we have discussed, this requires looking up our result in a critical value table to determine whether the calculated value is above threshold. In this case, the critical value for a chi-square with a 2 3 2 table 5 3.84, so we can feel confident in our value of 34.55— almost 10 times higher than the threshold value!

However, unlike correlation and regression coefficients, our chi-square results cannot tell us anything about the direction or magnitude of the relationship. A larger chi-square reflects a larger deviation from what we would expect by chance and is thus an index of statistical significance. In order to interpret the patterns of our data, we need to visually inspect the numbers in our data table. Better yet, we can create a bar graph like we did in Chapter 3 to visually display these frequencies.

new85743_04_c04_169-212.indd 204 6/18/13 12:02 PM

205

CHAPTER 4Section 4.5 Ethical Issues in Survey Research

As Figure 4.6 shows, the cell frequencies suggest a fairly clear interpretation: People who live in urban settings are much more likely than people who live in rural settings to sup- port a sales tax increase. In fact, urban residents support the increase by a 3-to-1 margin, while rural residents oppose the increase by a 3-to-1 margin.

Figure 4.6: Graph of chi-square results

4.5 Ethical Issues in Survey Research

Like all research, surveys should be carried out in ways that protect and avoid harm to the participants. Informed consent, as we discussed in Chapter 1 (Section 1.7, Ethics in Research), is a requirement for survey research and should include a clear description of the survey content and the purpose for conducting the research. Whether the survey is completed in person or online, a cover letter should be included that con- tains this information. Although not always required for surveys (as there is minimal risk of harm to participants taking questionnaires and surveys), some researchers also like to obtain signed consent forms from the participants. This is especially true for institutional review boards (IRBs), who want to ensure that participants were fully informed of any sensitive information that may be collected, any potential limits to the confidentiality of the data, or access to private records (such as medical records) that are being sought in addition to the survey. In any case, a researcher should never administer a survey if the participant has not provided verbal or written consent and should never use the data other than for reasons for which the participant provided consent.

Assurances of participant anonymity and confidentiality are also important, especially with respect to response rates to sensitive questions and response rates in general. Partici- pants are likely to feel more comfortable responding to sensitive subjects when they know that they are participating anonymously. Anonymity ensures that there is no way for their responses to be linked back to them. Conducting survey research in a completely anony- mous manner is much easier with online surveys; however, researchers can take steps

50

Rural Urban

45

40

35

30

25

20

15

5

10

0

Support

Don’t Support

new85743_04_c04_169-212.indd 205 6/18/13 12:02 PM

206

CHAPTER 4Summary

to protect the anonymity of participants in individual or group administrations as well. For example, anyone who has access to the surveys and completed data must commit in writing to preserving their confidentiality. Any links between the answers and the par- ticipants’ personal identifying information, such as names, email addresses, and phone numbers, should be minimized by removing the latter from the data collection and coding it (e.g., numbering each participant with an ID code rather than using his or her name). If personal information must be kept, it should be separated from the survey responses and destroyed as soon as the study is over. Any person who can identify the participant by looking at the pattern of responses, such as a supervisor, should not be permitted to view the survey responses. It is also important to be careful when reporting results for a small subpopulation of the sample, whose personal information may be identifiable. Finally, when the study is completed, researchers must ensure that they either destroy all survey responses and personal information or store them securely.

Another important consideration for researchers is to be aware of their own biases and the impact that those might have on the testing process. This is especially important in face-to-face inquiries. For example, if a researcher has a strong opinion or bias toward a particular topic, the questions might be worded in a way that persuades the participants to answer in a specific manner. Additionally, if the researcher has a strong belief about the topic, he or she may make this evident during the testing process, which may encourage the participant to respond in a way that is consistent with what the researcher believes. As we discussed in Chapter 3 (Section 3.1), participant-expectancy bias can occur during obser- vations and interviews, as well as during survey research, and needs to be considered when analyzing and interpreting data.

Finally, when using standardized questionnaires or surveys that have been developed by other researchers, professional training and competence in the tests being used are essen- tial. It is unethical for any researchers to administer, score, and interpret a test that they have not been trained on. In addition, it is unethical for researchers to select tests based on limited knowledge and experience and assume that these tests are reliable and valid for the purpose for which they are using them.

Summary

This chapter has covered the process of survey research from conceptualization through analysis. We first discussed the types of research questions that are best suited to survey research—essentially, those that can be answered based on people’s observations of their own behavior and characteristics. Survey research can involve either verbal reports (i.e., interviews) or written reports (i.e., questionnaires). In both cases, sur- veys are distinguished by their reliance on people’s self-reports of their attitudes, feelings, and behaviors.

This chapter covered several key points for writing survey items. The take-home point to our Five Rules for Designing Better Questionnaires is that your questions should be writ- ten as clearly and unambiguously as possible. This helps to minimize the error variance

new85743_04_c04_169-212.indd 206 6/18/13 12:02 PM

207

CHAPTER 4 Summary

that might result from participants imposing their own guesses and interpretations on the material. In designing survey items, you also have a broad choice between open-ended and fixed-format responses. The former provide richer and more extensive data but are harder to score and code; the latter are easier to code but can constrain people’s responses to your choice of categories. If and when you settle on a fixed-format response, you have another set of decisions to make regarding the response scaling, labels, and general format.

Once you have constructed the scale, it is time to begin collecting the data. This chap- ter discussed the concept of sampling, or choosing a portion of the population to use for your study. Broadly speaking, sampling can be either “probability” or “nonprobability,” depending on whether you have a known population size from which you sample ran- domly. Probability sampling is more likely to result in a representative sample, but this approach is not possible in all studies. In fact, a significant proportion of psychology research studies uses a form of nonprobability sampling called convenience sampling, meaning that the sample consists of those who show up for the study.

This chapter also covered three approaches to analyzing survey data and testing hypoth- eses about prediction. The first, correlational analysis, is a very popular way to analyze survey data. The correlation is a statistical test that gives an assessment of the linear relationship between two variables. The stronger the correlation between variables, the more we can accurately predict one based on knowing the other. Second, regression analyses allow us to expand our investigations into multiple predictors. The advantage of multiple regres- sion analysis is that it considers both the individual and the combined influence of the predictor variables. However, both correlation and regression require the variables to be quantitative—that is, measured on an ordinal, interval, or ratio scale. In cases where our survey produces nominal or categorical data, we use an alternative called the chi-square statistic, which determines whether two nominal variables are independent or related. The chi-square works by examining the extent to which our observed data deviate from the pattern we would expect if the variables were unrelated—that is, the null hypothesis. The common thread running through these analyses is that they measure the association between variables and do not tell us anything about the causal relationship between them. To make causal statements, we have to conduct experiments, which we will cover in the next chapter.

Finally, this chapter discussed the ethical issues that arise in survey research. In addition to the concepts discussed in Chapter 1 regarding informed consent and confidentiality, it is important that researchers conducting survey research ensure anonymity so that partic- ipants feel safe responding to sensitive topics and comfortable participating in the overall study. Participant response rates tend to be higher when participants know that there is no way to link them to their responses. Additionally, the researcher should be aware of his or her own biases and the impact that these might have on the collection, analysis, and interpretation of the data, as well as on how participants may respond. Also, when using standardized tests or tests that have been developed by other researchers, it is imperative that researchers be trained and competent in administering and scoring the results, as well as in interpreting those results.

new85743_04_c04_169-212.indd 207 6/18/13 12:02 PM

208

CHAPTER 4Key Terms

anchors Labels, or endpoints, for a rating scale.

beta weights See regression coefficients.

bipolar scale Rating scale that has polar opposites as its anchors.

branching schedule An interview format in which questions take different directions depending on participants’ answers.

chi-square statistic A statistical test similar to the correlation coefficient; deter- mines whether two nominal variables are independent or related.

cluster sampling A variation of simple random sampling that involves dividing the sample into groups based on more than one level of analysis.

contingency table A data summary table that shows the number of individuals in each combination of the nominal vari- ables; used as the first step in calculating chi-square.

convenience sampling A nonprobability sampling strategy that involves simply enrolling people who show up for the study.

correlation Statistical test that assesses the linear relationship between two variables; the stronger the correlation between vari- ables, the more accurately the prediction about one based on knowing the other.

correlation coefficient The number used to quantify a correlation; this coefficient (r) ranges from 21 to 11 and contains infor- mation about both the size and direction of the correlation.

criterion sampling A qualitative sampling method used to select cases that meet a predetermined criterion of importance.

double-barreled question A flawed sur- vey item that asks more than one question at a time.

error of estimation The degree to which the data from the sample are expected to deviate from the population as a whole.

error variance Variance from random sources that are irrelevant to the trait or ability that a questionnaire is purporting to measure.

expected frequency The frequency one would expect to see if the variables were not related.

expert sampling Sampling of a panel of individuals who have known expertise (knowledge and training) in a particular area.

extreme or deviant case sampling A form of sample selection used when the researcher wants information-rich data on unusual or special cases; the focus is on examining both ends of the spectrum of outcomes (e.g., outstanding successes and failures of the phenomenon being studied).

fixed-format response Answer to a lim- iting question or statement, involving choosing from a list of options; on a sur- vey, fixed-format responses are easier to code but can constrain the data into nar- row categories.

forced choice A rating scale that requires respondents to agree or disagree with a statement, usually through the use of an even number of scale points.

Key Terms

new85743_04_c04_169-212.indd 208 6/18/13 12:02 PM

209

CHAPTER 4Key Terms

heterogeneous (maximum variation) sampling A qualitative sampling method that may include both extreme and typi- cal cases; used to select a wide variety of cases in relation to the phenomenon being investigated.

interview A verbally administered survey.

interview schedule A plan, or script, for the progress of an interview, describing the list of questions and the order in which they should be asked.

leading question A flawed survey item worded in a way that suggests an answer.

Likert scale Format that uses anchors of “strongly agree” and “strongly disagree” to rate responses to a survey question.

linear schedule An interview format that asks the same questions in the same order for all participants.

multiple-choice format A fixed-format- response survey format that asks partici- pants to select from a set of predetermined responses.

multiple correlation coefficient A number that represents the combined association between the outcome variable and the full set of predictor variables.

multiple regression analysis A variation on correlational analysis in which more than one predictor variable is used to pre- dict a single outcome variable.

nonprobability sampling A group of sampling strategies used when the odds of any given individual’s being in the sample are unknown.

nonresponse bias Bias introduced when certain questions or entire surveys are not completed by participants, affecting the characteristics and size of the sample.

open-ended response Unstructured answer to a question or statement; on a survey, open-ended responses provide rich data but are difficult to code.

pilot testing A “test run” of a survey that involves giving the questionnaire to a small sample of people, getting their feed- back, and making any necessary changes.

population The entire collection of people who could possibly be relevant for a study.

probability sampling A group of data- collection strategies used when each per- son in the population has a known chance of being in the sample.

purposive sampling A qualitative sam- pling method that includes selecting relatively small samples in a purposeful manner.

questionnaire A survey that is adminis- tered in writing.

rating scale A fixed-format response that asks participants to place responses on a continuum.

regression coefficients (beta weights) Val- ues that represent the effects of each predictor variable while controlling for the effects of all the other predictors.

sampling bias The failure of the sample to represent the underlying distribution in the population.

new85743_04_c04_169-212.indd 209 6/18/13 12:02 PM

210

CHAPTER 4Apply Your Knowledge

Apply Your Knowledge

1. For each of the following poorly written questionnaire items, identify the major problem and then rewrite it so that the problem is resolved. a. How much do you like cats and ponies?

main problem: better item:

b. Do you think that John McCain’s complete lack of personality proved that he would have been a terrible president?

main problem: better item:

sampling error The degree to which the characteristics of the sample differ from the characteristics of the population.

sampling frame A list of all members of a particular population (e.g., a list of every resident of the United States) and a neces- sary requirement for probability sampling strategies.

self-reports Participants’ reports of their own attitudes, feelings, and behaviors.

self-selection bias Bias introduced into a survey study when the researcher receives responses only from those who are inter- ested in the topic.

simple random sampling A probability sampling strategy that involves randomly picking participants from a list of everyone in the population.

snowball sampling A nonprobability sampling strategy that involves recruiting by word-of-mouth referrals.

social desirability Participants’ reluctance to give unpopular answers to survey ques- tions; concern over how their attitude will be perceived.

stratified random sampling A variation of simple random sampling, used when subgroups of the population might be left out of a purely random sampling pro- cess; breaking the sampling frame into subgroups and then sampling a random number from each subgroup.

survey research Any method that relies on people’s observations of their own behavior.

theory-based sampling A qualitative sampling method (a version of criterion sampling) that focuses on obtaining cases that represent theoretical constructs.

true/false format A fixed-format survey response that asks participants to indicate whether they endorse a statement.

typical case sampling Research method that involves sampling the most frequent or “normal” case from a population.

unipolar scale Rating scale that assesses a single construct.

new85743_04_c04_169-212.indd 210 6/18/13 12:02 PM

211

CHAPTER 4Critical Thinking & Discussion Questions

c. Do you dislike not playing basketball?

main problem: better item:

d. Do you support SB 1070?

main problem: better item:

e. How often do you take drugs?

main problem: better item:

2. Dr. Truxillo is interested in Arizona residents’ thoughts and feelings about global warming. For each of the following examples, state the sampling method used by her research assistants. a. Reese sets up a table in the mall and hands a survey to people who

approach her. b. Catherine randomly chooses 5 cities, then chooses 3 neighborhoods in each,

then randomly samples 5,000 households for a phone survey. c. Jason starts with a list of the entire population of Arizona and selects partici-

pants by dialing random phone numbers. d. Anna gets the master list from Jason and divides the population according

to education level. She then randomly chooses 500 high school dropouts, 500 college graduates, and 500 people with some postgraduate education.

3. Based on each of the following study descriptions, choose whether the best analysis would be a correlation, a multiple regression, or a chi-square. a. Jim is interested in the relationship between annual income and self-

reported happiness. b. Shelia is interested in whether some ethnic groups are more likely to use

counseling services (a yes-or-no question). c. Angela is interested in knowing the best predictors of recovery from depres-

sion, comparing the influence of drugs, therapy, and family resources. d. Adam is interested in whether high school dropouts or college graduates are

more likely to vaccinate their children. e. Nicole is interested in understanding the best predictors of weight loss. f. Samantha is interested in the relationship between self-esteem and

prejudice.

Critical Thinking & Discussion Questions

1. In survey research, explain the trade-off between the “richness” of people’s responses and the ease of analyzing their responses.

2. When doing interviews, the researcher has a personal interaction with the subject. Why is this both good and bad?

new85743_04_c04_169-212.indd 211 6/18/13 12:02 PM

new85743_04_c04_169-212.indd 212 6/18/13 12:02 PM