Write a 2 page (double-spaced) paper addressing the following questions. Be sure to use information presented in the article to support your answers.
Fundamental Statistics for the Social and Behavioral Sciences
2
To my kids, Meagan and Will, and my parents, Katsumi and Grayce Tokunaga.
3
Fundamental Statistics for the Social and Behavioral Sciences
Howard T. Tokunaga San Jose State University
4
FOR INFORMATION:
SAGE Publications, inc.
2455 Teller Road
Thousand Oaks, California 91320
E-mail: [email protected]
SAGE Publications Ltd.
1 Oliver's Yard
55 City Road
London EC1Y 1SP
United Kingdom
SAGE Publications India Pvt. Ltd.
B 1/I 1 Mohan Cooperative Industrial Area
Mathura Road, New Delhi 110 044
India
SAGE Publications Asia-Pacific Pte. Ltd.
3 Church Street
#10-04 Samsung Hub
Singapore 049483
Copyright © 2016 by SAGE Publications, Inc.
All rights reserved. No part of this book may be reproduced or utilized in any form or by any means, electronic or mechanical, including photocopying, recording, or by any information storage and retrieval system, without permission in writing from the publisher.
All trademarks depicted within this book, including trademarks appearing as part of a screenshot, figure, or other image are included solely for the purpose of illustration and are the property of their respective holders. The use of the trademarks in no way indicates any relationship with, or endorsement by, the holders of said trademarks. SPSS is a registered trademark of International Business Machines Incorporated.
Printed in the United States of America
Cataloging-in-Publication Data is available for this title from the Library of Congress.
ISBN 978-1-4833-1879-0
5
This book is printed on acid-free paper.
15 16 17 18 19 10 9 8 7 6 5 4 3 2 1
Acquisitions Editor: Vicki Knight
Editorial Assistant: Yvonne McDuffee
Associate Editor: Katie Bierach
Production Editor: Kelly DeRosa
Copy Editor: Gillian Dickens
Typesetter: C&M Digitals (P) Ltd.
Proofreader: Jennifer Grubba
Indexer: Marilyn Augst
Cover Designer: Candice Harman
Marketing Manager: Nicole Elliott
6
Detailed Contents
Preface Acknowledgments About the Author Chapter 1. Introduction to Statistics Chapter 2. Examining Data: Tables and Figures Chapter 3. Measures of Central Tendency Chapter 4. Measures of Variability Chapter 5. Normal Distributions Chapter 6. Probability and Introduction to Hypothesis Testing Chapter 7. Testing One Sample Mean Chapter 8. Estimating the Mean of a Population Chapter 9. Testing the Difference between Two Means Chapter 10. Errors in Hypothesis Testing, Statistical Power, and Effect Size Chapter 11. One-Way Analysis of Variance (ANOVA) Chapter 12. Two-Way Analysis of Variance (ANOVA) Chapter 13. Correlation and Linear Regression Chapter 14. Chi-Square Tables Appendix: Review of Basic Mathematics Glossary References Index
7
Preface
It may surprise students to learn they have something in common with writers of books such as this one: When you get close to finishing a writing assignment, you get a bit tired and a bit lazy. The first attempt at this preface was written shortly after final drafts of chapters were sent to my editor at SAGE, Vicki Knight. After reading it, she said, “It's not bad, but it reads like the ‘typical’ Preface. I think it would be useful for the reader to have a sense of why you wrote this book and why you wrote it the way you did.”
In responding to my editor's plea for self-analysis, I found that this book's journey began in college. When I entered college, I thought my path would take me to law school; however, taking an Intro to Psych class my freshman year made me realize I enjoy the challenge of trying to understand the human mind. The school I attended, UC Santa Cruz, was a fairly unconventional university, but somehow in the midst of a sea of humanistic psychologists, I became attracted to the empirical and methodological aspects of psychology. This was a result of taking classes with instructors such as David Harrington and Dane Archer, who showed me that statistics could appeal to students if taught using a gentle, guiding approach that addresses questions relevant to students' lives. After graduating from college, I was able to get a job as a research assistant for a human resource consulting firm. Despite my lack of work experience, I was hired primarily as a result of having taken statistics and research methods courses, which taught me that learning statistics has benefits both inside and outside of the classroom.
Several years later, I started grad school at UC Berkeley, where two events critical to this book took place. First, serving as a teaching assistant, I found I really enjoyed helping students, particularly in statistics and research methods classes that were often viewed with fear and suspicion. Second, I took graduate classes from Geoff Keppel, who had developed his own unique method and system for analyzing experimental research designs. His lectures and books were instrumental in showing me that statistics can be taught in a systematic way that highlights similarities rather than differences between different research situations. Geoff managed to transform something as daunting sounding as a “3 × 2 × 4 research design” into the mathematical equivalent of playing with wooden toy alphabet blocks labeled “A,” “B,” and “C.” For a long time, I thought my gratitude to Geoff was an isolated occurrence. However, the appreciation others have for his approach to teaching statistics was made apparent to me several years later when I watched him receive an American Psychological Association (APA) Lifetime Achievement award.
After leaving Cal, I took on a teaching position at San Jose State, where my teaching responsibilities included an introductory statistics course aimed at students with a wide range of background, ability, and motivation. As I needed to select a textbook to use in this course, for the first time I looked carefully at the wide range of offerings. What I found
8
striking (and still find striking) was that the majority of books focused on providing formulas and very small sets of data designed to demonstrate how to correctly calculate the correct numbers from these formulas. Little emphasis, however, was given to what these numbers meant or implied. Given my own experiences learning statistics, I thought a book was needed that discusses statistics in a thematic manner, focusing on how they are used to answer questions and test ideas within the larger research process.
The primary purpose of this book is to not just teach students how to calculate statistics but how to interpret the results of statistical analyses in light of a study's research hypothesis and to communicate one's results and interpretations to a broader audience. Hopefully, this book will not only help students understand the purpose and use of statistics but also give them a greater understanding of how research studies are conceived, conducted, and communicated.
The 14 chapters of this book may be placed into three general categories. The first four chapters are designed to introduce students to the research process and how data that have been collected may be organized, presented, and summarized. Chapters 5 through 10 discuss the process of conducting statistical analyses to test research questions and hypotheses, as well as issues and controversies regarding this process. The final four chapters of this book, Chapters 11 to 14, discuss different statistical procedures used in research situations that vary in the number of independent variables in the study as well as how the independent and dependent variables have been measured.
9
A Few Tips for Students
To you, the college student about to read this book as part of taking a statistics course: “Welcome!” and “Great job!” I welcome you because you're about to embark on a semester-long journey that I hope will enhance your skills and widen your perspective; I congratulate you because it's a journey not everyone is willing to take.
At the present moment, I know my encouragement and appreciation may be of little comfort to you as you might be somewhat anxious about having to learn statistics. Some of you might be anxious about having to learn statistical concepts and formulas; some of you might be anxious because the research process seems pretty complicated. Talking with students who have taken my courses over the years has helped me assemble the following advice:
10
Master the Material Presented in the Early Chapters
Chapters 1 through 6 discuss how research is often conducted, how data are summarized and described, and the process by which researchers conduct statistical analyses to test their ideas. It is absolutely critical for you to have a firm understanding of these chapters as they lay the groundwork for later chapters that discuss a variety of statistical procedures used by researchers to test hypotheses.
What does it mean to master this material? First, read the chapters both before and after they're discussed in class. By reading the chapters beforehand, you'll be able to identify anything that's unclear to you and have your questions addressed by your instructor. Rereading the chapters after they're discussed in class will help confirm your understanding of the material. Next, be able to define and explain key concepts presented in these chapters. These concepts are highlighted in bold-faced type, and it is important to learn them when they're first introduced because they'll appear throughout the remainder of the book. Next, do the learning checks within the chapters and the exercises at the end of the chapters. The learning checks include both exercises and review questions you can ask yourself to assess your understanding of the material. Most important, do not miss class during the early part of the semester! It's been my experience that students who miss critical lectures at the beginning of the term often have difficulty keeping up as the semester continues.
11
Review High School Algebra
If you read newspapers or watch television, you might conclude that the most complicated data analysis people can comprehend is a pie chart. As frustrating as this is to researchers and statistics instructors, they understand that statistics can be confusing. If you happen to move beyond the introductory statistics course for which you are reading this book, you'll find that statistical procedures are often conducted using computers and statistical software rather than hand calculations. Consequently, some statistics textbooks no longer include mathematical formulas but instead have students conduct analyses using statistical software or Microsoft Excel. However, this book has chosen a different approach for two main reasons. First, I believe that to comprehend the results of data analysis, one must understand the underlying foundation of statistical procedures. Learning statistics via computer software can lead to a “brain-dead” approach in which students are at the mercy (rather than control) of their computers. I refer to this as the “I'm only as smart as my printout” method of learning statistics, also known as “Because this is what my computer told me.”
Second, my steadfast and somewhat stubborn adherence to an approach emphasizing mathematical formulas and calculations is based on a simple reason: The statistics in this book are not hard to calculate! To assess whether you're adequately prepared to read this book, take the following test:
1. Do you know how to add, subtract, multiply, and divide? 2. Do you know how to calculate the “square” or “square root” of a number? 3. Do you understand the “order of operations” and how to use parentheses within
mathematical calculations?
If your answers to these questions are yes, congratulations! You possess the ability to conduct every mathematical calculation in this book. None of the formulas in this book requires knowledge of geometry, trigonometry, or calculus. However, if you're unsure of your mathematical ability, you may want to review your high school algebra. At the end of this book is an appendix that includes the mathematical concepts and operations needed to conduct the statistical analyses in this book. I highly recommend you read this appendix to review and assess your mathematical skills.
12
Learn to Use a Statistical Calculator
Although learning the statistics in this book doesn't require the use of computers or software, you will need a hand calculator. Calculators with statistical capabilities are typically identified as statistical or scientific calculators. Although there's a wide selection of brands and models, one simple criterion to use in selecting a calculator is price. The statistical calculators most appropriate for this book cost somewhere in the $10 to $15 range at the time of this writing. I do not recommend you use or purchase a more expensive calculator! I prefer these simpler calculators because I've found students are able to conduct the calculations in this book more quickly and with fewer errors. I also recommend you purchase a calculator similar to the one your instructor uses as he or she will be more familiar with its features and idiosyncrasies.
In selecting a calculator, look closely at the keyboard; you'll need a calculator whose keys have labels such as “ X ¯ ,” “σ,” and “ΣX” (Chapters 3 and 4 discuss what these symbols represent). When you go calculator shopping, you'll find calculators that have keys with labels such as “ΣY” and “XY.” Although calculators with these “Y” symbols are only slightly more expensive than calculators with just the “X” symbols, I do not recommend their purchase because these features are not needed to perform the calculations in this book. In fact, I've found students with “Y” calculators have a more difficult time performing simpler calculations.
13
Ask Questions
Conducting research is the process of asking and answering questions. Accordingly, throughout this book, issues are often framed in the form of questions, which I hope encourages you to ask questions as well. The first person you should direct any questions to is yourself. As you work your way through the chapters, get into the habit of asking yourself, “What does this mean?” A good test of your level of comprehension is whether you can explain the material to yourself in a meaningful way.
Also, be sure to work with your instructor to confirm your level of understanding and clarify unanswered questions. Instructors often say to students, “The only ‘bad’ question is the one you don't ask.” Don't be hesitant about asking questions in class. Instructors will tell you they use one student's question to assess an entire class' level of understanding. This is what I call the “pencil test”—when one student asks a question, I look to see how many other pencils get raised, which reflects the number of students who had the same question and are ready to write down the response. Asking questions enhances the learning experience for yourself, your classmates, and your instructor.
14
Form a Study Group
No matter how much I encourage students to ask questions in class, it seems they hesitate for fear of drawing attention to themselves. As a result, I turn to another well-traveled saying: “Misery loves company.” Forming a study group to meet with other students on a regular basis will help you keep up with reading assignments, confirm your comprehension of the material, clarify any confusion you have, prepare for exams, and receive an invaluable source of social support. I do suggest that your group contain at least one person who will raise the group's questions and concerns to the instructor in class or office hours.
15
Practice, Practice, Practice
The best way to learn the topics covered by this book is to obtain as much practice as possible. Throughout the chapters are examples in which the calculations have been worked out for you; in the middle and end of each chapter are a number of exercises. The answers to some of the exercises are provided at the back of this book, and your instructor has access to the answers to the other exercises. Do the end-of-chapter exercises even if your instructor does not assign them as homework in order to identify any recurring mistakes you make. We find that the vast majority of errors students make are simple computational errors, committed when students do the calculations too quickly.
WELCOME!
GREAT JOB!
16
Acknowledgments
This book has traveled a long journey, with many people playing both transient and persevering roles in its development. First and foremost, I'd like to thank my kids, Meagan and Will, for providing me perspective and inspiration amid all of the twists and turns we faced while this book was being written. I'm so very blessed to have you as my “burden.” Away from home, I'm grateful to my graduate students at San Jose State—working on this book has given me much empathy for their thesis-related travails. However, they should be aware that the completion of this book ends the unspoken agreement that I not ask, “So how's your thesis?” if they don't ask, “So how's your book?” A number of friends have provided most welcome relief after long days staring at a computer screen, with particular thanks going to Ellie, Eileen, Susan, Bonnie, Ann, Lisa, and Terry. I would like to thank Charles Linsmeier and David Shirley for their many helpful comments and advice on earlier versions of this book, and Julie West for her help in developing chapter exercises.
I would like to express my appreciation to everyone at SAGE who has played a role in shepherding this book to completion. First and foremost, I must express my gratitude to my editor, Vicki Knight. It has been a joy and an honor to work with someone with such a truly exceptional combination of experience, knowledge, perspective, and humor. In production I was fortunate enough to work with others on the SAGE team who made the process a pleasure: Yvonne McDuffee, Editorial Assistant; Katie Guarino, Assistant Editor; Nicole Elliott, Executive Marketing Manager; Candice Harman, Cover Designer; and the copyeditor, Gillian Dickens. Furthermore, I am grateful to the reviewers, who took the time to read the chapters, raise questions, and provide constructive comments and suggestions to assist me in clarifying and strengthening the manuscript:
Holly R. Straub, The University of South Dakota
Shelly A. McGrath, University of Alabama at Birmingham
Jim Allen, State University of New York, Geneseo
David Schuster, San José State University
Hideki Morooka, Fayetteville State University
Andrea J. Sell, California Lutheran University
Christine D. MacDonald, Indiana State University
Robert G. LaChausse, California State University, San Bernardino
17
Melissa J Beers, Ohio State University
18
About the Author
Howard T. Tokunaga is Professor of Psychology at San Jose State University, where he serves as Coordinator of the MS Program in Industrial/Organizational (I/O) Psychology and teaches undergraduate and graduate courses in statistics, research methods, and I/O psychology. He received his bachelor's degree in psychology at UC Santa Cruz and his PhD in psychology at UC Berkeley. In addition to his teaching, he has consulted with a number of public-sector and private-sector organizations on a wide variety of management and human resource issues. He is coauthor (with G. Keppel) of Introduction to Design and Analysis: A Student's Handbook.
19
Chapter 1 Introduction to Statistics
20
Chapter Outline 1.1 What Is Statistics? 1.2 Why Learn Statistics? 1.3 Introduction to the Stages of the Research Process
Developing a research hypothesis to be tested Identifying a question or issue to be examined Reviewing and evaluating relevant theories and research Stating a research hypothesis: Independent and dependent variables
Collecting data Drawing a sample from a population Determining how variables will be measured: Levels of measurement Selecting a method to collect the data: Experimental and non-experimental research methods
Analyzing the data Calculating descriptive statistics Calculating inferential statistics
Drawing a conclusion regarding the research hypothesis Communicating the findings of the study
1.4 Plan of the Book 1.5 Looking Ahead 1.6 Summary 1.7 Important Terms 1.8 Exercises
In introducing this book to you, we assume you are a college student who is taking what is perhaps your first course in statistics to fulfill a requirement for your major or a general education requirement. If so, you may be asking yourself two questions:
What is statistics? Why learn statistics?
The ultimate goal of this book is to help you begin to answer these two questions.
21
1.1 What is Statistics?
Whether or not you are aware of it, you encounter a variety of “statistics” in your day-to- day activities: the typical cost of going to college, the yearly income of the average college graduate, the average price of a home, and so on. So what exactly is “statistics”? The Merriam-Webster dictionary defines statistics as a branch of mathematics dealing with the collection, analysis, interpretation, and presentation of masses of numerical data. When people think about statistics, they often focus on only the “analysis” aspect of the above definition—that is to say, they focus on numbers that result from analyzing data. However, statistics is not only concerned about how data are analyzed, it recognizes the importance of understanding how data are collected and how the results of analyses are interpreted and communicated. The purpose of this book is to introduce, describe, and illustrate the role of statistics within the larger research process.
22
1.2 Why Learn Statistics?
We believe there are a variety of reasons why you should learn statistics. First, not only do you currently encounter statistics in your daily activities, but throughout your life, you have been and will continue to be affected by the results of research and statistical analyses. Which college or graduate school you attend is based in part on test scores developed by psychologists. You may also have to take a personality or intelligence test to get a job. The choices of drugs and medicines available to you are based on medical research and statistical analyses. If you have children, their education may be affected by their scores on achievement or aptitude tests. Learning about statistics will help you become a more informed and aware consumer of research and statistical analyses that affect many aspects of your life.
A second reason for learning statistics is that you may be asked or required to read and interpret the results of statistical analyses. Many college courses require students to read academic research journal articles. Evaluating published research is complicated by the fact that different people studying the same topic may come up with diverse or even opposing conclusions. Understanding statistics and their role in the research process will help you decide whether conclusions drawn in research articles are appropriate and justified.
Another reason for learning statistics is that it will be of use to you in your own research. College courses sometimes have students design and conduct mini-research studies; undergraduate majors might require or encourage students to do senior honors theses; graduate research programs often require masters' theses and doctoral dissertations. Learning to collect and analyze data will help you address your own questions in an objective, systematic manner.
A final reason for learning statistics is that it may help you in your future career. The website Careercast.com conducts an annual survey in which they evaluate 200 professions on five dimensions: environment, income, employment outlook, physical demands, and stress. In 2013, the highest rated profession in this survey was “actuary,” defined as someone who “interprets statistics to determine probabilities of accidents, sickness, and death, and loss of property from theft and natural disasters.” Talking about his job, one actuary noted, “I can count on one hand the number of days I've said, ‘I don't want to go to work today’ … I've seen people come in to say thank you for the work I've done. That's pretty powerful.”
It is generally a good idea for students to maintain a healthy level of curiosity or even skepticism in regards to their education. However, we find that when it comes to learning statistics, the frame of mind of some students may be characterized as one of fear and anxiety. Although we understand these feelings, we hope the benefits associated with learning statistics will become clear to you and help you overcome any concerns you may
23
have.
24
1.3 Introduction to the Stages of the Research Process
Much of scientific research involves asking questions. Throughout this book, we will examine how contemporary researchers have asked and attempted to answer a broad range of questions regarding human attitudes and behavior. Below are research questions we will address in this chapter to introduce the stages of the research process:
Is students' performance on tests more influenced by their learning strategies (how they learn) or their motivation (why they learn)? Do college students and faculty differ in their beliefs about the prevalence of student academic misconduct such as cheating and plagiarism? Is the extent to which adolescents are exposed to violence in their community related to how they do in school? Is one method of disciplining ones children more effective than another? Does playing online computer games affect ones interpersonal relationships? Does providing substance abuse treatment to drug users have an effect on safety in the workplace?
How might you try to answer questions such as these? You could base your answers on your personal beliefs, or you could adopt the answers given to you by others. But rather than relying on subjective beliefs and feelings, researchers test their ideas using science and the scientific method. The scientific method is a method of investigation that uses the objective and systematic collection and analysis of empirical data to test theories and hypotheses.
At its simplest, this book will portray the scientific method as consisting of five main steps or phases:
developing a research hypothesis to be tested, collecting data, analyzing the data, drawing a conclusion regarding the research hypothesis, and communicating the findings of the study.
Accomplishing each of these five steps requires completing a number of tasks, as shown in Figure 1.1. Because this sequence of steps will be used throughout this book and will serve as the model for the wide assortment of research studies we will review and discuss, each step is briefly introduced below. It is important to understand that the research process depicted in Figure 1.1 represents an ideal way of doing research. The “real” way, as you may discover in your own efforts or from speaking with researchers, is often anything but a smooth ride but rather is filled with starts and stops, dead ends, and wrong turns.
25
Developing a Research Hypothesis to be Tested
The initial stage—and the first step—of the research process is to develop a research hypothesis to be tested. A research hypothesis is a statement regarding an expected or predicted relationship between variables. A variable is a property or characteristic of an object, event, or person that can take on different values. One example of a variable is “U.S. state,” a variable with 50 possible values (Alabama, Arkansas, etc.).
Figure 1.1 Steps in the Research Process within the Scientific Method
Research hypotheses are usually developed through the completion of several tasks:
identifying a question or issue to be examined, reviewing and evaluating relevant theories and research, and stating a research hypothesis.
Each of these three tasks is described below.
Identifying a Question or Issue to be Examined
Most research starts with a question posed by the researcher. These questions often come
26
from the researchers own ideas and daily observations. Although this may not seem terribly scientific, there is an advantage in using one's own experience as a starting point: People are generally much more motivated to explore a question or topic that concerns them personally. In teaching statistics, we frequently advise students developing their own research projects to study something that is of interest to them. Conducting research can be tedious, difficult, and frustrating. At various points during your research, you may ask yourself, “Why am I doing this?” Being able to provide a satisfactory answer to this question will help you overcome whatever obstacles you encounter along the way.
Reviewing and Evaluating Relevant Theories and Research
Beyond the researcher's own curiosity, research questions often arise from an examination of the theories, ideas, and research of others. A theory is a set of propositions used to describe or explain a phenomenon. The purpose of a theory is to summarize and explain specific facts using a logically consistent framework. Placing a question within a theoretical framework provides guidance and structure to research.
Reviewing and evaluating existing theories and research helps the researcher decide whether it is worth the time and energy required to conduct the study. By seeing what others have done, the researcher may decide that a particular idea has already been investigated and there is no reason to duplicate earlier efforts. On the other hand, the researcher may conclude that the current way of thinking is incomplete or mistaken. By doing this review, researchers are able to ensure that the studies they undertake add to and improve upon an existing body of knowledge.
Stating a Research Hypothesis: Independent and Dependent Variables
Understanding and evaluating an existing literature not only helps articulate a question of interest but also may lead to a predicted answer to that question. Within the scientific method, this answer is stated as a research hypothesis, defined earlier as a statement regarding an expected or predicted relationship between variables. Table 1.1 lists the research questions and research hypotheses for the studies listed at the beginning of this section. For example, the first research hypothesis states that “students who are taught effective learning skills will perform better on tests than students offered incentives to do well.”
One characteristic of research hypotheses such as those listed in Table 1.1 is that they identify the variables that are the focus of their research studies. As mentioned earlier, a variable is a property or characteristic with different values. Variables can be classified in several ways. In specifying a research hypothesis, researchers often speak in terms of “independent” and “dependent” variables. An independent variable may be defined as a
27
variable manipulated by the researcher. A dependent variable, on the other hand, is a variable measured by the researcher. Researchers are interested in examining the effect of the independent variable on the dependent variable.
Consider the first research hypothesis provided in Table 1.1: “Students who are taught effective learning skills will perform better on tests than students who are offered incentives to do well.” Here the independent variable is the instructional method by which students are taught, which consists of two values: learning skills and incentives. The dependent variable is the test performance that will be measured during the research. In this study, the effect of the independent variable on the dependent variable is that differences in students' test performance (the dependent variable) may “depend” upon which instructional method (learning skills or incentives) a student receives. Table 1.2 lists the independent and dependent variables for each of the research hypotheses in Table 1.1.
A second characteristic of research hypotheses, in addition to identifying variables, is that they specify with as much precision as possible the nature and direction of the relationship between variables. For example, the first research hypothesis in Table 1.1 states that “students who are taught effective learning skills will perform better on tests….” The word better indicates the nature and direction of the relationship between the independent variable (instructional method) and the dependent variable (test performance). The direction of the relationship would not have been stated if the hypothesis had included the less specific phrase “will perform differently on tests,” which simply indicates that the two groups are not expected to be the same. Table 1.3 provides directional and non-directional research hypotheses for the research studies from Table 1.1.
Table 1.1 Examples of Research Questions and Hypotheses Table 1.1 Examples of Research Questions and Hypotheses
Research Question Research Hypothesis
Is students' performance on tests more influenced by their motivation (why they learn) or their learning strategies (how they learn)?
Students who are taught effective learning skills will perform better on tests than students offered incentives to do well.
Do college students and faculty differ in their beliefs about the prevalence of student academic misconduct (i.e., cheating, plagiarism)?
Faculty members' beliefs about the frequency of student academic misconduct will be lower than students' beliefs.
Is there a relationship between adolescents' exposure to violence in their community and their academic achievement?
The more adolescents are exposed to violence in their community, the lower their levels of academic achievement.
Children will rate a disciplining strategy
28
Is one method of disciplining children more effective than another?
that emphasizes logic and reason as more effective than one based on rewards and punishment.
Does playing online games affect one's interpersonal relationships?
Heavy users of online games have less fulfilling interpersonal relationships than users spending little or no time playing online games.
Does providing substance abuse treatment to drug users have an effect on safety in the workplace?
Drug users are less likely to have work- related accidents after undergoing substance abuse treatment than before the treatment.
Table 1.2 Research Hypotheses and their Independent and Dependent Variables Table 1.2 Research Hypotheses and their Independent and Dependent Variables
Research Hypothesis Independent Variable
Dependent Variable
Students who are taught effective learning skills will perform better on tests than students offered incentives to do well.
Instructional method
Test performance
Faculty members' beliefs about the frequency of student academic misconduct will be lower than students' beliefs.
Member of college community
Beliefs about the frequency of student academic misconduct
The more adolescents are exposed to violence in their community, the lower their levels of academic achievement.
Exposure to community violence
Academic achievement
Children will rate a disciplining strategy that emphasizes logic and reason as more effective than one based on rewards and punishment.
Parental discipline strategy
Effectiveness of parental discipline strategy
Heavy users of online games have less fulfilling interpersonal relationships than users spending little or no time playing online games.
Online game-playing
Quality of interpersonal relationships
Drug users are less likely to have work-related accidents after undergoing substance abuse treatment than before the treatment.
Time Occurrence of a work-related accident
Table 1.3 Directional and Non-Directional Research Hypotheses Table 1.3 Directional and Non-Directional Research Hypotheses
29
Directional Research Hypothesis Non-directional Research Hypothesis
Students who are taught effective learning skills will perform better on tests than students offered incentives to do well.
Students who are taught effective learning skills will perform differently on tests than students offered incentives to do well.
Faculty members' beliefs about the frequency of student academic misconduct will be lower than students' beliefs.
Faculty members' beliefs about the frequency of student academic misconduct will be different than students' beliefs.
The more adolescents are exposed to violence in their community, the lower their levels of academic achievement.
The more adolescents are exposed to violence in their community, the more different their levels of academic achievement.
Children will rate a disciplining strategy that emphasizes logic and reason as more effective than one based on rewards and punishment.
Children will rate a disciplining strategy that emphasizes logic and reason differently than one based on rewards and punishment.
Heavy users of online games have less fulfilling interpersonal relationships than users spending little or no time playing online games.
The quality of interpersonal relationships is different for heavy users of online games than users spending little or no time playing online games.
Drug users are less likely to have work- related accidents after undergoing substance abuse treatment than before the treatment.
The likelihood of drug users having work- related accidents is different after undergoing substance abuse treatment than before the treatment.
The ability to form a directional research hypothesis is dependent on the state of the existing literature on the question of interest. If little or perhaps conflicting research has been conducted, researchers may not be able to form a directional hypothesis before they begin their research. One study, for example, examined the relationship between exercise deprivation (not allowing people to get their exercise) and tension, depression, and anger (Mondin et al., 1996). The researchers for the study reported, “We did not have a directional hypothesis, and when participants asked what we expected to find in this study, we replied: ‘We really are not sure since the results of earlier work on exercise deprivation are mixed’” (p. 1200).
30
Collecting Data
Once a research hypothesis has been formulated, researchers are ready to proceed to the second stage in the research process: collecting data relevant to this hypothesis. This step is seen as being composed of three tasks:
drawing a sample from a population, determining how the variables will be measured, and selecting a method by which to collect the data.
31
Learning Check 1: Reviewing what you've Learned So Far
1. Review questions a. What are the main steps involved in the research process? b. Why is it useful to review and evaluate theories and research before conducting a study? c. What are the two main characteristics of research hypotheses? d. What is the difference between an independent variable and a dependent variable?
2. Listed below are several research hypotheses from published studies. For each research hypothesis, identify the independent and dependent variable.
a. “College students will rate instructors who dress formally (i.e., business suit and tie) as having more expertise than instructors who dress casually (i.e., slacks and shirt)” (Sebastian & Bristow, 2008).
b. “It was expected that … greater amounts of television viewing … would predict greater … posttraumatic stress symptoms” (McLeish & Del Ben, 2008).
c. “We … hypothesized persons who estimated the HSAS level to be red (severe) or orange (high) … when the HSAS level was [in fact] yellow (elevated), would report greater worry about terrorism” (Eisenman et al., 2009).
d. “We hypothesized that prekindergarten children who participated in the 6-week intervention would perform better [on a test of literacy skills] than their peers in a control group who did not participate in the program” (Edmonds, O'Donoghue, Spano, & Algozzine, 2009, p. 214).
Drawing a Sample from a Population
The first step in collecting data is to identify the group of participants to which the research hypothesis applies. The group to which the results of a study may be applied or generalized is called a population. A population is the total number of possible units or elements that could potentially be included in a study. For example, researchers could variously define the population of interest for the first research hypothesis in Table 1.1 as “college students,” “college students in the United States,” “college students in Georgia,” or “college students at the University of Georgia.” Researchers typically try to define their populations as broadly as possibly (e.g., “college students in the United States” rather than “college students in Georgia”) to maximize the applications or implications of their research.
It is typically difficult to collect data from all members of a population. Imagine, for example, the time and money that would be needed to collect information from every college student in the United States. For this reason, researchers typically draw conclusions about populations based on information collected from a sample drawn from the population. A sample is a subset or portion of a population. Table 1.4 describes the samples used in the six studies introduced in Table 1.1. As you can see from Table 1.4, samples greatly vary in terms of their targeted population and size (the number of participants).
32
Determining how Variables will be Measured: Levels of Measurement
The research hypotheses described in Table 1.1 involve variables such as instructional method, test performance, online game playing, and quality of interpersonal relationships. To conduct a research study, the researcher must determine an appropriate way to measure the variables stated in the research hypothesis. Measurement is the assignment of categories or numbers to objects or events according to rules.
For example, to measure the variable “height,” a researcher might use “number of inches from the ground in bare feet” as a form of measurement. To measure a variable such as “success in college,” a student's grade point average (GPA) may be obtained from school transcripts. “Self-esteem” might be measured by having people complete a questionnaire and, on the basis of their responses, be categorized as having either “low” or “high” self- esteem. As these examples demonstrate, the result of measurement is an assignment of a number or a category to each participant in the research study. Different types of variables require different forms of measurement. In recognition of these differences, researchers have identified four distinct levels for measuring variables: nominal, ordinal, interval, and ratio.
The values of variables measured at the nominal level of measurement differ in category or type. The word nominal implies having to do with “names,” such that we use first names and surnames as ways of distinguishing between people. Gender is an example of a nominal variable, in that it consists of categories or types (male and female) rather than numeric values. In the first research hypothesis in Table 1.1, the independent variable, instructional method, is measured at the nominal level, consisting of two categories: learning skills and incentives.
Variables measured at the ordinal level of measurement have values that can be placed in an order relative to the other values. Rankings (such as finishing first, second, or third in a race) and size (small, medium, large, or extra large) are familiar examples of an ordinal scale. Ordinal scales allow researchers to demonstrate that one value represents more or less of a variable than do other values; however, it is not possible to specify the precise size or amount of the difference between values. For example, although you can say that a runner who finishes “first” in a race is faster than the runner who finishes “second,” you cannot specify the exact difference between the two runners' times.
Table 1.4 Research Hypotheses and Study Samples Table 1.4 Research Hypotheses and Study Samples
Research Hypothesis Study Sample
Students who are taught effective learning skills will perform better on tests than students offered
“The participants were 109 juniors and seniors in college … enrolled in three sections of an educational psychology course”
33
incentives to do well. (Tuckman, 1996, p. 200).
Faculty members' beliefs about the frequency of student academic misconduct will be lower than students' beliefs.
“At a medium-sized public university in the northeastern U.S … 166 undergraduate students … 157 members of the faculty” (Hard, Conway, & Moran, 2006, p. 1063).
The more adolescents are exposed to violence in their community, the lower their levels of academic achievement.
“118 adolescents (59 girls, 59 boys) enrolled in a longitudinal study of the effects of violence exposure” (Borofsky, Kellerman, Baucom, Oliver, & Margolin, 2013, p. 383).
Children will rate a disciplining strategy that emphasizes logic and reason as more effective than one based on rewards and punishment.
“663 students enrolled in public educational institutions in Manhattan, Kansas” (Barnett, Quackenbush, & Sinisi, 1996, p. 414).
Heavy users of online games have less fulfilling interpersonal relationships than users spending little or no time playing online games.
“180 students attending a college in northern Taiwan” (Lo, Wang, & Fang, 2005, p. 17).
Drug users are less likely to have work-related accidents after undergoing substance abuse treatment than before the treatment.
“334 drug-test positive workers who completed substance abuse treatment” (Elliot & Shelley, 2006, p. 132).
The values of variables measured at the interval level of measurement are equally spaced along a numeric continuum. One example of an interval variable is the Fahrenheit scale of temperature. Here, a difference of five degrees has the same meaning anywhere along the scale; for example, the difference between 45°F and 50°F is the same as the difference between 65°F and 70°F. Many variables studied in the behavioral sciences (e.g., personality characteristics or attitudes) are considered to be measured at the interval level of measurement. Interval variables not only provide more precise and specific information than do ordinal variables, but they also fulfill the requirements of the most commonly used statistical procedures.
Variables at the ratio level of measurement are identical to interval variables, with one exception: Ratio scales possess what is known as a true zero point, for which the value of zero (0) represents the complete absence of the variable. Variables that describe a physical dimension (such as height, weight, distance, and time duration) typically have a true zero point. In the first research hypothesis in Table 1.1, the researchers in this study measured the variable “test performance” by recording the number of correct answers to a test of reading comprehension. Test performance is a ratio variable because it has a true zero
34
point, where zero would indicate the complete absence of correct answers.
One advantage of ratio measurement versus interval measurement is that ratio variables allow for a greater number of comparisons among values. Consider, for example, the value 6 for the ratio variable “inches.” Not only is the difference between 4 and 6 inches the same as the difference between 6 and 8 inches (thereby involving addition and subtraction), 6 inches is also twice as long as 3 inches and half as long as 12 inches (involving multiplication and division). Multiplication and division comparisons cannot be made with interval variables. You cannot, for example, say that a temperature of 90°F is three times as much temperature as 30°F.
The bold-faced and italicized text in Table 1.5 illustrates how the authors of the six studies described in the earlier tables chose to measure their independent and dependent variables. As you can see, the variables in these studies have been measured in a variety of ways at different levels of measurement.
Table 1.5 Research Hypotheses and Measurement of Variables Table 1.5 Research Hypotheses and Measurement of Variables
Research Hypothesis Independent Variable
Dependent Variable
Students who are taught effective learning skills will perform better on tests than students offered incentives to do well.
Instructional method learning strategy, incentive motivation (nominal)
Test performance number of items correct (ratio)
Faculty members' beliefs about the frequency of student academic misconduct will be lower than students' beliefs.
Member of college community faculty, student (nominal)
Beliefs about the frequency of student academic misconduct 1–5 scale (1 = Never, 5 = Very often) (ordinal)
35
The more adolescents are exposed to violence in their community, the lower their levels of academic achievement.
Exposure to community violence 0 times, 1 time, 2 times, 3+ times (ratio)
Academic achievement Grade point average (GPA) (ratio)
Children will rate a disciplining strategy that emphasizes logic and reason as more effective than one based on rewards and punishment.
Parental discipline strategy power assertion, love withdrawal, induction (nominal)
Effectiveness of parental discipline strategy 1–5 scale (1 = Not at all effective, 5 = Very effective) (ordinal)
Heavy users of online games have less fulfilling interpersonal relationships than users spending little or no time playing online games.
Online game- playing heavy, light, nonplayer (nominal)
Quality of interpersonal relationships 1–6 scale (1 = Strongly agree, 6 = Strongly disagree) (ordinal)
Drug users are less likely to have work- related accidents after undergoing substance abuse treatment than before the treatment.
Time before treatment, after treatment (nominal)
Occurrence of a work- related accident Yes, No (nominal)
What Difference does the Level of Measurement Make?
36
A variable's level of measurement has important implications for researchers in that it influences how research hypotheses are stated as well as how data are analyzed. For example, imagine you are interested in studying the variable “success in college.” You could choose to measure this variable at the nominal level of measurement by having faculty members classify students into one of two groups: successful or unsuccessful. Researchers typically employ nominal variables when they are interested in questions involving differences between groups, such as, “Do successful and unsuccessful college students differ in their study habits?”
You could instead measure success in college as an ordinal variable by having faculty members rank their students from top to bottom. Ranking enables a researcher to study relative differences with less concern for the precise magnitude of these differences. For example, using ranked data, you could ask, “How similar are younger and older faculty members' rankings of their students?”
To study success in college using the interval level measurement, you could have faculty members rate each student on a scale from 1 (low) to 100 (high). You could then use these ratings to address a research question such as, “Is there a relationship between faculty members' ratings of their students and students' ratings of themselves?”
Finally, to measure success in college using the ratio level of measurement, you could use students' salary after graduation (in dollars) as an indicator of success. Salary is a ratio variable because it contains a true zero point. In practice, ratio variables can be used to ask many of the same questions as those involving interval variables. For example, you could ask, “Is there a relationship between faculty members' ratings of students and students' salaries after graduation?”
As you can see, how researchers choose to measure their variables influences how they state their research questions and research hypotheses. The level of measurement used for a variable has another important implication: It helps determine the statistical procedures researchers use to analyze the data. Certain statistical procedures can be applied only to variables measured at the interval or ratio levels of measurement, whereas other procedures are appropriate for variables measured at the nominal or ordinal level.
Selecting a Method to Collect the Data: Experimental and Non-Experimental Research Methods
In addition to drawing a sample from a population and determining how the variables in a study may be measured, the third step in collecting data is to determine the type of research method to use to collect the data on these variables. Research methods may be classified into two main types:
experimental research methods, and
37
non-experimental research methods.
Experimental Research Methods
Experimental research methods are methods designed to test causal relationships between variables—more specifically, whether changes in independent variables produce or cause changes in dependent variables. To make inferences about cause-effect relationships, researchers conducting an experiment must first eliminate all other possible causes or explanations for changes in the dependent variable besides the independent variable. If it can be claimed that a variable other than the independent variable created the observed changes in the dependent variable, the research results are said to be confounded. A confounding variable is a variable related to an independent variable that provides an alternative explanation for the relationship between the independent and dependent variables.
38
Learning Check 2: Reviewing what you've Learned So Far
1. Review questions a. What is the difference between a population and a sample? b. Within the research process, what is the relationship between populations and samples? c. What are the four levels of measurement? How do they differ?
2. For each of the following variables, name the scale of measurement (nominal, ordinal, interval, or ratio).
a. Type of school attended (public, private) b. Probability of graduating college in 4 years (0% to 100%) c. Rating of television (good, better, best) d. Number of computers in home
To understand how confounding variables work, consider the following question: What is the relationship between a mother's ethnicity and the birth weight of her baby? Although research has shown differences in the birth weights of babies of different ethnicities, one study noted that some of this research did not take into account the possibility that differences in birth weights may be partly due to differences between ethnic groups on such factors as the mother's average age at time of pregnancy, socioeconomic status, and behaviors such as smoking and drinking (Ma, 2008). These lifestyle characteristics are considered confounding variables in that they provide alternative explanations for any causal relationship between ethnicity and babies' birth weights.
Researchers minimize the influence of confounding variables in two main ways. First, researchers exercise experimental control by making the research setting (i.e., characteristics of the research participants, location of the experiment, the instruments or measures administered, the instructions given to the research participants) the same for all participants. In the birth weight study, for example, the research design might control for the effect of parental smoking on birth weight by only including nonsmokers in the study sample, excluding those who smoke.
Another way researchers exert control over the research setting is to include a condition known as a control group, which is a group of participants in an experiment not exposed to the independent variable the research is designed to test. For example, if you conducted an experiment designed to examine the effects of caffeine on one's health, one group (the experimental group) might be instructed to drink coffee while a second group (the control group) drinks decaffeinated coffee. By contrasting the results of the two groups, researchers can then assess the impact of caffeinated coffee on the health of those who drink it.
Researchers cannot possibly identify and control for the effects of all potential confounding variables. Consequently, a second strategy used to minimize the influence of confounding
39
variables is random assignment, which is assigning participants to each category of an independent variable in such a way that each participant has an equal chance of being assigned to each category. For example, to assign a participant to either the experimental or the control condition of a study, a researcher might simply flip a coin: heads for experimental, tails for control. The purpose of random assignment is to equalize or neutralize the effects of confounding variables by distributing them equally over all levels of the independent variable.
Non-Experimental Research Methods
Experimental research designs are one of the best tools researchers have for making causal inferences. However, the ability to make causal inferences requires a great deal of control over the situation, control that may not always be possible or desirable. For this reason, researchers often employ non-experimental research methods (sometimes referred to as correlational research methods). non-experimental research methods are research methods designed to measure naturally occurring relationships between variables without having the ability to infer cause-effect relationships. Some of the most common types of non- experimental research designs include quasi-experiments, survey research, observational research, and archival research.
Quasi-experimental research compares naturally formed or preexisting groups rather than employing random assignment to conditions. For example, suppose a researcher wanted to study the effects of different methods of teaching reading on children's verbal skills. Ideally, the researcher would randomly assign a sample of schoolchildren to receive the different teaching methods. However, because children are taught together in classes, it would be difficult to have children in the same classroom receive different methods. Implementing multiple methods in a single classroom would not only place an unreasonable burden on the teacher, but children would also see their classmates being treated differently, which might influence their behavior. To address these concerns, a researcher might assign entire classrooms of children to receive a particular method. Comparing the classrooms is an example of quasi-experimental research.
Survey research methods obtain information directly from a group of people regarding their opinions, beliefs, or behavior. The goal of survey research, which can involve the use of questionnaires and interviews, is to obtain information from a sample that can then be used to represent or estimate the views of a larger population. Because the researcher does not directly manipulate any variables, survey research is not conducted to make causal inferences but is instead used to describe a phenomenon or predict future behavior. As one example of survey research, political pollsters attempt to predict how people will vote in an election by asking a sample of voters about their preferences.
Observational research is the systematic and objective observation of naturally occurring behavior or events. The purpose of observational research is to study behavior or events,
40
with little, if any, intervention on the part of the researcher. Observational research is often used to study phenomena that the researcher either cannot or should not deliberately manipulate. For example, researchers studying aggressive behavior in children would never force children to push or hit each other. By observing playground behavior, however, researchers may be able to record acts of aggression if and when they occur.
Rather than observing or measuring behavior directly, archival research is the use of archives (records or documents of the activities of individuals, groups, or organizations) to examine research questions or hypotheses. One example of an archival research study was interested in studying criminal trials—more specifically, whether the race, age, or gender of an offender is related to the severity of the sentence they receive (Steffensmeier, Ulmer, & Kramer, 1998). To conduct their study, they obtained and analyzed state court records of more than 138,000 criminal trials, recording the severity of the sentences given to the offender as well as the offender's race, age, and gender.
An Example of Combined Experimental and Non-Experimental Research
Both experimental and non-experimental research methods have strengths and weaknesses. The strength of experimental research is the ability to demonstrate cause-effect relationships between independent and dependent variables; however, the control that experiments require creates situations that may not resemble the real world. Non-experimental research methods do not allow the researcher to make causal inferences because they do not involve experimental manipulation, experimental control, or random assignment; however, they have the advantage of allowing researchers to study variables as they naturally occur. Given the strengths and limitations of both research methods, one solution is to use and compare the findings from both methods for examining the same question.
Does playing violent video games lead to aggressive behavior? Two researchers studied this important question using both experimental and non-experimental research methods, saying, “We chose two different methodologies that have strengths that complement each other and surmount each others' weaknesses” (Anderson & Dill, 2000, p. 776).
For their experiment, participants were randomly assigned to one of two conditions, playing either a violent video game or a nonviolent game; “type of video game” was their independent variable. Next, they played another type of game in which they could punish their opponent (who was actually a computer) by delivering a loud blast of noise. The loudness and duration of the noise delivered by participants represented the dependent variable of “aggressive behavior.”
For their non-experimental method, the researchers used a survey methodology, asking students to fill out a questionnaire about the number of hours they played violent video games each week (the independent variable). For the dependent variable of aggressive behavior, students reported the number of times in the previous year that they had
41
performed eight different aggressive acts, such as hitting or stealing from other students.
What did the researchers find? In reporting their findings, they wrote, “In both a correlational investigation using self-reports of real-world aggressive behaviors and an experimental investigation using a standard, objective laboratory measure of aggression, violent video game play was positively related to increases in aggressive behavior” (Anderson & Dill, 2000, p. 787). In evaluating their study, they emphasized the advantages of using both types of research methods, explaining that the non-experimental method “measured video game experience, aggressive personality, and delinquent behavior in real life … [whereas] an experimental methodology was also used to more clearly address the causality issue” (p. 782).
42
Analyzing the Data
Once data have been collected, the next step in the research process involves analyzing them. This part of the research process addresses the primary topic of this book: statistics. Because this is the first chapter of this book, we will not describe analyzing data in detail. Instead, we introduce the notion that there are two main purposes of analyzing data: (1) to organize, summarize, and describe the data that have been collected and (2) to test and draw conclusions about ideas and hypotheses. These two purposes are met by calculating two main types of statistics: descriptive statistics and inferential statistics.
Calculating Descriptive Statistics
Descriptive statistics are statistics used to summarize and describe a set of data for a variable. For example, you have probably heard of crime statistics and unemployment statistics, statistics used to describe or summarize certain aspects of our society. Using a research-related example, Caitlin Abar, a researcher at Brown University, conducted a study examining increases in students' alcohol-related behaviors after entering college (Abar, 2012). To measure students' level of alcohol use, she created a variable called “typical weekend drinking,” which “was measured as the sum of drinks consumed on a typical Friday and Saturday within the past 30 days” (p. 22). One way to organize and summarize students' responses to the typical weekend drinking variable would be to calculate the mean, which is the mathematical average of a set of scores. The mean, described in Chapter 3, is an example of a descriptive statistic. The first part of this book describes a variety of descriptive statistics used by researchers to summarize data they have collected.
Calculating Inferential Statistics
Besides summarizing and describing data, a second purpose of statistics is to analyze data to test hypotheses and draw conclusions. Inferential statistics are statistical procedures used to test hypotheses and draw conclusions from data collected during research studies. By using inferential statistics, researchers are able to test the ideas, assumptions, and predictions on which their research is based. Using inferential statistics, researchers are able to make inferences about the existence of relationships between variables in populations based on information gathered from samples of the population.
As an example of an inferential statistic, let's return to the example of students' alcohol- related behaviors introduced above (Abar, 2012). This study was interested in seeing whether there was a relationship between these behaviors and students' perceptions regarding different aspects of their relationships with their parents. One of these aspects was “alcohol communications,” defined as “the extent that they discussed alcohol related topics with their parents at some point during the past several months” (p. 22). To test a
43
hypothesis regarding the relationship between alcohol communications and students' alcohol use, we may decide to use an inferential statistic known as the Pearson correlation coefficient, which will be discussed in Chapter 13. The chapters in the last half of this book discuss a broad variety of examples of inferential statistics, along with detailed instructions for analyzing and drawing conclusions from statistical data and presenting research findings.
44
Drawing a Conclusion regarding the Research Hypothesis
Once statistical analyses have been completed, the next step is to interpret the results of the analyses as they relate to the research hypothesis. More specifically, do the findings of the analyses support or not support the research hypothesis? The word support in the previous sentence is very important. Students and researchers are sometimes tempted to conclude that their findings “prove” their hypotheses are either true or false. However, as will be discussed in Chapter 5, because researchers typically do not collect data from the entire population, they cannot know with 100% certainty whether their hypothesis is in fact true or false. Also, given the complexity of the phenomena studied by researchers, it is extremely difficult for one research study to prove a hypothesis or theory is completely true or false.
45
Learning Check 3: Reviewing what you've Learned So Far
1. Review questions a. What are the main differences between experimental and Non-experimental research
methods? b. What are confounding variables? What do researchers do to minimize the effects of
confounding variables? c. What are the main types of Non-experimental research methods? d. What are the relative strengths and weaknesses of experimental and non-experimental
research? e. What are the main purposes of descriptive and inferential statistics?
46
Communicating the Findings of the Study
Conducting research requires a variety of different skills: conceptual skills to develop research hypotheses, methodological skills to collect data, and mathematical skills to analyze these data. Another integral part of the research process is the communication skills needed to inform others about a study. Researchers must be not only skilled scientists but also effective writers.
For many years, the American Psychological Association (APA), the professional association for psychologists, has recognized the need to provide researchers guidance on how to communicate the results of their research. In 1929, the APA published a seven-page article in the journal Psychological Bulletin entitled, “Instructions in Regard to Preparation of Manuscript.” In contrast, the sixth edition of the Publication Manual of the American Psychological Association, published in 2010, consists of 272 pages. This increase in length highlights the challenges faced by writers in communicating their research in an effective and efficient manner.
47
1.4 Plan of the Book
The primary purpose of this chapter was to introduce statistics and place it within the larger research process. The remainder of the book will discuss a number of different statistical procedures used by researchers to examine various questions of interest. Although it is critical for you to be able to correctly calculate statistics, it is equally important to understand the role of statistics within the research process and appreciate the conceptual and pragmatic issues related to the use (and sometimes misuse) of statistics.
The first half of the book (Chapters 2 through 6) will introduce you to conceptual and mathematical issues that are the foundation of statistical analyses. Chapters 2, 3, and 4 discuss how data may be examined, described, summarized, and presented in numeric, visual, and graphic form. Descriptive statistics are introduced and discussed in these chapters. Chapters 5 and 6 introduce critical assumptions about and characteristics of the inferential statistical procedures used to test research hypotheses. The remaining chapters of the book will introduce you to a number of different inferential statistical procedures, along with key issues surrounding the use of these procedures.
In most cases, the chapters in this book share a uniform format and structure, and are centered on the research process. In making our presentation, we will employ various examples of actual published research studies, addressing questions with which you may be familiar. In each case, we will clearly state the research hypothesis, briefly describe how the data in the study were collected, introduce and describe the calculation of both descriptive and inferential statistics, and discuss the extent to which the results of the statistical analyses support the research hypothesis. Finally, we will illustrate how to communicate one's research to a broader audience.
48
1.5 Looking Ahead
We began our presentation by defining statistics and providing reasons why learning about them may be of value to you. Next, we placed statistics within the stages of the larger research process, a process that will be the foundation of this textbook. As we have explained, the research process is centered on a research hypothesis, a predicted relationship between variables. Once a hypothesis has been stated, the next step is to collect data about the variables included in the research, after which the process of statistical analysis begins. Analyzing data involves the calculation of two main types of statistics, one designed to describe and summarize the data for a variable (descriptive statistics) and one designed to test research hypotheses (inferential statistics). However, before conducting analyses on a set of data, there is preliminary work that must be completed. The next chapter will focus on methods used by researchers to examine data.
49
1.6 Summary
Statistics may be defined as a branch of mathematics dealing with the collection, analysis, interpretation, and presentation of masses of numerical data. As such, statistics is not only concerned about how data are analyzed but also recognizes the importance of understanding how data are collected and how the results of analyses are interpreted and communicated.
There are a variety of reasons why students should learn statistics: Statistics are encountered in a wide variety of daily activities, students may be asked or required to read and interpret the results of statistical analyses in their courses, students may use statistics in conducting their own research, and statistics may help one's career.
Researchers conduct research using the scientific method of inquiry, a method of investigation that uses the objective and systematic collection and analysis of empirical data to test theories and hypotheses. The research process used within the scientific method of inquiry consists of five main steps: developing a research hypothesis to be tested, collecting data, analyzing the data, drawing a conclusion regarding the research hypothesis, and communicating the findings to the study.
A research hypothesis is a statement regarding an expected or predicted relationship between variables. Developing a research hypothesis involves identifying a question or issue to be examined, reviewing and evaluating relevant theories and research, and stating the research hypothesis to be tested in the study. A theory is a set of propositions used to describe or explain a phenomenon. A research hypothesis contains variables, which are properties or characteristics of some object, event, or person that can take on different values. More specifically, a research hypothesis states the nature and direction of a proposed relationship between an independent variable (a variable manipulated by the researcher) and a dependent variable (a variable measured by the researcher).
Collecting data involves drawing a sample from a population, determining how the variables will be measured, and selecting a method to collect the data. A population is the total number of possible units or elements that could potentially be included in a study; a sample is a subset or portion of a population. Variables can be measured at one of four levels of measurement: nominal (values differing in category or type), ordinal (values placed in an order relative to the other values), interval (values equally spaced along a numeric continuum), or ratio (values equally spaced along a numeric continuum with a true zero point).
There are two main types of methods used to collect data: experimental research methods, designed to test whether changes in independent variables produce or cause changes in dependent variables, and Non-experimental research methods, designed to examine the
50
relationship between variables without having the ability to infer cause-effect relationships.
In experimental research, researchers are concerned about possible confounding variables, which are variables related to independent variables that provide an alternative explanation for the relationship between independent and dependent variables. To minimize the impact of confounding variables, a research study may exercise experimental control over various aspects of the situation, including the use of a control group, which is a group of participants not exposed to the independent variable the research is designed to test, or random assignment, which involves assigning participants to each category of an independent variable in such a way that each participant has an equal chance of being assigned to each category.
Examples of non-experimental research methods are survey research, in which information is directly obtained from a group of people regarding their opinions, beliefs, or behavior; observational research, which involves the systematic and objective observation of naturally occurring events; and archival research, which uses archives (records or documents of the activities of individuals, groups, or organizations) to examine research questions or hypotheses.
Once data have been collected, the next step is to analyze them using two main types of statistics: descriptive statistics, which summarize and describe a set of data, and inferential statistics, which are statistical techniques used to test hypotheses and draw conclusions from data collected during research studies.
Once the statistical analyses have been completed, the results of the analyses are interpreted regarding whether they support or not support the study's research hypothesis. The final step in the research process is to communicate the study to a broader audience.
51
1.7 Important Terms
statistics (p. 1) scientific method (p. 3) research hypothesis (p. 3) variable (p. 3) theory (p. 5) independent variable (p. 5) dependent variable (p. 5) population (p. 9) sample (p. 9) measurement (p. 9) level of measurement (nominal, ordinal, interval, ratio) (p. 9–10) experimental research methods (p. 12) confounding variable (p. 13) control group (p. 13) random assignment (p. 14) non-experimental research methods (p. 14) quasi-experimental research (p. 14) survey research (p. 14) observational research (p. 14) archival research (p. 14) descriptive statistics (p. 16) inferential statistics (p. 16)
52
1.8 Exercises
1. Listed below are a number of hypothetical research hypotheses. For each hypothesis, identify the independent and dependent variable.
a. Male drivers are more likely to exhibit “road rage” behaviors such as aggressive driving and yelling at other drivers than are female drivers.
b. The more time a student takes to finish a midterm examination, the higher his or her score on the examination.
c. Men are more likely to be members of the Republican political party than are women; women are more likely to belong to the Democratic political party than are men.
d. Students who receive a newly designed method of teaching reading will display higher scores on a test of comprehension than the method currently used.
e. The more time a child spends in daycare outside of the home, the less he or she will be afraid of strangers.
2. Listed below are a number of research questions and hypotheses from actual published articles. For each hypothesis, identify the independent and dependent variable.
a. “The use of color in a Yellow Pages advertisement will increase the perception of quality of the products for a particular business when compared with noncolor advertisements” (Lohse & Rosen, 2001, p. 75).
b. “We hypothesized that parents who use more frequent corporal and verbal punishment … will report more problem behaviors in their children” (Brenner & Fox, 1998, p. 252).
c. “It was hypothesized that adolescents with anorexia nervosa would … be more respectful when compared with peers with bulimia nervosa” (Pryor & Wiederman, 1998, p. 292).
d. “The purpose of the present research was to assess brand name recognition as a function of humor in advertisements…. It was predicted that participants would recognize product brand names that had been presented with humorous advertisements more often than brand names presented with nonhumorous advertisements” (Berg & Lippman, 2001, p. 197).
e. “The purpose of this study was to explore the relation between time spent in daycare and the quality of exploratory behaviors in 9-month-old infants … it was hypothesized that … infants who spent greater amounts of time in center- based care would demonstrate more advanced exploratory behaviors than infants who did not spend as much time in center-based care” (Schuetze, Lewis, & DiMartino, 1999, p. 269).
f. “It was hypothesized that students would score higher on test items in which the narrative contains topics and elements that resonate with their daily experiences versus test items comprised of material that is unfamiliar” (Erdodi,
53
2012, p. 172). 3. Listed below are additional research questions and hypotheses from actual published
articles. For each hypothesis, identify the independent and dependent variable. a. “It is expected that achievement motivation will be a positive predictor of
academic success” (Busato, Prins, Elshout, & Hamaker, 2000, p. 1060). b. “Men are expected to employ physical characteristics (particularly those that are
directly related to sex) more often than women in selecting dating candidates” (Hetsroni, 2000, p. 91).
c. “The purpose of our study was to gain a better understanding of the relationship between social functioning and problem drinking…. We predicted that problem drinkers would endorse more social deficits than nonproblem drinkers” (Lewis & O'Neill, 2000, pp. 295–296).
d. “We predicted that people in gain-framed conditions would show greater intention to use sunscreen … than those people in loss-framed conditions” (Detweiler et al., 1999, p. 190).
e. “Students with learning disabilities who received self-regulation training would obtain higher reading comprehension scores than students (with learning disabilities) in the control group” (Miranda, Villaescusa, & Vidal-Abarca, 1997, p. 504).
f. “We hypothesized that consumers would think that a ‘sale price’ presentation would generate a greater monetary savings than an ‘everyday low price’ presentation” (Tom & Ruiz, 1997, p. 403).
g. “We hypothesize that physical coldness (vs. warmth) would activate a need for psychological warmth, which in turn increases consumers' liking of romance movies” (Hong & Sun, 2012, p. 295).
4. Name the scale of measurement (nominal, ordinal, interval, ratio) for each of the following variables:
a. The amount of time needed to react to a sound b. Gender c. Score on the Scholastic Aptitude Test (SAT) d. Political orientation (not at all conservative, conservative, very conservative) e. Political affiliation (Democrat, Republican, Independent)
5. Name the scale of measurement (nominal, ordinal, interval, ratio) for each of the following variables:
a. One's age (in years) b. Size of soft drink (small, medium, large, extra large) c. Voting behavior (in favor vs. against) d. IQ score e. Parent (mother vs. father)
6. A faculty member wishes to assess the relationship between students' scores on the Scholastic Aptitude test (SAT) and their performance in college.
a. What is a possible research hypothesis in this situation?
54
b. What are the independent and dependent variables? c. How could you measure the variable “performance in college” at each of the
four levels of measurement? 7. A researcher hypothesizes that drivers who use cellular phones will get into a greater
number of traffic accidents than drivers who do not use these phones. For this example,
a. What are the independent and dependent variables? b. How could you measure the dependent variable? c. How could you conduct this study using an experimental research method?
How could you conduct this study using a non-experimental method?
55
Answers to Learning Checks
Learning Exercise 1
2.
a. IV: Instructors' dress; DV: Expertise
b. IV: Amount of television viewing; DV: Posttraumatic stress symptoms
c. IV: Estimation of HSAS level; DV: Worry about terrorism
d. IV: Use of humor; DV: Product recognition
e. IV: Participation in intervention program;
DV: Literacy skills
Learning Exercise 2
2. a. Nominal b. Ratio c. Ordinal d. Ratio
56
Answers to Odd-Numbered Exercises
1.
a. Independent variable (IV): Gender;
Dependent Variable (DV): “Road rage” behaviors
b. IV: Time; DV: Exam score
c. IV: Gender; DV: Political affiliation
d. IV: Method of teaching; DV: Comprehension test scores
e. IV: Time in daycare; DV: Fear of strangers
3.
a. IV: Achievement motivation; DV: Academic success
b. IV: Gender; DV: Emphasis on physical characteristics
c. IV: Alcohol drinking; DV: Social deficits
d. IV: Framing; DV: Intention to use sunscreen
e. IV: Training program; DV: Reading comprehension scores
f. IV: Price presentation; DV: Perception of savings
g. IV: Physical temperature; DV: Liking of romance movies
5. a. Ratio b. Ordinal c. Nominal d. Interval e. Nominal
7. a. IV: Cellular phone use; DV: Number of traffic accidents b. Example: Determine the number of traffic accidents each person has
experienced in the past year.
c. Example of experimental: Use a driving simulation with the “driver” talking on a phone and measure the number of potential driving mistakes or accidents.
Example of non-experimental: Take a survey of the number of accidents people have been in and whether or not they were using a phone during the accident.
57
Sharpen your skills with SAGE edge at edge.sagepub.com/tokunaga
SAGE edge for students provides a personalized approach to help you accomplish your coursework goals in an easy-to-use learning environment. Log on to access:
eFlashcards Web Quizzes Chapter Outlines Learning Objectives Media Links
58
Chapter 2 Examining Data: Tables and Figures
59
Chapter Outline 2.1 An Example From the Research: Winning the Lottery 2.2 Why Examine Data?
Gaining an initial sense of the data Detecting data coding or data entry errors Identifying outliers Evaluating research methodology Determining whether data meet statistical criteria and assumptions
2.3 Examining Data Using Tables What are the values of the variable? How many participants in the sample have each value of the variable? What percentage of the sample has each value of the variable?
2.4 Grouped Frequency Distribution Tables Cumulative percentages
2.5 Examining Data Using Figures Displaying nominal and ordinal variables: bar charts and pie charts Displaying interval and ratio variables: histograms and frequency polygons Drawing inappropriate conclusions from figures
2.6 Examining Data: Describing Distributions Modality Symmetry Variability
2.7 Looking Ahead 2.8 Summary 2.9 Important Terms 2.10 Formulas Introduced in This Chapter 2.11 Using SPSS 2.12 Exercises
Chapter 1 described a process by which scientific research may be conducted. This process begins by stating a research hypothesis regarding an expected relationship between variables. The next step involves collecting data, data that will ultimately be used in statistical analyses designed to test the research hypothesis. Researchers do not conduct statistical analyses, however, until they have first examined the data to ensure appropriate conclusions can be drawn from the analyses. This chapter describes why researchers examine data and how data may be examined using tables, charts, and graphs. As with many of the chapters in this book, this discussion will be guided by the findings from a published research study.
60
2.1 An Example from the Research: Winning the Lottery
A friend of yours flips a coin and asks you to guess whether the coin has landed “heads” or “tails.” You guess “tails,” but the answer is “heads.” Your friend flips the coin a second time, and again you predict “tails.” But again the answer is “heads.” On the third coin flip, you guess “tails” once more—but once again you are wrong. As your friend prepares to flip the coin again, you tell yourself there is little chance the coin could land on “heads” a fourth consecutive time. But the reality is that the fourth coin flip is equally likely to land on “heads” or “tails.” This is because the outcome of the first three coin flips has absolutely no impact on the outcome of the fourth flip. The outcome of a coin flip is a random event that cannot be controlled or predetermined. But do you really believe this?
The impact of random events can take on greater significance than the outcome of a coin flip. For example, some government-run lotteries provide the opportunity to win huge amounts of money through the random selection of numbers. In 2001, Dr. Karen Hardoon, a researcher at McGill University in Canada, studied the thought processes people might use when purchasing lottery tickets (Hardoon, Baboushkin, Derevensky, & Gupta, 2001). As previous research had found some gamblers believe they can control random events such as the rolling of dice or the spinning of a roulette wheel, Dr. Hardoon and her colleagues were interested in determining “whether individuals with gambling problems perceive the purchase of lottery tickets in a similar manner as non-problem gamblers” (p. 752). On the basis of the earlier findings, Dr. Hardoon hypothesized that problem gamblers are less likely to believe lottery outcomes are random than are non– problem gamblers.
To test their research hypothesis, Dr. Hardoon and her associates located a group of gamblers and asked whether they had ever experienced problems as a result of their gambling. The research team then classified the participants into two groups: problem gamblers and non–problem gamblers.
Each participant in the study, which we will refer to as the lottery study, was shown four lottery tickets and was asked, “If you were to buy a ticket to play in the lottery, which one would you select?” Each ticket was composed of six numbers selected from among the numbers 1 through 49; the four tickets were given a particular label. The sequence ticket had numbers in consecutive order (30–31–32–33–34–35); numbers on the pattern ticket increased by 5s (5–10–15–20–25–30); the nonequilibrated numbers (35–37–40–43–44– 49) were clustered at the upper end of the 49 numbers, and the random ticket (7–8–23– 34–36–42) shared none of the characteristics of the other three tickets. We will call the variable in this study “lottery ticket,” a variable with four values or categories: sequence, pattern, nonequilibrated, and random.
The researchers recorded which of the four tickets was chosen by each participant. Table
61
2.1 lists the data for the lottery ticket variable for the non–problem gamblers (the data for the problem gamblers in the study are provided in the exercises at the end of this chapter). Looking at this table, we see that the problem gamblers differed in their lottery ticket choices; however, it is difficult to fully understand or describe the extent or nature of these differences given the unorganized nature of the collected data. Consequently, before conducting any statistical analyses on their data, researchers typically organize and examine it. In the following sections, we will discuss why and how data are organized and examined.
62
2.2 Why Examine Data?
Before statistically analyzing their data, there are a variety of reasons why researchers first examine them. These reasons include the following:
Table 2.1 The Lottery Ticket Choices of 22 Non–Problem Gamblers Table 2.1 The Lottery Ticket Choices of 22 Non–Problem
Gamblers
Participant Lottery Ticket
1 Random
2 Random
3 Nonequilibrated
4 Pattern
5 Random
6 Sequence
7 Random
8 Random
9 Pattern
10 Pattern
11 Random
12 Pattern
13 Random
14 Nonequilibrated
15 Random
16 Pattern
17 Nonequilibrated
18 Random
19 Random
20 Pattern
21 Random
22 Random
63
to gain an initial sense of the data, to detect data coding or data entry errors, to identify outliers, to evaluate research methodology, and to determine whether data meet statistical criteria and assumptions.
Each of these reasons is discussed below.
64
Gaining an Initial Sense of the Data
Researchers spend a great amount of time and energy designing their research studies; consequently, they are eager to analyze the data they collect so that they may test their research hypotheses. Students taking courses in statistics, on the other hand, face a different challenge: They are given sets of data with which they have no prior experience and have only a short amount of time to learn and apply statistical formulas to the data. Although both researchers and students may be motivated to analyze their data as quickly as possible, it is important to first examine the data in order to gain an initial sense of them.
Imagine you conduct a study regarding the fuel efficiency of automobiles in which you collect data on the miles per gallon (MPG) of different types of cars. As such, you might expect the scores in your data set to range from 10 to 40 MPG, with the majority of scores between 15 and 25 MPG. Examining data helps researchers gain an initial sense of the data they have collected and whether the data are in line with their expectations. Examining data also provides a way of confirming the results of later statistical analyses. We have found students often make errors in their calculations that could have been avoided had they first looked at the set of data. For example, examining data will help you avoid making statements such as, “The MPG for the cars in my sample ranged from 10 to 300 MPG” or “The average GPA in my sample of college students was 6.36.”
65
Detecting Data Coding or Data Entry Errors
Another reason for examining data before conducting statistical analyses is to detect any errors made in the coding of data that have been collected. Imagine, for example, that a group of students is asked to complete a 20-item test of mathematical calculations; each student's score on the test is the number of items answered correctly. Once the test scores are calculated, they are entered into a computer file. In this situation, it is crucial to ensure that no mistakes are made in calculating the test scores or in data entry. Data coding and data entry errors are not at all uncommon. Left unattended, seemingly minor errors can lead to a great loss of time and energy—including the need, in some cases, to repeat the statistical analysis of the data.
66
Identifying Outliers
Researchers also examine their data to identify outliers, defined as rare, extreme scores that lie outside of the range of the majority of scores in a set of data. Returning to our example of the 20-item test of mathematical calculations, suppose the test is given to 15 students. In scoring the tests, a researcher finds that 14 students correctly answered between 8 and 17 questions. However, one student provided the wrong answer to every question on the test, for a score of 0. In this case, the score of 0 would be considered an outlier.
If they are not identified, outliers may distort conclusions researchers draw about their sample and their data, particularly when the sample is relatively small. In the above example, the single score of 0 might lead an observer to conclude the class as a whole performed more poorly on the test than they actually did.
67
Evaluating Research Methodology
Examining data is also useful in assessing the effectiveness of the specific methods researchers use to collect their data. For the 20-item math test, for example, the possible scores range from 0 to 20 questions answered correctly. Having administered the test to 15 students, it might be reasonable to expect the majority of the data should fall in the middle of the 0 to 20 range. But what if every student answered at least 16 of the 20 questions correctly? Based on such a finding, we might conclude that the test was too easy and that a more difficult set of questions is needed to accurately assess students' level of knowledge. By comparing the collected data for a variable with the range of possible values, researchers are able to evaluate and modify their measurement tools.
68
Determining Whether Data Meet Statistical Criteria and Assumptions
Researchers also examine their data to see whether they meet certain statistical criteria and assumptions. Many of the statistical procedures discussed in this book are based on specific assumptions regarding the shape and nature of a set of data that has been collected. These procedures generally assume, for example, that in most data sets, the majority of the data points will fall in the middle of the range of possible values, with a relatively small amount of data at the highest and lowest possible values. In the example of the 20-item math test, it might be assumed that most scores will be between 7 and 13, with smaller number of scores either less than 7 or greater than 13. When assumptions such as this are not met, it is more likely a researcher may draw inappropriate conclusions from any statistical analyses that are conducted.
69
2.3 Examining Data Using Tables
One part of examining data that have been collected for a variable involves answering a series of questions:
What are the values of the variable? How many participants in the sample have each value of the variable? What percentage of the sample has each value of the variable?
As it would be difficult to answer questions such as these just by looking at the data in Table 2.1, researchers typically organize their data by creating a table known as a frequency distribution table, a table that summarizes the number and percentage of participants for the different values of a variable. This section discusses the steps involved in creating frequency distribution tables using the questions listed above.
70
What are the Values of the Variable?
The first step in organizing the data for a variable is to identify all of the possible values for the variable. Table 2.2 begins the construction of a frequency distribution table by listing the four values for the lottery ticket variable.
71
How Many Participants in the Sample have Each Value of the Variable?
Once the values of the variable have been identified, the next step in organizing data that have been collected is to determine the frequency of each value, which is the number of participants in a sample corresponding to a value of a variable. These frequencies provide an indication of the nature and shape of the distribution of values for a variable in a sample.
Table 2.2 Identifying the Values of the Lottery Ticket Variable Table 2.2
Identifying the Values of the
Lottery Ticket Variable
Lottery Ticket
Sequence
Pattern
Nonequilibrated
Random
One way to determine the frequencies for a variable is to sort the data into the different values and then count the number of participants having each value using a “tick mark method.” Using this method, a tick mark is placed next to a value of the variable each time that value appears in the sample. As shown in Table 2.3, the tick marks are then counted to determine the frequency for each value. For example, the number of tick marks and therefore the frequency (f) for the sequence ticket is 1 because it was picked by one participant (Participant 6 in Table 2.1). Similarly the tick marks for each of the other three types of tickets can be recorded and counted until all of the data in the sample have been accounted for.
Table 2.3 Using the Tick Mark Method to Determine the Frequencies of the Lottery Ticket Variable Table 2.3 Using the Tick Mark
Method to Determine the Frequencies of the Lottery Ticket
Variable
Lottery Ticket Tick Marks f
Sequence 1
Pattern 6
72
Nonequilibrated 3
Random 12
Total 22
73
What Percentage of the Sample has Each Value of the Variable?
Although it is important to determine the frequency for each value of a variable, these frequencies by themselves are not always particularly informative. For example, the statement “There are 35 Democrats in my sample” may be interpreted differently depending on whether the total sample consists of 40 voters versus 400 voters.
Because the interpretation of frequencies depends on the size of the sample, it is useful to calculate the percentage of the sample having each value of a variable. These percentages can be calculated using the following formula:
(2-1) % = f total number of scores ∗ 100
where f is the frequency for a value of a variable.
From Table 2.3, we found that 1 of the 22 non-problem gamblers chose the sequence ticket. Therefore, the percentage of the sample choosing the sequence ticket is %= f t o t a l n u m b e r o f s c o r e s ∗ 100 = 1 22 ∗ 100 = .05 ∗ 100 = 5 %
After calculating the percentage for each of the four tickets, Table 2.4 provides the frequency distribution table for the lottery ticket variable.
Once a frequency distribution table has been constructed for a variable, researchers can begin to draw conclusions about the data in their sample. For the lottery ticket variable, the random ticket had the highest frequency and percentage (f = 12, 54%) compared to the pattern (f = 6, 27%), non-equlibrated (f = 3, 14%), and sequence (f = 1, 5%) tickets. As a result, the researchers in the lottery study made the following observation:
The results of the present study indicated that, for the entire sample, the most commonly cited reason for selecting a lottery ticket was perceived randomness. Furthermore, with respect to actual ticket selections, irrespective of explanations, the greatest percentage of tickets selected by the entire sample … were random tickets. (Hardoon et al., 2001, p. 760)
Table 2.4 Frequency Distribution Table for the Lottery Ticket Variable Table 2.4 Frequency
Distribution Table for the Lottery Ticket Variable
74
Lottery Ticket f %
Sequence 1 5%
Pattern 6 27%
Nonequilibrated 3 14%
Random 12 54%
Total 22 100%
75
Learning Check 1: Reviewing what you've Learned So Far
1. Review questions a. What are some reasons for examining data before conducting statistical analyses? b. Why are outliers of concern to researchers? c. What questions do you address about a set of data for a variable in constructing frequency
distribution tables? d. Why is it useful to calculate percentages in addition to frequencies for each value of a
variable? 2. For each of the following situations, create a frequency distribution table.
a. A college counselor asks a group of 116 seniors what their plans are after graduating from college; 71 say they are going to “work,” 34 are planning on going to “graduate school,” and 11 say they are “not sure” what their plans are.
b. A high school instructor teaching a course finds that of the students she teaches, 69 are Freshmen, 18 are Sophomores, 12 are Juniors, and 9 are Seniors.
c. A pollster stops 12 people and asks if they are either “in favor” or “against” a local proposition:
Person Position
1 In favor
2 In favor
3 Against
4 In favor
5 Against
6 Against
7 In favor
8 Against
9 In favor
10 Against
11 In favor
12 In favor
d. The company that makes M&Ms candy (www.mms.com) conducted a survey asking people which new color they would like to have added: purple, aqua, or pink. Below are the votes of a hypothetical sample:
Person Color
1 Pink
2 Purple
3 Pink
4 Aqua
5 Purple
6 Purple
76
7 Purple
8 Aqua
9 Purple
10 Pink
11 Aqua
12 Pink
13 Pink
14 Purple
15 Aqua
16 Pink
17 Purple
18 Purple
19 Aqua
20 Purple
21 Purple
22 Pink
23 Pink
24 Purple
77
2.4 Grouped Frequency Distribution Tables
The frequency distribution table for the lottery ticket variable in Table 2.4 was sufficient for presenting the frequencies for each of the four lottery ticket choices. However, other situations may involve numeric variables with a large number of possible values. Consider, for example, the variable grade point average (GPA). The values for GPA typically range from 0.00 to 4.00, with hundreds of possible values in between. Creating a frequency distribution table for a variable such as GPA would result in an extremely large table, with many values having frequencies of zero.
To examine a variable with a large number of values, it is often useful to construct a grouped frequency distribution table, a table that groups the values of a variable measured at the interval or ratio level of measurement into a small number of intervals and then provides the frequency and percentage within each interval. To illustrate how grouped frequency distribution tables are constructed, Table 2.5(a) provides a sample of 42 hypothetical GPAs. In Table 2.5(b), each of these 42 GPAs has been placed into one of eight intervals that comprise a grouped frequency distribution table.
To illustrate how to read and interpret a grouped frequency distribution table, the lowest interval in Table 2.5(b) consists of students with GPAs less than (<) 2.00. The frequency in this interval (f = 1) consists of one student: Student 29 (GPA = 1.96). The next interval consists of students with GPAs ranging from 2.00 to 2.29 and contains three students (4, 22, and 38). As in our previous example, the grouping continues until all of the GPAs in the sample have been accounted for.
To create a grouped frequency distribution table, the real limits of each interval need to be identified. The real limits of an interval are the values of the variable that fall halfway between the top of one interval and the bottom of the next interval. For example, because of how values of GPA are rounded, the lowest value of GPA that would lead someone to be placed in the interval 2.30–2.59 is not 2.30 but rather 2.295, which is halfway between the top of the 2.00–2.29 interval and the bottom of the 2.30–2.59 interval. In this case, 2.295 represents the real lower limit for the 2.30–2.59 interval. The real lower limit is the smallest value of a variable that would be grouped into a particular interval. On the other hand, the real upper limit is the largest value of a variable that would be grouped into a particular interval. For the interval 2.30–2.59, the real upper limit is 2.594. This is the real upper limit because it falls halfway between the top of one interval (2.59) and the bottom of the next interval (2.60).
What preliminary conclusions might be drawn about these students from examining the GPAs in Table 2.5(b)? First, because the interval with the greatest frequency is 2.90–3.19 (f = 10), you might conclude that the typical student in this sample has a GPA close to 3.0. Second, moving in both directions away from the 2.90–3.19 interval, the frequencies of the
78
different intervals grow progressively smaller, with fewer and fewer students receiving either very high or very low GPAs.
Table 2.6 provides a list of guidelines for creating grouped frequency distribution tables. Grouped frequency distribution tables are useful in summarizing variables that have a large number of values. However, by combining values of a variable, these tables provide little detail and specificity about individual values. For example, although Table 2.5(b) reveals that 10 students have GPAs in the interval of 2.90–3.19, it does not provide the exact GPA of any particular student. For example, there is no way of determining how many students had a GPA of exactly 3.00.
79
Cumulative Percentages
In addition to determining the percentage of the sample that has each individual value or interval for a variable, it is sometimes useful to combine these percentages. For the GPA example, you may be interested in knowing, “What percentage of the sample had GPAs less than 2.60?” To answer this question, the percentages in the “%” column of the frequency distribution table must be combined. This type of percentage is referred to as a cumulative percentage, defined as the percentage of a sample at or below a particular value of a variable.
Table 2.5 Grade Point Average (GPA), 42 Students Table 2.5 Grade
Point Average (GPA), 42 Students
a. Individual GPAs
Student GPA
1 3.05
2 2.83
3 3.26
4 2.19
5 3.52
6 3.34
7 3.02
8 2.97
9 3.71
10 2.50
11 2.77
12 3.10
13 3.06
14 3.83
15 2.39
16 3.16
17 3.92
18 2.37
19 3.65
80
20 3.70
21 3.00
22 2.26
23 2.65
24 3.40
25 2.87
26 2.92
27 2.89
28 3.41
29 1.96
30 3.77
31 3.20
32 2.86
33 2.94
34 2.71
35 3.43
36 3.95
37 3.32
38 2.28
39 3.08
40 3.25
41 2.40
42 3.22
Table 2.5 Grade Point Average (GPA), 42
Students
b. Grouped Frequency Distribution Table
GPA f %
3.80+ 3 7%
3.50–3.79 5 12%
3.20–3.49 9 21%
2.90–3.19 10 24%
81
2.60–2.89 7 17%
2.30–2.59 4 10%
2.00–2.29 3 7%
< 2.00 1 2%
Total 42 100%
Table 2.6 Guidelines for Creating Grouped Frequency Distribution Tables Table 2.6 Guidelines for Creating Grouped Frequency Distribution Tables
1. Variables are grouped into approximately 10 intervals.
The exact number of intervals depends on the data in each sample. For example, data that fall within a small range of values may be accurately represented with a relatively small number of intervals, whereas data extending across a wide range may require a larger number of intervals.
2. The number of intervals should accurately represent the data.
The intervals should represent the nature of the data as accurately as possible. For example, there should not be so large a number of intervals that many intervals have frequencies of zero (for example, 3.00–3.05) or so few intervals that a large majority of the sample falls into only one or two intervals (for example, 3.00–3.99).
3. Intervals should be of equal size.
To make the intervals comparable, they should be of equal size or width. For example, you would not want one interval of 2.01–2.50 (a width of .50) and another of 2.51–2.70 (a width of .20) because the first interval includes a wide range of GPAs and the second contains a much narrower range. One exception to this guideline is the lowest or highest interval (for example, less than 2.00 or greater than 3.70), which is typically wider than the others because it contains a relatively small frequency.
4. Intervals should not overlap.
Each interval should be distinct from the other intervals. For example, you would not want to have one interval be 2.30–2.60 and the next be 2.60–2.90 because the GPA 2.60 appears in both intervals. This would lead to confusion regarding which category to place someone with a GPA of 2.60. Each score should be included in one, and only one, interval.
82
Table 2.7 includes the cumulative percentages for the GPA example. As an example of a cumulative percentage, the cumulative percentage for the 2.30–2.59 interval (19%) has been calculated by combining the percentages in the < 2.00, 2.00–2.29, and 2.30–2.59 intervals (2% + 7% + 10% = 19%). As a result, we could conclude that 19% of this sample had GPAs less than 2.60. The accumulating of percentages continues until all (100%) of the scores in the sample are combined. As shown in Table 2.7, cumulative percentages are typically located in the final column of a frequency distribution table and labeled “Cum %.” (Although we have used a grouped frequency distribution table to illustrate the concept of cumulative percentages, keep in mind that cumulative percentages may also be calculated for frequency distribution tables in which the values for the variable are not grouped together.)
83
2.5 Examining Data Using Figures
In addition to organizing data into frequency distribution tables, researchers also examine data visually using figures such as charts or graphs. There are four types of figures often used to visually portray data for variables:
Table 2.7 Frequency Distribution Table with Cumulative Percentages for the GPA Variable
Table 2.7 Frequency Distribution Table with Cumulative Percentages
for the GPA Variable
GPA f % Cum %
3.80+ 3 7% 100%
3.50–3.79 5 12% 93%
3.20–3.49 9 21% 81%
2.90–3.19 10 24% 60%
2.60–2.89 7 17% 36%
2.30–2.59 4 10% 19%
2.00–2.29 3 7% 9%
< 2.00 1 2% 2%
Total 42 100%
bar charts, pie charts, histograms, and frequency polygons.
Which type of figure a researcher may use will depend on the variables level of measurement (nominal, ordinal, interval, ratio). Data for variables measured at the nominal or ordinal level of measurement are typically displayed using bar charts or pie charts. The values of interval and ratio variables are graphed using histograms and frequency polygons. Each of these types of visual illustration is discussed below.
84
Displaying Nominal and Ordinal Variables: Bar Charts and Pie Charts
The values of variables measured at the nominal level of measurement differ in category or type; one example of a nominal variable is the lottery ticket variable discussed earlier, which consisted of four categories: sequence, pattern, nonequilibrated, and random. The values of variables measured at the ordinal level of measurement can be placed in an order relative to the other values; size (small, medium, large) is one example of an ordinal variable.
One way to display nominal and ordinal variables is to use a bar chart. A bar chart is a figure that uses bars to represent the frequency or percentage of a sample corresponding to each value of a variable, with the different bars not touching each other. Figure 2.1 shows a bar chart for the data from the lottery ticket variable. In this figure, each bar corresponds to one of the four types of lottery tickets; the height of each bar corresponds to the frequency (f) of non-problem gamblers who selected that type of ticket.
Figure 2.1 Bar Chart Showing the Frequencies of the Lottery Ticket Variable for the Non–Problem Gamblers
In general, when creating a bar chart for a variable, the values of the variable are placed along the horizontal (X) axis and the frequency (f) or percentage (%) for each value is placed along the vertical (Y) axis. Note that the different bars in the bar chart in Figure 2.1 do not touch or intersect. The gap between bars indicates that the values of the variable represent distinct groups or categories that cannot be connected together along a numeric continuum.
A second type of figure used to display nominal or ordinal variables is a pie chart. A pie chart is a figure that uses a circle divided into proportions to represent the percentage of the sample corresponding to each value of a variable. Pie charts are created by dividing the 360 degrees of a circle into pieces (or proportions) corresponding to the percentage of each value.
85
Figure 2.2 shows a pie chart for the data for the lottery ticket variable. To illustrate how pie charts are constructed, the frequency distribution table in Table 2.4 shows that 5% of non– problem gamblers chose the sequence ticket. Given that a pie chart is a circle consisting of 360 degrees, 5% of the 360 degrees are needed to represent the sequence ticket. As 5% ∗
360 = 18, 18 degrees of the pie is assigned to the sequence ticket. Similarly, 97 degrees of the pie chart (27% of 360) are used to represent the pattern ticket. This process continues until all 360 degrees of the pie chart are accounted for.
86
Displaying Interval and Ratio Variables: Histograms and Frequency Polygons
Interval and ratio variables have values equally spaced along a numeric continuum and are identical to each other with one exception: Ratio scales possess a true zero point, for which the value of zero represents the complete absence of the variable. Examples of ratio variables include length (measured in inches) and speed (measured in miles per hour). The values of interval and ratio variable may be illustrated by using histograms and frequency polygons, the choice of which depends on the number of values for the variable.
Figure 2.2 Pie Chart Showing the Percentages of the Lottery Ticket Variable for the Non–Problem Gamblers
A histogram is a figure in which bars are used to represent the frequency of each value of a variable, with the different bars touching each other. The values of the variable are located on the X-axis and the frequency or percentage for each value is on the Y-axis. Histograms are similar to bar charts, with one important difference: The bars in a histogram touch each other, indicating that the values of the variable are connected to each other along a numeric continuum.
Imagine a researcher is interested in studying aggressive behavior in children. She decides to measure this variable by counting the number of fights each student is involved in over a given time period. “Number of fights” is a ratio variable with a small number of values; it is a ratio rather than interval variable because the value of 0 represents the complete absence of fights. Figure 2.3 provides a frequency distribution table and histogram for this variable.
A histogram is used when an interval or ratio variable consists of a relatively small number of values. But what if, for example, the number of fights could be anywhere from 0 to 20 rather than 0 to 5? In order to have the histogram fit on one page, the 21 bars would need to be very thin and crowded together, making the histogram confusing and difficult to read. In situations such as this, the values of the variable can be more clearly illustrated
87
using a figure known as a frequency polygon. A frequency polygon is a line graph that uses data points to represent the frequency of each value of a variable, with lines connecting the data points. In a frequency polygon, the X-axis represents the values of the variable, while the Y-axis represents the frequency for each value.
Figure 2.3 Frequency Distribution Table and Histogram for the Number of Fights Variable
Figure 2.4 Frequency Polygon Showing the Frequency of Students in Each Interval of the Grade Point Average (GPA) Variable
Figure 2.4 shows a frequency polygon based on the grouped frequency distribution table of the earlier example of students' GPA (Table 2.5(b)). In this frequency polygon, the data points representing the frequencies of the eight GPA intervals are connected using straight, sharp lines. These lines indicate that the values of the variable can be connected together along a numeric continuum.
Although a frequency polygon like the one in Figure 2.4 may look jagged as it moves from one data point to the next, it provides a general sense of the shape of the distribution of scores for a variable. For example, looking at Figure 2.4, you might conclude that the majority of GPAs in this sample are in the middle of the 1.00 to 4.00 range, with relatively low frequencies of students with either low or high GPAs.
88
Learning Check 2: Reviewing what you've Learned So Far
1. Review questions a. What is the difference between a frequency distribution table and a grouped frequency
distribution table? b. For what levels of measurement would you create a pie chart, bar chart, histogram, or a
frequency polygon? c. What is the difference between a bar chart and a histogram? d. Why would it be inappropriate to create a histogram or frequency polygon for a nominal or
ordinal variable? e. Under what circumstances might you use a frequency polygon rather than a histogram to
graph an interval or ratio variable? 2. For each of the following variables, identify whether you could use a bar chart, pie chart, histogram,
and/or a frequency polygon to visually display the data. a. Miles per gallon (MPG) b. Living situation (on-campus, off-campus) c. Number of brothers and sisters d. Marital status (single, married, divorced)
89
Drawing Inappropriate Conclusions from Figures
A well-constructed pie chart, bar chart, histogram, or frequency polygon can greatly facilitate the understanding of complex information. Conversely, a poorly designed figure may not only confuse the reader but also lead to inappropriate or misleading conclusions. Table 2.8 provides guidelines for constructing figures based on recommendations from the American Psychological Association (APA).
A simple example may demonstrate how altering the features of a figure can lead to different interpretations of the same set of data. Imagine, a week before an election, a pollster is interested in assessing how voters with different political affiliations plan to vote on a ballot measure. She asks a sample of Democrats, Republicans, and Independents the following question: “Do you plan on voting in favor of Measure A?” In examining her data, the pollster finds 56% of Democrats, 47% of Republicans, and 52% of Independents intend to vote “yes” on the measure.
Figure 2.5 graphs these three percentages using two bar charts; bar charts are used to display the data because the variable of interest (political affiliation) is a nominal variable (Democrat, Republican, or Independent). The bars in each chart do not touch each other because the three political affiliations are distinct from each other and cannot be placed along a numeric continuum.
Although the basic structure of the two bar charts is similar, there is one important difference in how they are constructed. Following the guidelines in Table 2.8, the bar chart in Figure 2.5(a) starts the Y-axis with the value 0% and ends with 100%. The bar chart in Figure 2.5(b), however, starts its Y-axis with the value of 45% and ends with 65%. By limiting the range of values for the percentages, the bar chart in Figure 2.5(b) exaggerates differences in the heights of the three bars. As a result, someone looking at Figure 2.5(b) may interpret the differences between the three political affiliations as being larger than someone looking at Figure 2.5(a).
The saying, “a picture is worth a thousand words,” is particularly relevant for the presentation of statistical data. Figures provide a vivid representation of data. It is the responsibility of researchers to portray their data in an accurate manner and to avoid creating distorting images that may mislead or confuse the reader.
Table 2.8 Some APA Guidelines for Creating Figures Table 2.8 Some APA Guidelines for Creating Figures
1. The figure should be formatted in a portrait rather than landscape page orientation (the reader should not turn the page sideways to read the table).
2. The figure is numbered using Arabic numbers (Figure 1, not Figure I or
90
Figure A). 3. The figure includes a title that describes the contents of the figure. For
example, it is not enough to name a figure simply “Figure 2.1” or “Figure 2.1. Bar chart.”
4. Figures have a rectangular shape, with the height of the vertical (Y) axis two- thirds to three-fourths the length of the horizontal (X) axis.
5. The horizontal (X) axis contains the values of the variable; the vertical (Y) axis contains the frequency (f) or percentage (%) of the sample having each value of the variable.
6. The X- and Y-axes include descriptive labels written in both upper- and lowercase letters (as opposed to all capital letters).
7. The labels for the X- and Y-axes are short enough to fit on one line along the axes.
8. The label of the Y-axis is placed parallel to the Y-axis (the letters face sideways rather than vertically).
9. The Y-axis starts with the value zero (0) and is divided into equally spaced values or intervals.
10. The Y-axis is divided into 6 to 10 values (just enough values to accommodate bar length). The bars should neither be flattened along the bottom of the Y- axis nor bumped up against the top of the Y-axis.
Figure 2.5 Different Displays of the Same Data
91
2.6 Examining Data: Describing Distributions
We have discussed how researchers construct tables and figures to organize and examine data they have collected. Looking at a table or figure enables researchers to describe how values or scores for a variable are distributed. Describing a distribution of data for a variable involves addressing questions such as, “Which values of the variable are the most common or occur most often?”, “How do the frequencies of the values of the variable change in relation to the most common values?”, and “What is the amount or nature of differences among the different values?” These questions pertain to three aspects of distributions, aspects that are introduced below and covered in greater detail in the next two chapters of this book:
modality, symmetry, and variability.
92
Modality
The first feature of a distribution of data for a variable is its modality. The modality of a distribution refers to the value or values of the variable that have the highest frequency or occur most often in a set of data. Using the lottery study as an example, identifying the modality of the lottery ticket variable involves answering the question, “Which of the four lottery tickets was chosen the most often?”
The modality of a distribution can be identified by looking at a frequency distribution table; the modality is the value(s) of the variable with the highest frequency (f). The modality may also be determined by examining a figure and locating the tallest bar in a bar chart or histogram, the highest point in a frequency polygon, or the biggest “slice” in a pie chart.
In identifying the modality of a distribution, a researcher may discover that one value or a limited number of values adjacent to one another clearly have the highest frequency. In such cases, the distribution is referred to as a unimodal distribution. A unimodal distribution is a distribution where one value occurs with the greatest frequency. Looking at the frequency distribution table (Table 2.4) or figure (Figure 2.1) for the lottery ticket variable, this distribution would be considered unimodal as one value (the random ticket) clearly occurred more often than the other three ticket choices.
In contrast to a unimodal distribution, it is possible that two values clearly distinct from each other may both have relatively high frequencies; if so, the distribution is referred to as a bimodal distribution. A bimodal distribution is a distribution where two values occur with the greatest frequency. A bimodal distribution sometimes resembles the two humps on a camel's back. And, in addition to unimodal and bimodal distributions, it is also possible for a distribution to have more than two values with relatively equally high frequencies; this type of distribution is known as a multimodal distribution—a distribution where more than two values have the greatest frequency. Figure 2.6 presents an example of a unimodal, bimodal, and multimodal distribution.
93
Symmetry
The second feature of a distribution of data is its symmetry. The symmetry of a distribution refers to how the frequencies of values of a variable change in relation to the most common or frequently occurring values. Describing the symmetry of a distribution means answering the question, “What is the shape of the distribution?”
Figure 2.6 Distributions with Different Modality (Unimodal, Bimodal, and Multimodal)
The symmetry of a distribution can be determined by looking at the frequencies in a frequency distribution table or the bars, points, or pieces of a figure. By doing so, a distribution may be described as being either symmetric or asymmetric. A symmetric distribution is one in which the frequencies change in a similar manner moving away in both directions from the most frequently occurring values. An example of a symmetric distribution is illustrated in Figure 2.7. One way to determine whether a distribution is symmetric is to draw the distribution on a piece of paper and then fold the distribution vertically in half down the middle so that the two halves are on top of each other. A distribution is symmetric when the two halves are similar in shape such that they resemble mirror images of each other.
However, a distribution may also be asymmetric. An asymmetric distribution (also known as a skewed distribution) is one in which the frequencies change in a different manner moving away in both directions from the most frequently occurring values. In asymmetric distributions, the highest frequencies are located at one end of the distribution rather than in the middle, such that if you fold an asymmetric distribution in half down the middle, the shapes of the two halves are not the same and do not resemble each other.
Figure 2.7 Distributions with Different Symmetry (Symmetric, positively Skewed, and Negatively Skewed)
94
In an asymmetric distribution, the frequencies take on the shape of a long tail as the values move from the values with the highest frequencies toward the other end of the distribution; the scores in the tail of the distribution are typically considered outliers. Furthermore, depending on the location of this tail, asymmetric distributions are referred to either as positively skewed or negatively skewed. A positively skewed distribution is one in which the higher frequencies are at the lower end of the distribution, with the tail on the upper (right) end of the distribution; a negatively skewed distribution is one in which the higher frequencies are at the upper end of the distribution, with the tail on the lower (left) end of the distribution. Figure 2.7 provides an illustration of a symmetric distribution and two asymmetric distributions (one positively skewed and one negatively skewed).
95
Variability
The third way in which distributions may be described is in terms of the amount or nature of differences among the different values for a variable—what is known as variability. The variability of a distribution refers to the amount of differences in a distribution of data for a variable. The variability in a distribution is described by answering the question, “To what extent are the scores in the distribution similar or different from each other?”
The amount of variability in a distribution can be estimated by looking at a frequency distribution table or figure. Imagine, for example, an instructor gives an 11-item test of reading comprehension to her class and finds almost all of the students answered six, seven, or eight of the items correctly. In this situation, the distribution would resemble the shape of a tall mountain: thin and tall, with many of the scores residing in a few values of the variable. A distribution that has a small amount of variability such as this one is referred to as a peaked distribution, defined as a distribution where much of the data is in a small number of values of a variable.
On the other hand, what would the distribution have looked like if she had found her students' scores were distributed relatively evenly across the entire range of possible scores? In this situation, the distribution would not be tall and thin like a mountain but rather would be short and wide like a plateau. This type of distribution is referred to as a flat distribution, which is a distribution where the data are spread evenly across the values of a variable.
In discussing modality, symmetry, and variability, distributions may have more than one mode, may be asymmetric, and may be peaked or flat. But what if a distribution has none of these qualities? A distribution that is in essence unimodal, symmetric, and neither peaked nor flat may be described as “bell-shaped.” If you look at a bell, you will notice that it is the highest in the middle, and as you move away from the top of a bell in both the left and right directions, the shape of the bell slopes downward in the same manner. A bell- shaped distribution for a variable is given a special name: a normally distributed variable.
A normally distributed variable is a distribution for a variable that is considered unimodal, symmetric, and neither peaked nor flat. Note that the word normal is a statistical concept and not a value judgment. In other words, asymmetric, peaked, or flat distributions are not considered “abnormal.” The concept of normal distributions will be discussed in greater detail in Chapter 5. Figure 2.8 provides an illustration of three distributions with different amounts of variability: peaked, flat, and normally distributed.
Figure 2.8 Distributions with Different Variability (Peaked, Flat, and Normally Distributed)
96
97
Learning Check 3: Reviewing what you've Learned So Far
1. Review questions a. What are the three aspects used to describe distributions of variables? b. What is the difference between unimodal, bimodal, and multimodal distributions? c. What is the difference between symmetric and asymmetric (skewed) distributions? d. What is the difference between positively and negatively skewed distributions? e. What is the difference between peaked and flat distributions?
2. For each of the following sets of data, determine whether the distribution of scores is unimodal, bimodal, or multimodal.
a. 4, 2, 11, 6, 10, 2, 6, 13, 2, 10, 6, 10, 1 b. 15, 17, 13, 18, 17, 14, 20, 18, 17, 14, 17 c. 20, 10, 50, 10, 80, 30, 70, 10, 60, 10, 80, 0, 30, 80, 10, 100, 80
3. For each of the following sets of data, determine whether the distribution of scores is symmetric or asymmetric (skewed).
a. 4, 5, 1, 4, 2, 4, 3, 5, 3, 4, 4, 3 b. 220, 230, 250, 230, 200, 240, 210, 230, 270, 220, 240, 230, 240, 230 c. 3, 1, 9, 2, 3, 6, 2, 7, 3, 5, 2, 6, 3, 2, 2
4. For each of the following sets of data, determine whether the distribution of scores is peaked or flat. a. 16, 10, 19, 12, 13, 15, 14, 17, 15, 20, 11, 13, 19, 16, 15 b. 12, 15, 12, 11, 13, 11, 12, 12, 14, 10, 12, 11, 16, 12, 13, 12, 11 c. 14, 12, 16, 14, 13, 15, 12, 14, 14, 15, 14, 14, 13, 14
98
2.7 Looking Ahead
This chapter has discussed why and how researchers examine data. Examining data by creating tables and figures requires a number of different skills, including organizational skills (the ability to take a set of data and prepare it for statistical analysis), communication skills (the ability to present data in tabular or graphic form), and conceptual skills (the ability to describe a distribution). The next two chapters begin to focus on mathematical skills related to the statistical analysis of data, describing how researchers examine and describe data and distributions numerically rather than visually. Given that this is a statistics textbook, we will focus on calculating statistics to describe and analyze data. However, and we cannot say this strongly enough, it is extremely important to examine a set of data before conducting statistical analyses as this will prevent errors both in your calculations of statistics as well as your interpretation of the results of these calculations.
99
2.8 Summary
Before conducting statistical analyses, researchers typically examine their data to gain an initial sense of the data, detect data coding or data entry errors, identify outliers (rare, extreme scores that lie outside of the range of the majority of scores in a set of data), evaluate research methodology, and determine whether the data meet statistical criteria and assumptions.
One common way of examining variables involves creating tables known as frequency distribution tables (tables that summarize the number and percentage of participants that express the different values for a variable) and grouped frequency distribution tables (tables that group the values of a variable measured at the interval or ratio level of measurement into a small number of intervals and then provide the frequency and percentage within each interval). To create a grouped frequency table, the real limits of each interval (the values of the variable that fall halfway between the top of one interval and the bottom of the next interval) need to be identified.
Data for a variable may also be examined by creating figures such as bar charts, pie charts, histograms, and frequency polygons. Bar charts and pie charts are used for variables measured at the nominal or ordinal level of measurement; a bar chart uses bars to represent the frequency or percentage of a sample corresponding to each value of a variable, with the different bars not touching each other; a pie chart is a circle divided into proportions that represent the percentage of the sample corresponding to each value of a variable. Histograms and frequency polygons are used for variables measured at the interval or ratio level of measurement; a histogram is a figure in which bars are used to represent the frequency of each value of a variable, with the different bars touching each other; a frequency polygon is a line graph that uses data points to represent the frequency of each value of a variable, with lines connecting the data points.
From examining their data, researchers can describe the distribution of values or scores for a variable. There are three main aspects of distributions: modality, symmetry, and variability.
The modality of a distribution is the value(s) of the variable that have the highest frequencies or occur most often in a set of data. Unimodal distributions are distributions where one value occurs with the greatest frequency, bimodal distributions have two values with the greatest frequency, and multimodal distributions have two values with the greatest frequencies.
The symmetry of a distribution refers to how the frequencies of values of a variable change in relation to the most common or frequently occurring values of the variable. In symmetric distributions, the frequencies change in a similar manner, moving away in both directions from the most common values of the variable. In asymmetric (skewed) distributions, the
100
frequencies change in a different manner, moving away in both directions from the most common values of the variable.
The variability of a distribution refers to the amount of differences in a distribution of data. In peaked distributions, much of the data are in a few values of a variable; in flat distributions, the data are spread evenly across the values of a variable. A normally distributed variable has a symmetric distribution that is neither peaked nor flat.
101
2.9 Important Terms
outlier (p. 27) frequency distribution table (p. 28) frequency (f) (p. 28) grouped frequency distribution table (p. 32) real limits (p. 32) real lower limit (p. 32) real upper limit (p. 32) cumulative percentage (p. 34) bar chart (p. 35) pie chart (p. 36) histogram (p. 37) frequency polygon (p. 37) modality (p. 42) unimodal distribution (p. 42) bimodal distribution (p. 42) multimodal distribution (p. 42) symmetry (p. 42) symmetric distribution (p. 43) asymmetric (skewed) distribution (p. 43) positively skewed distribution (p. 44) negatively skewed distribution (p. 44) variability (p. 44) peaked distribution (p. 44) flat distribution (p. 44) normally distributed variable (p. 44)
102
2.10 Formulas Introduced in this Chapter
103
Percentage (%)
(2-1) % = f total number of scores ∗ 100
104
2.11 Using IBM® SPSS® Software∗
* IBM® SPSS® software. SPSS is a registered trademark of International Business Machines Corporation.
105
Creating Data Files: Defining Variables and Entering Data
1. Open a new data file.
2. Define and name variable(s).
How? (1) Click Variable View tab, (2) click below Name, and (3) type name of variable in the cell.
3. Determine number of decimals for your variable(s).
How? Click desired number of decimal places in the Decimals box.
106
4. Provide labels for your variable(s).
How? (1) Click below Label, and (2) type the label for the variable in the cell.
5. Provide labels for values of nominal (categorical) variable(s).
How? (1) Click below Values, and (2) type labels for coded values of the variable.
6. Enter data.
How? (1) Click Data View tab and (2) enter data for the variable in the cells.
107
Examining Data: Frequency Distribution Tables and Figures: The Lottery Study (2.1)
1. Select the frequency distribution procedure within SPSS.
How? (1) Click Analyze menu, (2) click Descriptive Statistics, and (3) click Frequencies.
2. Select the variable to be analyzed.
How? (1) Click variable and .
108
3. Select the appropriate type of figure for the variable.
How? (1) click Charts, (2) select the desired type of figure, (3) click Continue, and (4) click OK .
4. Examine output.
109
2.12 Exercises
1. For each of the following variables, identify whether you could use a pie chart, bar chart, histogram, or frequency polygon to visually display the data:
a. Gender (male, female) b. Weight (in pounds) c. Number of children in families d. Baseball batting averages e. Eye color f. Temperature
2. For each of the following variables, identify whether you could use a pie chart, bar chart, histogram, or frequency polygon to visually display the data:
a. Shoe size b. College major c. Favorite radio station d. Midterm exam score e. Zip code f. Marital status
3. The owner of an ice cream store asks 75 people which flavor of ice cream they prefer. Thirteen of them say strawberry, 11 say chocolate, 24 say vanilla, and 27 provide a flavor other than strawberry, chocolate, or vanilla. Create a frequency distribution table for these data.
4. A political pollster approaches people on the street and asks them to describe their political affiliation. Twenty-eight people describe themselves as Democrats, 25 as Republicans, 8 people provide a political party other than Democrat or Republican, 13 label themselves as Independent, and 10 people say they do not belong to a political party. Create a frequency distribution table for these data.
5. Below is a frequency distribution table for a hypothetical variable:
Value f %
100 3 10%
90 8 26%
80 5 17%
70 6 20%
60 3 10%
50 2 7%
40 0 0%
30 2 7%
110
20 1 3%
10 0 0%
Total 30 100%
a. How many of the scores for this variable have the value 70? b. What percentage of the scores has the value of 30?
6. Below is a frequency distribution table for a hypothetical variable: a. How many of the scores for this variable have the value 2? b. What percentage of the scores has the value of 7?
Value f %
7 5 15%
6 8 24%
5 10 29%
4 6 18%
3 3 9%
2 2 5%
1 0 0%
Total 34 100%
(Exercises 7–10 pertain to the following situation.) You're standing in line to see a movie with two of your friends (who also happen to be taking a statistics class). You are not sure whether you want to see this particular movie, so you decide to stop people coming out of the theater and ask them for their opinion about the movie.
7. You stop 10 people and ask them whether they would or would not recommend the movie to others. They give you the following answers:
Person Recommend?
1 No
2 Yes
3 No
4 No
5 Yes
6 No
7 Yes
8 Yes
9 No
111
10 No
a. Construct a frequency distribution table of these data (be sure to include columns representing the frequency and percentages).
b. What level of measurement is this variable? c. Construct an appropriate figure for these data.
8. One of your friends feels the responses of your 10 people did not result in a clear recommendation or nonrecommendation of the movie. She decides to ask 15 people to give the movie one of three ratings: above average, average, or below average. Their ratings are listed below:
Person Rating
1 Average
2 Below average
3 Average
4 Above average
5 Average
6 Above average
7 Average
8 Average
9 Average
10 Above average
11 Average
12 Above average
13 Average
14 Average
15 Below average
a. Construct a frequency distribution table of these data. b. What level of measurement is this variable? c. Construct an appropriate figure for these data.
9. Your other friend asks 20 people to rate the movie using a 1- to 5-star rating: the higher the number of stars, the higher the recommendation. Their ratings are listed below:
Person # Stars
112
1 ∗∗∗
2 *****
3 ∗∗
4 ∗∗∗∗
5 ∗∗∗
6 ∗∗∗
7 ∗∗∗∗
8 ∗∗
9 ∗∗∗
10 *****
11 ∗
12 ∗∗∗∗
13 ∗∗∗∗
14 ∗∗∗
15 ∗∗∗
16 ∗∗
17 ∗∗∗
18 ∗∗
19 ∗∗∗∗
20 ∗∗∗
a. Construct a frequency distribution table of these data. b. What level of measurement is this variable? c. Construct an appropriate figure for these data.
10. Another friend asks 30 people to rate the movie by their likelihood of seeing the movie again, ranging from 0% to 100%. Their ratings are listed below:
Person Likelihood
1 50%
2 10%
3 66%
4 60%
5 50%
6 90%
7 33%
113
8 95%
9 75%
10 5%
11 99%
12 50%
13 35%
14 75%
15 60%
16 25%
17 75%
18 90%
19 30%
20 33%
21 60%
22 50%
23 80%
24 70%
25 50%
26 65%
27 20%
28 30%
29 70%
30 33%
a. Construct a grouped frequency distribution table of these data, with intervals of 10 (0%–9%, 10%–19%, 20%–29%, etc.).
b. What level of measurement is this variable? c. Construct an appropriate figure for these data. d. Construct a grouped distribution of these data with intervals of 33 (0%–33%,
34%–66%, and 67%–100%). How does this table compare with the table with one having intervals of 10?
11. Given the growing popularity of video games among young people, one study examined how women are portrayed in these games (Dietz, 1998). Dr. Dietz was particularly interested in whether video games maintain gender role stereotypes or contain violence against women. She viewed 25 video games with human characters popular at that time and defined the primary role of women in each of them. Each
114
video game was placed into one of five groups: no female characters, female characters portrayed as sex objects, females as the victim, females as the hero, or females in traditional feminine roles.
Video Game Role of Women
1 Victim
2 No females
3 No females
4 Victim
5 No females
6 Sex object
7 Sex object
8 Sex object
9 No females
10 Hero
11 Sex object
12 No females
13 No females
14 No females
15 No females
16 Sex object
17 No females
18 Hero
19 Hero
20 Feminine
21 No females
22 Victim
23 Sex object
24 Hero
25 Victim
a. Create a frequency distribution table for these data.
115
b. What level of measurement is this variable? c. Construct an appropriate figure for these data.
12. An instructor administers a 27-item quiz to her class of 25 students. Each student's score on the quiz is the number of items answered correctly. These scores are listed below:
a. Construct a grouped frequency distribution table of these data, following the guidelines presented in this chapter.
b. What level of measurement is this variable? c. Construct an appropriate figure for these data.
Person Quiz Score
1 22
2 15
3 11
4 19
5 12
6 21
7 22
8 19
9 23
10 17
11 18
12 14
13 10
14 6
15 16
16 20
17 21
18 19
19 20
20 21
21 17
22 18
23 15
24 24
116
25 22
d. How would you describe the shape of the distribution in terms of its symmetry?
e. How would you characterize the students' performance? Did they do well or poorly? Would you say the quiz was easy or hard?
13. Imagine you are in your local department store to buy a sweater. Which of the following signs would make you think you're getting a good deal: “Everyday Low Price $15.00” or “Regularly $20.00/Sale $15.00”? Because the actual price in both signs is the same, you should be equally likely to pick either sign. One study investigated whether the presentation of a price influences customers' perceptions of savings (Tom & Ruiz, 1997). They asked 49 college students to select one of the two signs. Twelve students selected the everyday low price sign, and 37 selected the sale price sign.
a. Create a frequency distribution table for these data. b. The researchers “hypothesized that consumers would think that a sale price
presentation would generate a greater monetary savings than an everyday low price presentation” (p. 403). Do the percentages in your frequency distribution table provide preliminary support for this hypothesis?
14. The lottery ticket variable discussed in this chapter (Hardoon et al., 2001) examined the ticket preferences of non–problem gamblers. One of the purposes of this study was to compare the preferences of non–problem gamblers and problem gamblers (people experiencing personal, professional, financial, or legal problems as a result of their gambling). The lottery ticket choices for the 24 problem gamblers in this study are listed below:
a. Construct a frequency distribution table of these lottery ticket choices. b. Construct an appropriate figure for these data. c. Looking at your table and figure, how would you summarize the lottery ticket
choices of the problem gamblers?
Problem Gambler Ticket Choice
1 Pattern
2 Random
3 Nonequilibrated
4 Pattern
5 Random
6 Random
117
7 Nonequilibrated
8 Random
9 Sequence
10 Random
11 Random
12 Pattern
13 Nonequilibrated
14 Random
15 Random
16 Sequence
17 Pattern
18 Nonequilibrated
19 Random
20 Random
21 Nonequilibrated
22 Pattern
23 Random
24 Pattern
d. Compare the percentages in this table with the table for the non–problem gamblers (Table 2.4). In what ways are the lottery ticket choices of the two groups similar and in what ways are they different?
15. Another published study of government-sponsored lotteries examined reasons why people purchase lottery tickets (Miyazaki, Langenderfer, & Sprott, 1999). The researchers approached a number of people in different parts of one city and asked them, “At the times that you decide to go ahead and purchase tickets, what are the reasons why you personally do buy them?” (p. 6). The researchers grouped these people's responses into the following categories:
Reason f
Desire to win money 70
Impulse purchase 28
Feeling lucky 18
Enjoyment and fun 15
To help schools 7
Miscellaneous/other reasons 14
118
Total 152
a. What level of measurement is this variable? b. Construct an appropriate figure for these data. c. Looking at the table and figure, how would you summarize the reasons why
people purchase lottery tickets? 16. How information is presented may influence people's decision-making processes; this
was examined in a study of the doctor-patient relationship (Gurm & Litaker, 2000). This study examined how the manner in which the potential risks and benefits of a medical procedure are presented affect patients' willingness to undergo the procedure. In one condition of this study, 63 participants were told that “99% of patients undergoing the procedure do not have any of these complications.” When asked to describe the likelihood they would undergo the procedure, 16 said “definitely,” 36 said “probably,” 9 said “probably not,” and 2 said “definitely not.”
a. Create a frequency distribution table for these data. b. How would you summarize the likelihood these participants would undergo
the procedure? 17. In a second condition of the study described in Exercise 16, 53 participants were
told, “These complications are seen in 1 out of 100 who undergo the procedure.” Note that both conditions contain the same degree of risk (99% chance of no complications is equal to 1% chance of complications). When participants in the second condition were asked whether they would undergo the procedure, 4 said “definitely,” 23 said “probably,” 24 said “probably not,” and 2 said “definitely not.”
a. Create a frequency distribution table for these data. b. How would you summarize the responses of the participants in this second
condition? c. How would you describe the difference between the responses of the
participants in the two conditions? What implications might these findings have for physicians?
18. For each of the following sets of data, determine whether the distribution of scores is unimodal or bimodal.
a. 17, 10, 8, 7, 14, 7, 6, 10, 9, 10, 7, 8 b. 12, 10, 9, 11, 9, 20, 9, 4, 9, 16, 9, 10 c. 9, 6, 8, 5, 8, 7, 9, 10, 8, 7, 8 d. 160, 197, 170, 115, 158, 170, 175, 115, 112, 115, 104, 170, 136, 115, 102,
181, 170 19. For each of the following sets of data, determine whether the distribution of scores is
symmetric or asymmetric (skewed). a. 9, 7, 6, 8, 6, 13, 7, 6, 5, 6 b. 10, 17, 16, 14, 16, 21, 19, 15, 18, 15, 16, 17, 16 c. 22, 8, 17, 3, 19, 14, 19, 6, 19, 23, 16, 19, 22, 17, 19, 20 d. 101, 109, 99, 112, 109, 110, 109, 107, 109, 108, 117, 109, 108, 110, 109
119
20. For each of the following sets of data, determine whether the distribution of scores is peaked or fat.
a. 18, 15, 19, 18, 23, 18, 17, 19, 18, 17, 18 b. 15, 8, 12, 6, 9, 11, 7, 17, 9, 10, 8, 5, 4, 8, 9, 7 c. 11, 13, 15, 14, 17, 16, 13, 9, 13, 12, 7, 13, 12, 14, 13 d. 72, 56, 63, 66, 49, 63, 60, 56, 60, 63, 53, 63, 70, 60, 44, 60, 56
120
Answers to Learning Checks
Learning Check 1
2. a.
Plans After Graduation f %
Work 71 61%
Graduate school 34 29%
Not sure 11 10%
Total 116 100%
b.
Year in College f %
Freshman 69 64%
Sophomore 18 17%
Junior 12 11%
Senior 9 8%
Total 108 100%
c.
Voting Intention f %
In favor 7 58%
Against 5 42%
Total 12 100%
d.
New M&M Color f %
Aqua 5 21%
Pink 8 33%
Purple 11 46%
Total 24 100%
121
Learning Check 2
2. a. Frequency polygon or histogram b. Bar chart or pie chart c. Histogram or frequency polygon d. Bar chart or pie chart
Learning Check 3
2. a. Multimodal b. Unimodal c. Bimodal
3. a. Asymmetric (negatively skewed) b. Symmetric c. Asymmetric (positively skewed)
122
Answers to Odd-Numbered Exercises
1. a. Pie chart or bar chart b. Frequency polygon or histogram c. Histogram or frequency polygon d. Frequency polygon or histogram e. Bar chart or pie chart f. Frequency polygon or histogram
2. a. Histogram or frequency polygon b. Bar chart or pie chart c. Bar chart or pie chart d. Histogram or frequency polygon e. Bar chart or pie chart f. Bar chart or pie chart
3.
Flavor f %
Strawberry 13 17%
Chocolate 11 15%
Vanilla 24 32%
Other 27 36%
Total 75 100%
4.
Political Affiliation f %
Democrat 28 33%
Republican 25 30%
Other 8 10%
Independent 13 15%
No affiliation 10 12%
Total 84 100%
5. a. 6 b. 7%
123
6. a. 2 b. 15%
7. a.
Recommend? f %
Yes 4 40%
No 6 60%
Total 10 100%
b. Nominal c. Bar chart
8. a.
Rating f %
Below average 2 13%
Average 9 60%
Above average 4 27%
Total 15 100%
b. Ordinal c. Bar chart
124
9. a.
# Stars f %
5 2 10%
4 5 25%
3 8 40%
2 4 20%
1 1 5%
Total 20 100%
b. Interval c. Histogram
10. a.
Likelihood f %
90%–100% 4 13%
80%–89% 1 3%
125
70%–79% 5 17%
60%–69% 5 17%
50%–59% 5 17%
40%–49% 0 0%
30%–39% 6 20%
20%–29% 2 7%
10%–19% 1 3%
0%–9% 1 3%
Total 30 100%
b. Ratio c. Frequency polygon
d.
Likelihood f %
67%–100% 10 33%
34%–66% 10 33%
0%–33% 10 33%
Total 30 100%
e. Grouping the ratings into only three categories gives the impression people's ratings are equally distributed when in fact there is a fair amount of variability.
11. a.
Role of Women f %
No females 10 40%
126
Sex object 6 24%
Victim 4 16%
Hero 4 16%
Traditional 1 4%
Total 25 100%
b. Nominal c. Bar graph
12. a.
Score f %
25–27 0 0%
22–24 5 20%
19–21 8 32%
16–18 5 20%
13–15 3 12%
10–12 3 12%
7–9 0 0%
4–6 1 4%
0–3 0 0%
Total 25 100%
b. Ratio c. Frequency polygon
127
d. The distribution is negatively skewed. e. The students generally did well on the quiz, suggesting that the quiz was easy.
13. a.
Sign f %
“Everyday Low Price $15.00” 12 24%
“Regularly $20.00/Sale $15.00” 37 76%
Total 49 100%
b. Yes, the table provides preliminary support for this hypothesis in that a majority of the sample (76%) picked the sale price presentation (“Regularly $20.00/Sale $15.00”).
14. a.
Lottery Ticket f %
Sequence 2 8%
Pattern 6 25%
Nonequilibrated 5 21%
Random 11 46%
Total 24 100%
b. Bar chart
128
c. The problem gamblers chose the random pattern (46%) much more often than the other three types; the sequential pattern (8%) was chosen the least often.
d. The problem gamblers had a slightly lower preference for the random pattern and a higher preference for the nonequilibration pattern than did the non– problem gamblers, suggesting that problem gamblers are less likely to believe lottery numbers are selected randomly.
15. a. Nominal b. Bar chart
c. In this sample, the desire to win money is clearly the most frequently given reason for buying lottery tickets; helping schools is the least often given reason.
16. a.
Likelihood f %
Definitely 16 26%
Probably 36 57%
Probably not 9 14%
129
Definitely not 2 3%
Total 63 100%
b. When told that 99% of patients do not have any complications, most of the participants (83%) would either “definitely” or “probably” undergo the medical procedure.
17. a.
Likelihood f %
Definitely 4 8%
Probably 23 43%
Probably not 24 45%
Definitely not 2 4%
Total 53 100%
b. When told that 1% of patients experience complications, the participants are almost equally likely to say they would “definitely” or “probably” (51%) undergo the medical procedure as they would “probably not” or “definitely not” (49%).
c. The wording doctors use in presenting the risk factors influences the likelihood the participant will undergo medical treatments such that emphasizing the possibility of complications may lower the likelihood a patient will undergo a medical procedure.
18. a. Bimodal b. Unimodal c. Unimodal d. Bimodal
19. a. Positively skewed b. Symmetric c. Negatively skewed d. Symmetric
20. a. Peaked b. Flat c. Peaked d. Flat
130
Sharpen your skills with SAGE edge at edge.sagepub.com/tokunaga
SAGE edge for students provides a personalized approach to help you accomplish your coursework goals in an easy-to-use learning environment. Log on to access:
eFlashcards Web Quizzes Chapter Outlines Learning Objectives Media Links SPSS Data Files
131
Chapter 3 Measures of Central Tendency
132
Chapter Outline 3.1 An Example From the Research: The 10% Myth 3.2 Understanding Central Tendency 3.3 The Mode
Determining the mode 3.4 The Median
Calculating the median: odd number of scores Calculating the median: even number of scores Calculating the median when multiple scores have the median value
3.5 The Mean Calculating the mean
Reporting calculated numbers Rounding calculated numbers
Calculating the mean from frequency distribution tables The mean as a balancing point in a distribution The sample mean vs. the population mean: statistics vs. parameters
3.6 Comparison of the Mode, Median, and Mean Strengths and weaknesses of the mode Strengths and weaknesses of the median Strengths and weaknesses of the mean Choosing a particular measure of central tendency
3.7 Measures of Central Tendency: Drawing Conclusions 3.8 Looking Ahead 3.9 Summary 3.10 Important Terms 3.11 Formulas Introduced in This Chapter 3.12 Exercises
Chapter 2 discussed the importance of examining data using tables and figures. Researchers examine frequency distribution tables and figures to draw preliminary conclusions regarding three aspects of distributions of data: modality, symmetry, and variability. In addition to visually examining data, researchers typically describe data numerically by calculating statistics. Recall from Chapter 1 that there are two main types of statistics: descriptive and inferential. The purpose of descriptive statistics is to numerically describe or summarize data for a variable. This chapter discusses one of the main types of descriptive statistics, known as measures of central tendency. The next section of this chapter introduces a published research study that we'll use to help introduce and illustrate our discussion of these measures.
133
3.1 An Example from the Research: The 10% Myth
One reason for conducting research is to test ideas systematically and objectively rather than rely on one's personal perceptions or beliefs. One example of a popular belief related to the field of psychology is the commonly heard statement, “Most people use only 10% of their brains.” This statement implies that people are not living up to their full potential and that we could all have phenomenal mental abilities if only we were able to use more of our brains. This belief has received widespread attention through the popular media, appearing in movies, advertisements, and magazines. For all its influence and popularity, however, there is one problem with the statement: It is not true. In his study, “Whence Cometh the Myth That We Only Use Ten Percent of Our Brains?” (Beyerstein, 1999), psychologist Barry Beyerstein demonstrated there is no scientific evidence to support this popular belief, referring to it as the “ten-percent myth.”
The origin of the 10% myth is not precisely known. It may be a misrepresentation of a statement made by William James, one of the pioneers of psychology. In his 1911 book, The Energies of Men, James stated that “as a rule men habitually use only a small part of the powers which they actually possess…. We are making use of only a small part of our possible mental and physical resources” (pp. 11–12).
Researchers Kenneth Higbee and Samuel Clay decided to investigate the potential impact of higher education on dispelling misconceptions such as the 10% myth. “We expected that the kind of academic training psychology students receive … would give them a healthy skepticism” (Higbee & Clay, 1998, p. 471). More specifically, they hypothesized that “psychology majors who had a significant amount of training in psychology would be less likely to believe the ten-percent myth than would non-majors who had no training in psychology” (p. 471).
To test their research hypothesis, Higbee and Clay located two groups of college students: 38 psychology majors and 39 non–psychology majors. The students in the study, which we will refer to as the 10% myth study, were asked the following question: “About what percentage of their potential brain power do you think most people use?” The students were asked to choose from among 21 scores, ranging from 0% to 100%, with scores increasing in 5% increments (0%, 5%, 10%, etc.); we will refer to this variable as “estimated brain power.”
The example from the 10% myth study included in this chapter will focus on the data provided by the psychology majors. (The data for the non–psychology majors is included in the exercises at the end of this chapter.) The estimated brain power chosen by the 38 psychology majors is listed in Table 3.1(a); to organize these data, Table 3.1(b) arranges the data into a frequency distribution table.
134
Figure 3.1 illustrates the data from the frequency distribution table using a frequency polygon (a frequency polygon was used because estimated brain power is measured at the ratio level of measurement, with the values of the variable equally spaced along a numeric continuum with a true zero point). As we can see from the table and figure, the majority of the data are at the low end of the 0% to 100% scale, with many psychology majors reporting the belief that most people use 5% or 10% of their potential brain power. This distribution is best described as asymmetric; more specifically, it is positively skewed because the higher frequencies are at the lower end of the distribution, and the tail is at the upper end of the distribution.
Table 3.1 The Estimated Brain Power by 38 Psychology Majors Table 3.1 The Estimated Brain Power by 38
Psychology Majors
(a) Raw Data
Psychology Major Estimated Brain Power
1 45
2 40
3 10
4 30
5 10
6 40
7 10
8 10
9 5
10 15
11 75
12 10
13 10
14 5
15 5
16 10
17 10
18 15
19 20
20 10
21 5
135
22 35
23 5
24 5
25 10
26 20
27 30
28 40
29 5
30 30
31 35
32 75
33 10
34 5
35 10
36 90
37 20
38 55
Table 3.1 The Estimated Brain Power by 38
Psychology Majors
(b) Frequency Distribution Table
Estimated Brain Power f %
100% 0 0%
95% 0 0%
90% 1 3%
85% 0 0%
80% 0 0%
75% 2 5%
70% 0 0%
65% 0 0%
60% 0 0%
55% 1 3%
50% 0 0%
136
45% 1 3%
40% 3 8%
35% 2 5%
30% 3 8%
25% 0 0%
20% 3 8%
15% 2 5%
10% 12 31%
5% 8 21%
0% 0 0%
Total 38 100%
Figure 3.1 Frequency Polygon for the Estimated Brain Power Variable
Examining the frequency distribution table and frequency polygon for the estimated brain power variable provides an initial indication of the extent to which the sample of psychology majors believed in the 10% myth. However, to test the study's research hypothesis, the researchers needed to statistically analyze the data. A common first step in statistical analyses is to summarize a set of data by calculating descriptive statistics. The next section introduces one type of descriptive statistic, measures of central tendency, which summarize a set of data by identifying the most common or frequently occurring values of variables.
137
3.2 Understanding Central Tendency
For many variables studied by researchers, the majority of data cluster or center on the middle of the distribution. For this reason, the first statistical index or measure used to describe data numerically is a measure of central tendency, defined as a statistic that identifies the center of a distribution; a measure of central tendency is a single score that is the most typical, common, or frequently occurring value for a variable.
The following sections describe three measures of central tendency:
the mode, the median, and the mean.
After describing and illustrating the different measures, we will compare them in terms of their relative strengths and weaknesses and the conditions under which it may be preferable to use a particular measure.
138
3.3 The Mode
One way to describe or summarize a set of data for a variable is to ask, “What values of the variable are the most common or frequently occurring?” The first measure of central tendency that we will consider, the mode, sometimes referred to as the modal score, is the score or value of a variable that appears most frequently in a set of data.
139
Determining the Mode
The simplest way to determine the mode is to examine a frequency distribution table or figure for a variable. For the estimated brain power variable, the mode may be identified by looking at the frequency distribution table presented earlier in Table 3.1(b). By scanning down the frequency (f) column in this table, we see that 12 of the 38 psychology majors believed that most people use 10% of their potential brain power. Because it has the highest frequency, 10% is the mode or modal score for this set of data. Following American Psychological Association (APA) style, the abbreviation Mo is used in presenting the mode, such as “Mo = 10%.” (Note that APA style requires the word Mo be italicized.)
To locate the mode in a figure such as a bar chart, histogram, or frequency polygon, we identify the tallest bar in the bar chart or histogram or the highest data point in the polygon. Looking at the frequency polygon for the estimated brain power variable in Figure 3.2, because the highest data point corresponds to the value of 10%, it is the mode for this data set.
140
3.4 The Median
In addition to the mode, another way of describing a distribution of data for a variable involves the question, “What value of the variable is located in the center of the distribution?” A measure of central tendency designed to answer this question is the median, defined as the value of a variable that splits a distribution of scores in half, with the same number of scores above the median as below.
The calculation of the median varies depending on whether the set of data consists of an odd number of scores or an even number of scores. These two situations are discussed in turn below.
141
Calculating the Median: Odd Number of Scores
Because the data for the estimated brain power variable consist of an even number of scores (N = 38), we will illustrate how to calculate the median when the set of data consists of an odd number of scores using the following set of nine scores: 2 3 1 4 3 1 4 6 3
Figure 3.2 Identifying the Mode for the Estimated Brain Power Variable
142
Learning Check 1: Reviewing what you've Learned So Far
1. Review questions a. What is the main purpose of a measure of central tendency? b. How do you identify the mode from a frequency distribution table or figure?
2. Determine the mode for each of the following sets of data. a. 2, 1, 4, 2, 7, 2, 3 b. 6, 7, 6, 4, 5, 4, 2, 6 c. 12, 5, 9, 12, 10, 11, 12, 8, 15, 12, 7 d. 4, 2, 2, 1, 5, 4, 5, 2, 3, 2, 4, 2 e. 3.0, 2.8, 3.2, 3.6, 3.7, 3.2, 3.5, 3.6, 3.1, 3.2, 3.0, 3.2, 3.8, 3.3
There are three steps involved in calculating the median for a set of data: (1) sort the scores from lowest to highest, (2) calculate the median score in the set of data, and (3) determine the value of the median score. To accomplish the first step, the nine scores are sorted from lowest to highest below: 1 lowest 1 2 3 3 3 4 4 6 highest
The second step is to calculate the median score in the set of data. Formula 3-1 presents the formula for calculating the median with an odd number of scores:
(3-1) Median = N + 1 2 th score
where N is the total number of scores. Because there are nine scores (N = 9) in this set of data, the median is calculated as follows: Median = N + 1 2 t h s c o r e = 9 + 1 2 th score = 10 2 th score = 5th score
Therefore, when N = 9, the median is located at the fifth score in the set of data. It is important to note that the median for this variable is not 5 but is rather the value of the fifth score.
The third step is to determine the value of the median score. Looking at the nine sorted scores, we find the value of the fifth score is equal to 3: 1 1 st 1 2 nd 2 3 rd 3 4 th 3 5 t h 3 6 th 4 7 th 4 8 th 6 9 th
The median in the set of nine scores is equal to 3; it is the median because the same number of scores (four) is above this score as below it. When reporting the median following APA format, use the abbreviation Mdn and write “Mdn = 3.”
143
Calculating the Median: Even Number of Scores
If a data set has an odd number of scores, the median is the value of the one score located in the center of the sorted scores. However, if a set of data consists of an even number of scores, the formula for the median changes because no single score has the same number of scores above and below it. In this situation, the median is determined by identifying the two scores in the middle of the distribution and then calculating the number halfway between the values of the two scores; this is represented in Formula 3-2:
(3-2) Median = N 2 th score + N 2 + 1 th score 2
where N is the total number of scores.
As a simple example, let's say we had the following eight scores: 39, 15, 11, 24, 45, 31, 42, and 19. To calculate the median, the first step is to sort the scores from lowest to highest: 11 lowest 15 19 24 31 39 42 45 highest
The second step is to calculate the median scores in the set of data using Formula 3-2: Median = N 2 th score + ( N 2 + 1 ) th score 2 = 8 2 th score + ( 8 2 + 1 ) th score 2 = 4 th score + 5th score 2
Therefore, for N = 8, the value of the median falls halfway between the values of the fourth and fifth scores.
The third step is to calculate the value of the median. In this example, we need to identify the values of the fourth and fifth scores: 11 1 st 15 2 nd 19 3 rd 24 4 th 31 5 th 39 6 th 42 7 th 45 8 th
Because the fourth score in this distribution is 24 and the fifth score is 31, the value for the median is the number halfway between these two scores: Median = 4 th score + 5th score 2 = 24 + 31 2 = 55 2 = 27.50
In this set of eight scores, the median is 27.50, or “Mdn = 27.50.” There are two things to note about this value of the median. First, the same number of scores falls below it (11, 15, 19, 24) as above it (31, 39, 42, 59). Second, because 27.50 is the same distance from the score below it (24) as the score above it (31), it lies precisely in the center of the set of scores.
144
As a second example, let's calculate the median for the 38 scores for the estimated brain power variable. Once the 38 scores have been ranked from lowest to highest (see below), the two median scores are calculated using Formula 3-2: Median = N 2 th score + ( N 2 + 1 ) th score 2 = 38 2 th score + ( 38 2 + 1 ) th score 2 = 19 th score + 20th score 2
For the estimated brain power variable, the median falls between the 19th and 20th scores:
As both the 19th and 20th scores are equal to 10, the value for the median is the following: Median = 19 th score + 20th score 2 = 10 + 10 2 = 20 2 = 10.00
So, for this sample of 38 psychology majors, the median for the estimated brain power variable is 10.00 (Mdn = 10.00). Figure 3.3 identifies the median on the frequency polygon for the estimated brain power variable.
145
Calculating the Median when Multiple Scores have the Median Value
The median of 10.00 for the estimated brain power variable suggests that, in the sample of 38 psychology majors, an equal number of the scores fall below and above the value of 10.00. This is not precisely true, however, as more than two scores have the median value of 10.00. Returning to Table 3.1(b), we can see that, because 12 of the 38 psychology majors gave the response of 10%, only 8 of the 38 scores actually fall below the median value of 10.
Figure 3.3 Identifying the Median for the Estimated Brain Power Variable
146
Learning Check 2: Reviewing what you've Learned So Far
1. Review questions a. Why are there different formulas for calculating the median depending on whether there are
an odd or even number of scores for a variable? 2. Calculate the median for each of the following sets of data.
a. 2, 3, 3, 4, 5, 6, 8 b. 21, 14, 13, 17, 30, 17, 14, 11, 13, 14, 19 c. 5, 6, 7, 9, 12, 15, 16, 17 d. 73, 66, 91, 84, 69, 87, 62, 79, 82, 90 e. 11, 9, 12, 9, 10, 7, 9, 8, 9, 10, 9, 8
When there are multiple scores with the median value, the method for calculating the median presented in this book does not provide the exact value for the median but rather an approximation. This is because the method does not consider all of the scores with the median value, but only one (in samples with an odd number of scores) or two (in even- numbered samples) of the scores. Calculating the exact median when multiple scores have the median value requires using a mathematical operation known as interpolation. This book will not cover interpolation, however, because the differences between the approximate and exact values of the median are typically small.
147
3.5 The Mean
A third measure of central tendency used to represent a distribution concerns the question, “What is the average value of the variable in this set of data?” The descriptive statistic that answers this question is the mean, which is defined as the arithmetic average of a set of scores. The mean is the measure of central tendency with which you are most likely familiar. You may have heard or read, for example, about the “average cost of a college education,” “average life expectancy,” “average miles per gallon,” or “earned run average (ERA).” Because this book will refer to scores for a variable using the symbol X, the mean of scores in a set of data is represented by the letter X with a bar above it X ¯ . This symbol is referred to as the “mean of X” or “X bar.”
148
Calculating the Mean
Mathematically the formula for the mean (Formula 3-3) is the following:
(3-3) X ¯ = ∑ X N
where X is a score for a variable and N is the total number of scores. To calculate the mean, we first calculate the sum of the scores, which is represented by ΣX (the Σ [sigma] symbol represents “the sum of”). The mean is then calculated by dividing the sum of scores by the total number of scores (N).
To introduce the calculation of the mean, we return the set of nine scores presented earlier: 2,3, 1, 4, 3, 1, 4, 6, and 3. The mean for this set of data is calculated as follows: X ¯ = ∑ X N = 2 + 3 + 1 + 4 + 3 + 1 + 4 + 6 + 3 9 = 27 9 = 3.00
In reporting the mean of a variable in a paper or journal article, rather than use X ¯ , APA format uses the italicized letter M, as in “M = 3.00.”
For the 10% myth study, let's calculate the mean for the 38 scores for the estimated brain power variable in Table 3.1(a): X ¯ = ∑ X N = 45 + 40 + 10 + ⋯ + 90 + 20 + 55 9 = 870 38 = 22.89
The ellipses (…) included in the middle of this calculation (45 + 40 + 10 + … + 90 + 20 + 55) represent mathematical operations conducted to correctly carry out the formula but have not been printed to save space.
On the basis of the above calculations, we would conclude that the mean amount of brain power that psychology majors believe people use is 22.89% (M = 22.89). The mean for the estimated brain power variable is illustrated in Figure 3.4.
Reporting Calculated Numbers
Figure 3.4 Identifying the Mean for the Estimated Brain Power Variable
149
As statistics uses a large number of formulas involving mathematical calculations, this section discusses several issues concerning the reporting and rounding of numbers resulting from these calculations. In reporting calculated numbers, the general guideline is to round calculated numbers to two decimal places. Including two decimal places indicates that the reported number is the result of a mathematical calculation. For the sake of consistency, calculated numbers are rounded to two decimal places even when the result does not require two decimal places (i.e., 21 3 = 7.00 [rather than 7], 38 10 = 3.80 [rather than 3.8]).
In some situations, more than two decimal places may be needed to accurately describe a variable. For example, a study looking at how long it takes people to respond to a stimulus presented on a computer screen may produce scores such as .007 and .014 seconds. Here you would want to report the results of any calculations using these data with three or perhaps four decimal places (e.g., “the sum of scores is equal to .021”).
Rounding Calculated Numbers
When rounding numbers to two decimal places, we will use the method employed by most computers and calculators:
If the number in the thousandths column is 0, 1, 2, 3, or 4, the number in the hundredths column remains unchanged. If the number in the thousandths column is 5, 6, 7, 8, or 9 the number in the hundredths column is increased by 1.
Using these rounding rules, the number 5.263 would be rounded down to 5.26 because the number in the thousandths column is 3. The number 3.417 would be rounded up to 3.42 because the number in the thousandths column is 7.
Students sometimes become concerned about a possible loss of accuracy due to rounding, choosing instead to report numbers with many decimal places. For example, a student may wish to report a number as 4.2316 rather than 4.23. However, for most situations, rounding has very little impact on the conclusions drawn from statistical analyses.
150
Calculating the Mean from Frequency Distribution Tables
When a data set is large, an alternative way to calculate the mean involves first organizing the data into a frequency distribution table. The mean can then be calculated using the following formula:
(3-4) X ¯ = ∑ fX N
where X is a score for the variable, f is the frequency of the score in the set of data, and N is the total number of scores. To calculate the mean using Formula 3–4, each score (X) is multiplied by its frequency (f) within the set of data; the results of these multiplications (fX) are then summed, and this sum (Σ(fX)) is divided by the total number of scores (N).
Using the frequency distribution table for the estimated brain power variable in Table 3.1(b), the mean for this data set may be calculated using Formula 3–4 as follows: X ¯ = ∑ ( f X ) N = ( 0 ∗ 0 ) + ( 8 ∗ 5 ) + ( 12 ∗ 10 ) + … + ( 1 ∗ 90 ) + ( 0 ∗ 95 ) + ( 0 ∗ 100 ) 38 = 0 + 40 + 120 + … + 90 + 0 + 0 38 = 870 38 = 22.89
The value of the mean (M = 22.89%) is the same as the one earlier obtained using Formula 3-3.
In the above calculations, parentheses are used to indicate the order in which calculations are made. Mathematical operations placed within parentheses are calculated before operations outside of parentheses. In Formula 3–4, the parentheses in the quantity Σ(fX) indicate that each score (X) is multiplied by its frequency (f) (such as (0 ∗ 0) and (8 ∗ 5)) before any summing is performed. Without the parentheses, the quantity ΣfX might incorrectly suggest that the frequencies should be summed (Sigma;f) before performing any multiplication.
151
The Mean as a Balancing Point in a Distribution
Table 3.2 Example of the Mean as a Balancing Point in a Distribution Table 3.2 Example of the Mean
as a Balancing Point in a Distribution
Score Mean Score – Mean
2 3.00 –1.00
3 3.00 .00
1 3.00 –2.00
4 3.00 1.00
3 3.00 0.00
1 3.00 –2.00
4 3.00 1.00
6 3.00 3.00
3 3.00 .00
One way to think of the mean as a measure of central tendency is as a balancing point in a distribution of scores for a variable. To illustrate this concept, let's return to the simple example of nine scores (2, 3, 1, 4, 3, 1, 4, 6, 3) used earlier. The first column of Table 3.2 lists the nine scores, the second column lists the mean of these scores (M = 3.00), and the numbers in the third column represent the difference between each score and the mean. For example, by subtracting the mean from the first score (2 – 3.00), we obtain a value of −1.00. It is noteworthy that some of these differences are positive (1.00, 3.00) and some are negative (–1.00, −2.00).
The mean is a balancing point because the sum of the negative differences of scores from the mean is always equal to the sum of the positive differences of scores from the mean (ignoring the minus sign). Using the third column in Table 3.2: Negative differences from mean = Positive differences from mean − 1.00 + − 2.00 + − 2.00 = 1.00 + 1.00 + 3.00 − 5.00 = 5.00
Figure 3.5 plots the three negative differences (–1.00, −2.00, −2.00) and the three positive differences (1.00, 1.00, 3.00) across the length of a balance beam. Notice that the balance beam does not lean or tilt to the left or the right; this is because the sums of the weights on each side of the balance beam (the negative and positive differences from the mean) are equal to each other. This simple example illustrates the mean as a measure of central tendency: mathematically, the mean is in the center of the distribution of any set of data.
152
Figure 3.5 Illustration of the Mean as a Balancing Point
153
Learning Check 3: Reviewing what you've Learned So Far
1. Review questions a. Why do we say that the mean is the balancing point in a distribution of scores?
2. Calculate the mean using Formula 3-3 for each of the following sets of data. a. 2, 3, 3, 5, 7 b. 8, 6, 4, 8, 3, 7 c. 71, 84, 65, 78, 89, 72, 60, 85 d. 2.5, 1.0, 3.5, 3.0, 2.0, 2.5, 4.0, 1.0, 2.0 e. 10, 8, 12, 8, 9, 7, 23, 6, 8, 11, 10, 6
3. Calculate the mean using Formula 3–4 for the data in each of the following frequency distribution tables.
a.
Score f %
5 1 10%
4 1 10%
3 2 20%
2 4 40%
1 2 20%
Total 10 100%
b.
Score f %
6 3 11%
5 7 25%
4 5 18%
3 8 29%
2 0 0%
1 4 14%
0 1 4%
Total 28 100%
4. For the sets of data in 2(a) and 2(b), calculate the difference between each score and the mean (use Table 3.2 as an example). Next, compare the sum of the positive and negative differences to see if they balance each other.
154
The Sample Mean vs. the Population Mean: Statistics vs. Parameters
For the estimated brain power variable, a mean of 22.89% was calculated for the sample of 38 psychology majors. A mean calculated from data for a sample may be referred to as a sample mean, which we have represented using the symbol X ¯ . Instead of a sample drawn from the population, what if we were able to collect data from the entire population of psychology majors? A mean associated with an entire population is called the population mean, represented by the Greek symbol μ (pronounced “mew”). Because researchers typically collect data from samples rather than entire populations, when we refer to “the mean” in this book, we will be referring to the sample mean.
A numeric characteristic of a sample (such as X ¯ ) is called a statistic; a numeric characteristic of a population (such as μ) is called a parameter. This relationship between statistics and parameters is summarized below:
Target Numeric Characteristic Mean
Sample Statistic X ¯
Population Parameter μ
As we discussed in Chapter 1, one goal of research is to test hypotheses regarding what is believed to exist in the population. To test their hypotheses, researchers collect data from samples drawn from the population; the statistics calculated from these samples may be used to estimate population parameters. For example, a sample mean ( X ¯ ) might be used to estimate a population mean (μ). For the 10% myth study example, the sample mean of 22.89% could be used to estimate the mean for the entire population of psychology majors. In doing so, researchers are able to test the hypothesis that data collected from the population of psychology majors would result in a mean of 22.89%.
The relationship between samples and populations and the use of statistics to estimate parameters is critical to testing research hypotheses. The process of hypothesis testing and statistics used to test research hypotheses, known as inferential statistics, will be illustrated in greater detail starting in Chapter 6 of this book.
155
3.6 Comparison of the Mode, Median, and Mean
Each of the three measures of central tendency presented in this chapter has strengths and weaknesses, depending on the nature of the data and the objectives of the research. This section evaluates and compares the mean, median, and mode, identifying situations in which the use of one may be preferable to the others.
156
Strengths and Weaknesses of the Mode
The first comparative advantage of the mode is that it can be determined relatively easily and quickly because it does not require any mathematical calculations. Unlike the median and mean, the mode is identified simply by looking at a frequency distribution table or figure. Another advantage of the mode is that it may be used to summarize categorical variables, which are variables measured at the nominal level of measurement. Imagine, for example, 50 people are asked to indicate their favorite flavor of ice cream and their data are organized into the following frequency distribution table:
Flavor f %
Vanilla 18 36%
Chocolate 23 46%
Strawberry 9 18%
Total 50 100%
At a single glance at this table, we could state that “the modal flavor was chocolate.”
Another strength of the mode is that it is always an actual score for a variable, something that may not be true for the mean or median. You may have heard, for example, that “the average American family has two and a half children.” If these data had been summarized using the mode rather than the mean, you would not lose sleep trying to figure out what “half” a child looks like.
An additional advantage of the mode is that it accurately describes distributions that have more than one mode. Below is a small set of data that has been rank ordered from lowest to highest: 3 4 4 4 6 8 10 10 10
A histogram has been created for this set of data in Figure 3.6. Looking at the histogram, we see that this distribution has two modes: 4 and 10. In Chapter 2, we referred to a distribution with two modes as a bimodal distribution. As we will demonstrate shortly, the mode is a more appropriate measure of central tendency for bimodal distributions than the median or the mean.
Finally, because the mode is the single most common value for a variable, determining the mode is not affected by outliers in the distribution (as is, for example, the mean). In Chapter 2, outliers were defined as rare, extreme scores for a variable lying outside the range of the majority of the scores.
Figure 3.6 Identifying the Modes in a Bimodal Distribution
157
Despite these strengths, the mode is of limited use to researchers. First, because the mode is based on just one value of a variable, it may not adequately represent the entire set of data. Consider the following two sets of data, each of which consists of eight scores: Variable 1: 5 7 10 11 11 13 15 18 Variable 2: 11 11 14 18 19 20 22 25
Although the mode for the two variables is the same (Mo = 11), the two sets of data are clearly different in nature. For Variable 1, the modal value of 11 lies near the center of the distribution of scores and provides a useful representation of the distribution. For Variable 2, however, the modal value is at one end of the distribution and does not accurately represent the set of data. Because the mode does not take into account all of the scores in the distribution, it cannot be used in statistical analyses that are designed to test hypotheses about distributions.
158
Strengths and Weaknesses of the Median
The primary strength of the median is that it provides a more accurate description of skewed (asymmetric) distributions than does the mean. Recall from Chapter 2 that a skewed distribution is one in which the majority of data for a variable are located at one end of the range of values. Skewed distributions are different from symmetric distributions, where the highest frequencies are in the center of the distribution, with the frequencies decreasing in a similar manner as we move away from the center in both the left and right directions.
A good example of a variable with a skewed distribution is income. As a simple yet representative example of income, Table 3.3 lists the annual salaries for the players on the starting offensive unit of the San Francisco 49ers professional football team, rank ordered from lowest to highest. The distribution of salaries in this table is positively skewed because most are at the lower end of the distribution, with the tail at the upper end consisting of the particularly high salaries of $4,675,000 and $11,800,000.
Suppose we decided to describe the distribution of salaries in Table 3.3 using the mean as the measure of central tendency. The mean of the 11 salaries is calculated below: X ¯ = ∑ X N = 449800 + 791720 + … + 4675000 + 11800000 11 = 34354800 11 = 3123163.64
Table 3.3 Salaries of the Starting Offensive Unit of the San Francisco 49ers Table 3.3 Salaries of
the Starting Offensive Unit of the San Francisco 49ers
Player Salary ($)
1 449,800
2 791,720
3 1,250,000
4 1,575,520
5 2,000,000
6 2,330,000
7 2,662,000
8 2,820,760
9 4,000,000
10 4,675,000
159
11 11,800,000
Although this mean of $3,123,163.64 may be a balancing point of the distribution, it is not located in the center of the distribution because the salaries of 8 of the 11 players (73%) are below the mean. The mean is affected by outliers, such as the one players salary of $11,800,000, and hence does not accurately portray a skewed distribution such as this one.
An alternative way to describe the salaries of these players would be to use the median. Because there are an odd number of scores in the variable (N = 11), Formula 3-1 is used to calculate the median as follows: Median = N + 1 2 th score = 11 + 1 2 th score = 12 2 th score = 6 th score
Starting from the lowest salary, the sixth players salary is $2,330,000. Table 3.4 illustrates the mean and median of the salaries of the 11 players. As we can see, the median ($2,330,000) is located closer to the center of the distribution than the mean ($3,123,163.64).
Table 3.4 Salaries of the Starting Offensive Unit of the San Francisco 49ers Football Team, With the Mean and Median Indicated
Although the median is useful for describing skewed distributions, it does have several limitations. First, like the mode, the median is not based on all of the scores in a distribution of data but rather is based only on the one or two scores in the center of the distribution. Consequently, the median cannot be used in statistical analyses that are designed to test hypotheses about distributions in the population.
A second limitation of the median is that it does not accurately describe bimodal distributions. Let's return to the set of nine numbers that comprise the bimodal distribution in Figure 3.6. What if we were to calculate the median for this set of data? Because there are an odd number of scores (N = 9), we apply Formula 3-1 to the data as follows: Median = N + 1 2 th score = 9 + 1 2 th score = 10 2 th score = 5 th
160
Finding the fifth score, the median of the set of nine scores is equal to 6. But as we can see from the histogram in Figure 3.7, the two modes of 4 and 10 more accurately represent the most common values of this bimodal distribution than does the median value of 6.
161
Strengths and Weaknesses of the Mean
Figure 3.7 Identifying the Modes and the Median in a Bimodal Distribution
Compared with the mode and median, the mean has one critical strength: It is calculated from all of the scores in a distribution of data (ΣX) rather than just the most frequently appearing scores or the one or two scores in the center of the distribution. As such, the mean possesses important mathematical properties that allow it to be used in statistical analyses to test hypotheses regarding populations. Because testing hypotheses is a critical aspect of research, the mean will be used as the primary measure of central tendency for the remainder of this book.
However, the mean has several limitations. As illustrated in the football salary example in Table 3.4, the mean does not accurately represent skewed distributions because it is affected by outliers. In the 10% myth study, the sample mean for the estimated brain power variable was calculated to be 22.89%; however, the mode and the median were both 10%, which was in fact the most frequently occurring value in the distribution. Why was the mean so much higher than the mode and median? As the frequency distribution table in Table 3.1(b) reveals, the value of the mean was influenced by several outliers, particularly one student's score of 90%.
The mean is also limited in its ability to provide an accurate representation of bimodal distributions. Returning once more to the bimodal distribution in Figure 3.6, below we calculate the mean for the set of nine scores: X ¯ = ∑ X N = 3 + 4 + 4 + 4 + 6 + 9 + 10 + 10 + 10 9 = 6.67 = 60 9
As was the case with the median, the mean of 6.67 represents the distribution less accurately than the two modes of 4 and 10. These comparisons illustrate an important principle for statistical analysis: It is often helpful to calculate more than one measure of central tendency to fully understand the shape and nature of a distribution.
162
Choosing a Particular Measure of Central Tendency
Table 3.5 summarizes the relative strengths and weaknesses of the mode, median, and mean. Under what conditions might one use a particular measure of central tendency? As the previous examples have demonstrated, the choice will be largely based on the shape of the distribution.
Figure 3.8 displays the mode, median, and mean for three distributions: symmetric, asymmetric (skewed), and bimodal. The distribution in Figure 3.8(a) is symmetric and unimodal. In this situation, the mode, median, and mean are all in the center of the distribution. When the distribution is symmetric and unimodal, the mean is used as the measure of central tendency to conduct statistical analyses designed to test hypotheses. The two distributions in Figure 3.8(b) are both skewed. In portraying skewed distributions, the median is the preferred measure of central tendency because it lies closer to the center of the distribution of scores than does the mean or the mode. Finally, for the bimodal distribution illustrated in Figure 3.8(c), the mode is a more useful choice as a measure of central tendency than either the median or the mean because only the mode highlights the two scores with the highest frequencies.
Table 3.5 Strengths and Weaknesses of the Mode, Median, and Mean Table 3.5 Strengths and Weaknesses of the Mode, Median, and Mean
Measure of Central Tendency
Strength Weakness
Mode
Quick and easy to identify (no calculations necessary) Can be used with categorical variables An actual score for the variable Describes bimodal or multimodal distributions Not affected by outliers
Not based on all of the scores in a set of data Cannot be used in statistical analyses to test hypotheses
Median
Describes skewed
distributions
Not based on all of the scores in a set of data Cannot be used in
statistical analyses to test
163
Not affected by outliers hypotheses May not describe bimodal distributions
Mean
Based on all of the scores in a set of data Can be used in statistical analyses to test hypotheses
Affected by outliers May not describe skewed distributions May not describe bimodal distributions
164
3.7 Measures of Central Tendency: Drawing Conclusions
One of the last steps in conducting statistical analyses is to draw conclusions that interpret the results of statistical analyses in light of a study's research questions or hypotheses. Drawing conclusions from analyses involves answering the question, “So what does it mean?” Answering this question accurately and appropriately is perhaps one of the most difficult aspects of the research process.
What conclusions can be drawn from a measure of central tendency? Recall from Chapter 2 that researchers examine the data they have collected to gain an understanding of three aspects of distributions: modality, symmetry, and variability. Measures of central tendency provide a numerical description of the modality and, to a lesser degree, the symmetry of a variable.
In terms of modality, the most common value for a variable, the information provided by a measure of central tendency can be used to draw conclusions about the sample from whom the data were collected. For example, the finding that “the mean GPA of a sample of students was 3.80” may lead to the conclusion that this is a very capable group of students, given that the highest possible value of GPA is 4.00.
Figure 3.8 The Mode, Median, and Mean in Different Distributions
165
In addition to modality, conclusions about the symmetry of the distribution can be drawn from a measure of central tendency. For example, the real estate sections of newspapers often report the “median price of a home.” Using the median (rather than the mean) to describe housing prices helps us understand that the distribution of home prices is skewed rather than symmetric.
The study introduced in the beginning of this chapter examined the relationship between studying psychology and a lowered belief in the 10% myth. In identifying the mode and the median of the data collected from the sample of 38 psychology majors, however, we found they were more likely to say that people use 10% of their brain power than any other amount, with the majority of these students' beliefs at the low end of the 0% to 100% range. On the basis of their calculation and examination of measures of central tendency, the researchers for the study proposed the following implication:
Students trained in psychology may not be more critical of the ten-percent claim than are other students…. The results of this study suggest that for this popular misconception, additional psychology courses taken by psychology majors may not have much effect. (Higbee & Clay, 1998, p. 472)
166
Learning Check 4: Reviewing what you've Learned So Far
1. Review questions a. What is the difference between the sample mean ( X ¯ ) and the population mean (μ)? b. What is the difference between statistics and parameters? c. What are the relative strengths and weaknesses of the mode, median, and mean as measures
of central tendency? For what types of situations might you use one rather than another? d. For which of the three aspects of distributions (modality, symmetry, and variability) do
measures of central tendency provide information?
167
3.8 Looking Ahead
Measures of central tendency summarize data in terms of the most frequent or central score, a score that is based on what the scores have in common with each other. However, to accurately portray a set of data, we must also describe differences among scores. For example, your instructor might tell your class, “The average score on the midterm was 34.75.” Let's say you received a midterm score of 39. How do you evaluate your performance? To answer this question, you must examine the third aspect of distributions: variability, which refers to the amount of differences in a distribution of data. How much variability was there in midterm scores? In Chapter 4, you will learn how to numerically describe variability.
168
3.9 Summary
A measure of central tendency is a descriptive statistic of the most typical, common, or frequently occurring value for a variable. Three common measures of central tendency are the mode, the median, and the mean.
The mode is the score or value of a variable that appears most frequently in a set of data. The simplest way to determine the mode is to examine a frequency distribution table or figure.
The median is the number, located at the center of a set of numbers that has been arranged in ascending or descending order, that splits a distribution of scores in half, with the same number of scores above the median as below. The median is determined by rank ordering the scores and uses the appropriate formula based on whether the total number of scores is odd or even.
The mean is the arithmetic average of a set of scores. The mean is determined by calculating the sum of a set of scores and dividing this sum by the total number of scores. The mean may also be calculated from data that have been organized into a frequency distribution table. The mean is a balancing point because the sum of the negative differences of scores from the mean is always equal to the sum of the positive differences of scores from the mean.
A mean calculated from data for a sample may be referred to as a sample mean ( X ¯ ); a mean associated with a population is called the population mean (μ). A numeric characteristic of a sample (such as X ¯ ) is called a statistic; a numeric characteristic of a population (such as μ) is called a parameter.
The mode, median, and mean each have relative strengths and weaknesses. As a result, the mode is the preferred measure of central tendency when a distribution of scores for a variable is bimodal, the median is preferred when the distribution is skewed, and the mean is preferred when the distribution is symmetric and unimodal.
169
3.10 Important Terms
measure of central tendency (p. 73) mode (modal score) (p. 74) median (p. 74) mean (p. 80) sample mean (p. 86) population mean (p. 86) statistic (p. 86) parameter (p. 86)
170
3.11 Formulas Introduced in this Chapter
171
Median (Odd Number of Scores)
(3-1) Median = N + 1 2 th score
Median (Even Number of Scores)
(3-2) Median = N 2 th score + N 2 + 1 th score 2
Mean
(3-3) X ¯ = ∑ X N
Mean (Frequency Distribution Table)
(3-4) X ¯ = ∑ fX N
172
3.12 Exercises
1. Identify the mode for each of the following sets of data: a. 7, 3, 7, 1, 7 b. 3, 1, 1, 2, 3, 3 c. 8, 22, 8, 10, 9, 8, 12, 7, 8 d. 5, 3, 6, 4, 5, 3, 1, 5, 2, 5, 6, 5 e. 11, 9, 13, 6, 14, 12, 7, 12, 10, 8, 15, 5, 12 f. 7, 3, 14, 9, 7, 6, 4, 7, 4, 10, 4, 2, 8, 5, 11, 12
2. Identify the mode for each of the following sets of data: a. 6, 2, 7, 6, 6, 4 b. 13, 19, 12, 13, 7, 13, 20, 13, 15 c. 4, 2, 5, 1, 2, 2, 5, 5, 3, 1, 5, 2, 6, 3 d. 40, 10, 35, 30, 10, 25, 5, 10, 15, 10, 5, 35, 10, 40, 20 e. 1, 4, 2, 4, 4, 1, 5, 1, 4, 2, 3, 1, 4, 2, 3, 5, 1
3. Calculate the median for each of the following sets of data: a. 3, 4, 6, 7, 10 b. 14, 16, 19, 21, 22, 27, 36 c. 6, 6, 9, 9, 9, 10, 13, 16, 16, 21, 24 d. 4, 2, 6, 3, 8, 3, 8 e. 27, 24, 35, 30, 41, 32, 36, 35, 27 f. 3, 5, 5, 6, 9, 10 g. 11, 13, 13, 16, 17, 22, 24, 27, 28, 30 h. 12, 10, 9, 11, 9, 15, 10, 9, 8, 13, 9, 10
4. Calculate the median for each of the following sets of data: a. 2, 2, 3, 4, 6, 9, 9 b. 7, 7, 8, 12, 13, 14, 17, 21, 23 c. 11, 6, 3, 14, 12, 15, 8, 10, 7 d. 9, 6, 7, 5, 8, 7, 8, 9, 10, 8, 7 e. 10, 13, 13, 14, 15, 16, 17, 24 f. 6.0, 8.5, 7.0, 3.0, 4.5, 2.5, 3.5, 7.5, 3.5, 8.0
5. Calculate the mean for each of the following sets of data: a. 2, 2, 3, 5 b. 3, 5, 5, 7, 7, 8, 8, 11 c. 9, 6, 7, 5, 8, 7, 8, 9, 10, 8, 7 d. 11, 13, 14, 15, 16, 16, 16, 17, 18, 20, 22, 25 e. 2.57, 3.62, 2.84, 1.90, 3.15, 3.41, 3.03 f. 5, 2, 3, 6, 3, 4, 7, 2, 2, 8 g. 12.50, 10.00, 9.75, 11.50, 9.00, 15.00, 10.25, 9.50, 8.75, 13.75, 9.50, 10.25
6. Calculate the mean for each of the following sets of data: a. 1, 1, 2, 3, 3
173
b. 1, 4, 4, 5, 6, 7, 8 c. 47, 56, 62, 69, 70, 73 d. .75, .23, .48, .60, .98, .65, .08, .12., .39 e. 11.02, 13.67, 17.39, 14.74, 16.10, 9.28, 12.31, 8.66 f. 12, 10, 9, 11, 9, 15, 10, 9, 8, 13, 9, 10
7. Identify the mode in each of the following frequency distribution tables. a.
Score f %
4 5 31%
3 3 19%
2 6 38%
1 2 12%
Total 16 100%
b.
Score f %
5 4 16%
4 8 32%
3 6 24%
2 3 12%
1 2 8%
0 2 8%
Total 25 100%
c.
Score f %
8 0 0%
7 2 6%
6 3 8%
5 6 17%
4 10 28%
3 8 21%
2 5 14%
174
1 2 6%
Total 36 100%
d.
Score f %
100 0 0%
90 1 3%
80 2 7%
70 0 0%
60 2 7%
50 3 10%
40 6 20%
30 5 17%
20 8 26%
10 3 10%
Total 30 100%
8. Using Formula 3–4, calculate the mean for the variables in the frequency distribution tables in Exercise 7.
9. If a variable with N = 25 has a mean of 3.25, what is the value of ΣX? 10. For the sets of data in Exercise 5(a) and (b), calculate the difference between each
score and the mean (use Table 3.2 as an example). Next, compare the sum of the positive and negative differences to see if they balance each other.
11. What makes you happy? Participants in a research experiment reported their levels of happiness and the amount of time spent alone versus with friends or family. The researchers found that people who were reportedly “very happy” spent more time with friends and family and concluded that sociability may be related to happiness (Diener & Seligman, 2002). The following data represent the number of hours eight “very happy people” spend with friends and family per day (the data below are representative of the study's findings):
Participant # Hours
1 5
2 3
3 7
4 5
175
5 4
6 5
7 5
8 7
a. Calculate the number of scores (N) and the sum of scores (ΣX). b. Calculate the mode, median, and the mean.
12. (This example was introduced in Chapter 2.) A friend of yours asks 20 people to rate a movie using a 1- to 5-star rating: the higher the number of stars, the higher the recommendation. Their ratings are listed below:
Person # Stars
1 ***
2 *****
3 **
4 ****
5 ***
6 ***
7 ****
8 **
9 ***
10 *****
11 *
12 ****
13 ****
14 ***
15 ***
16 **
17 ***
18 **
19 ****
20 ***
a. Calculate the number of scores (N) and the sum of scores (ΣX). b. Calculate the mode, median, and the mean. c. Looking at the mode, median, and mean for the data in Exercise 15, how
176
would you describe the shape of the distribution of star ratings? 13. One study stated the research hypothesis, “Violence behavior in children may be
reduced by teaching them conflict resolution skills” (DuRant et al., 1996). The variable “violence behavior” was measured by the number of fights in which each student was involved. Below is a frequency distribution table for the number of fights for 13 students in this study.
# Fights f %
4 1 8%
3 2 15%
2 3 23%
1 3 23%
0 4 31%
Total 13 100%
a. Calculate the mode, median, and the mean for these data. b. Comparing the mode, median, and the mean, how would you describe the
shape of the distribution? 14. The 10% myth study discussed in this chapter measured the beliefs of both
psychology majors and non–psychology majors. The 39 non–psychology majors in this study provided the following values for the estimated brain power variable:
Non–Psychology Major Estimated Brain Power
1 15
2 45
3 10
4 40
5 5
6 10
7 50
8 45
9 10
10 10
11 35
12 5
13 15
177
14 10
15 20
16 50
17 25
18 45
19 40
20 25
21 10
22 5
23 15
24 5
25 40
26 30
27 15
28 25
29 10
30 20
31 10
32 5
33 15
34 10
35 5
36 60
37 10
38 5
39 15
a. Calculate the number of scores (N) and the sum of scores (ΣX). b. Calculate the mode, median, and the mean. c. Looking at the mode, median, and mean, how would you summarize the
beliefs of the non–psychology majors? How would you compare their beliefs to those of the psychology majors discussed earlier in this chapter?
15. At an ice skating competition, the score for each skater is the mean of the different judges' scores. However, before this mean is calculated, the lowest and the highest scores among the judges are discarded.
178
a. Why are the highest and lowest scores not included in the calculation of the mean?
b. Which measure(s) of central tendency would eliminate the need to discard scores?
16. Given the following values for the mode, median, and mean, determine whether you believe the distribution is symmetrical, positively skewed, or negatively skewed.
a. Mode = 4, median = 5, mean = 8 b. Mode = 9, median = 8, mean = 4 c. Mode = 6, median = 6, mean = 6
17. On your own, generate a set of data where the mean, median, and mode are the same, and then create a figure for these data. How would you describe the shape of this distribution?
18. On your own, generate a set of data that is positively skewed, and then create a figure for these data. If you calculate the mean and median of these data, which is larger?
179
Answers to Learning Checks
Learning Check 1
2. a. Mo = 2 b. Mo = 6 c. Mo = 12 d. Mo = 2 e. Mo = 3.2
Learning Check 2
2. a. Mdn = 4 b. Mdn = 14 c. Mdn = 10.50 d. Mdn = 80.50 e. Mdn = 9
Learning Check 3
2. a. M = 4.00 b. M = 6.00 c. M = 75.50 d. M = 2.39 e. M = 9.83
3. a. M = 2.50 b. M = 3.61
4.
Score Mean Score – Mean
2 4.00 –2.00
3 4.00 –1.00
3 4.00 –1.00
5 4.00 1.00
7 4.00 3.00
180
a. Negative differences from mean = Positive differences from mean − 2.00 + − 1.00 + − 1.00 = 1.00 + 3.00 − 4.00 = 4.00
Score Mean Score – Mean
8 6.00 2.00
6 6.00 .00
4 6.00 –2.00
8 6.00 2.00
3 6.00 –3.00
7 6.00 1.00
b. Negative differences from mean = Positive differences from mean − 2.00 + − 3.00 = 2.00 + 2.00 + 1.00 − 5.00 = 5.00
181
Answers to Odd-Numbered Exercises
1. a. Mo = 7 b. Mo = 3 c. Mo = 8 d. Mo = 5 e. Mo = 12 f. Mo = 4 and 7
3. a. Mdn = 6 b. Mdn = 21 c. Mdn = 10 d. Mdn = 4 e. Mdn = 32 f. Mdn = 5.50 g. Mdn = 19.50 h. Mdn = 10.00
5. a. X ¯ = 3.00 b. X ¯ = 6.75 c. X ¯ = 7.64 d. X ¯ = 16.92 e. X ¯ = 2.93 f. X ¯ = 4.20 g. X ¯ = 10.81
7. a. Mo = 2 b. Mo = 4 c. Mo = 4 d. Mo = 20
9. ΣX = (3.25 ∗ 25) = 81.25 11.
a. N = 8, ΣX = 41 b. Mode = 5, median = 5, mean = 5.13
13. a. Mode = 0, median = 1, mean = 1.46 b. The fact that the mode (Mo = 0) is less than the median (Mdn = 1), which in
turn is less than the mean (M = 1.46), suggests that the distribution of the number of fights is positively skewed, with the majority of the sample having relatively few fights.
182
15. a. The highest and lowest scores are not included to prevent the mean from being
affected by possible outliers. b. If they used the median or the mode rather than the mean, they would not
need to discard any scores because the median and mode are not affected by outliers.
17. Note: A distribution with equal values for the mode, median, and mean should be symmetrical.
183
Sharpen your skills with SAGE edge at edge.sagepub.com/tokunaga
SAGE edge for students provides a personalized approach to help you accomplish your coursework goals in an easy-to-use learning environment. Log on to access:
eFlashcards Web Quizzes Chapter Outlines Learning Objectives Media Links
184
Chapter 4 Measures of Variability
185
Chapter Outline 4.1 An Example From the Research: How Many “Sometimes” in an “Always”? 4.2 Understanding Variability 4.3 The Range
Strengths and weaknesses of the range 4.4 The Interquartile Range
Strengths and weaknesses of the interquartile range
4.5 The Variance (s2) Definitional formula for the variance Computational formula for the variance Why not use the absolute value of the deviation in calculating the variance? Why divide by N – 1 rather than N in calculating the variance?
4.6 The Standard Deviation (s) Definitional formula for the standard deviation Computational formula for the standard deviation
4.7 Measures of Variability for Populations
The population variance (σ2) The population standard deviation (σ) Measures of variability for samples vs. populations
4.8 Measures of Variability: Drawing Conclusions 4.9 Looking Ahead 4.10 Summary 4.11 Important Terms 4.12 Formulas Introduced in This Chapter 4.13 Using SPSS 4.14 Exercises
Chapter 3 introduced descriptive statistics, the purpose of which is to numerically describe or summarize data for a variable. As we learned in Chapter 3, a set of data is often described in terms of the most typical, common, or frequently occurring score by using a measure of central tendency such as the mode, median, or mean. Measures of central tendency numerically describe two aspects of a distribution of scores, modality and symmetry, based on what the scores have in common with each other. However, researchers also describe the amount of differences among the scores of a variable in analyzing and reporting their data. This chapter introduces measures of variability, which are designed to represent the amount of differences in a distribution of data. In this chapter, we will present and evaluate different measures of variability in terms of their relative strengths and weaknesses. As in the previous chapters, our presentation will depend heavily on the findings and analysis from a published research study.
186
4.1 An Example from the Research: How Many “Sometimes” in an “Always”?
Throughout our daily lives, we are constantly asked to fill out questionnaires and surveys. Restaurants ask us to evaluate the quality of the food and service they provide; companies want to know how often we shop online; political organizations are interested in knowing what we think about our elected officials. In responding to these types of scenarios, we are often asked to describe our feelings, behaviors, or beliefs by choosing among comparative terms such as excellent, good, or poor. To what degree do words such as these have the same meaning for different people? Researchers who collect data using questionnaires are concerned about the clarity of the words they use. If research participants do not agree about the meaning of key terms and phrases, it is difficult to combine or compare their responses. The following example demonstrates how important it is to understand and address these differences when research is conducted with people of different backgrounds, cultures, or languages.
As members of an international research team, two British researchers, Suzanne Skevington and Christine Tucker, conducted a study assessing differences in interpretation of labels for questionnaire items measuring the frequency of health-related behaviors, such as, “How often do you suffer pain?” (Skevington & Tucker, 1999). As a part of their study, they examined how British people define frequency-related terms such as seldom, usually, and rarely. To this end, the researchers assigned numeric values to participants' evaluation of these words in order to measure the degree to which people differ in their interpretation of these words.
In their study, Skevington and Tucker (1999) presented 20 British adults a series of frequency-related words. For each word, they included a line 100 millimeters (about 4 inches) in length. At opposite ends of the line were attached the labels Never and Always, representing the lowest and highest possible values for frequency. After reading each word, participants were asked to place an X on the point in the line that in their opinion best represented the frequency of the word, relative to the end points of Never (0%) and Always (100%). An example of their methodology is provided in Figure 4.1.
After the participants had completed their responses, the researchers used a ruler to measure, for each word, the distance from the word Never to the “X” provided by the participant. This distance, which was measured in millimeters, represented the perceived “frequency” of the word. For example, in Figure 4.1, the X for the word sometimes is in the exact middle of the line; consequently, this word is given a frequency rating of 50% because it is 50 mm from the left end. In this study, the variable, which we will call “frequency rating,” had possible values ranging from 0 to 100.
187
In this example, we will focus on the study's findings regarding the word sometimes. The frequency ratings for the word sometimes for the 20 participants are listed in Table 4.1(a). Table 4.1(b) organizes the raw data into a grouped frequency distribution table, and Figure 4.2 illustrates the distribution using a frequency polygon. A frequency polygon is used for this variable (rather than a histogram or pie chart) because the variable “frequency rating” is measured at the ratio level of measurement, with the value of 0 representing the complete absence of distance from the left end of the 100-mm line.
As was discussed in Chapter 3, the data for a variable can be summarized using a measure of central tendency. Using Formula 3-3 from Chapter 3, the mean frequency rating of the word sometimes for the 20 participants is calculated as follows: X ¯ = ∑ X N = 46 + 15 + 21 + … + 22 + 66 + 26 20 = 666 20 = 33.30
Figure 4.1 Example of Ratings of Frequency-Related Words Along a 100-mm Line
From this calculation, we conclude that the participants in the study believed, on average, the word sometimes represents a 33.30% frequency. However, the grouped frequency distribution table in Table 4.1(b) reveals that participants provided a wide range of ratings, some of which were far from the average rating of 33.30%. For example, the ratings provided by 11 of the 20 participants were lower than 30%, while 6 of the 20 participants provided ratings of 40% or higher.
Given the variety of ratings these participants gave the word sometimes, in order to accurately describe the distribution of responses for this variable, one must not only include a measure of central tendency (which focuses on what the scores have in common with each other) but also represent the degree to which the scores differ, or vary. The next section introduces the concept of variability.
188
4.2 Understanding Variability
We begin our discussion of variability by examining the three distributions presented in Figure 4.3. Chapter 2 discussed three aspects of distributions: modality, symmetry, and variability. The modality and symmetry of a distribution can be described using measures of central tendency such as those introduced in Chapter 3. However, because all three distributions in Figure 4.3 are unimodal and symmetric, they would be described as having the same mode, median, and mean. It is obvious, however, that the shapes of the three distributions are very different. As the remainder of this chapter will illustrate, these differences can be portrayed using a statistic that describes the third aspect of distributions: variability.
The word variability typically evokes words such as differences, dispersion, or changes. In a statistical sense, variability refers to the amount of spread or scatter of scores in a distribution. The concept of variability is a critical issue in the behavioral sciences, where research frequently examines differences in such things as characteristics, attitudes, and cognitive abilities. Different people, for example, express different levels of extroversion, different attitudes toward capital punishment, and different learning styles. Ultimately, the primary goal of a science such as psychology is to describe, understand, explain, and predict variability.
Table 4.1 Frequency Rating of the Word Sometimes by 20 Participants Table 4.1 Frequency Rating of the
Word Sometimes by 20 Participants
(a) Raw Data
Participant Frequency Rating
1 46
2 15
3 21
4 49
5 23
6 39
7 25
8 27
9 20
10 23
189
11 58
12 24
13 36
14 20
15 49
16 32
17 45
18 22
19 66
20 26
Table 4.1 Frequency Rating of the Word Sometimes by
20 Participants
(b) Grouped Frequency Distribution Table
Frequency Rating f %
70–100 0 0%
60–69 1 5%
50–59 1 5%
40–49 4 20%
30–39 3 15%
20–29 10 50%
10–19 1 5%
0–9 0 0%
Total 20 100%
Researchers have developed statistics designed to measure variability. A measure of variability is a descriptive statistic of the amount of differences in a set of data for a variable. The purpose of measures of variability is to numerically represent a set of data based on how the scores differ or vary from each other. Similar to measures of central tendency, there are multiple measures of variability. The next part of this chapter presents and discusses four measures of variability:
Figure 4.2 Frequency Polygon, Frequency Ratings of Sometimes
190
the range, the interquartile range, the variance, and the standard deviation.
Each of these measures of variability will be defined, illustrated in terms of their necessary calculations, and evaluated based on their relative strengths and weaknesses.
191
4.3 The Range
One way to describe the amount of variability in a distribution of data for a variable is to focus on the two ends of the distribution. Therefore, the first measure of variability we will discuss is the range, defined as the mathematical difference between the lowest and highest scores in a set of data:
(4-1) Range = highest score − lowest score
Figure 4.3 Three Distributions with the Same Modality and Symmetry but Different Variability
The range is computed by identifying the lowest and highest scores in a set of data and then subtracting the lowest score from the highest score to compute the difference between the two scores.
To illustrate how to calculate the range, let's return to the small set of data introduced in Chapter 3 to discuss measures of central tendency: 2, 3, 1, 4, 3, 1, 4, 6, and 3. Among the nine scores in this set of data, the lowest score is 1 and the highest score is 6. Using Formula 4-1, the range for this set of data is the following: Range = highest score − lowest score = 6 − 1 = 5
As a second example, to calculate the range for the frequency rating variable in Table 4.1(a), the ratings of the 20 participants are ranked from lowest to highest:
Among the 20 participants, the lowest and highest ratings are 15 and 66, respectively. Therefore, the range for the frequency rating variable is Range = highest score − lowest score = 66 − 15 = 51
192
When researchers report the range for a variable, they generally provide the actual values of the lowest and highest scores. For the frequency rating variable, for example, the range might be reported as the following: “In this sample of 20 participants, the frequency ratings for the word sometimes had a range of 51% (low = 15%, high = 66%).” Including the lowest and highest scores may provide information about the sample from which the data were collected. For example, even though the range ($35,000) may be the same in two samples, a sample in which the lowest income is $5,000 and the highest income is $40,000 would be considered very differently from a sample in which the lowest and highest incomes are $175,000 and $210,000.
193
Strengths and Weaknesses of the Range
As a measure of variability, the range has comparative strengths and weaknesses. The primary strength of the range is that it is easy and quick to compute, particularly if the sample is small or if computer software is used to sort the scores from lowest to highest. A second strength of the range, as mentioned earlier, is that indicating the lowest and highest scores for a variable provides information about the sample from which the data were collected.
However, because the range is calculated from only two scores in the distribution (the lowest and highest), it may not accurately reflect the amount of variability in the entire distribution of scores. Consider, for example, the following two sets of scores sorted from highest to lowest: Set 1: 1 2 2 3 3 3 4 4 5 Set 2 : 1 3 4 5 5 5 5 5 5
Although the range for both sets of data is the same (5 – 1 = 4), there is much more variability in the first set of data than in the second. The data in Set 2 illustrates another weakness of the range: It is affected by outliers (in this case, the one score of 1). Because the range does not take into account all of the scores in the distribution, a fundamental weakness of the range is that it cannot be used in statistical analyses designed to test hypotheses about distributions.
194
4.4 The Interquartile Range
One of the limitations of the range as a measure of variability is that it is affected by extreme scores known as outliers. One way to overcome the possible influence of outliers on the range is to calculate the interquartile range, which is the range of the middle 50% of the scores in a set of data. The interquartile range is calculated using Formula 4-2:
(4-2) Interquartile range = N − N 4 th score − N 4 + 1 th score
where N is the total number of scores. The interquartile range is calculated by removing the highest and lowest 25% of the distribution and then calculating the range of the remaining scores. The primary purpose of the interquartile range is to decrease the influence of outliers in representing the variability in a set of data.
As a simple example, consider the following eight scores (N = 8), ranked from lowest to highest:
Using Formula 4-2, the first step is to identify the N − N 4 and the N 4 + 1 th scores. Because N = 8 in this example, Interquartile range = ( N − N 4 ) th score − ( N 4 + 1 ) th score = ( 8 − 8 4 ) th score - ( 8 4 + 1 ) th score = ( 8 − 2 ) th score - ( 2 + 1 ) th score = 6 th score-3rd score=39 - 19 = 20
So, for this small set of data, the interquartile range is the difference between the values of the sixth and the third scores. Among the ranked scores, the sixth score is 39 and the third score is 19; therefore, the interquartile range is equal to (39 – 19), or 20. Note that the interquartile range of 20 is much smaller than the range of this set of data, which is (89 – 11), or 78.
To calculate the interquartile range for the frequency rating variable, because N = 20 in this example, we start by entering the value 20 into Formula 4-2: I n t e r q u a r t i l e r a n g e = ( N − N 4 ) t h s c o r e − ( N 4 + 1 ) t h s c o r e = ( 20 − 20 4 ) t h s c o r e − ( 20 4 + 1 ) t h s c o r e = ( 20 − 5 ) t h s c o r e − ( 5 + 1 ) t h s c o r e = 15 t h s c o r e − 6 t h s c o r e = 45 − 23 = 22
To calculate the interquartile range for the frequency rating variable, the 15th and 6th scores must be identified. Earlier, the 20 ratings were ranked from lowest to highest to calculate the range; looking at the ranked ratings, we find the 15th score is equal to 45 and the 6th score is equal to 23. Therefore, the interquartile range for this set of data is equal to (45 – 23), or 22.
195
Strengths and Weaknesses of the Interquartile Range
Compared with the range, the primary purpose of the interquartile range is to represent the variability in a set of data while lessening the influence of outliers. However, using the interquartile range has the potential to misrepresent a set of data by ignoring half of the scores (the top and bottom 25%) in the data. It is, after all, somewhat counterintuitive to measure the variability in a set of data with only half of the data. Also, because the interquartile range, like the range, does not take into account all of the scores in the distribution, it cannot be used in statistical analyses designed to test hypotheses about distributions.
Given its limitations, under what conditions is the interquartile range most appropriately used? Concern about the impact of outliers on a set of data was discussed in Chapter 3, where one advantage of the median as a measure of central tendency is that it is not affected by outliers. Therefore, the interquartile range is typically reported along with the median to represent the variability and central tendency in distributions that either are skewed or have outliers.
196
Learning Check 1: Reviewing what you've Learned So Far
1. Review questions a. Why is it important to calculate measures of variability for a variable in addition to
measures of central tendency? b. What are the relative advantages and disadvantages of the range as a measure of variability? c. What is the difference between the range and the interquartile range?
2. Calculate the range for each of the following sets of data: a. 2, 1, 4, 2, 7, 3 b. 15, 6, 17, 9, 12, 5, 16, 7 c. 21, 14, 13, 17, 30, 17, 14, 11, 13, 14, 19 d. 3.0, 2.8, 3.2, 3.6, 3.7, 3.2, 3.5, 3.6, 3.1, 3.2, 3.0, 3.2, 3.8, 3.3
3. Calculate the interquartile range for each of the following sets of data: a. 9, 3, 2, 7, 15, 10, 14, 8 b. 15, 6, 17, 9, 12, 5, 16, 7 c. 13, 9, 9, 16, 10, 9, 5, 7, 9, 8, 10, 9 d. 3.0, 2.8, 3.2, 3.6, 3.7, 3.2, 3.5, 3.2, 3.0, 3.2, 3.8, 3.3
197
4.5 The Variance (s2)
The range and interquartile range describe the variability of a distribution of data for a variable based on two scores at or near the ends of the distribution. However, other measures of variability are based on all of the scores in a set of data and do so based on the relationship between each score and the mean of all of the scores in the sample.
To illustrate measures of variability based on all of the scores in a set of data, let's return to the simple data set of nine scores used earlier to illustrate the range: 2,3, 1, 4, 3, 1, 4, 6, and 3. The mean ( X ¯ ) of this sample of nine scores is calculated below: X ¯ = ∑ X N = 2 + 3 + 1 + 4 + 3 + 1 + 4 + 6 + 3 9 = 27 9 = 3.00
Given the mean is designed to represent the center of a distribution, one way to measure the variability in a set of data is based on the extent to which each score differs from the mean. In this example, the variability in this set of data could be represented by the average difference between each score (X) and the mean of 3.00 X ¯ = 3 . 00 . Referring to this difference as a “deviation,” the third column in Table 4.2 calculates the deviation of each of the nine scores from the mean, symbolized by X − X ¯ .
Calculating the average of these deviations consists of summing the deviations and then dividing the summed deviations by the number of deviations, which is equal to N. The formula for the average deviation from the mean is provided below: Average deviation from the mean = ∑ X − X ¯ N
The numerator of this formula is “the sum of deviations from the mean.” Note that parentheses are placed around X – X − X ¯ to indicate that this deviation should be calculated for each score before adding the deviations together; if we did not include the parentheses, the notation ∑ X − X ¯ would imply that the first step would be to calculate the sum of the scores (SX) and then subtract the mean ( X ¯ ) from this sum. As in all mathematical calculations, the placement of parentheses indicates the order in which mathematical operations are performed.
Using the deviations calculated in Table 4.2, the average deviation from the mean for the set of nine scores is calculated below: Average deviation = ∑ ( X − X ¯ ) N = − 1.00 + .00 + … 3.00 + .00 9 = 0 9 .00
Here, the average deviation from the mean is equal to zero (0). From this, we would conclude there is zero variability among the nine scores. However, concluding there is “zero” variability among the scores implies that all nine scores are exactly the same, when in fact we know this is not the case. How did we reach the erroneous conclusion that all the scores are identical?
198
The average deviation from the mean is always equal to zero. This is because, as we explained in Chapter 3, the sum of the positive deviations from the mean is always equal to the sum of the negative deviations from the mean, which in turn leads the sum of the deviations to be equal to zero. Because the sum of the deviations is always equal to zero, the average deviation from the mean is also always equal to zero, regardless of the actual amount of variability among the scores in a set of data.
199
Definitional Formula for the Variance
Calculating a measure of variability based on the deviations from the mean is complicated by the fact that the sum of the negative and positive deviations will always be equal to each other. One way to overcome this “balancing act” is to eliminate the negative deviations by squaring each deviation: X − X ¯ 2 ). The square of any number, negative or positive, is always a positive number. So, rather than calculating the average deviation from the mean, the amount of variability in a sample of data may be measured by calculating the variance (s2), defined as the average squared deviation from the mean. Formula 4-3 provides what is known as the definitional formula for the variance:
(4-3) s 2 = ∑ X − X ¯ 2 N − 1
where X is a score for the variable, X ¯ is the mean of the sample, and N is the total number of scores.
Two important aspects of the formula for the variance should be considered here. First, in the numerator, parentheses are placed around the deviation between each score and the mean X − X ¯ ) to separate the deviation from the squaring and the summing. Second, in the denominator, the sum of the squared deviations is not divided by the total number of scores (N) but instead is divided by the total number of scores minus 1 (N – 1). The rationale for dividing by N – 1 rather than by N will be discussed after we illustrate the calculations for the variance.
Table 4.3 begins the calculation of the variance for the set of nine scores. The first three columns of this table are identical to the same columns in Table 4.2—the fourth column squares each of the deviations. For the first score, for example, the squared deviation is equal to (–1.00)2 or 1.00. Note that all of the squared deviations are positive numbers.
Using the squared deviations calculated in Table 4.3, the variance for the set of nine scores can be calculated using Formula 4-3: s 2 = ∑ ( X − X ¯ ) 2 N − 1 = 1.00 + .00 + 4.00 + … + 1.00 + 9.00 + .00 9 − 1 = 20.00 8 = 2.50
Table 4.2 Calculation of the Deviation of Each Score from the Mean ( X − X ¯ ) , Simple Example
Table 4.2 Calculation of the Deviation of Each Score from the Mean ( X − X ¯ ) , Simple
Example
Score (X) Mean X ¯ Score – Mean X − X ¯
200
2 3.00 –1.00
3 3.00 .00
1 3.00 –2.00
4 3.00 1.00
3 3.00 .00
1 3.00 –2.00
4 3.00 1.00
6 3.00 3.00
3 3.00 .00
Table 4.3 Calculation of Squared Deviation from the Mean X − X ¯ 2 , Simple example
Table 4.3 Calculation of Squared Deviation from the Mean X − X ¯ 2 , Simple example
Score (X) Mean X ¯ Score – Mean X − X ¯ (Score – Mean)2 X − X ¯ 2
2 3.00 –1.00 1.00
3 3.00 .00 .00
1 3.00 –2.00 4.00
4 3.00 1.00 1.00
3 3.00 .00 .00
1 3.00 –2.00 4.00
4 3.00 1.00 1.00
6 3.00 3.00 9.00
3 3.00 .00 .00
From this calculation, we may conclude that the variance in this set of data is equal to 2.50. This means that, for the sample of nine scores, the average squared deviation of a score from the mean is 2.50.
In calculating the variance, because a squared deviation is always a positive number, the sum of the squared deviations ∑ X − X ¯ 2 ) and subsequently the variance (s2) must also always be positive numbers. This is an important point that helps students identify calculation errors in homework and paper assignments: Obtaining a negative value for the sum of squared deviations or the variance means you have made a mistake in your calculations.
As a second example, let's calculate the variance for the frequency rating variable. Using the sample mean ( X ¯ ) of 33.30 calculated earlier this chapter, the last column in Table 4.4
201
shows the squared deviation for each of the 20 ratings in this sample.
Using the squared deviations, the variance for the frequency rating variable is calculated using Formula 4-3 as follows: s 2 = ∑ ( X − X ¯ ) 2 N − 1 = 161.29 + 334.89 + … 1069.29 + 53.29 20 − 1 = 3940.20 19 = 207.38
For this sample of 20 participants, the variance, which is to say the average squared deviation of a frequency rating from the mean of 33.30, is equal to 207.38.
202
Computational Formula for the Variance
The formula for the variance presented in Formula 4-3 represents the literal definition of variance: the average squared deviation from the mean. For this reason, it is referred to as a definitional formula, which is a formula based on the actual or literal definition of a concept. However, using a definitional formula to analyze a set of data can be tedious (because it requires calculating the deviation of each score from the mean) and complicated (because the mean often possesses decimal places [e.g., 33.30]). Because using the definitional formula in a large or complicated data set increases the chances of making computational errors, the sum of squared deviations can be algebraically manipulated to create what is known as a computational formula, defined as a formula not based on the definition of a concept but is designed to simplify mathematical calculations.
Formula 4-4 provides the computational formula for the variance:
(4-4) s 2 = ∑ X 2 − ∑ X N 2 N − 1
where ΣX2 is the sum of squared scores, (ΣX)2 is the sum of scores squared, and N is the total number of scores. Note that the sum of scores squared ((ΣX)2) requires squaring each score and then summing the squared scores (first squaring, then summing), whereas the sum of scores squared ((X)2) involves summing a set of scores and then squaring this sum (first summing, then squaring).
Table 4.5 begins the process of calculating the variance for the set of nine scores used earlier using the computational formula. The bottom of the first column contains the sum of the nine scores (ΣX = 27); the sum of scores squared ((ΣX)2) is equal to (27)2 or 729. The second column of Table 4.5 provides the squared values of the scores—the sum of squared scores is located at the bottom of this column (ΣX2 = 101). The results of these calculations are then inserted into Formula 4-4 as follows:
Table 4.4 Calculation of Squared Deviation from the Mean X − X ¯ 2 , Frequency Rating Variable
Table 4.4 Calculation of Squared Deviation from the Mean X − X ¯ 2 , Frequency Rating Variable
Frequency Rating (X)
Mean ( X ¯ )
Frequency Rating – Mean X − X ¯
(Frequency Rating – Mean)2 X − X ¯ 2 )
46 33.30 12.70 161.29
15 33.30 –18.30 334.89
21 33.30 –12.30 151.29
203
49 33.30 15.70 246.49
23 33.30 –10.30 106.09
39 33.30 5.70 32.49
25 33.30 –8.30 68.89
27 33.30 –6.30 39.69
20 33.30 –13.30 176.89
23 33.30 –10.30 106.09
58 33.30 24.70 610.09
24 33.30 –9.30 86.49
36 33.30 2.70 7.29
20 33.30 –13.30 176.89
49 33.30 15.70 246.49
32 33.30 –1.30 1.69
45 33.30 11.70 136.89
22 33.30 –11.30 127.69
66 33.30 32.70 1069.29
26 33.30 –7.30 53.29
Table 4.5 Calculating the Sum of Squared Scores (ΣX2) and Sum of Scores Squared ((ΣX)2), Simple Example Table 4.5 Calculating the
Sum of Squared Scores (ΣX2) and Sum of Scores Squared ((ΣX)2), Simple Example
Score (X) (Score)2 (X2)
2 4
3 9
1 1
4 16
3 9
1 1
4 16
6 36
3 9
204
ΣX = 27 ΣX2 = 101 (ΣX)2 = 729
s 2 = ∑ X 2 − ( ∑ X ) N 2 N − 1 = 101 − 729 9 9 − 1 = 101 − 81.00 8 = 20.00 8 = 2.50
The same value for the variance (s2 = 2.50) was obtained whether we use the computational formula in Formula 4-4 or the definitional formula in Formula 4-3.
To calculate the variance for the frequency rating variable using the computational formula, Table 4.6 calculates the sum of squared scores and the sum of scores squared for the 20 participants. Once these two quantities have been calculated, the variance may be calculated as follows: s 2 = ∑ X 2 − ∑ X 2 N N − 1
Table 4.6 Sum of Squared Scores (ΣX2) and Sum of Scores Squared ((ΣX)2), Frequency Rating Variable
Table 4.6 Sum of Squared Scores (ΣX2) and Sum of Scores Squared ((ΣX)2), Frequency Rating Variable Frequency Rating (X) (Frequency Rating)2 (X2)
46 2116
15 225
21 441
49 2401
23 529
39 1521
25 625
27 729
20 400
23 529
58 3364
24 576
36 1296
20 400
49 2401
32 1024
45 2025
205
22 484
66 4356
26 676
ΣX = 666 ΣX2 = 26,118 (ΣX)2 = 443,556
= 26118 − 443556 20 20 − 1 = 26118 − 22177.80 19 = 3940.20 19 = 207.38
The value of 207.38 for the variance is again identical using either the definitional or the computational formula.
206
Why Not Use the Absolute Value of the Deviation in Calculating the Variance?
To compute the variance using the definitional formula provided in Formula 4-3, the deviation between each score and the mean must be squared to eliminate negative deviations. Instead of doing all of this squaring, you may be asking yourself, wouldn't it be easier to eliminate the negative deviations simply by using the absolute value of each deviation? If so, one could simply calculate the average of the absolute values. The absolute value of a number, symbolized by parallel vertical lines, ignores the sign (+/–) of the number. For example, for the first score in Table 4.2, the absolute value of the deviation |2 – 3.00| is equal to 1.00.
The logic behind using absolute values of the deviations may be intuitively appealing. However, more advanced statistical procedures pertaining to variability require algebraic manipulations that cannot be carried out using absolute values. For this reason, it is necessary to use the formulas for the variance that are based on the squaring of the deviations.
207
Why Divide by N – 1 Rather than N in Calculating the Variance?
In calculating the variance, students often wonder why the numerator is divided by N – 1, rather than by N. One reason for dividing by N – 1 is to estimate the variability in a population using data collected from a sample of the population. To illustrate this concept, consider an example in which data are collected from a sample of 100 first-year students at a local university in order to represent the entire population of first-year university students nationwide. It is reasonable to suspect that the smaller sample of local university students would be less diverse than the entire population. That is, the sample may not accurately represent all of the possible values for ethnicity, economic background, age, attitudes, and so on that exist in the entire population. Consequently, the amount of variability in the sample will be less than what is believed to exist in the population.
Because the amount of variability in a sample is less than the variability in the population from which the sample is drawn, the sample variance underestimates the variance in the population. As a result, the sample variance is a biased estimate of the population variance; a biased estimate is a statistic based on a sample that systematically underestimates or overestimates the population from which the sample was drawn. What is needed, therefore, is to correct the sample variance to make it an unbiased estimate of the population variance, which is a statistic based on a sample that is equally likely to underestimate or overestimate the population from which the sample was drawn.
Given that the sample variance systematically underestimates the population variance, we need to increase the value of the sample variance to more accurately estimate the variability in the population. This sample variance is increased by dividing the numerator, the sum of squared deviations, by N – 1 rather than by N. Dividing by a smaller number makes the result larger.
208
4.6 The Standard Deviation (s)
The variance is the average squared deviation of a score from the mean. However, researchers typically want to represent the variability in a set of data not in terms of the average squared deviation but rather simply the average deviation. This is accomplished by calculating a measure of variability known as the standard deviation (s), defined as the square root of the variance. Mathematically, the standard deviation is the square root of the average squared deviation from the mean, and it represents the average deviation of a score from the mean.
209
Definitional Formula for the Standard Deviation
The standard deviation, represented by the symbol s, is calculated by computing the square root of the variance. The purpose of calculating the square root of the variance is to “undo” the effect of squaring the deviations. Formula 4–5 provides the definitional formula for the standard deviation—this formula places the definitional formula for the variance (Formula 4-3) under a square root symbol:
210
Learning Check 2: Reviewing what you've Learned So Far
1. Review questions a. Why does calculating the variance involve squaring the deviation of each score from the
mean? b. What is the purpose of computational formulas? c. In calculating a measure of variability, why can't you use the absolute value of the deviation
of each score from the mean rather than the squared deviation? d. In calculating the variance, why is the sum of squared deviations divided by N – 1 rather
than N? e. What is the difference between a biased estimate and an unbiased estimate?
2. Calculate the variance (s2) using the definitional and computational formulas for each of the following data sets.
a. 2, 3, 3, 5, 7 b. 5, 4, 7, 5, 10, 5, 6 c. 10, 13, 13, 14, 15, 16, 17, 24 d. 73, 66, 91, 84, 69, 87, 62, 79, 82, 90 e. 11, 9, 9, 12, 10, 9, 7, 8, 9, 8, 10, 9
(4-5) s = ∑ X − X ¯ 2 N − 1
For the simple data set of nine scores used throughout this chapter, using the squared deviations calculated in Table 4.3, the standard deviation may be calculated as follows: s = ∑ ( X - X ¯ ) 2 N − 1 = 1.00 + .00 + 4.00 + … + 1.00 + 9.00 + .00 9 − 1 = 20 8 = 2.50 = 1.58
From this value of the standard deviation, we conclude that the average deviation of the nine scores in this sample from the mean of 3.00 is equal to 1.58.
For the frequency rating variable, using the calculations in Table 4.4, the standard deviation is equal to the following: s = ∑ ( X - X ¯ ) 2 N − 1 = 161.29 + 334.89 + 151.29 + … + 127.69 + 1069.29 + 53.29 20 − 1 = 3940.29 19 = 207.38 = 14.40
Here, we can conclude that, in this sample of 20 participants, in rating the frequency of the word sometimes, the average difference between a participants frequency rating and the sample mean of 33.30% was 14.40%.
Chapter 3 mentioned that descriptive statistics are often provided within the body or text of a paper. For the frequency rating example,
211
The average frequency rating for the word sometimes was approximately one-third the distance between Never and Always, representing a 33% frequency (M = 33.30, SD = 14.40).
In this example, the symbols M and SD represent the mean and standard deviation, respectively. Both measures of central tendency and variability are reported because they provide different pieces of information about the nature and shape of the distribution of scores for a variable.
212
Computational Formula for the Standard Deviation
The computational formula for the standard deviation (Formula 4–6) simply places the computational formula for the variance (Formula 4-4) within the square root symbol:
(4-6) s = ∑ X 2 − ∑ X 2 N N − 1
The standard deviation for the simple example of nine scores and the frequency rating variable are calculated using Formula 4–6 in Table 4.7 (the values for the sum of scores squared [(ΣX)2] and the sum of squared scores [ΣX2] for the two examples are found in Tables 4.5 and 4.6). Similar to the variance, note the same value for the standard deviation is obtained regardless of whether the definitional or computational formula is used.
Table 4.7 Calculation of the Standard Deviation Using the Computational Formula, Simple Example and Frequency Rating Variable
Table 4.7 Calculation of the Standard Deviation Using the Computational Formula, Simple Example and Frequency Rating Variable
Simple Example Frequency Rating
s = ∑ X 2 − ( ∑ X ) 2 N N − 1 = 101 − 729 9 9 − 1 = 101 − 81.00 8 8 = 20.00 8 = 2.50 = 1.58
s = ∑ X 2 − ( ∑ X ) 2 N N − 1 = 26118 − 443556 20 20 − 1 = 26118 − 22177.80 19 19 = 3940.20 19 = 207.38 = 14.40
213
Learning Check 3: Reviewing what you've Learned So Far
1. Review questions a. What is the relationship between the variance and the standard deviation? b. Is it possible to get a negative value for the variance or standard deviation? Why or why not?
2. Calculate the standard deviation (s) using the definitional and computational formula for each of the following data sets.
a. 5, 3, 9, 2, 6 b. 11, 19, 8, 10, 9, 7, 13 c. 7, 9, 11, 12, 13, 14, 17, 21, 23 d. 10, 8, 12, 8, 9, 7, 23, 6, 8, 11, 10, 6 e. 11, 9, 13, 6, 14, 12, 7, 12, 10, 8, 15, 5, 12
214
4.7 Measures of Variability for Populations
The formulas for the variance and standard deviation described thus far are used for samples drawn from populations. But what if you were to collect data from the entire population? This section discusses two measures of variability for populations: the population variance and the population standard deviation.
215
The Population Variance (σ2) The variance for data collected from a population is called the population variance (σ2) (σ is the lowercase Greek letter sigma), which is defined as the average squared deviation of a score from the population mean. The definitional formula for the population variance is provided in Formula 4–7:
(4-7) σ 2 = ∑ X − μ 2 N
where X is a score for a variable, μ is the population mean, and N is the total number of scores.
The formula for the population variance differs from the formula for the variance s2
(Formula 4-4) in two important ways. First, the mean in the formula is the population mean μ rather than the sample mean X ¯ . Second, the sum of the squared deviations (Σ(X – μ)2) is divided by N rather than N – 1; this is because you are no longer estimating the variability in the population from data collected from a sample but instead have collected data from the entire population.
Imagine, for a moment, that the nine scores in the first column of Table 4.3 represent a population rather than a sample. Using the squared deviations for these scores calculated in the last column of Table 4.3, the population variance would be equal to the following: σ 2 = ∑ ( X − μ ) 2 N = 1.00 + .00 + 4.00 + … + 1.00 + 9.00 + .00 9 = 20.00 9 = 2.22
From this calculation, you would conclude that, assuming these nine scores represent the entire population, the average squared deviation of a score from the population mean of 3.00 is 2.22.
216
The Population Standard Deviation (σ) Calculating the population variance involves squaring the deviation of each score from the population mean μ. In order to obtain a measure of variability that represents the average deviation (rather than the average squared deviation) from the population mean, the square root of the population variance may be calculated. This is the population standard deviation (σ), defined as the square root of the population variance. Mathematically, the population standard deviation is the square root of the average squared deviation from the population mean, and it represents the average deviation of a score from the population mean. Formula 4–8 provides the definitional formula for the population standard deviation.
(4-8) σ = ∑ X − μ 2 N
Relying again on Table 4.3, the population standard deviation for the set of nine scores is calculated below. σ = ∑ ( X - μ ) 2 N = 1.00 + .00 + 4.00 + … + 1.00 + 9.00 + .00 9 = 20 9 = 2.22 = 1.49
From these calculations, it may be concluded that the average deviation of the nine scores from the population mean of 3.00 is equal to 1.49.
217
Measures of Variability for Samples vs. Populations
The population variance and standard deviation are examples of parameters, which were introduced in Chapter 3. Parameters are numeric characteristics of populations and are distinguished from statistics, which are numeric characteristics of samples. Parameters are calculated when data are collected from the entire population, whereas statistics are calculated when data are collected from a sample drawn from a larger population. This is summarized below for both a measure of central tendency (the mean) and a measure of variability (the standard deviation):
Target Numeric Characteristic Mean Standard Deviation
Sample statistic X ¯ s
Population parameter μ σ
Because researchers rarely collect data from entire populations, statistics such as the sample mean and standard deviation are much more likely to be calculated than population parameters. Throughout this book, when you see the words mean, variance, or standard deviation, you should assume they refer to samples rather than populations. In fact, numeric values of parameters such as μ or σ are typically values believed or hypothesized to be true rather than values based on the actual collection of data.
218
Learning Check 4: Reviewing what you've Learned So Far
1. Review questions a. What is the main difference between the population standard deviation and the standard
deviation? b. Under what conditions would you calculate the population standard deviation for a set of
data rather than the standard deviation? 2. Calculate the population standard deviation (σ) using the definitional and computational formula
for each of the following data sets: a. 2, 4, 5, 6, 8 b. 8, 6, 4, 8, 3, 7 c. 71, 84, 65, 78, 89, 72, 60, 85 d. 3.0, 1.5, 3.5, 3.0, 2.0, 2.5, 4.0, 1.0, 2.0 e. 10, 8, 12, 8, 9, 7, 23, 6, 8, 11, 10, 6
219
4.8 Measures of Variability: Drawing Conclusions
These early chapters of this book have discussed the importance of examining and drawing appropriate conclusions about three aspects of distributions: modality, symmetry, and variability. Measures of central tendency such as the mean, median, and mode provide a numerical description of the modality and, to a lesser degree, the symmetry of a variable by focusing on the center of the distribution and what scores have in common with each other. However, it is equally important to consider the variability in a distribution, which is the degree to which scores differ from the center of the distribution and from each other. Measures of variability such as the variance and standard deviation provide valuable information, for example, regarding the degree to which participants in a sample agree when asked the same survey question or respond in a similar manner to the same experimental manipulation.
The beginning of this chapter introduced a research study designed to measure the degree to which people differ in their interpretation of words such as sometimes, often, and seldom. Based on the variability of the frequency ratings, what did these researchers conclude regarding the degree to which people agree in their perceptions of these words? Comparing the results of their study with those conducted in other countries, they concluded the following:
There are subtle variations in how those sharing the same language describe the intermediate points … so it cannot be assumed that rating scales developed in one culture can be automatically used in another, even where there is a common language. (Skevington & Tucker, 1999, p. 59)
220
4.9 Looking Ahead
Several critical issues have been identified and discussed in the first four chapters of the book. The first is the importance of examining data before conducting statistical analyses on the data. Appropriate and meaningful conclusions cannot be drawn from statistical analyses until the accuracy of the data has been confirmed and the distribution of scores has been understood. Second, there are many different types of distributions; you have seen distributions labeled as symmetric, skewed, peaked, flat, unimodal, and bimodal. Third, distributions can be described numerically using measures of central tendency and variability. The next chapter discusses yet another type of distribution: “normal distributions.” Like the other distributions discussed so far, normal distributions can be described using descriptive statistics such as measures of central tendency and variability. What makes normal distributions unique, as we will discuss in greater detail going forward, is that they possess characteristics that enable them to be the basis of a second type of statistics, known as “inferential statistics,” whose goal is to test hypotheses about populations based on the information from samples.
221
4.10 Summary
Variability is a third aspect of distributions and refers to the amount of spread or scatter of scores in a distribution. One goal of a science such as psychology is to describe, understand, explain, and predict variability in the phenomena studied by researchers.
A measure of variability is a descriptive statistic of the amount of differences in a set of data for a variable. Common measures of variability are the range, the interquartile range, the variance, and the standard deviation.
The range is the mathematical difference between the lowest and highest scores in a set of data. The interquartile range is the range of the middle 50% of the scores for a variable, calculated by removing the highest and lowest 25% of the distribution. The variance (s2) is the average squared deviation of a score from the mean. The standard deviation (s) is the square root of the variance, and it represents the average deviation of a score from the mean. Because the variance and standard deviation (unlike the range and interquartile range) are based on all of the scores in the distribution, they can be used in statistical analyses designed to test hypotheses about distributions.
The variance and the standard deviation may be calculated either using a definitional formula, which is a formula based on the actual or literal definition of a concept, or a computational formula, defined as a formula not based on the definition of a concept but is designed to simplify mathematical calculations. Computational formulas are used because using the definitional formula in a large or complicated data set increases the chances of making computational errors.
The variance and standard deviation measure the variability in data collected from samples drawn from populations; the population variance (σ2) (the average squared deviation of scores from the population mean) and the population standard deviation (σ) (the square root of the population variance) are calculated when data are collected from the entire population. Because data are rarely collected from entire populations, numeric values of parameters such as σ are typically based on beliefs and hypotheses rather than the actual collection of data.
222
4.11 Important Terms
measure of variability (p. 107) range (p. 108) interquartile range (p. 110) variance (s2) (p. 114) definitional formula (p. 116) computational formula (p. 116) biased estimate (p. 120) unbiased estimate (p. 120) standard deviation (s) (p. 120) population variance (σ2) (p. 124) population standard deviation (σ) (p. 124)
223
4.12 Formulas Introduced in this Chapter
224
Range
(4-1) Range = highest score − lowest score
Interquartile Range
(4-2) Interquartile range = N − N 4 th score − N 4 + 1 th score
Variance (s2) (Definitional Formula)
(4-3) s 2 = ∑ X − X ¯ 2 N − 1
Variance (s2) (Computational Formula)
(4-4) s 2 = ∑ X 2 − ∑ X 2 N N − 1
Standard Deviation (s) (Definitional Formula)
(4-5) s = ∑ X − X ¯ 2 N − 1
Standard Deviation (s) (Computational Formula)
(4-6) s = ∑ X 2 − ∑ X 2 N N − 1
Population Variance (σ2) (Definitional Formula) (4-7) σ 2 = ∑ X − μ 2 N
Population Standard Deviation (σ) (Definitional Formula) (4-8) σ = ∑ X − μ 2 N
225
4.13 Using SPSS
226
Calculating Measures of Central Tendency and Variability: The Frequency Rating Study (4.1)
1. Define variable (name, # decimal places, label for the variable) and enter data for the variable.
2. Select the descriptive statistics procedure within SPSS.
How? (1) Click Analyze menu, (2) click Descriptive Statistics, and (3) click Descriptives.
3. Select the variable to be analyzed.
How? (1) Click variable and , (2) click .
227
4. Examine output.
228
4.14 Exercises
1. Calculate the range for each of the following sets of data: a. 4, 6, 9 b. 14, 17, 11, 19, 12 c. 25, 22, 27, 30, 21, 26, 29 d. 10, 8, 4, 16, 9, 7, 9, 13, 6, 11 e. 73, 66, 91, 84, 69, 87, 62, 79, 82, 90 f. 3.50, 4.21, 3.95, 2.27, 3.06, 4.58, 2.74, 3.89, 2.65, 2.03, 4.41, 3.76, 2.35
2. Calculate the range for each of the following sets of data: a. 9, 10, 13 b. 5, 3, 9, 12 c. 6, 2, 8, 12, 9, 5, 7 d. 8, 4, 1, 6, 14, 9, 12, 5, 11, 7, 4 e. 16.65, 12.98, 31.74, 18.80, 27.31, 29.92, 34.65, 23.68, 28.20, 20.77
3. Calculate the interquartile range for each of the following sets of data: a. 3, 6, 7, 12, 15, 17, 23, 28 b. 8, 4, 1, 6, 13, 10, 12, 5 c. 3, 8, 14, 11, 16, 7, 14, 15, 11, 9, 12, 6 d. 15, 22, 17, 13, 31, 25, 22, 19, 26, 30, 27, 19, 23, 21, 27, 29 e. 450, 560, 340, 510, 390, 670, 540, 420, 720, 480, 560, 510 f. 29, 27, 26, 15, 28, 26, 32, 27, 26, 25, 30, 26, 27, 18, 23, 35
4. Calculate the interquartile range for each of the following sets of data: a. 1, 2, 2, 3, 4, 5, 7, 10 b. 21, 8, 17, 7, 12, 19, 5, 12 c. 6, 9, 5, 14, 3, 15, 19, 7, 13, 6, 8, 5 d. 6.78, 7.81, 6.35, 9.65, 5.43, 8.62, 5.90, 7.13, 6.56, 3.27, 8.92, 4.49 e. 78, 55, 82, 64, 93, 69, 71, 82, 59, 71, 76, 52, 89, 75, 81, 78
5. For each of the following sets of data, (1) calculate the mean of the scores ( X ¯ ), (2) calculate the deviation of each score from the mean X − X ¯ , and (3) check to see if the sum of the deviations equals zero ∑ X − X ¯ = 0 .
a. 2, 4, 6 b. 5, 6, 13 c. 4, 7, 8, 9 d. 5, 2, 7, 13, 11 e. 3, 8, 14, 11, 16, 7, 14, 15, 11
6. For each of the following sets of data, (1) calculate the mean of the scores ( X ¯ ), (2) calculate the deviation of each score from the mean X − X ¯ , and (3) check to see if the sum of the deviations equals zero ∑ X − X ¯ = 0 .
a. 7, 5, 12 b. 6, 7, 9, 12
229
c. 6, 8, 4, 11, 3, 7 d. 5, 7, 3, 1, 4, 8, 5, 2, 7, 9, 12, 4, 15
7. For each of the sets of data in Exercise 5, calculate the variance (s2) using the definitional formula and the computational formula.
8. For each of the sets of data in Exercise 5, calculate the population variance (σ2) using the definitional formula. Comparing your calculations for each data set with those done in Exercise 7, which is larger, the variance or the population variance? Why?
9. For each of the sets of data in Exercise 5, calculate the standard deviation (s) and the population standard deviation (σ). For each data set, which is larger, the standard deviation or the population standard deviation?
10. (This example was discussed in Chapters 2 and 3.) A friend of yours asks 20 people to rate a movie using a 1- to 5-star rating: the higher the number of stars, the higher the recommendation. Their ratings are listed below:
Person # Stars
1 ***
2 *****
3 **
4 ****
5 ***
6 ***
7 ****
8 **
9 ***
10 *****
11 *
12 ****
13 ****
14 ***
15 ***
16 **
17 ***
18 **
19 ****
20 ***
230
a. Calculate the variance (s2) and standard deviation (s) of the ratings using either the definitional or the computational formulas. (Note: From Chapter 3, you may have already calculated the number of scores (N), sum of scores squared ((ΣX)2), the sum of squared scores (ΣX2), and the mean X ¯ .)
b. Based on your value of the standard deviation, what would you conclude regarding the degree to which these people agree or disagree about this movie?
11. One study stated the research hypothesis, “Violence behavior in children may be reduced by teaching them conflict resolution skills” (DuRant et al., 1996). The variable “violence behavior” was measured by the number of fights in which each student was involved. Below is a frequency distribution table for the number of fights for 12 students.
# Fights f %
4 1 8%
3 2 17%
2 5 42%
1 1 8%
0 3 25%
Total 12 100%
a. Calculate the variance (s2) and standard deviation (s) of the number of fights using either the definitional or the computational formulas.
b. Based on your value of the standard deviation, what is the average difference between the number of fights a student got involved in and the mean?
12. (This example was introduced in Chapter 3). The 10% myth study discussed in this chapter measured the beliefs of both psychology majors and non–psychology majors. The 39 non– psychology majors in this study provided the following values for the brain power variable:
Non–Psychology Major Estimated Brain Power
1 15
2 45
3 10
4 40
5 5
6 10
7 50
8 45
231
9 10
10 10
11 35
12 5
13 15
14 10
15 20
16 50
17 25
18 45
19 40
20 25
21 10
22 5
23 15
24 5
25 40
26 30
27 15
28 25
29 10
30 20
31 10
32 5
33 15
34 10
35 5
36 60
37 10
38 5
39 15
a. Calculate the variance (s2) and standard deviation (s) of the estimates using
232
either the definitional or the computational formulas. (Note: From Chapter 3, you may have already calculated the number of scores (N), sum of scores squared ((ΣX)2), the sum of squared scores (ΣX2), and the mean ( X ¯ ).)
13. (This example was introduced in Chapter 2.) An instructor administers a 27-item quiz to her class of 25 students. Each student's score on the quiz is the number of items answered correctly. These scores are listed below:
Person Quiz Score
1 22
2 15
3 11
4 19
5 12
6 21
7 22
8 19
9 23
10 17
11 18
12 14
13 10
14 6
15 16
16 20
17 21
18 19
19 20
20 21
21 17
22 18
23 15
24 24
25 22
233
a. Calculate the variance and standard deviation of the ratings using the computational formulas.
14. On your own, create two distributions of 10 scores that have the same mean but differ in modality: one unimodal and one bimodal. Calculate the standard deviation off the two distributions. Is there greater variability in the unimodal distribution or the bimodal distribution?
15. On your own, (a) create two distributions of 10 scores that have the same mean but differ in their amount of variability, and (b) create two distributions of 10 scores that have different means but have the same amount of variability. (c) What does this imply about the information that is necessary to accurately describe a distribution of scores for a variable?
16. What if someone told you, “There is very little variability in the scores for my variable.” From this statement, would you be more likely to describe the shape of the distribution as peaked or flat?
17. What if someone told you, “There is a great deal of variability in the scores for my variable.” From this statement, can you tell whether the shape of the distribution is symmetrical or skewed? Why or why not? If not, what would you do to determine the symmetry of the distribution?
18. What if someone told you, “The mean age of the participants in this sample was 35.50 and the standard deviation was 2.53.” Why would you interpret the sample differently if you had been told the standard deviation was 14.71?
19. An employment survey in 2009 was conducted to find the average starting salaries of people receiving doctorate (PhD) degrees in psychology (Michalski, Kohout, Wicherski, & Hart, 2011). The mean starting salary for psychologists in assistant professor positions in universities was approximately $60,000 with a standard deviation of $11,000. Clinical psychologists in their first year of practice also report a mean of approximately $60,000 but with a standard deviation of $16,000.
a. Which type of psychologist has more variability in income? b. For which type of psychologist could you more accurately predict a salary for,
and why?
234
Answers to Learning Checks
Learning Check
2. a. Range = 6 b. Range = 12 c. Range = 19 d. Range = .90
3. a. Interquartile range = 3 b. Interquartile range = 8 c. Interquartile range = 1 d. Interquartile range = .30
Learning Check 2
2. a. s2 = 4.00 b. s2 = 4.00 c. s2 = 17.07 d. s2 = 105.79 e. s2 = 1.84
Learning Check 3
2. a. s = 2.74 b. s = 4.04 c. s = 5.33 d. s = 4.55 e. s = 3.12
Learning Check 4
2. a. σ = 2.00 b. σ= 1.91 c. σ= 9.58
235
d. σ= .91 e. σ= 4.36
236
Answers to Odd-Numbered Exercises
1. a. Range = 5 b. Range = 8 c. Range = 9 d. Range = 12 e. Range = 29 f. Range = 2.55
3. a. Interquartile range = 10 b. Interquartile range = 5 c. Interquartile range = 6 d. Interquartile range = 8 e. Interquartile range = 110 f. Interquartile range = 2
5. a.
Score (X) Mean ( X ¯ ) Mean X − X ¯
2 4.00 –2.00
4 4.00 .00
6 4.00 2.00
∑ X − X ¯ = 0
b.
Score (X) Mean ( X ¯ ) Mean X − X ¯
5 8.00 –3.00
6 8.00 –2.00
13 8.00 5.00
∑ X − X ¯ = 0
c.
Score (X) Mean ( X ¯ ) Score - Mean X − X ¯
4 7.00 –3.00
7 7.00 .00
237
8 7.00 1.00
9 7.00 2.00
∑ X − X ¯ = 0
d.
Score (X) Mean ( X ¯ ) Score - Mean X − X ¯
5 7.60 –2.60
2 7.60 –5.60
7 7.60 –.60
13 7.60 5.40
11 7.60 3.40
∑ X − X ¯ = 0
e.
Score (X) Mean ( X ¯ ) Score - Mean X − X ¯
3 11.00 –8.00
8 11.00 –3.00
14 11.00 3.00
11 11.00 .00
16 11.00 5.00
7 11.00 –4.00
14 11.00 3.00
15 11.00 4.00
11 11.00 .00
∑ X − X ¯ = 0
7. a. s2 = 4.00 b. s2 = 19.00 c. s2 = 4.67 d. s2 = 19.80 e. s2 = 18.50
9.
238
a. s = 2.00, σ = 1.63 b. s = 4.36, σ = 3.56 c. s = 2.16, σ = 1.87 d. s = 4.45, σ = 3.98 e. s = 4.30, σ = 4.06
Note: In all data sets, the standard deviation (s) is larger than the population standard deviation (σ) because you are dividing by a smaller number (N – 1 rather than N). 11.
a. s2 = 1.66, s = 1.29 b. The average difference between the number of fights of a student and the mean
is 1.29 fights. 13.
a. s2 = 19.89, s = 4.46 15.
a. Sample Data Set 1: 4, 4, 5, 5, 5, 5, 5, 6, 6, 6 X ¯ = 5.10 , s = .74
Sample Data Set 1: 2, 2, 3, 4, 5, 5, 6, 6, 9, 9 X ¯ = 5.10 , s = 2.51
b. Sample Data Set 3: 3, 3, 4, 4, 5, 5, 5, 6, 7, 7 X ¯ = 4.90 , s = 1.45
Sample Data Set 4: 8, 8, 9, 9, 10, 10, 10, 11, 12, 12
X ¯ = 9.90 , s = 1.45
c. Distributions with the same mean or the same standard deviation can be describing very different information; therefore, it is important to provide both measures of central tendency and variability to describe a set of data.
17. There can be a great deal of variability in both symmetrical and skewed distributions. The simplest way to determine the symmetry of a distribution is to look at a figure of the distribution. 19.
a. Clinical psychologists have more variability in their starting salaries. b. You could more accurately predict the income for a psychologist going into an
assistant professor position because there is less variability in their starting salaries.
239
Sharpen your skills with SAGE edge at edge.sagepub.com/tokunaga
SAGE edge for students provides a personalized approach to help you accomplish your coursework goals in an easy-to-use learning environment. Log on to access:
eFlashcards Web Quizzes Chapter Outlines Learning Objectives Media Links SPSS Data Files
240
Chapter 5 Normal Distributions
241
Chapter Outline 5.1 Example: SAT Scores 5.2 Normal Distributions
What are normal distributions? Why are normal distributions important? Is there only one normal distribution?
5.3 The Standard Normal Distribution What is the standard normal distribution? Interpreting z-scores: distance from the mean Interpreting z-scores: relative position within the distribution Finding the area between the mean and a z-score Finding the area less or greater than a z-score Finding the area between two z-scores
5.4 Applying z-Scores to Normal Distributions Transforming normal distributions into the standard normal distribution Understanding the z-score formula Example: transforming SAT scores into z-scores Interpreting z-scores: distance from the mean Interpreting z-scores: relative position within distribution Finding the area between the mean and a score Finding the area less or greater than a score Finding the area between two scores Transforming scores into z-scores: drawing conclusions
5.5 Standardizing Frequency Distributions Example: measuring politicians' behavior Standardizing a frequency distribution Interpreting standardized scores Using standardized scores to compare or combine variables Characteristics of standardized distributions When can you apply the normal curve table to a standardized distribution?
5.6 Looking Ahead 5.7 Summary 5.8 Important Terms 5.9 Formulas Introduced in This Chapter 5.10 Exercises
The first chapters of this book have introduced a process by which research may be conducted. This process starts by stating a research hypothesis to be tested; next, data relevant to the research hypothesis are collected. The first part of analyzing data is to examine them (using tables and figures) and describe them (using descriptive statistics such as measures of central tendency and variability). But given the goal of research is to determine whether or not a research hypothesis is supported, data must not only be examined and described but also evaluated. The preceding chapters have emphasized the importance of understanding distributions because they are the basis of statistical analyses designed to test research hypotheses. The purpose of this chapter is to discuss one type of these distributions, known as normal distributions. To illustrate the characteristics and uses
242
of normal distributions, we will use an example familiar to many college students: the SAT.
243
5.1 Example: Sat Scores
Each year, many thousands of high school seniors apply for college. One of the more anxiety-provoking aspects of this process may involve taking the SAT Reasoning Test, known by most people as the SAT. The SAT is a 3-hour, 45-minute examination divided into three sections (math, reading, and writing), each scored on a scale ranging from 200 to 800. Although educators have long disagreed regarding the appropriate role of the SAT and other standardized tests in the college admissions process, it continues to be used by many universities.
244
5.2 Normal Distributions
Imagine you have recently taken the SAT and received a score of 660 on the math section. Based on this score, how could you evaluate your performance on this section of the test? Your first step might be to compare your score with the scores of other students in the student population at large. This will require locating your score relative to the other scores within the distribution.
245
What are Normal Distributions?
Evaluating a score relative to the other scores in a population may require an understanding of a specific type of distribution known as a normal distribution, defined as a distribution based on a population of an infinite number of scores calculated from a mathematical formula. Note that a normal distribution differs from the types of distributions discussed earlier in this book (what we have referred to as frequency distributions) in that scores in a normal distribution are generated from a mathematical formula rather than data collected from research participants. An example of a normal distribution is illustrated in Figure 5.1.
There are four distinguishing features of normal distributions. First, because they consist of an infinite number of scores generated from mathematical formulas, the left and right ends of normal distributions continue to infinity (symbolized by the Greek symbol ∞) and do not touch the horizontal (X) axis. The ends, or tails as they are sometimes called, of normal distributions do not touch the X-axis because to do so would signify the end of the distribution (a score with a frequency [f] of zero). Second, in terms of their shape, normal distributions are unimodal and symmetric and are often referred to as “bell-shaped.” The symmetrical nature of normal distributions implies that if you were to fold one down the middle, the left and right halves of the distribution would exactly match one another. Third, in terms of central tendency, because normal distributions are based on a population of scores, the mean of a normal distribution is the population mean μ. Finally, in terms of variability, the standard deviation of a normal distribution is the population standard deviation σ.
246
Why are Normal Distributions Important?
Normal distributions differ from frequency distributions in that frequency distributions are constructed using data collected from samples, whereas normal distributions consist of an infinite number of scores generated by mathematical formulas. Because data collected by researchers often do not closely resemble the smooth, curved shape of a normal distribution, you may wonder why normal distributions are the basis of statistical techniques designed to analyze collected data. The answer to this question is threefold.
Figure 5.1 Example of a Normal Distribution
First, normal distributions are used because many variables (e.g., height or weight) are generally believed by researchers to be normally distributed in the population. Therefore, researchers apply the characteristics of normal distributions to data collected from samples, particularly when the sample is large. The discovery that such variables are not normally distributed in a particular sample is typically attributed to limitations in how the data were collected. For example, findings that a variable is not normally distributed commonly result from samples that are relatively small or not representative of the entire population.
A second reason for using normal distributions, related to the first, is that many statistical procedures designed to test research hypotheses are based mathematically on the assumption that variables are normally distributed. This assumption reinforces the importance of examining the distribution of data collected from samples to assess the degree to which it resembles a normal distribution.
Last, because normal distributions are based on mathematical formulas, it is possible to determine the proportion of the distribution associated with any score. For example, because a normal distribution is symmetric, exactly 50% of the distribution is located below the mean and 50% is located above the mean. As a second example, assuming SAT scores are normally distributed in the population, one can evaluate individual scores such as the score of 660 mentioned earlier, answering questions such as, “What percentage of scores on the math section of the SAT are less than 660?”
247
Is There Only One Normal Distribution?
The concept of normal distributions is a very important one to researchers. However, there is one problem associated with using these distributions: There is not just one normal distribution. The mathematical formula for the normal distribution involves stating a value for the population mean (μ) and standard deviation (σ). As such, changing either of these values results in a change in the shape of the distribution.
Consider, for example, the variables height and weight. It is assumed both of these variables are normally distributed in the adult population. The shapes of the two distributions are not the same, however, because they are measured using different units of measurement (inches vs. pounds). As a result, how a score is evaluated depends on how the variable is measured. For example, a score of 80 is interpreted very differently for these two variables. For height, 80 inches (6′ 8″) is near the upper end of the distribution for adults; for weight, 80 pounds is at the lower end.
As further illustration of different types of normal distributions, Figure 5.2(a) displays two normal distributions that have different values for the population mean μ but the same standard deviation σ. Figure 5.2(b) displays two normal distributions that have the same value for σ but a different value for σ. Because there are an infinite number of possible normal distributions, how you interpret a score depends on its particular distribution.
248
5.3 The Standard Normal Distribution
To avoid confusion resulting from using normal distributions with different units of measurement, researchers rely on a single normal distribution that can be applied to any normally distributed variable, regardless of how the variable has been measured. This distribution is known as the standard normal distribution, which is a normal distribution measured in standard deviation units with a mean equal to 0 and a standard deviation equal to 1. Figure 5.3 is a visual illustration of the standard normal distribution.
Figure 5.2 Different Types of Normal Distributions
249
What is the Standard Normal Distribution?
In terms of describing the standard normal distribution, it is important to understand it shares many of the features of normal distributions. There are an infinite number of scores in this distribution, with the left and right tails never touching the horizontal axis. In terms of its shape, the standard normal distribution is unimodal and symmetric.
However, the standard normal distribution is distinguished from other normal distributions in three important ways. First, scores for the standard normal distribution, known as z-scores, are measured in standard deviation units, which is the number of standard deviations from the mean. This is in contrast to other normal distributions, where measurements are a function of how each particular variable is measured (such as “inches” or “pounds”). Second, the mean of the standard normal distribution is equal to zero (0), in contrast to other normal distributions, which have the population mean μ. Third, in terms of variability, the standard deviation of the standard normal distribution is equal to 1 rather than the population standard deviation σ.
The next section will discuss the distinctive features of the standard normal distribution. In presenting these features, our discussion will focus on two ways that z-scores may be interpreted: a scores distance from the mean and a scores location relative to the entire distribution.
250
Interpreting z-Scores: Distance from the Mean
Figure 5.3 The Standard Normal Distribution
Because z-scores are measured in standard deviation units, the numeric value of a z-score represents its distance from the mean in standard deviations. For example, a z-score of zero (.00) is located zero standard deviations away from the mean z-score of 0. For any z-score other than zero, the sign (+/–) of the z-score indicates its location relative to the mean. z- scores greater than the mean are positive (+) numbers; z-scores less than the mean have negative (ITC Berkeley Oldstyle Std) values. Note that the sign of a z-score is typically only included when the z-score is negative; for example, a z-score of +.35 is reported as .35. Figure 5.4 displays three z-scores in the standard normal distribution.
In Figure 5.4(a), the z-score .00 is located at the mean of the distribution. (Although there are an infinite number of z-scores, researchers typically round them to two decimal places.) In Figure 5.4(b), the z-score of 1.00 is located exactly one standard deviation to the right of (i.e., greater than) the mean. Finally, in Figure 5.4(c), the z-score −1.24 is located a distance of 1.24 standard deviations below (or less than) the mean. Because the standard normal distribution is symmetric, the z-score −1.24 is the same distance from the mean as the z- score 1.24.
251
Interpreting z-Scores: Relative Position within the Distribution
In addition to its distance from the mean in standard deviation units, a z-score in the standard normal distribution may also be interpreted in terms of its location relative to the entire distribution. For example, it is possible to answer questions such as, “What percentage of the standard normal distribution is less than the z-score +1.00?”, “What percentage are greater than a z-score of −2.24?” or “What percentage is between the z-scores –.50 and +.50?” Later in this chapter, you will see how z-scores can be used to answer the question, “What percentage of SAT Math scores are less than a score of 660?”
To determine the area of the standard normal distribution corresponding to a particular z- score, we will use the normal curve table, a table containing the proportion of the standard normal distribution associated with different z-scores. Selected parts of the normal curve table have been listed in Table 5.1; the entire table may be found in the back of this book in Table 1.
The first column listed in Table 5.1, “z,” provides different values of z-scores. You may notice that all of the values of z in this column are positive numbers. Because the standard normal distribution is symmetric, it is only necessary for the normal curve table to include the area of the distribution to the right of the mean of .00. This is because the area of the distribution between the mean and a positive z-score (i.e., between z = .00 and z = .75) is exactly the same as the area between the mean and the negative z-score of the same value (i.e., between z = .00 and z = –.75). To locate the area associated with a negative z-score, you simply ignore the negative sign.
Figure 5.4 Three z-Scores in the Standard Normal Distribution
Table 5.1 Portions of the Appendix (Proportions of Area under the Standard Normal Distribution)
Table 5.1 Portions of the Appendix (Proportions of Area under the Standard Normal Distribution)
z Area Between Mean and z Area Beyond z
.00 .0000 .5000
.01 .0040 .4960
252
.02 .0080 .4920
.03 .0120 .4880
.04 .0160 .4840
.05 .0199 .4801
.06 .0239 .4761
.07 .0279 .4721
.08 .0319 .4681
.09 .0359 .4641
—– —— —–
.85 .3023 .1977
.86 .3051 .1949
.87 .3078 .1922
.88 .3106 .1894
.89 .3133 .1867
—– —— —–
1.00 .3413 .1587
1.01 .3438 .1562
1.02 .3461 .1539
1.03 .3485 .1515
1.04 .3508 .1492
—– —— —–
1.95 .4744 .0256
1.96 .4750 .0250
1.97 .4756 .0244
1.98 .4761 .0239
1.99 .4767 .0233
2.00 .4772 .0228
2.01 .4778 .0222
2.02 .4783 .0217
2.03 .4788 .0212
2.04 .4793 .0207
—– —— —–
2.95 .4984 .0016
253
2.96 .4985 .0015
2.97 .4985 .0015
2.98 .4986 .0014
2.99 .4986 .0014
3.00 .4987 .0013
4.00 .4999 .0001
The symmetric nature of the standard normal distribution also implies that exactly 50% of z-scores are located in each of the two halves of the distribution. In other words, 50% of z- scores have positive values greater than the mean of .00, and 50% have negative values less than the mean of .00. Adding the two halves together (50% + 50% = 100%) accounts for all possible z-scores.
Finding the Area between the Mean and a z-Score
What if you wanted to know, “How much of the standard normal distribution is located between z = .00 and z = 3.00?” Answering this question begins by examining the first column of Table 5.1—the column labeled “z.” This column lists z-scores in increasing order, starting with the mean of the distribution, z = .00. After moving down the “z” column until you reach the value 3.00, move to the right until you reach the corresponding number under the column labeled “Area Between Mean and z.” The number .4987 reveals that 49.87% (of a possible 50%) of the standard normal distribution lies between z = .00 and z = 3.00. This area is visually illustrated by the lighter shaded area in Figure 5.5.
As we mentioned earlier, the area of the distribution between the mean and a positive z- score is the same as the area between the mean and the same negative z-score. In other words, if we had asked, “How much of the standard normal distribution is located between z = .00 and z = −3.00?” we would have arrived at the same answer of 49.87%. Although the z-score may be a negative number, please note that the percentage cannot be a negative number because the smallest possible value of a percentage is zero (0%).
Figure 5.5 Percentage of z-Scores between Mean and z = 3.00
254
Now let's determine how much of the standard normal distribution is greater than the z- score 3.00. Returning to Table 5.1, move from our z-score of 3.00 to the corresponding number in the column labeled “Area Beyond z.” Here we find the value .0013, from which we learn that .13% (or less than 1%) of z-scores in the standard normal distribution are greater than (“beyond”) 3.00. This area is visually illustrated by the darker shaded portion located to the right of the value 3.00 in Figure 5.5. (We have found that drawing and shading figures such as the one in Figure 5.5 is a useful way to help check your calculations of these areas.)
We can also verify our findings by noting that the two percentages (49.87% and .13%) add up to 50%, which is the entire right half of the standard normal distribution: Area under the right half of = ( area between mean and z = 3.00 ) + the standard normal distribution ( area beyond z = 3.00 ) + 50 % = 49.87 % + .13 % 50 % = 50 %
Finding the Area Less or Greater than a z-Score
In addition to determining the area of the standard normal distribution between the mean z-score of .00 and a particular z-score, you can also evaluate a z-score in terms of its relationship to the entire standard normal distribution. For example, what if you asked, “What percentage of z-scores in the standard normal distribution are less than z = 1.00?” This has been represented by the shaded area in Figure 5.6.
Figure 5.6 Percentage of z-Scores Less than 1.00
255
The first step in determining the percentage of z-scores less than z = 1.00 is to determine the area between the mean of .00 and the stated z-score. For this example, move down the “z” column in Table 5.1 until you reach z = 1.00. Under the column labeled “Area Between Mean and z,” you will find associated with z = 1.00 the number .3413. This implies that 34.13% of the standard normal distribution lies between the mean and z = 1.00. However, this only determines the area of z-scores less than 1.00 that are located in the right half of the distribution. To find the percentage of all z-scores less than 1.00, you need to add the left half of the distribution to the percentage you just identified. Given that half, or 50%, of the standard normal distribution is less than .00, you add 50% to 34.13%: Area less than z = + 1.00 = ( area between mean and z = 1.00 ) + ( area less than mean ) = 34.13 % + 50 % = 84.13 %
From these calculations, we can conclude that 84.13% of the total distribution of z-scores is less than 1.00.
What percentage of the z-scores in the standard normal distribution is greater than (to the right of) z = 1.00? To answer this question, turn to Table 5.1 and look under the column “Area Beyond z.” Associated with z = 1.00 is the number .1587. This means that 15.87% of the z-scores are greater than z = 1.00. Note that the sum of the two percentages you have just calculated (84.13 + 15.87) is equal to 100%, which is the entire distribution.
Finding the Area between Two z-Scores
It is also possible to identify the percentage of the standard normal distribution that falls between two z-scores. For example, what percentage of z-scores is between z = –.85 and z = 1.95? This area is shaded in Figure 5.7. Calculating the percentage of scores within this area requires three steps. The first step is to divide the area into two sections: the area between the negative z-score and the mean of .00, as well as the area between the mean and the positive z-score. In this example: Area between z = − . 85 and z = 1.95 = area between z = − . 85 and mean + area between mean and z = 1.95
Figure 5.7 Percentage of z-Scores between z = –.85 and z = 1.95
256
Because the distribution is symmetric, the area between the negative z-score and the mean of .00 should be treated as if it were above the mean. In this example, the area of the distribution between the mean and z = –.85 is the same as the area between the mean and z = .85.
The second step is to determine the area between the mean and each of the two z-scores. To do this, we must return to the “Area Between Mean and z” column in Table 5.1. Here the area between the mean and z = –.85 is 30.23%, and the area between the mean and z = 1.95 is 47.44%.
The third and last step is to add these two percentages together: Area between z = − .85 and z = 1.95 = 30.23 % + 47.44 % = 77.67 %
On the basis of these three steps, we can conclude that 77.67% of z-scores in the standard normal distribution fall between z = –.85 and z = 1.95.
Figure 5.8 Percentage of z-Scores within 1, 2, and 3 Standard Deviations of the Mean
The above example shows how to calculate the area between a negative z-score and a positive z-score. Several of these areas, illustrated in Figure 5.8, are particularly important to remember. First, what percentage of z-scores is between the mean and one standard deviation in both directions (z = ±1.00)? (The ±symbol means “plus or minus.”) Looking at Table 5.1, under the “Area Between Mean and z” column, you find that 34.13% of z- scores are between the mean and z = 1.00. To cover the area both left and right of the mean (z = ±1.00), you simply multiply 34.13% by 2 and conclude that 68.26%, or roughly two thirds, of the standard normal distribution lies within ±1 standard deviation of the mean. Furthermore, for z = ±2.00, 95.44% (47.72% + 47.72%) of the distribution falls within 2 standard deviations of the mean. Finally, slightly less than 100% of z-scores, 99.74%, are within 3 standard deviations of the mean (z = ±3.00).
In addition to these examples, it is also possible to determine the area of the distribution
257
that is greater or less than a positive or negative z-score or that is between or beyond two different z-scores. Table 5.2 lists some simple rules regarding how the area under the standard normal distribution may be determined for stated z-scores.
258
5.4 Applying Z-Scores to Normal Distributions
Table 5.2 Rules for Determining Area of Distribution Associated with Stated z-Scores
Table 5.2 Rules for Determining Area of Distribution Associated with Stated z-Scores To find the percentage of z-scores that is …
between the mean and a positive or negative z-score, find the “Area Between Mean and z” greater than a positive z-score, find the “Area Beyond z” greater than a negative z-score, find the “Area Between Mean and z” and add 50% less than a positive z-score, find the “Area Between Mean and z” and add 50% less than a negative z-score, find the “‘Area Beyond z” between two positive or two negative z-scores, find the “Area Between Mean and z” for each of the two z-scores and subtract the smaller area from the larger area between one positive and one negative z-score, find the “Area Between Mean and z” for each of the two z-scores and add the two areas together beyond two z-scores, calculate the area between the two z-scores and subtract it from 100%
The previous section spent a great deal of time defining and explaining z-scores. However, researchers do not typically collect data in the form of z-scores. For example, if you ask people how tall they are, you would get responses such as 5′ 5″, 6′ 4″, or 5′ 9″ rather than −1.33, 2.33, or .00. Fortunately, the information provided by z-scores and the standard normal distribution can be applied to variables not measured in standard deviation units. This serves a valuable purpose: the ability to interpret and evaluate scores on variables. This section discusses how the standard normal distribution may be applied to variables assumed to be normally distributed in the population.
259
Transforming Normal Distributions into the Standard Normal Distribution
Normal distributions and the standard normal distribution share many features, such as their bell shape and infinite number of scores. The primary way in which the two types of distributions differ is the unit of measurement. The unit of measurement for normal distributions is specific to each particular variable (i.e., inches, pounds, miles per hour); consequently, it is difficult to evaluate scores for a variable as well as compare scores across different variables. However, the unit of measurement in the standard normal distribution is common to all variables: distance from the mean in standard deviations. Scores for variables that are assumed to be normally distributed in the population can be evaluated by transforming the normal distribution into the standard normal distribution—that is, transforming scores into z-scores.
260
Learning Check 1: Reviewing what you've Learned So Far
1. Review questions a. What are the main features and characteristics of normal distributions? b. Why are normal distributions important to researchers? c. What is a problem or concern about using normal distributions? d. What are the main features of the standard normal distribution? e. What distinguishes the standard normal distribution from normal distributions? f. What are the characteristics of z-scores? g. What are the two main ways z-scores may be interpreted? h. Why does the normal curve table only cover half of the standard normal distribution? i. How do you determine the area of the standard normal distribution that is between the
mean and a stated z-score? Less or greater than a z-score? Between two z-scores? j. What percentage of the standard normal distribution lies within one, two, and three
standard deviations of the mean? 2. Using the normal curve table, determine the area of the standard normal distribution that is
between the mean of .00 and the following z-scores: a. z = .46 b. z = 1.30 c. z = 2.87 d. z = –.19 e. z = −1.53
3. Using the normal curve table and the rules in Table 5.2, determine the area of the standard normal distribution that is either less (<) or greater (>) than the following z-scores:
a. z < .66 b. z < 2.55 c. z < –.75 d. z > 1.73 e. z > −2.17
4. Using the normal curve table and the rules in Table 5.2, determine the area of the standard normal distribution that is between the following z-scores.
a. z = .25 and z = .50 b. z = 1.33 and z = 2.25 c. z = –.67 and z = −1.50 d. z = –.44 and z = .82 e. z = −1.17 and z = 2.60
261
Understanding the z-Score Formula
Let's return to the example of the SAT introduced earlier in this chapter. In taking the math section of the SAT, you receive a score of 660 and want to know how “good” it is— that is, how a score of 660 relates to the scores of other students. Answering this question involves applying the properties of the standard normal distribution to the distribution of SAT scores in the population.
To apply the standard normal distribution to a normal distribution, the scores in the normal distribution must be transformed into a z-score (z) using the following formula:
(5-1) z = X − μ σ
where X is a score for the variable, μ is the population mean, and σ is the population standard deviation.
How does Formula 5-1 transform a score into a z-score? To answer that question, let's first examine the numerator: X – μ. Here we see that the value of a z-score is a function of the difference between the score and the population mean. Whether the score is greater (or less) than the mean will determine whether the corresponding z-score is a positive (or negative) number. Next, the denominator for the formula is a standard deviation; dividing the numerator by a standard deviation implies that a z-score is measured in standard deviation units. Combining the numerator and the denominator of Formula 5-1 illustrates that a z- score is the difference between the score and the mean in standard deviation units.
Converting a score into a z-score requires two pieces of information: the population mean (μ) and standard deviation (σ) of the variable to be transformed. There are two main ways to determine the values for μ and σ. The first way is to collect data from the entire population and then calculate the mean and standard deviation from these data. However, as it is extremely difficult to collect data from an entire population, a second, and more common, method is to examine data collected from a large sample or large number of samples and then to estimate the population mean and standard deviation from these samples.
262
Example: Transforming SAT Scores into z-Scores
Each year, the Education Testing Service (ETS) administers more than two million SAT examinations. The extremely large size of the sample increases confidence in the ETS's claim that SAT scores are normally distributed with a mean (μ) of 500 and standard deviation (σ) of 100. Therefore, SAT scores may be transformed into z-scores using the following formula: z = X − 500 100
To illustrate how to transform scores to z-scores, let's say you have the SAT scores of three students: 630, 450, and 500. Substituting each score for “X” in the above formula, the three scores are transformed to z-scores as follows:
SAT = 630 SAT = 450 SAT = 500
z = 630 − 500 100 = 130 100 = 1.30
z = 450 − 500 100 = − 50 100 = − .50
z = 500 − 500 100 = 0 100 = .00
263
Interpreting z-Scores: Distance from the Mean
In the previous example, the SAT scores of 630, 450, and 500 were transformed into the z- scores 1.30, –.50, and .00, respectively. As you have already learned, there are two main ways to interpret z-scores: the distance of the score from the mean of .00 and the location of the score relative to the entire distribution.
First, let's look at how these three SAT scores can be interpreted in terms of their distance from the mean. The first z-score of 1.30 indicates that an SAT score of 630 is located 1.30 standard deviations above the population mean of 500. The second z-score (z = –.50) implies that a score of 450 is .50 standard deviations below the mean. Finally, the third z- score of .00 means that an SAT score of 500 is equal to the population mean. In summary, an SAT score of 630 would be considered above average, a score of 450 is below average, and a score of 500 is average performance.
264
Interpreting z-Scores: Relative Position within Distribution
In addition to their distance from the mean, scores transformed into z-scores can be interpreted relative to the entire distribution. Or, to put it differently, it is possible for researchers to determine the area of the distribution between specified z-scores (e.g., the percentage of SAT scores less than 660) by using the guidelines provided earlier in Table 5.2.
Finding the Area between the Mean and a Score
As a first example, what if you wanted to know the percentage of SAT scores between the mean of 500 and a score of 700? The first step in answering this question is to transform the stated score (700) into a z-score using Formula 5-1: z = 700 − 500 100 = 200 100 = 2.00
The above calculation tells us that an SAT score of 700 is equal to a z-score of 2.00.
Now that we have calculated the z-score, the next step is to determine the “Area Between Mean and z” in the normal curve table. Looking at Table 5.1, for z = 2.00, we see the value .4772. We may conclude, therefore, that 47.72% (or almost half) of SAT scores are between 500 and 700. This is illustrated by the shaded portion of Figure 5.9.
Finding the Area Less or Greater than a Score
Earlier in this chapter, we posed the situation where you have taken the SAT math section and received a score of 660. What if you wanted to compare your score to others' scores by determining the percentage of SAT scores less than 660? Again, you start by converting the score into the corresponding z-score: z = 660 − 500 100 = 160 100 = 1.60
Figure 5.9 Percentage of SAT Scores between 500 and 700
Figure 5.10 Percentage of SAT Scores Less than 660
265
Because you are looking for the area less than a positive z-score, Table 5.2 informs you to find the “Area Between Mean and z” and add 50%. Looking at the Appendix, move down the “z” column until you reach z = 1.60. Under the “Area Between Mean and z” column, you find the value .4452. By adding 50% to 44.52%, you can conclude that 94.52% of SAT scores are less than 660. This has been represented by the shaded area in Figure 5.10.
What percentage of SAT scores are greater than 660? Looking under the column “Area Beyond z,” you find associated with z = 1.60 the number .0548, implying that 5.48% of SAT scores (100% – 94.52%) are greater than 660.
Finding the Area between Two Scores
What if you asked, “What percentage of SAT scores is between 470 and 540?” (see Figure 5.11). You begin by converting each of these two scores to z-scores:
SAT = 470 SAT = 540
z = 470 − 500 100 = − 30 100 = − .30 z = 540 − 500 100 = 40 100 = .40
Because this example involves one positive and one negative z-score, the next step is to divide this area into two sections: the area between the negative z-score and the mean of .00, as well as the area between the mean and the positive z-score. Ignoring the sign of the negative z-score, you determine the “Area Between Mean and z” for the two z-scores and add these two percentages together. In this example,
Figure 5.11 Percentage of SAT Scores between 470 and 540
266
Area between z = − .30 and z = .40 = ( area between z = − .30 and mean ) + ( area between mean and z = .40 ) = 11.79 % + 15.54 % = 27.33 %
From these calculations, you would conclude that 27.33% of SAT scores are between 470 and 540.
267
Transforming Scores into z-Scores: Drawing Conclusions
The previous section discussed how the standard normal distribution may be applied to variables assumed to be normally distributed in the population. Although we emphasized the presentation and illustration of mathematical formulas and calculations, it is equally important to develop the ability to interpret and evaluate the results of these calculations. Transforming a score into a z-score provides the ability to draw conclusions about the score. As an example, we posed the situation where you wanted to evaluate your SAT score of 660. Transforming this score into a z-score of 1.60 revealed that 94.52% of SAT scores are less than 660, whereas only 5.48% of scores are greater than 660. Consequently, it can be concluded that only a very small percentage of the population receives scores higher than 660 and that you should feel confident in your abilities.
268
Learning Check 2: Reviewing what you've Learned So Far
1. Review questions a. Why might one want to apply the standard normal distribution to a normal distribution? b. What's the first step in applying the standard normal distribution to a normal distribution?
(Exercises 2 and 3 are based on the following example): A study of more than 1,000 pregnant women in India found the length of pregnancy was normally distributed, with a mean (μ) of 272 days (about 9 months) and a standard deviation (σ) of 9 days (Bhat & Kushtagi, 2006).
2. Transform the following lengths of pregnancy into z-scores: a. 275 days b. 280 days c. 265 days d. 286 days e. 260 days
3. What percentage (%) of pregnancies … a. is shorter than 277 days? b. is shorter than 262 days? c. is longer than 270 days? d. is longer than 283 days? e. is between 279 and 286 days? f. is between 268 and 276 days?
(Exercise 4 is based on the following example): A survey conducted in 2005 (www.macintouch.com) asked a large number of iPod owners whether their iPod had ever failed (completely stopped working). Based on their results, it was estimated that the likelihood of experiencing a failure was normally distributed, with an average likelihood of failure (μ) of 13.70%, with a standard deviation (σ) of 7.54%.
4. What percentage (%) of iPods … a. have a likelihood of failure greater than 25%? b. have a likelihood of failure greater than 10%? c. have a likelihood of failure less than 20%? d. have a likelihood of failure less than 5%? e. have a likelihood of failure between 15% and 30%? f. have a likelihood of failure between 11% and 18%?
269
5.5 Standardizing Frequency Distributions
The previous section discussed how the standard normal distribution may be applied to normal distributions, which are distributions based on populations with a known population mean and standard deviation. The primary purpose of applying the standard normal distribution to normally distributed variables is to interpret and evaluate scores for a variable by transforming them into z-scores. As you will see in later chapters of this book, this transformation, interpretation, and evaluation are the foundation of statistical procedures designed to test research hypotheses.
In addition to transforming normal distributions, z-scores may also be calculated for scores within a frequency distribution (i.e., a distribution of scores based on data collected from a sample). For example, imagine you take a midterm exam in a class of 23 students and want to answer the question, “How well did I do?” You might want to evaluate your score relative to the other students in the class but have no reason to believe that the scores in your class are normally distributed. This section uses a published research study to show how and why researchers transform scores in a frequency distribution into z-scores.
270
Example: Measuring Politicians' Behavior
Voters have a vested interest in understanding the thoughts, feelings, and behaviors of their elected officials, including the impact of politicians' personal goals on their behavior. Researchers might choose to examine, for example, whether the behavior of a state governor who wishes to be president differs from the behavior of a governor who prefers serving at a lower level of government.
One study, conducted by Dr. Rebekah Herrick from Oklahoma State University, investigated the political ambition of members of the U.S. House of Representatives (Herrick, 2001). More specifically, “Members' responses to questions about their desires for future political office were compared to their legislative behaviors to see if ambition affects their behavior” (p. 471). The example in this chapter will focus on the legislative behavior of these members.
One aspect of legislative behavior examined in this study was “floor activity,” measured using two variables: the number of remarks made on the floor and the number of amendments proposed to other members' legislation. The higher the number of remarks or amendments, the higher the level of floor activity. Dr. Herrick recorded the number of remarks and amendments made by a sample of members of the House of Representatives during their 2-year term in office. For the purposes of our discussion, this example will be limited to the data from 23 of the 47 members in her study. (The findings reported here closely resemble the overall findings from the published study.)
Table 5.3(a) provides the number of remarks and the number of amendments for the 23 members to be considered in this example. Table 5.3(b) provides frequency distribution tables for the data presented in Table 5.3(a).
Looking at the frequency distribution tables, the two variables differ greatly in their range of possible values. The number of remarks ranges from 0 to 157; however, the number of amendments proposed is much smaller because developing an amendment requires more time and effort than what is typically required for public speaking. These differences are reinforced by Table 5.4(a) and (b), which calculate the mean and standard deviation for the two variables using the formulas introduced in Chapters 3 and 4. As you can see, a comparison of the two means, M = 65.22 for the number of remarks and M = 1.74 for the number of amendments, clearly reflects the difference in the nature of the two variables.
Dr. Herrick's objective was to combine each member's scores on the two variables into a single measure of floor activity. But because of the difference in the distributions of the two variables, she felt it would be inappropriate to simply add the two scores together. For instance, although a score of 0 is an extremely low number of remarks, it is actually the modal score for the number of amendments. Consequently, to compare and combine
271
variables that have different units of measurement, the variables must first be transformed to the same unit of measurement. This is accomplished by transforming each distribution of scores. As she wrote, “Since the number of amendments were significantly smaller than the number of remarks, z-scores were calculated and used” (Herrick, 2001, p. 471).
Table 5.3 Number of Remarks and Number of Amendments Variables Table 5.3 Number of Remarks and Number of Amendments Variables
(a) Raw Data
Member # Remarks # Amendments
1 49 0
2 82 1
3 66 3
4 25 0
5 44 0
6 113 13
7 52 1
8 61 1
9 70 2
10 54 2
11 84 1
12 70 0
13 57 2
14 105 0
15 78 0
16 36 1
17 68 7
18 29 1
19 157 0
20 55 1
21 61 0
22 0 0
23 84 4
Table 5.3 Number of Remarks and Number of
272
Amendments Variables
(b) Frequency Distribution Tables
# Remarks f % # Amendments f %
140+ 1 4% 7+ 2 9%
120–139 0 0% 6 0 0%
100–119 2 9% 5 0 0%
80–99 3 13% 4 1 4%
60–79 7 31% 3 1 4%
40–59 6 26% 2 3 13%
20–39 3 13% 1 7 31%
0–19 1 4% 0 9 39%
Total 23 100% Total 23 100%
273
Standardizing a Frequency Distribution
When scores in a frequency distribution are transformed into z-scores, the distribution of z- scores is called a standardized distribution. A standardized distribution is different from the standard normal distribution in that a standardized distribution does not assume the variable is normally distributed with a known population mean and standard deviation. The z-scores within a standardized distribution are referred to as standardized scores.
A frequency distribution of scores may be transformed into a standardized distribution of standardized scores using Formula 5-2:
(5-2) z = X − X ¯ s
where X is a score for the variable, X ¯ is the sample mean and s is the standard deviation.
Formula 5-2 is very similar to Formula 5-1, which is used to transform a normal distribution into the standard normal distribution. The numerator in both formulas is the deviation of a score from a mean, and both require dividing the numerator by a standard deviation to measure z-scores in standard deviation units. There are, however, two critical differences between the two formulas. First, creating a standardized distribution using Formula 5-2 does not assume the knowledge of a population mean (μ) or population standard deviation (σ). Instead, z-scores in a standardized distribution are based on the sample mean ( X ¯ ) and sample standard deviation (s). As such, you are relating the scores for a variable to the other scores in the sample rather than to scores in the population. Second, the scores in the frequency distribution are not assumed to be normally distributed. Regardless of its shape (modality, symmetry, and variability), any frequency distribution may be transformed into a standardized distribution.
274
Interpreting Standardized Scores
Member 1 (# Remarks = 49)
Member 2 (# Remarks = 82)
Member 3 (# Remarks = 66)
z = 49 − 65.22 32.16 = − 16.22 32.16 = − .50
z = 82 − 65.22 32.16 = 16.78 32.16 = .52
z = 66 − 65.22 32.16 = .78 32.16 = .02
To illustrate how a frequency distribution may be transformed into a standardized distribution, we will transform the first three scores for the number of remarks variable into their corresponding standardized scores. Inserting the sample mean ( X ¯ = 65.22 = 65.22) and standard deviation (s = 32.16) calculated in Table 5.4 into Formula 5-2, the three z- scores are calculated on page 160, in which the number of remarks of the first three members (49, 82, and 66) have been transformed into standardized scores of –.50, .52 and .02.
Table 5.4 Descriptive Statistics, Number of Remarks, and Amendments Variables Table 5.4 Descriptive Statistics, Number of Remarks, and Amendments Variables
(a) Calculation of Means ( X ¯ )
# Remarks # Amendments
X ¯ = ∑ X N = 49 + 82 + ⋯ + 0 + 84 23 = 65 . 22 = 1500 23
X ¯ = ∑ X N = 0 + 1 + ⋯ + 0 + 4 23 = 40 23 = 1.74
Table 5.4 Descriptive Statistics, Number of Remarks, and Amendments Variables
(b) Calculation of Standard Deviations (s)
# Remarks # Amendments
s = ∑ ( X − X ¯ ) 2 N-1 = ( 49 − 65.22 ) 2 + ⋯ + ( 84 − 65.22 ) 2 23 − 1 = 263.09 + ⋯ + 352.69 22 = 22751.95 22 = 1034.18 = 32.16
s = ∑ ( X − X ¯ ) 2 N − 1 = ( 0 − 1.74 ) 2 + ⋯ + ( 4 − 1.74 ) 2 23 − 1 = 3.03 + ⋯ + 5.11 22 = 192.43 22 = 8.75 = 2.96
Standardized scores are typically interpreted in terms of their distance from the sample mean in standard deviation units. For example, the first members z-score of –.50 indicates that 49 remarks is .50 standard deviations, or half a standard deviation, below the sample mean of 65.22. On the other hand, the second members 82 remarks is .52 standard deviations above the mean. Finally, the 66 remarks made by the third member is very close to the mean for this sample. Transforming scores into standardized scores helps researchers evaluate and interpret their data. For example, we can say that the 49 remarks made by the first member is below the average of the sample, whereas the second member's 82 remarks
275
is above average.
276
Using Standardized Scores to Compare or Combine Variables
Perhaps the most common reason for creating standardized distributions is to compare and combine scores on variables measured on different units of measurement. To illustrate this use of standardized scores, the information from Tables 5.3 and 5.4 is used to transform the scores for the first three members on the other variable in this study, the number of amendments introduced by each member:
Member 1 (# Amendments = 0)
Member 2 (# Amendments = 1)
Member 3 (# Amendments = 3)
z = 0 − 1.74 2.96 = − 1.74 2.96 = − .59
z = 1 − 1.74 2.96 = − .74 2.96 = − .25
z = 3 − 1.74 2.96 = 1.26 2.96 = .43
Transforming scores into standardized scores provides a way of comparing scores on the two variables. Looking, for example, at the two standardized scores for the first member, a score of 49 on the number of remarks variable (z = –.50) is comparable to a score of 0 on the number of amendments variable (z = -.59). As such, we can conclude that this members level of floor activity on both of these variables was below the average of this sample of representatives. Furthermore, because both variables are now measured on the same unit of measurement, scores on the two variables may be combined. In this example, by transforming the number of remarks and number of amendments variables into standardized scores, Dr. Herrick was able to combine them into a single measure of floor activity.
277
Characteristics of Standardized Distributions
This section illustrates the similarities and differences between standardized distributions and the standard normal distribution. To begin this discussion, Table 5.5(a) provides the scores and z-scores for the number of amendments variable in the Herrick (2001) study. Table 5.5(b) calculates the mean and standard deviation of the standardized scores for this variable.
Looking at the calculations for the mean and standard deviation, we discover that standardized distributions share several of the characteristics of the standard normal distribution. Specifically, the mean and standard deviation of both distributions are equal to 0 and 1, respectively. But is the shape of the two distributions also the same?
Table 5.6 presents frequency distribution tables for the number of amendments variable in both original and standardized score form. Notice that although the scores and z-scores differ, the numbers in the “f” and “%” columns are the same in both tables. As we can see from this example, standardizing a frequency distribution does not change the shape of the distribution. The shape of the distribution is the same, regardless of whether a variable is measured in score or standardized score form. We can also assert from this example that standardizing a distribution does not transform a frequency distribution into a normal distribution. If the scores for a variable are not normally distributed, they will not be normally distributed after being transformed into standardized scores.
Table 5.5 Standardized Scores for the Number of Amendments Variable Table 5.5 Standardized Scores for the Number of
Amendments Variable
(a) Listing of Scores and z-Scores for Each Member
Member # Amendments z-Score
1 0 –.59
2 1 –.25
3 3 .43
4 0 –.59
5 0 –.59
6 13 3.81
7 1 –.25
8 1 –.25
9 2 .09
10 2 .09
278
11 1 –.25
12 0 –.59
13 2 .09
14 0 –.59
15 0 –.59
16 1 –.25
17 7 1.78
18 1 –.25
19 0 –.59
20 1 –.25
21 0 –.59
22 0 –.59
23 4 .76 Table 5.5 Standardized Scores for the Number of Amendments Variable
(b) Descriptive Statistics of Standardized Scores
Mean Standard Deviation
∑ z N = − . 59 + − 25 + … + − . 59 + . 76 23 = . 00 23 = . 00
∑ z − z ¯ 2 N − 1 = . 59 − . 00 2 + … + . 76 − . 00 2 23 − 1 = . 35 + … + . 58 22 = 22.00 22 = 1.00 = 1.00
Table 5.6 Frequency Distribution Tables, Number of Amendments Variable Table 5.6 Frequency Distribution Tables,
Number of Amendments Variable
Scores z-Scores
Score f % z-Score f %
5+ 2 9% 1.78+ 2 9%
4 1 4% .76 1 4%
3 1 4% .43 1 4%
2 3 13% .09 3 13%
1 7 31% –.25 7 31%
0 9 39% –.59 9 39%
Total 23 100% Total 23 100%
The fact that standardizing a frequency distribution does not alter the shape of the distribution has an important implication for how standardized scores can be interpreted.
279
For example, we know that exactly 50% of the z-scores in the standard normal distribution lie on either side of the mean. However, in the standardized distribution in Table 5.6, only 17% (4 of 23) of the standardized scores for the number of amendments variable are greater than .00. Students sometimes mistakenly believe that converting scores to standardized scores changes the shape of the distribution into a normal distribution. This is simply not the case.
Using Formula 5-2 to standardize a frequency distribution does not alter the shape of the distribution because this formula is a linear transformation, which is a mathematical transformation of a variable comprising addition, subtraction, multiplication, or division. Linear transformations do not change the shape of the distribution because the relative distances between scores in the distribution do not change. Another example of a linear transformation is temperature. If, for example, you wanted to convert a temperature of 80° Fahrenheit to its Celsius equivalent, you would use the formula of Celsius = 5 9 Fahrenheit − 32 . For a temperature of 80° Fahrenheit, its Celsius equivalent is 5 9 80 − 32 , or 5 9 48 , or 26.67°. If one person says today's temperature is 80° Fahrenheit and another person says it's 27° Celsius, you would not say that the temperature has changed. Similarly, transforming scores to z-scores to standardize a frequency distribution does not change the shape of the distribution.
280
When can You Apply the Normal Curve Table to a Standardized Distribution?
Given that standardizing a frequency distribution does not change the shape of the distribution, is it ever possible to apply the characteristics of the standard normal distribution to a standardized distribution? The answer to this question depends on the shape of the frequency distribution. The features of the standard normal distribution can be applied to standardized distributions if and when the shape of the distribution of scores is approximately normal. But this answer leaves us with yet another question: What is meant by “approximately” normal?
281
Learning Check 3: Reviewing what you've Learned So Far
1. Review questions a. What are the differences between calculating z-scores for normal distributions versus
frequency distributions? b. What are common reasons for standardizing frequency distributions? c. What are the main characteristics of standardized distributions? d. Does standardizing a frequency distribution create a normal distribution? Why or why not? e. Under what conditions could you apply the normal curve table to a standardized
distribution? 2. Using the descriptive statistics calculated in Table 5.4, calculate standardized scores for the number
of remarks variable for the following members. a. Member 4 (# remarks = 25) b. Member 5 (# remarks = 44) c. Member 6 (# remarks = 113) d. Member 7 (# remarks = 61) e. Member 8 (# remarks = 70)
3. Imagine you take a quiz in one of your classes and learn that the class mean ( X ¯ ) was 15.53 and the standard deviation (s) was 4.22. Calculate standardized scores for the following quiz scores.
a. 14 b. 18 c. 7 d. 17 e. 23
4. You play on a softball team and at the end of the season find the batting averages of the players on your team are as follows:
Player Batting Average
1 .375
2 .250
3 .289
4 .321
5 .264
6 .298
7 .311
8 .232
9 .211
a. Calculate the mean ( X ¯ ) and standard deviation (s) of the batting averages. b. Transform each batting average into a standardized score. c. Calculate the mean and standard deviation of the standardized scores.
Because the standard normal distribution comprises an infinite number of scores based on a mathematical formula, data collected from a sample cannot be expected to completely resemble a normal distribution. However, the larger the sample, the more likely it is that
282
the data will begin to resemble a normal distribution. How large does your sample need to be for it to be assumed that the data approach what researchers typically describe as a normal distribution? While there is no definitive answer to this question, some researchers have come to rely on the formula, N = 30. This means that a sample size of at least 30 is necessary for the normal curve to be applied to a standardized distribution.
283
5.6 Looking Ahead
This chapter began by reviewing a process used to conduct research. A central part of this process is conducting statistical analyses designed to test research hypotheses. As you will see in the next chapter, these statistical analyses depend heavily on different types of distributions. As such, the main purpose of this chapter was to introduce two critical types of distributions: normal distributions and the standard normal distribution. Several aspects of these distributions will be built upon in future chapters. First, these distributions are unimodal and symmetric, with an infinite number of scores. Also, it is possible to evaluate a score in terms of the proportion of the distribution associated with that score. However, rather than evaluating individual scores, researchers analyze data collected from samples to test research hypotheses about populations. Because data are not collected from the entire population, researchers must rely on probability to evaluate their data and draw appropriate conclusions about their hypotheses. The next chapter describes the role of probability in the research process and introduces steps taken by researchers to conduct statistical analyses to test their hypotheses.
284
5.7 Summary
Normal distributions are distributions based on a population of an infinite number of scores calculated from mathematical formulas. Some of the features of normal distributions are that they consist of an infinite number of scores, are unimodal and symmetric, have a mean that is equal to the population mean μ, and have a standard deviation equal to the population standard deviation σ.
Normal distributions are important to researchers because many variables in populations are generally believed by researchers to be normally distributed, many statistical techniques designed to test research hypotheses are based on the assumption that variables are normally distributed, and, because normal distributions are based on mathematical formulas, it is possible to determine the proportion of the distribution associated with any score.
One problem with using normal distributions is that the shape and nature of a particular normal distribution depend on how each particular variable is measured. To address this problem, the standard normal distribution, which is a normal distribution with a mean equal to 0 and a standard deviation equal to 1, may be applied to any normal distribution. The standard normal distribution has several distinguishing characteristics: Its scores (known as z-scores) are measured in standard deviation units, the mean of this distribution is equal to 0, and the standard deviation is equal to 1.
It is possible to evaluate z-scores in two ways: distance from the mean in standard deviation units and the percentage of the distribution associated with a stated z-score or scores (using a table known as the normal curve table).
Scores in normally distributed variables may be evaluated by transforming its scores into z- scores. Scores in a frequency distribution (data collected by researchers) may be transformed into z-scores known as standardized scores; the distribution of standardized scores is known as a standardized distribution. A common reason for creating standardized distributions is to compare and combine scores on variables measured on different scales of measurement. Because a standardized distribution does not assume the variable is normally distributed, if the scores in a frequency distribution are not normally distributed, they will not be normally distributed after being transformed into standardized scores. However, the larger the sample, the more likely it is a frequency distribution will resemble a normal distribution.
285
5.8 Important Terms
normal distribution (p. 140) standard normal distribution (p. 142) z-scores (p. 143) normal curve table (p. 144) standardized distribution (p. 160) standardized scores (p. 160) linear transformation (p. 164)
286
5.9 Formulas Introduced in this Chapter
287
z-Score
(5-1) z = X − μ σ
Standardized Score
(5-2) z = X − X ¯ s
288
5.10 Exercises
NOTE: In Exercises 1 to 4, it may be helpful to draw a figure such as Figure 5.5.
1. Using the normal curve table, determine the area of the standard normal distribution that is between the mean of .00 and the following z-scores:
a. z = .69 b. z = 1.45 c. z = 2.01 d. z = –.25 e. z = −1.86
2. Using the normal curve table, determine the area of the standard normal distribution that is less than the following z-scores:
a. z = .30 b. z = 1.75 c. z = 2.42 d. z = –.68 e. z = −1.11
3. Using the normal curve table, determine the area of the standard normal distribution that is greater than the following z-scores:
a. z = .78 b. z = 1.22 c. z = 2.30 d. z = –.90 e. z = −1.34
4. Using the normal curve table, determine the area of the standard normal distribution that is between the following z-scores:
a. z = .30 and z = 1.00 b. z = 1.56 and z = 2.45 c. z = –.50 and z = –.80 d. z = –.90 and z = −2.67 e. z = −1.34 and z = .10 f. z = –.75 and z = .75
(Exercises 5 through 9 are based on the following example): IQ scores are normally distributed in the population with a mean (μ) of 100 and a standard deviation (σ) of 16. Use the normal curve table to complete these exercises.
5. Transform the following IQ scores into z-scores. a. 116 b. 84 c. 124
289
d. 95 e. 130
6. What percentage of IQ scores is between the mean of 100 and … a. 116 b. 68 c. 105 d. 90 e. 140
7. What percentage of IQ scores is … a. less than 122 b. less than 77 c. less than 104 d. greater than 135 e. greater than 73 f. greater than 112
8. What percentage of IQ scores is … a. between 115 and 130 b. between 105 and 120 c. between 75 and 85 d. between 80 and 95 e. between 90 and 110 f. between 80 and 120
9. What percentage of IQ scores is … a. less than 110 or greater than 130 b. less than 80 or greater than 90 c. less than 95 or greater than 105 d. less than 84 or greater than 116
10. A researcher constructs a test to measure self-esteem. In doing so, she calculates a mean for her sample of 10 with a standard deviation of 2. Assuming that scores on this test in the larger population are normally distributed, what percentage of scores on the test is …
a. greater than 12 b. greater than 16 c. greater than 8 d. less than 11 e. between 13 and 15 f. between 6 and 14
(Exercises 11 through 13 are based on the following example): Imagine you are the president of a toy company that builds video games. You have determined that the time needed to build these games is normally distributed with a mean of 20 minutes and a standard deviation of 5 minutes.
290
11. What percentage of the video games is built … a. in less than 15 minutes b. in less than 10 minutes c. in more than 20 minutes d. in more than 30 minutes e. between 26 and 32 minutes f. between 12 and 24 minutes
12. Five percent of the video games (the bottom 5%) take less than ____ minutes to build.
13. Ninety percent of the video games (the top 90%) are built in ____ minutes. 14. How many base hits can a baseball team expect to get in a game? Frohlich (1994)
recorded the number of hits by the 28 major league baseball teams for all of the games played from 1989 to 1993 (each team plays 162 games a year). He found that the number of hits the teams made in the games was normally distributed, with a mean of 8.72 and a standard deviation of 1.10.
a. In what percentage of games would you expect a team to get more than nine hits?
b. In what percentage of the games would you expect a team to get less than six hits?
15. In another baseball-related story, between 1980 and 2000, there was a remarkable increase in the number of home runs hit in the major leagues. A recent article stated that the average number of home runs per game rose from 1.47 in 1980 to 2.34 in 2000 (Rist, 2001). A factor thought to be a contributing factor to this increase is the ball itself. One of the responsibilities of a university-based research center sponsored by Major League Baseball is to check the weight of baseballs. In a recent report, this research center found that the weight of baseballs used in the 2000 season had a mean of 5.07 ounces with a standard deviation of .06 ounces.
a. Major League Baseball specifies that the weight of baseballs should be between 5.00 and 5.25 ounces. What percentage of the baseballs falls outside of the specified range? That is, what percentage of the 2000 baseballs was less than 5.00 ounces; what percentage weighed more than 5.25 ounces?
b. A sample of baseballs weighed in 1999 had a mean weight of 5.09 ounces with a standard deviation of .05 ounces. What percentage of the 1999 baseballs fell outside of 5.00 to 5.25 range?
c. Comparing the results of the two seasons, what conclusions might you reach regarding the weight of baseballs?
16. What if you took midterms in two different courses and happened to get the same grade of 68 on both of them. In the first course, the mean was 63 and the standard deviation was 10. In the second course, the mean was 72 and the standard deviation was 8.
a. Calculate the z-score for your score in both courses. b. Relative to the other students in the courses, in which course would you
291
consider your performance to be better? Why? 17. What if you took final examinations in two different courses and happened to get the
same grade of 72 on both of them? The mean in the two courses turns out to be the same X ¯ = 60 . However, the standard deviation in the two courses differs: s = 12 in the first course and s = 6 in the second.
a. Calculate the z-score for your score in both courses. b. Even though your score on the two tests was the same and the mean in the two
courses was the same, why are your two z-scores different from each other? c. Relative to the other students in the courses, in which course would you
consider your performance to be better? Why? 18. The mean of the Graduate Record Examination (GRE), for the verbal section, is 500
with a standard deviation of 100. Your friend scores a 529. a. What percentage of test takers did she score above? b. How many standard deviations above or below the mean is she?
19. Two students are comparing their recent midterm scores. The first student received a 92 and the second student received an 86. However, they are in different classes, and their tests were slightly different. The mean score in the first student's class was 82, with a standard deviation of 10 points. The mean score in the second student's class was a 75, with a standard deviation of 10. Both tests were normally distributed.
a. Convert both students' test scores to z-scores. b. What percentage of test scores was below each score? c. Which student did better on the statistics test?
20. You collect the grade point average (GPA) from 10 students:
Student GPA
1 3.45
2 2.40
3 2.80
4 2.55
5 3.80
6 2.55
7 2.30
8 2.80
9 3.05
10 2.30
a. Calculate the mean ( X ¯ ) and standard deviation (s) of these GPAs. b. Calculate the z-score for each GPA in order to standardize the distribution. c. Calculate the mean and standard deviation of these z-scores.
292
d. What percentage of z-scores is greater than the mean? What percentage is less than the mean? Are these two percentages the same? Why or why not?
293
Answers to Learning Checks
Learning Check 1
2. a. 17.72% b. 40.32% c. 49.79% d. 7.53% e. 4.37%
3. a. 74.54% b. 99.46% c. 22.66% d. 4.18% e. 98.50%
4. a. 9.28% b. 7.95% c. 18.46% d. 46.39% e. 87.43%
Learning Check 2
2. a. z = .33 b. z = .89 c. z = –.78 d. z = 1.56 e. z = −1.33
3. a. 71.22% (z = .56) b. 13.35% (z = −1.11) c. 58.71% (z = –.22) d. 11.12% (z = 1.22) e. 15.83% (z = .78 and z = 1.56) f. 34.01% (z = –.44 and z = .44)
4. a. 6.69% (z = 1.50)
294
b. 68.79% (z = –.49) c. 79.95% (z = .84) d. 12.51% (z = −1.15) e. 41.71% (z = .17 and z = 2.16) f. 35.62% (z = –.36 and z = .57)
Learning Check 3
2. a. z = −1.25 b. z = –.66 c. z = 1.49 d. z = –.13 e. z = .15
3. a. z = –.36 b. z = .59 c. z = −2.02 d. z = .35 e. z = 1.77
4. a. X ¯ = . 284 , s = .051 b.
Player Batting Average z
1 .375 1.78
2 .250 –.67
3 .289 .10
4 .321 .73
5 .264 –.39
6 .298 .27
7 .316 .63
8 .232 –1.02
9 .207 –1.51
c. Mean of z scores = 0, standard deviation of z scores = 1.00.
295
Answers to Odd-Numbered Exercises
1. a. 25.49% b. 42.65% c. 47.78% d. 9.87% e. 46.86%
3. a. 21.77% b. 11.12% c. 1.07% d. 81.59% e. 90.99%
5. a. z = 1.00 b. z = −1.00 c. z = 1.50 d. z = –.31 e. z = 1.88
7. a. 91.62% (z = 1.38) b. 7.49% (z = −1.44) c. 59.87% (z = .25) d. 1.43% (z = 2.19) e. 95.45% (z = −1.69) f. 22.66% (z = .75)
9. a. 76.58% (z =.63 and z = 1.88) b. 84.13% (z = −1.25 and z = –.63) c. 75.66% (z = –.31 and z = .31) d. 31.74% (z = −1.00 and z = 1.00)
11. a. 15.87% (z = −1.00) b. 2.28% (z = −2.00) c. 50.00% (z = .00) d. 2.28% (z = 2.00) e. 10.69% (z = 1.20 and z = 2.40) f. 73.33% (z = −1.60 and z = .80)
13. 26.40 minutes (z = 1.28) 15.
296
a. 12.10% weigh less than 5 oz (z = −1.17); .13% weigh more than 5.25 oz (z = 3.00).
b. 3.59% weigh less than 5 oz (z = −1.80); .07% weigh more than 5.25 oz (z = 3.20).
c. It appears baseballs that weighed outside the allowed range were being used more frequently in 2000 than in 1999.
17. a. First course: z = 1.00, second course: z = 2.00 b. The students in the second course scored more similarly; therefore, a deviation
from the mean is a greater change in relation to the other students. c. In relation to other students, the score in Course B is better (better than
97.72% of the class) than the score in Course A (better than 84.13% of the class).
19. a. First student: z = 1.00; second student: z = 1.10 b. 84.13% of the class scored below the first student; 86.85% scored below the
second student. c. The second student did better on his or her midterm, even though the score
was lower.
297
Sharpen your skills with SAGE edge at edge.sagepub.com/tokunaga
SAGE edge for students provides a personalized approach to help you accomplish your coursework goals in an easy-to-use learning environment. Log on to access:
eFlashcards Web Quizzes Chapter Outlines Learning Objectives Media Links
298
Chapter 6 Probability and Introduction to Hypothesis Testing
299
Chapter Outline 6.1 A Brief Introduction to Probability
What is probability? Why is probability important to researchers? Applying probability to normal distributions Applying probability to binomial distributions
6.2 Example: Making Heads or Tails of the Super Bowl 6.3 Introduction to Hypothesis Testing
State the null and alternative hypotheses (H0 and H1) The null hypothesis (H0) The alternative hypothesis (H1)
Make a Decision About the Null Hypothesis Defining a “low” probability of a statistic: setting alpha (α) Identifying the values of the statistic with a low probability Stating a decision rule Calculating a value of a statistic Making the decision whether to reject the null hypothesis
Draw a conclusion from the analysis Relate the result of the analysis to the research hypothesis Summary of steps in hypothesis testing
6.4 Issues Related to Hypothesis Testing: An Introduction The issue of “proof” in hypothesis testing Errors in decision making Factors influencing the decision about the null hypothesis
Sample size Alpha (α) Directionality of the alternative hypothesis
6.5 Looking Ahead 6.6 Summary 6.7 Important Terms 6.8 Formulas Introduced in This Chapter 6.9 Exercises
In Chapter 5, we discussed normal distributions and the standard normal distribution. One critical feature of these distributions is the ability to evaluate a score in terms of its relationship to the other scores in the distribution. Understanding these types of distributions and how to evaluate scores in these distributions is the foundation of this chapter, which introduces the process of conducting statistical analyses to test research hypotheses regarding what is believed to exist in a population. But rather than evaluate individual scores, these statistical analyses evaluate data collected from samples. Because data are typically collected from samples of a population rather than the entire population, researchers cannot test hypotheses with complete certainty; instead, researchers must determine the likelihood (or probability) of obtaining their results. Therefore, before introducing the process of hypothesis testing, the next section provides a very brief introduction to probability.
300
6.1 A Brief Introduction to Probability
Although you may not be aware of it, you constantly come in contact with and use probability. You hear commercials on the radio saying that a particular device makes it “five times less likely” your car will be stolen. If you play card games, you keep or discard cards based on the likelihood of making the best possible hand. At a grocery store, you look at five checkout lines and pick the one you believe is most likely to move quickly. Probability is also an area of study given a great deal of attention by researchers, with entire textbooks written and courses taught on the subject. For the purposes of our discussion, however, we will focus on aspects of probability that are relevant to the research process portrayed in this book.
301
What is Probability?
On the most basic conceptual level, probability may be defined as the likelihood of the occurrence of a particular outcome of an event given all possible outcomes. For the sake of our discussion, an event may be defined as an action performed by a person, and the possible results of this event are referred to as outcomes. For example, in the very simple event of flipping a coin, there are two possible outcomes: heads and tails. In the event of rolling a die (half a pair of dice), there are six possible outcomes: the numbers 1 through 6.
Expressed mathematically, the probability (p) of an outcome is the number of ways an outcome can occur divided by the total number of possible outcomes:
(6-1) p outcome = number of ways an outcome can occur total number of possible outcomes
Returning to the grocery store example presented earlier, if you randomly choose one of the five lines, the probability that this checkout line will be the fastest is equal to 1/5 or .20 or 20%. Because a single number (your one checkout line) is divided by a larger number (the five checkout lines from which you are choosing), the probability of an outcome is a percentage ranging from 0% to 100%; in this book, we will report probabilities using decimal places such as .00 and 1.00. At the extremes of the range of probabilities, a probability of .00 for an outcome means the outcome is certain not to occur, whereas a probability of 1.00 indicates the outcome is certain to occur.
Moving to a second example, for the event of flipping a coin, the probability of the outcome of “heads” is calculated by dividing the number of ways this outcome can occur (1) by the total number of possible outcomes (2: heads or tails). Using Formula 6-1, the probability of getting heads is 1/2 or .50; this can be represented by “p(heads) = .50.” Given that the probability of the other outcome (tails) is exactly the same, we may state that “p(tails) = .50.”
The simple event of flipping a coin illustrates two important aspects of probability. First, the sum of the probabilities of all possible outcomes of an event equals 1.00 (100%). In flipping a coin, p(heads) + p(tails) = .50 + .50 = 1.00. Second, according to the addition rule, the combined probability of mutually exclusive outcomes is the sum of their individual probabilities. Outcomes are “mutually exclusive” when they cannot occur at the same time. For example, in flipping a coin, we can get either heads or tails, but we cannot get both heads and tails. Because the two outcomes are mutually exclusive, the two probabilities can be added together, and we can say “the probability of getting either a heads or a tails from flipping a coin is .50 + .50, or 1.00.” Another example of the addition rule is determining the probability of rolling a 1, 2, 3, or 4 on a single roll of a die. Because these four outcomes are mutually exclusive, their combined probability is (1/6 + 1/6 + 1/6
302
+ 1/6) or 4/6 or 2/3 or .67; in other words, p(1 or 2 or 3 or 4) = .67.
303
Why is Probability Important to Researchers?
The previous section focused on mathematical characteristics of probability. However, conducting research is very different from waiting in grocery store lines, flipping coins, or rolling dice. How is probability a part of the research process? As we stated at the beginning of this chapter, one of the goals of research is to test hypotheses regarding what is believed to be true in the population. In most cases, however, collecting data from an entire population is difficult and impractical. Consequently, researchers base their conclusions about their hypotheses on data provided by a sample of a population. Because they have not collected data from the entire population, researchers must instead rely on probability to evaluate the data collected from their samples.
The concept of sampling error refers to differences between statistics calculated from a sample and statistics pertaining to the population from which the sample is drawn; these differences are attributed to random, chance factors. For example, assume the average height of adult males in the population is 5′ 9″. However, when we draw a random sample of males from the population in this sample, we calculate a mean height of 6′ 1″. We would attribute the difference between the mean of our sample and the mean of the population to sampling error. Given the many ways in which the members of a population may differ from each other, researchers cannot be certain that any particular sample is completely representative of the larger population.
In addition to the difference between a sample and a population, sampling error also refers to differences between samples drawn from the same population. For example, if we were to draw 10 random samples of adult males from the population, because the samples contain different people, we may calculate 10 different values of the mean height. These differences across samples may be defined as sampling error.
The possibility of sampling error exists in any study in which data are collected from a sample rather than the entire population. As a result, researchers can never be certain that their sample completely represents the population that is the basis or target of their research hypotheses. This implies that researchers cannot test their research hypotheses with absolute certainty but must instead rely on probability to evaluate the data collected from their samples.
304
Applying Probability to Normal Distributions
Research involves developing hypotheses about the relationship between variables in a population. As such, it is important to understand how probability may be applied to the distribution of values for these variables in the population. The principles of probability may be applied to distributions such as normal distributions and the standard normal distribution. In Chapter 5, for example, we asked questions regarding such things as the percentage of z-scores between z = –.85 and z = 1.95 and the percentage of SAT scores less than 660. As it turns out, these questions can be restated in terms of probability. For example, determining the percentage of z-scores between z = –.85 and z = 1.95 is the same as determining the probability a z-score is between z = –.85 and z = 1.95. Similarly, our finding in Chapter 5 that 94.52% of SAT scores are less than 660 implies there is a .9452 probability that any randomly selected SAT score will be less than 660.
The ability to apply the concepts of probability to normal distributions and the standard normal distribution is important to researchers because it enables researchers to evaluate data collected from samples. For instance, using our earlier example of the height of adult males, we could determine the probability of obtaining a mean of 6′ 1″ for a sample drawn from a population with a hypothesized mean of 5′ 9″. This probability could then be used to test a hypothesis regarding the height of adult males.
305
Applying Probability to Binomial Distributions
In addition to normal distributions and the standard normal distribution, which involve numeric variables measured at the interval or ratio level of measurement, probability may also be applied to variables measured at the nominal level of measurement, with values consisting of two or more distinct categories. For example, a binomial variable is a variable consisting of exactly two categories; examples of binomial variables include gender (male, female), test response (correct, incorrect), and jury verdict (guilty, not guilty).
When probability is applied to binomial variables, the probabilities of the two categories are labeled p and q. In our example of coin flipping, p = p(heads) and q = p(tails). The sum of the probabilities of the two categories of a binomial variable is equal to 1.00; when flipping a coin, the probability of heads (p) and tails (q) is both .50, which sum to 1.00.
If a coin is flipped once, the two possible outcomes (heads or tails) each have a probability of .50. But what if we flip a coin two times? To help us answer this question, let's first look at the list below, which provides all of the possible outcomes from two coin flips:
First Coin Flip Second Coin Flip
Heads Heads
Heads Tails
Tails Heads
Tails Tails
Focusing on the number of heads in two coin flips, there are four possible outcomes. The four possible outcomes have been organized into the distribution table in Table 6.1; this table is an example of a binomial distribution, which is the distribution of probabilities for a binomial variable. Looking at Table 6.1 and Figure 6.1, we see that in two coin flips, there is a .25 probability of getting 0 heads, a .50 probability of getting one head, and a .25 probability of getting two heads. Given that the outcomes in the event of two coin flips are mutually exclusive, the addition rule may be applied to this event such that the probabilities may be added together (.25 + .50 + .25), which sum to 1.00.
As a second example of a binomial distribution, let's consider the probabilities involved in a sample of four coin flips. Table 6.2(a) lists 16 possible outcomes that can occur, and Table 6.2(b) calculates the probability of the number of heads across these outcomes. Looking at these tables and at Figure 6.2, we see that the number of heads in four coin flips can range from 0 to 4, with their corresponding probabilities ranging from .06 to .38.
Table 6.1 Binomial Distribution, Number of Heads in Two Coin Flips Table 6.1 Binomial
306
Table 6.1 Binomial Distribution, Number of Heads in Two Coin Flips
# Heads f Probability
2 1 .25
1 2 .50
0 1 .25
Total 4 1.00
Figure 6.1 Binomial Distribution, Number of Heads in Two Coin Flips
We've introduced the binomial distribution to illustrate the probabilities of different possible outcomes for different samples of data. In the next section, we will introduce a research situation that involves applying probability to a binomial distribution to test a research hypothesis about a population.
Table 6.2 Possible Outcomes and Binomial Distribution Table, Four Coin Flips Table 6.2 Possible Outcomes and Binomial Distribution Table, Four Coin Flips
(a) Possible Outcomes
Outcome First Coin Flip
Second Coin Flip
Third Coin Flip
Fourth Coin Flip
1 Heads Heads Heads Heads
2 Heads Heads Heads Tails
3 Heads Heads Tails Heads
4 Heads Heads Tails Tails
5 Heads Tails Heads Heads
6 Heads Tails Heads Tails
307
7 Heads Tails Tails Heads
8 Heads Tails Tails Tails
9 Tails Heads Heads Heads
10 Tails Heads Heads Tails
11 Tails Heads Tails Heads
12 Tails Heads Tails Tails
13 Tails Tails Heads Heads
14 Tails Tails Heads Tails
15 Tails Tails Tails Heads
16 Tails Tails Tails Tails Table 6.2 Possible Outcomes and
Binomial Distribution Table, Four Coin Flips
(b) Binomial Distribution Table
# Heads f Probability
4 1 .06
3 4 .25
2 6 .38
1 4 .25
0 1 .06
Total 16 1.00
308
Learning Check 1: Reviewing what you've Learned So Far
1. Review questions a. What is probability (conceptually and mathematically)? b. Why is probability important to researchers? c. What is the addition rule of probability? d. What is the relationship between sampling error and probability? e. How can probability be applied to distributions? f. What is the difference between normal distributions and binomial distributions?
2. Here is a set of scores: 7 2 6 1 9 3 5 2. If we were to randomly select one of these scores, what is that probability this score will be …
a. equal to 6? b. equal to 2? c. greater than 4? d. less than 8? e. an even number?
3. Using the standard normal distribution and the normal curve table, what is the probability of a z- score …
a. greater than 1.25? b. less than .40? c. greater than –.50? d. less than −1.33? e. between .50 and 1.50?
4. According to the National Center for Health Statistics, in 2005 the average birth weight of a newborn baby was approximately normally distributed with a mean of 119 ounces (7 pounds, 7 ounces) and a standard deviation of 19 ounces. What is the probability a newborn baby will weigh …
a. more than 128 ounces (8 pounds)? b. less than 112 ounces (7 pounds)? c. more than 144 ounces (9 pounds)? d. less than 124 ounces (7 pounds, 12 ounces)? e. between 96 ounces (6 pounds) and 136 ounces (8.5 pounds)?
309
6.2 Example: Making Heads or Tails of the Super Bowl
The Super Bowl, the game that determines the champion of the National Football League, has become an integral part of American society. Each year, the Super Bowl is watched by more people than any other televised event. Fans, gamblers, and curious onlookers spend countless hours before the game debating the merits of the two competing teams in attempting to predict which team will win. Although some factors that influence the outcome of a football game (such as physical ability and strategy) are within the control of the players and coaches, other factors such as weather and injuries also play a role. As a result, the winner of the Super Bowl cannot be predicted with 100% certainty. It is possible, however, that understanding and analyzing the potential impact of different factors may increase the probability of picking the winning team.
Figure 6.2 Binomial Distribution, Number of Heads in Four Coin Flips
A coin flip is held at the beginning of every Super Bowl game to determine which team gets to decide who will receive and possess the ball at the start of each half of the game. Assuming this decision plays a role in determining who will win the game, it's plausible that the team that wins the pregame coin flip has a greater probability of winning the game. Let's imagine we decide to conduct a research study to test the relationship between winning the pregame coin flip and winning the game.
As discussed in Chapter 1, the first step in the research process is to develop a research hypothesis, which is an expected outcome or relationship between variables. Imagine that, based on an examination of the literature regarding sporting events, we believe that the team that wins the coin flip has a greater likelihood of scoring first than the team that loses the coin flip. As history has shown that the team scoring first is more likely to win the game, our research hypothesis might be stated as the following: “It is hypothesized that teams are more likely to win the Super Bowl if they win the pregame coin flip than if they
310
lose the coin flip.”
The second step in the research process is to collect data, which involves defining the target population, drawing a sample from this population, and determining what data to collect from the sample and how they will be collected. In the case of the Super Bowl example, the relevant population consists of the teams that have won the pregame coin flip. The example discussed below is based on a sample of the teams competing in 12 Super Bowls between 2002 and 2013. In this example, the relevant data to be collected would consist of determining whether the team that won the coin flip went on to win the game. Because the National Football League keeps records of many aspects of each Super Bowl, this information can be easily obtained by examining books and documents produced by the league. For each of the 12 Super Bowls between 2002 and 2013, whether the team that won the coin flip went on to win or lose the game is provided in Table 6.3.
The third step in the research process is to analyze the data that have been collected. As we saw in Chapter 2, it is useful to examine data before proceeding with any statistical analyses. A frequency distribution table has been created in Table 6.4 for the sample of 12 Super Bowl coin flip winners. Looking at the table, the data for the Super Bowl example may be summarized as follows: “Of the 12 Super Bowls from 2002 to 2013, the team that won the coin flip won five (42%) of the games.” (Simple examples like the current one do not require calculating descriptive statistics such as the mean and standard deviation discussed in Chapters 3 and 4; research situations requiring the calculation of descriptive statistics will be discussed starting in Chapter 7.)
Table 6.3 Examples of Research Questions and Hypotheses Table 6.3 Examples of Research
Questions and Hypotheses
Super Bowl Outcome of Game
2002 Lose
2003 Win
2004 Lose
2005 Lose
2006 Lose
2007 Lose
2008 Win
2009 Lose
2010 Win
2011 Win
2012 Lose
311
2013 Win
Table 6.4 Frequency Distribution Table, Outcome of Super Bowl for Team Winning Coin Flip, 2002–2013
Table 6.4 Frequency Distribution Table, Outcome of Super Bowl for Team Winning Coin Flip, 2002–
2013
Outcome of Game f %
Win 5 42%
Lose 7 58%
Total 12 100%
The fourth step in the research process is to draw a conclusion regarding whether the research hypothesis has been supported based on the results of the statistical analyses. Does finding that five (42%) of the 12 teams that won the coin flip in the Super Bowls between 2002 and 2013 went on to win the game enable us to draw a conclusion regarding whether winning the coin flip affects a team's chance of winning the Super Bowl? No—because in this example we have collected data from a sample rather than the entire population, we need to conduct statistical analyses specifically designed to test research hypotheses by evaluating data collected from samples of the population. The next section introduces the steps in hypothesis testing; as we will see, hypothesis testing relies heavily on probability.
312
6.3 Introduction to Hypothesis Testing
Conducting statistical analyses to test research hypotheses consists of a number of steps, steps that will be used in many of the remaining chapters of this book. The purpose of this section is to use the Super Bowl example to illustrate why and how these steps are completed. The steps in hypothesis testing may be stated and summarized as follows:
State the null and alternative hypotheses (H0 and H1).
State two mutually exclusive conclusions regarding the existence of a hypothesized change, difference, or relationship in the population.
Make a decision about the null hypothesis.
Make a decision to reject or not reject the null hypothesis based on the probability of a calculated value of a statistic.
Draw a conclusion from the analysis.
Based on the decision about the null hypothesis, draw a conclusion about the hypothesized change, difference, or relationship.
Relate the result of the analysis to the research hypothesis.
Interpret the results of the statistical analysis in terms of whether it supports or does not support the study's research hypothesis.
Each of these steps is discussed in detail in the sections below.
313
State the Null and Alternative Hypotheses (H0 and H1)
The first step in hypothesis testing is to represent two mutually exclusive conclusions that can be made based on the results of analyzing data collected from a sample of the population: (1) a conclusion regarding what is believed to exist in the population and (2) a mutually exclusive alternative to this conclusion. Ultimately, as a result of conducting statistical analyses, a decision is made about which of these conclusions is considered to be true.
Each of these two conclusions is represented by a statistical hypothesis: a statement about an expected outcome or relationship involving population parameters. As stated in Chapter 3, a parameter, such as the population mean μ, is a numeric characteristic of a population. Statistical hypotheses differ from research hypotheses in that research hypotheses involve concepts and are expressed using words, whereas statistical hypotheses involve mathematical terms and are expressed with numbers. Below we define the two statistical hypotheses and develop hypotheses applicable to our example of the number of heads in 12 Super Bowl coin flips.
The Null Hypothesis (H0)
The first statistical hypothesis is the null hypothesis (H0) (pronounced “H sub zero,” where sub stands for “subscript”), defined as a statistical hypothesis that states that a hypothesized change, difference, or relationship among groups or variables does not exist in the population. The word null and the subscript 0 are used to designate the absence or lack of existence of a change, difference, or relationship.
In the Super Bowl example, the null hypothesis represents the conclusion that winning the coin flip does not affect a team's chances of winning the Super Bowl. Stating the null hypothesis requires representing this conclusion mathematically for our sample of 12 Super Bowls. If we were to draw random samples of 12 games from the population of Super Bowls, on average, how many of the 12 teams that won the coin flip should win the game? If winning the coin flip does not affect a team's chances of winning the game, assuming the probability of winning the game is .50, we could expect an average of half (or 6) of the 12 teams to win the game. Therefore, for the population of 12 Super Bowl games, the null hypothesis would be stated the following way: H 0 : μ = 6
This null hypothesis states that, in the population of 12 Super Bowls, the mean number of wins for the teams winning the coin flip is equal to 6.
Before moving on, we want to emphasize that statistical hypotheses state an expected
314
outcome or relationship involving population parameters rather than sample statistics. For example, the null hypothesis in the Super Bowl example does not state that our one sample of 12 Super Bowl coin flip winners will produce six winning teams. Instead, the null hypothesis asserts that in the population of 12 Super Bowl coin flip winners, a population based on an infinite number of random samples, the mean number of winning teams (μ) will be equal to 6.
The Alternative Hypothesis (H1)
The alternative hypothesis (H1) (pronounced “H sub one”) is the statistical hypothesis that a hypothesized change, difference, or relationship among groups or variables does exist in the population. The alternative hypothesis provides a mutually exclusive alternative to the null hypothesis in that only one of the two hypotheses may be true.
In the Super Bowl example, the alternative hypothesis represents the conclusion that winning the coin flip does in fact affect a team's chances of winning the Super Bowl. One way to represent this conclusion is to state that winning the coin flip makes the probability of winning the game not equal to .50; if so, the mean number of victories by the teams that wins the coin flip in the population of 12 Super Bowls will not be equal to 6. Consequently, the alternative hypothesis could be stated as follows: H 1 : μ ≠ 6
Therefore, the null hypothesis that μ is equal to 6 will be rejected if the number of wins is either less than or greater than the population mean of 6.
Because the research hypothesis in the Super Bowl example implies that winning the coin flip increases a team's chances of winning the game in the population of 12 Super Bowls, the number of wins by teams winning the coin flip should be greater than the population mean of 6. However, we have chosen to state the alternative hypothesis as H1: μ ≠ 6, which allows for the possibility that winning the coin flip could either increase or decrease a team's chance of winning the game. Later in this chapter, we illustrate choices available to researchers regarding how the alternative hypothesis may be stated.
315
Make a Decision about the Null Hypothesis
Having stated two mutually exclusive statistical hypotheses, the next step in hypothesis testing is to decide which of the two statistical hypotheses is supported by the data. More specifically, we will decide whether or not to reject the null hypothesis (which implies that the change, difference or relationship does not exist) in favor of the alternative hypothesis.
We could make a decision about the null hypothesis simply by comparing the value of the statistic calculated on our sample with that of the parameter specified in the null hypothesis. In the Super Bowl example, the statistic of five winning teams is not the same as the parameter μ = 6. But should this observation necessarily lead to the conclusion that the alternative hypothesis is true? The answer to this question is no; this is because we have learned from our discussion of sampling error that a sample may differ from the population parameter because of random, chance factors.
Because of the existence of sampling error, it would be inappropriate to require every sample of 12 Super Bowl teams to contain precisely 6 winning teams for the null hypothesis to be true. Instead, the decision about the null hypothesis is made based on the likelihood, or probability, of obtaining the value of the statistic calculated from the sample of data. In essence, if the statistic has a low probability of occurring under the assumption that the null hypothesis is true, we will conclude that the change, difference, or relationship found in the sample is not due to random, chance factors. Subsequently, the decision will be made to reject the null hypothesis in favor of the alternative hypothesis.
Defining a “Low” Probability of a Statistic: Setting Alpha (α) In hypothesis testing, it is necessary to state how “low” the probability of a statistic must be to make a decision to reject the null hypothesis. This probability is represented by the Greek letter alpha (α), defined as the probability of a statistic used to make a decision whether to reject the null hypothesis.
In many academic disciplines, it is customary to define a “low” probability at .05 (α = .05). A strategy based on this probability would be as follows:
The null hypothesis will be rejected if the probability of the statistic is less than .05.
In the Super Bowl example, we need to determine the probability of winning 5 out of 12 games, assuming the average is equal to 6. If this probability is less than .05, we will reject the null hypothesis and conclude that winning the coin flip does in fact affect a team's
316
chances of winning the Super Bowl. However, if this probability is not less than .05, we will not reject the null hypothesis and conclude that winning the coin flip does not affect a team's chances of winning the Super Bowl.
Although researchers traditionally set alpha at .05, this is by no means used by every researcher in every situation. Reading reports of research findings in journal articles, we may encounter situations where a low probability has been defined as .10 (10%) or perhaps as .01 (1%). Different values for alpha may be chosen based on such things as the size of the sample or the researcher's preferred level of confidence that a hypothesized effect does in fact exist. The reasons for and consequences of using a particular value of alpha will be further discussed later in this chapter and in subsequent chapters.
Identifying the Values of the Statistic with a Low Probability
Once alpha has been set, the next step in making the decision about the null hypothesis is to identify values of the statistic that have the defined low probability of occurring. Calculating one of these values from the data we have collected will lead to the decision to reject the null hypothesis in favor of the alternative hypothesis.
Throughout this book, we will discuss a variety of statistics; the specific type of statistic calculated for a particular research situation is a function of the type of variable used in the analysis. The Super Bowl example uses a binomial variable (win, lose); earlier in this chapter, we determined the probabilities for the number of heads that can occur in two or four coin flips. For our example of N = 12, determining all of the possible outcomes and the probability of each outcome would be extremely time-consuming; fortunately, this information can be obtained by consulting existing tables of binomial distributions.
In the back of this book, Table 2 (Binomial Probabilities) presents a table of binomial probabilities for distributions of different sample sizes. Looking first at the columns of the table, these columns represent a wide range of probabilities of the two values of the variable. Consider the example of an item on a multiple-choice test that contains four alternatives (a through d); if the student were to randomly choose one of the four alternatives, the probability of answering the item correctly (p) is ¼ or .25. If, on the other hand, the item consisted of five alternatives (a through e), the probability of answering the item correctly (p) would be ⅕ or .20. In the Super Bowl example, because we have stated that the probability (p) of winning a game is .50, we will rely on the .50 column.
The rows of the binomial probabilities table represent different distributions of probabilities depending on the size of the sample (N). To illustrate how to read this table, let's return to our earlier example of two coin flips. Looking under the N column, move down to the number 2 (N = 2). Here, we find three rows, labeled 0, 1, and 2, that correspond to the possible number of heads that may result from two coin flips. Based on the assumption that the probability of heads is .50, move to the right until we reach the .50
317
column. Here, the numbers .2500, .5000, and .2500 correspond to the probabilities (calculated in Table 6.1) of the outcomes of 0 (.25), 1 (.50), and 2 (.25) heads, respectively.
For the Super Bowl example, we need to determine the number of wins in 12 games whose combined probability is low enough to lead to the decision to reject the null hypothesis. Looking at the Appendix, move down the N column until we reach N = 12. Here the 13 rows inform us of the possible outcomes for our example (0 to 12 wins). Assuming the null hypothesis is true and the probability (p) of winning a game is .50, move to the column labeled .50 to determine the probability for each outcome. These probabilities are illustrated in Figure 6.3.
Figure 6.3 Binomial Distribution, Number of Wins in 12 Super Bowls
Within hypothesis testing, distributions of values of a statistic such as the one in Figure 6.3 are divided into two parts or regions: the region of rejection and the region of non- rejection. The region of rejection represents the values of a statistic whose combined probability is low enough that obtaining one of these values results in the decision to reject the null hypothesis. The region of non-rejection represents the values of a statistic whose combined probability is high enough that obtaining one of these values results in the decision to not reject the null hypothesis. A critical value is the value of a statistic that separates the region of rejection from the region of non-rejection.
Figure 6.4 indicates the region of rejection, region of non-rejection, and critical values for a hypothetical distribution. If alpha is set at .05, the region of rejection contains the 5% of values of the statistic whose probability is low enough to lead to the rejection of the null hypothesis; the region of non-rejection contains the remaining 95% (100% – 5%) of the values of the statistic. Note that the shaded 5% region of rejection in Figure 6.4 consists of two parts: the 2½% of the distribution at each end of the distribution. The region of rejection is divided in this way when the alternative hypothesis includes the ≠ symbol (e.g., H1: μ ≠ 6). The ≠ symbol implies that the null hypothesis will be rejected if the value of the statistic is either sufficiently greater than or less than the population parameter;
318
consequently, the 5% region of rejection must be split into the two tails of the distribution. Later in this chapter, we will discuss different ways the alternative hypothesis may be stated.
For the Super Bowl example, we need to determine the number of wins in 12 games that lie in the region of rejection. First, let's look at the probabilities of the different number of wins at the lower end of the distribution. Starting from zero wins, the probabilities of zero, one, and two wins are added together below:
Figure 6.4 Regions of Rejection and Non-rejection for Alpha (α) = .05
# Wins Probability
0 .0002
1 .0029
2 .0161
Sum .0192
Note that the sum of these three probabilities, 1.92%, is less than the stated cutoff of 2½%. Therefore, the region of rejection at the lower end of the distribution consists of zero, one, and two wins, and the critical value at this end of the distribution is three wins.
Turning our attention to the upper end of the distribution, the probabilities of 12, 11, and 10 wins are added together as follows:
# Wins Probability
12 .0002
11 .0029
10 .0161
Sum .0192
As a result of these calculations, the region of rejection at the upper end of the distribution consists of 12, 11, and 10 wins, and the critical value at this end of the distribution is 9 wins.
319
Based on these observations, the two critical values for the Super Bowl example may be stated as follows: For α = .05 and N = 12 games, critical values = 3 wins and 9 wins.
The regions of rejection and non-rejection and the critical values for the Super Bowl example are illustrated in Figure 6.5.
Stating a Decision Rule
Now that the regions of rejection and non-rejection and the critical values have been identified, it is useful to provide a rule that explicitly states the logic used to make a decision about the null hypothesis. A decision rule is a rule that specifies the values of the statistic that result in the decision to reject or not reject the null hypothesis. The general decision rule in hypothesis testing may be stated as follows:
If the value of the statistic calculated for the sample lies beyond the critical values, reject the null hypothesis; otherwise, do not reject the null hypothesis.
Figure 6.5 Regions of Rejection and Non-rejection, Super Bowl Example (N = 12)
The decision rule for the Super Bowl example is stated below:
If the number of wins in 12 games is < 3 or > 9, reject H0; otherwise, do not reject H0.
This decision rule implies that the null hypothesis of μ = 6 will be rejected if the number of wins in 12 games is either less than 3 or more than 9, as any number of wins beyond 3 or 9 are located in the region of rejection. On the other hand, if the number of wins in 12
320
games is not less than 3 or more than 9, the null hypothesis will not be rejected because the number of wins falls in the region of non-rejection.
Calculating a Value of a Statistic
Once the decision rule for rejecting the null hypothesis has been stated, the next step in hypothesis testing is to conduct statistical analyses on the data collected from the sample. The purpose of these analyses is to calculate a value of a statistic that will be compared with the critical values to make a decision about the null hypothesis. These statistical analyses typically involve the calculation of inferential statistics, defined in Chapter 1 as statistical procedures used to test research hypotheses.
Many of the remaining chapters of this book will introduce different inferential statistical procedures; the specific inferential statistic calculated in each case will depend on the research hypothesis of interest, as well as the nature of the variables involved in the analysis. For example, a different statistical procedure will be used to compare the means of two groups (Chapter 9) versus comparing the means of three or more groups (Chapter 11).
Unlike future chapters, the Super Bowl example is unique in that we have already calculated the statistic relevant to the research hypothesis: the number of times the teams that won the coin flip also won the game (5 out of 12 teams). Due to the relative simplicity of this example, the purpose of which is to introduce the steps in hypothesis testing, no additional statistical analyses or calculations are necessary.
Making the Decision Whether to Reject the Null Hypothesis
Following the logic stated in the decision rule, the decision about the null hypothesis is made by comparing the value of the statistic calculated from the sample data with the critical values; if the value of the statistic exceeds a critical value, the null hypothesis will be rejected; otherwise, the null hypothesis will not be rejected. In the sample of 12 games in our Super Bowl example, there were five winning teams. By comparing this statistic with the critical values, the decision about the null hypothesis may be stated as the following:
Because 5 wins is neither less than 3 nor greater than 9, the decision is made to not reject the null hypothesis because the probability of winning 5 out of 12 games is greater than .05.
Below is a more concise way of stating the above decision:
5 wins is not < 3 or > 9; therefore, do not reject H0 (p > .05).
321
In this example, the null hypothesis is not rejected because 5 wins in 12 games falls in the region of non-rejection, meaning that the probability of 5 wins in 12 games is greater than the .05 alpha level (p > .05).
322
Draw a Conclusion from the Analysis
The next step in hypothesis testing is to draw a conclusion based on the decision to reject or not reject the null hypothesis. Whenever a statistical analysis is conducted, drawing as specific a conclusion as possible regarding the results of the analysis reduces the likelihood that others will misunderstand or misinterpret a study's findings. In the Super Bowl example, based on our analysis, we could conclude the following:
For the 12 teams that won the coin flips in the Super Bowls between 2002 and 2013, 5 of the 12 went on to win the game. Consequently, the null hypothesis of μ = 6 was not rejected because the probability of 5 wins in 12 games is greater than α = .05.
Stating our conclusion in this manner informs the reader about several critical aspects of the analysis:
The sample from whom data were collected: “the 12 teams that won the coin flips in the Super Bowls between 2002 and 2013” The value of the statistic calculated from the data: “5 of the 12 went on to win the game” The decision about the null hypothesis: “the null hypothesis of μ = 6 was not rejected” The probability of the value of the statistic: “the probability … is greater than α = .05”
In later chapters, as we discuss different types of inferential statistical procedures, we will expand on this discussion regarding how to draw conclusions from statistical analyses and communicate these conclusions to others.
323
Relate the Result of the Analysis to the Research Hypothesis
We are now prepared to relate the results of the analysis back to the original research hypothesis. More specifically, it is important to determine whether the results of a statistical analysis support or do not support the research hypothesis.
For the Super Bowl example, the research hypothesis was stated as follows: “Teams are more likely to win the Super Bowl if they win the pregame coin flip than if they lose the coin flip.” Do the results of our analysis support or not support this research hypothesis? Because in this analysis, the null hypothesis was not rejected, we could state the following:
The results of the statistical analysis do not support the research hypothesis that teams are more likely to win the Super Bowl if they win the pregame coin flip than if they lose the coin flip.
324
Summary of Steps in Hypothesis Testing
Using the Super Bowl example, the steps involved in hypothesis testing are summarized in Table 6.5. This process will be referred to and expanded on throughout the remainder of this book. Although the statistical procedures discussed in future chapters may seem very different from each other, because they share the same goal of testing hypotheses about populations based on data collected from samples, the steps used to test these hypotheses will remain the same.
Table 6.5 Summary of Steps in Hypothesis Testing, Super Bowl Example Table 6.5 Summary of Steps in Hypothesis Testing, Super Bowl Example
State the null and alternative hypotheses (H0 and H1).
H 0 : μ = 6 H 1 : μ ≠ 6 Make a decision about the null hypothesis.
Set alpha, identify the critical values, and state a decision rule. For α = .05, if the number of wins is < 3 or > 9, reject H0; otherwise, do not reject H0. Calculate a statistic: number of wins in 12 Super Bowls. 5 wins in 12 games Make a decision whether to reject the null hypothesis. 5 wins is not < 3 or > 9; therefore, do not reject H0 (p > .05). Draw a conclusion from the analysis.
For the 12 teams that won the coin flips in the Super Bowls between 2002 and 2013, 5 of the 12 went on to win the game. Consequently, the null hypothesis of μ = 6 was not rejected because the probability of 5 wins in 12 games is greater than α = .05. Relate the result of the analysis to the research hypothesis.
The results of the statistical analysis do not support the research hypothesis that teams are more likely to win the Super Bowl if they win the pregame coin flip than if they lose the coin flip.
325
Learning Check 2: Reviewing what you've Learned So Far
1. Review questions a. What are the main steps in hypothesis testing? b. What are the differences between research hypotheses and statistical hypotheses? c. What are the null and the alternative hypotheses? d. What are the main steps in making the decision about the null hypothesis? e. What is the difference between the region of rejection and the region of non-rejection? f. Does the value of a statistic that exceeds the critical value have a low or high probability of
occurring? g. In stating the conclusion of a statistical analysis, what information is useful to provide?
2. Imagine you decide to conduct a study to test each of the following statements. For each one, state a null and alternative hypothesis (H0 and H1).
a. A survey of iPod owners conducted in 2005 estimated that the likelihood of having ones iPod completely stop working is 13.70%.
b. The average American adult female weighs 164 pounds. c. In 2004, the California Student Public Interest Group (CALPIRG) published a report
stating college students in California spend an average of $450 a semester on textbooks. d. One study estimated that the average college freshman owes $900 on his or her credit cards
(Arano, 2006). e. A nationwide survey of college students estimated that 50% of college students have cheated
at least once. 3. Returning to the Super Bowl example, imagine you draw a sample of 20 teams (N = 20) rather than
12. For this situation, a. State the null and alternative hypotheses (H0 and H1). b. Identify the critical values for α = .05 and state a decision rule. c. What decision would you make about the null hypothesis if you found 16 of the 20 teams
who won the coin flip also won the game? d. If you had found that 13 of the 20 teams that won the coin flip also won the game, would
this support or not support the research hypothesis? 4. A friend of yours claims to have extrasensory perception (ESP). You test this by creating a deck of
16 cards, half of which are red and the other half green. Sitting behind a partition, you hold up each card and ask her to guess the color. Assuming the probability of guessing the color correctly is 50%, you find she correctly guesses the color on 12 of the 16 cards.
a. State the null and alternative hypotheses (H0 and H1). b. Identify the critical values for α = .05 and state a decision rule. c. Make a decision about the null hypothesis. d. Does the result of your analysis support or not support your friends claim?
326
6.4 Issues Related to Hypothesis Testing: An Introduction
The previous section described the steps involved in conducting statistical analyses designed to test research hypotheses. Because the process of hypothesis testing is centered on collecting data from samples rather than populations and therefore relies on probability rather than certainty this process raises a number of issues and concerns that researchers must understand. The purpose of this section is to briefly introduce several issues related to hypothesis testing:
the issue of “proof” in hypothesis testing, errors that can be made in making the decision about the null hypothesis, and factors that affect the decision about the null hypothesis.
327
The Issue of “Proof” in Hypothesis Testing
In hypothesis testing, when the decision is made to reject the null hypothesis, does this “prove” the null hypothesis is false and the alternative hypothesis is true? On the other hand, if the null hypothesis is not rejected, does this “prove” the null hypothesis is true and the alternative hypothesis is false? And regardless of which decision is made, is it appropriate to conclude a research hypothesis has been “proven” to be true or false? The issue of “proof” as it pertains to hypothesis testing may be summarized as follows:
Hypothesis testing cannot provide proof (or lack of proof) for a research hypothesis; instead it can only provide support (or lack of support) for a research hypothesis.
Hypothesis testing cannot prove or disprove a research hypothesis because proof implies establishing the accuracy of a statement with absolute certainty. Proof cannot be provided for a research hypothesis because data have been collected from a sample rather than the entire population. Regardless of what decision is made about the null hypothesis, the concept of sampling error implies that the sample may not be representative of the population.
Because decisions and conclusions about hypotheses are based on probability, we can never be certain that the results of statistical analyses and the findings of research represent what would have happened if data had been collected from the entire population. Consequently, the best we can say is that our findings either support or do not support a research hypothesis, with the word support implying that our findings have increased or decreased the likelihood that a particular statement is accurate.
328
Errors in Decision Making
Collecting data from samples rather than populations is one reason why researchers exercise restraint in drawing conclusions about their hypotheses. There is another reason why relying on probability results in carefully worded conclusions:
It is possible that the decision made about the null hypothesis may either be correct or incorrect.
For example, when the decision is made to reject the null hypothesis, we can never completely eliminate the possibility that this was the wrong decision and that the hypothesized effect does not actually exist and the decision to reject the null hypothesis was based on chance factors. In flipping coins, even though there is less than a .0001 probability of getting 30 heads in 30 flips of a fair coin, the possibility of this happening does in fact exist.
It is also possible to make a second type of error: making the decision to not reject the null hypothesis when we should. In the case of the Super Bowl, for example, we made the decision to not reject the null hypothesis and conclude that winning the coin flip does not affect the outcome of the game. However, because we did not collect data from the entire population, we must recognize and accept the possibility that the hypothesized effect may in fact exist in the population of Super Bowl games even though we did not find it in our particular sample.
We introduce these two types of errors at this point in the book to further emphasize the importance of appropriately interpreting the results of statistical analyses. Chapter 10 discusses each type of error in greater detail, the factors that lead to these errors, and how the likelihood of making each of these errors can be lowered.
329
Factors Influencing the Decision about the Null Hypothesis
Although the process of making the decision to reject or not reject the null hypothesis may be fairly straightforward, a number of factors directly influence which of these two decisions are eventually made. This section introduces three factors that influence the decision about the null hypothesis:
the size of the sample, the value of alpha (the probability of the statistic used to make a decision about the null hypothesis), and the directionality of the alternative hypothesis.
Because each of these factors is under the control of the researcher, it is important to gain an understanding of their role in hypothesis testing.
Sample Size
The relationship between the size of the sample (i.e., the number of research participants from whom data are collected) and the decision about the null hypothesis may be summarized as follows:
The larger the sample size, the greater the likelihood of rejecting the null hypothesis.
Let's illustrate this relationship using the example of the Super Bowl. In the example described earlier, a sample size of 12 Super Bowls required relatively extreme outcomes (less than 3 or more than 9 wins) for us to decide to reject the null hypothesis. But what if the sample size had been larger? For a sample size of 30 games, how many games must the teams win to reject the null hypothesis that μ = 15? Again dividing the alpha level of .05 in half, Table 6.6 uses the binomial probabilities table in the back of the book to calculate a 2½% region of rejection at each of the two ends of the distribution. (Although some of the probabilities are listed as .0000, this does not mean the probability is equal to zero but rather that it is less than .0001.) These calculations identify the critical values to be 10 wins or 20 wins. Therefore, for N = 30, the decision rule is, “If the number of wins is < 10 or > 20, reject H0; otherwise, do not reject H0.”
Table 6.6 Calculation of Critical Values for Binomial Distribution of N = 30 Table 6.6 Calculation of Critical Values for Binomial
Distribution of N = 30
Lower End of Distribution Upper End of Distribution
330
# Wins Probability # Wins Probability
0 .0000 30 .0000
1 .0000 29 .0000
2 .0000 28 .0000
3 .0000 27 .0000
4 .0000 26 .0000
5 .0001 25 .0001
6 .0006 24 .0006
7 .0019 23 .0019
8 .0055 22 .0055
9 .0133 21 .0133
Sum .0214 Sum .0214
The critical values for samples of 12 and 30 games are illustrated in Figure 6.6. Comparing the vertical lines that separate the regions of rejection and non-rejection, the critical values for N = 30 are closer to the center of the distribution than the critical values for N = 12. For example, when N = 12 (Figure 6.6(a)), the teams must win at least 10 of the 12 games (10/12 = 83%) to reject the null hypothesis. However, the teams need only win at least 21 of the 30 games (21/30 = 70%) to reject the null hypothesis when N = 30 (Figure 6.6(b)). As the size of the sample increases, less extreme values of the statistic are needed to reject the null hypothesis. Consequently, the larger the sample size, the greater the likelihood of rejecting the null hypothesis.
Alpha (α) The null hypothesis is rejected when the calculated value of a statistic has a “low” probability of occurring; this probability is represented by alpha (α). Although it is traditional in many academic disciplines to use an alpha level of .05, other values of alpha may be chosen. The relationship between alpha and the decision about the null hypothesis may be stated the following way:
The smaller the value of alpha, the lower the likelihood of rejecting the null hypothesis.
Figure 6.6 Critical Values, Binomial Distribution for N = 12 and N = 30
331
As we mentioned earlier in this chapter, researchers may require the probability of a statistic to be less than .01 (rather than .05) to reject the null hypothesis. One reason why alpha may be set to .01 is to require stronger, more convincing evidence before reaching the conclusion that the proposed change, difference, or relationship exists. This is analogous to what may happen in jury trials, in which a jury may be required to reach a unanimous decision rather than a two-thirds majority to find someone guilty of committing a crime.
Let's return to the Super Bowl example to illustrate the impact of using α = .01 rather than .05 on the decision about the null hypothesis. For a sample of 30 games, Table 6.7 calculates the critical values for α = .01 (the .01 alpha level has been divided into two halves of .005). Based on Table 6.7, for α = .01, the decision rule may be stated as follows: “If the number of wins is < 8 or > 22, reject H0; otherwise, do not reject H0.”
The critical values for α = .05 and α = .01 for N = 30 are illustrated in Figure 6.7. Looking at this figure, we see that the critical values are farther from the center of the distribution for α = .01 (Figure 6.7(b)) than for α = .05 (Figure 6.7(a)). This indicates that more extreme values of the statistic are needed to reject the null hypothesis when a “low”
332
probability is defined as .01 rather than .05. Consequently, it is more difficult to reject the null hypothesis when alpha is set at .01 rather than .05.
Directionality of the Alternative Hypothesis
Table 6.7 Calculation of Critical Values for Binomial Distribution of N = 30, Alpha = .01 Table 6.7 Calculation of Critical Values for Binomial
Distribution of N = 30, Alpha = .01
Lower End of Distribution Upper End of Distribution
# Wins Probability # Wins Probability
0 .0000 30 .0000
1 .0000 29 .0000
2 .0000 28 .0000
3 .0000 27 .0000
4 .0000 26 .0000
5 .0001 25 .0001
6 .0006 24 .0006
7 .0019 23 .0019
Sum .0026 Sum .0026
Hypothesis testing begins by stating two mutually exclusive statistical hypotheses: a null hypothesis (H0) and an alternative hypothesis (H1) that states that a hypothesized change, difference, or relationship in the population does not exist or does exist, respectively. As it turns out, the alternative hypothesis may be stated in one of two ways: non-directional or directional. This section describes the following relationship between the directionality of the alternative hypothesis and the decision to reject the null hypothesis:
Figure 6.7 Critical Values, Binomial Distribution, α =.05 and α = .01 (N = 30)
333
There is a greater likelihood of rejecting the null hypothesis when the alternative hypothesis is directional than non-directional.
In the Super Bowl example, the alternative hypothesis was stated as H1: μ ≠ 6. This is an example of a non-directional alternative hypothesis, which is an alternative hypothesis that does not indicate the direction of the change, difference, or relationship between groups or variables. By including the ≠ symbol, this alternative hypothesis implies that the null hypothesis is rejected if the number of wins is either less than or greater than the population mean of 6.
A researcher may instead choose to use a directional alternative hypothesis, an alternative hypothesis that indicates the direction of the change, difference, or relationship by including the > or < (greater than or less than) symbol. Returning to the Super Bowl example, suppose we had reason to believe that winning the coin flip can increase (but not decrease) a team's chances of winning the game. If so, we may have stated the alternative hypothesis as H1: μ > 6, which implied that the null hypothesis is rejected only if the
334
number of wins is greater than the population mean of 6.
In terms of deciding which type of alternative hypothesis to use, researchers often choose non-directional hypotheses to allow for the possibility of a change or difference in both directions from the population mean μ. In the Super Bowl example, even though the research hypothesis predicts the number of wins by teams winning the coin flip should be greater than the population mean of 6, we have chosen to use a non-directional alternative hypothesis to allow for the possibility that winning the coin flip could either increase or decrease a team's chance of winning the game.
Let's return to the Super Bowl example of N = 30 to demonstrate the implications of using a directional alternative hypothesis. Setting alpha at .05, using the non-directional alternative hypothesis of H1: μ ≠ 15, we identified critical values of 9 and 20. Because its region of rejection is located in both ends (or tails) of the distribution, a non-directional alternative hypothesis is often referred to as “two-tailed.” But suppose we decide to state the alternative hypothesis as H1: μ > 15. When the alternative hypothesis is directional, the region of rejection and critical value are located at only one end of the distribution (“one- tailed”). Because H1: μ > 15 includes the “>” symbol, the critical value is located at the upper end of the distribution—in Table 6.8, we find this critical value to be equal to 19. As such, the decision rule may be stated as follows: “If the number of wins is > 19, reject H0; otherwise, do not reject H0.”
Figure 6.8 illustrates the difference between a non-directional (two-tailed) and a directional (one-tailed) alternative hypothesis for the N = 30 Super Bowl example. Comparing the 5% region of rejection in Figure 6.8(b) with the 2½% region on the right end of the distribution in Figure 6.8(a), we see that the region of rejection is slightly larger when the alternative hypothesis is directional rather than non-directional. That illustrates that a less extreme result will lead to rejection of the null hypothesis when a directional rather than a non-directional alternative hypothesis is chosen.
335
6.5 Looking Ahead
The majority of this chapter focused on describing the steps, logic, and process of conducting statistical analyses to test research hypotheses; in doing so, critical concepts were defined and illustrated. It is important to gain an understanding of these steps and concepts as they will be repeated throughout many of the remaining chapters of this book. We also defined and emphasized the role of probability in conducting and interpreting the results of statistical analyses, ending the chapter with a brief introduction to issues that arise in testing hypotheses. Given the numerous steps and concepts that are part of hypothesis testing, we chose to minimize mathematical calculations in this chapter by using the relatively simple example of a binomial variable. The next chapter will be the first of several chapters that discuss testing research hypotheses by conducting statistical procedures that involve calculating and comparing sample means with population means. Although the mathematical calculations in future chapters may become increasingly complicated, keep in mind that the steps used in hypothesis testing will remain the same.
Figure 6.8 Critical Values, Binomial Distribution, Non-Directional and Directional Alternative Hypothesis (N = 30)
336
Table 6.8 Calculation of Critical Values for Binomial Distribution of N = 30, Directional Alternative Hypothesis (H1: μ > 15)
Table 6.8 Calculation of Critical Values for
Binomial Distribution of N = 30, Directional Alternative Hypothesis
(H1: μ > 15) # Wins Probability
30 .0000
29 .0000
28 .0000
27 .0000
26 .0000
25 .0001
24 .0006
23 .0019
22 .0055
21 .0133
20 .0280
Sum .0494
337
Learning Check 3: Reviewing what you've Learned So Far
1. Review questions a. Within hypothesis testing, why do we say a research hypothesis has been “supported” rather
than “proved”? b. What are the two types of errors you can make regarding the two statistical hypotheses? c. What are the three factors discussed in this chapter that affect the decision about the null
hypothesis? In what specific ways does each of these influence this decision? d. What is the difference between a directional (one-tailed) and a non-directional (two-tailed)
alternative hypothesis? Under what research situations might you use one versus the other? 2. Using the Super Bowl example, imagine you draw the following three samples. Use the binomial
probabilities table to determine the critical values (assume the alternative hypothesis is non- directional [two-tailed] and alpha = .05). For each sample, also calculate the percentage of wins needed to reject the null hypothesis at the upper end of the distribution.
a. N = 14 b. N = 16 c. N = 22
3. Returning to the Super Bowl example, imagine you draw the following three samples. Use the binomial probabilities table to determine the critical values for α = .05 and α = .01 (assume the alternative hypothesis is non-directional [two-tailed]). For each sample, also calculate the percentage of wins needed to reject the null hypothesis at the upper end of the distribution.
a. N = 10 b. N = 16 c. N = 24
4. Returning to the Super Bowl example, imagine you draw the following three samples. Use the binomial probabilities table to determine the critical values for a non-directional alternative hypothesis and a directional alternative hypothesis (assume alpha = .05).
a. N = 10 (H1: μ ≠ 5 vs. H1: μ > 5) b. N = 18 (H1: μ ≠ 9 vs. H1: μ < 9) c. N = 26 (H1: μ ≠ 13 vs. H1: μ > 13)
338
6.6 Summary
Conceptually, probability may be defined as the likelihood of occurrence of a particular outcome of an event given all possible outcomes. Mathematically, the probability (p) of an outcome is the number of ways the outcome can occur divided by the total number of possible outcomes. Two important aspects of probability are that the sum of the probabilities of all possible outcomes of an event equals 1.00 (100%) and that the combined probability of mutually exclusive outcomes is the sum of their individual probabilities. This second observation is known as the addition rule.
Researchers cannot be certain that their sample is completely representative of the larger population. Sampling error refers to differences between statistics calculated from a sample and statistics pertaining to the population from which the sample is drawn that are due to random, chance factors. Because of sampling error, researchers cannot test research hypotheses with absolute certainty and instead rely on probability to evaluate the data collected from their samples.
Probability may be applied to distributions such as normal distributions, the standard normal distribution, and binomial distributions (distributions of probabilities for variables consisting of exactly two categories) to determine the probability of a particular score or outcome.
Conducting statistical analyses to test research hypotheses consists of several steps. The first step is to state two statistical hypotheses (i.e., statements regarding expected outcomes or relationships involving population parameters). The first statistical hypothesis is the null hypothesis (H0), which states that in the population, there exists no change, difference, or relationship among groups or variables. The second statistical hypothesis, the alternative hypothesis (H1), states that the hypothesized change, difference, or relationship among groups or variables does exist in the population.
The second step in hypothesis testing is to make a decision about which of the two statistical hypotheses, the null or the alternative, is believed to be true. This step requires several steps of its own. First, alpha (α) (i.e., the probability of a statistic used to make a decision whether to reject the null hypothesis) is stated. In many academic disciplines, it is customary to set an alpha level of .05 or 5% (α = .05). Setting alpha divides the distribution of values of a statistic into two regions: the region of rejection (values of the statistic whose probability is low enough to lead to the decision to reject the null hypothesis) and the region of non-rejection (values of the statistic whose probability is high enough to lead to the decision to not reject the null hypothesis). A value of the statistic that separates the regions of rejection and non-rejectionis known as a critical value. Second, a decision rule is stated specifying the values of the statistic that result in the decision to reject
339
the null hypothesis. Third, statistical analyses are conducted on the data collected from the sample to calculate a value of a statistic. Fourth, the decision whether to reject the null hypothesis is then made by comparing the value of the statistic calculated from the sample with the critical values stated in the decision rule. If the value of the statistic exceeds a critical value, the null hypothesis will be rejected; otherwise, the null hypothesis will not be rejected.
The third step in hypothesis testing is to draw a conclusion based on the decision to reject or not reject the null hypothesis, which typically involves communicating several aspects of the analysis, including the sample from whom data were collected, the value of the statistic calculated from the data, the decision about the null hypothesis, and the probability of the value of the statistic.
The fourth step in hypothesis testing is to relate the results of the analysis back to the original research hypothesis, more specifically, to determine whether the results of the analysis support or do not support the research hypothesis.
Because hypothesis testing is centered on probability rather than certainty, it raises a number of issues and concerns. First, hypothesis testing cannot provide “proof” (or a lack of proof) for a research hypothesis; instead it can only provide “support” (or a lack of support) for a research hypothesis. Second, it is possible that the decision made about the null hypothesis may either be correct or incorrect. When the decision is made to reject the null hypothesis, we can never completely eliminate the possibility that the hypothesized effect does not actually exist and the decision to reject the null hypothesis was based on chance factors. It is also possible to make a second type of error: making the decision to not reject the null hypothesis when we should. Third, several factors (sample size, alpha, and the directionality of the alternative hypothesis) under the control of the researcher influence the decision about the null hypothesis: The larger the sample size, the greater will be the likelihood of rejecting the null hypothesis; the smaller the value of alpha, the lower will be the likelihood of rejecting the null hypothesis; and there is a greater likelihood of rejecting the null hypothesis using a directional (one-tailed) alternative hypothesis (which predicts the direction of the change, difference, or relationship) than using a non-directional (two- tailed) alternative hypothesis (which does not predict the direction of the change, difference, or relationship).
340
6.7 Important Terms
probability (p. 176) addition rule (p. 176) sampling error (p. 177) binomial variable (p. 178) binomial distribution (p. 178) statistical hypothesis (p. 184) null hypothesis (H0) (p. 185) alternative hypothesis (H1) (p. 185) alpha (α) (p. 186) region of rejection (p. 187) region of non-rejection (p. 188) critical value (p. 188) decision rule (p. 190) non-directional (two-tailed) alternative hypothesis (p. 201) directional (one-tailed) alternative hypothesis (p. 201)
341
6.8 Formulas Introduced in this Chapter
342
Probability
(6-1) p outcome = number of ways an outcome can occur total number of possible outcomes
343
6.9 Exercises
1. You have collected the following data: 3 8 2 6 11
If you place these five numbers in a bag and randomly select one, what is the probability the number (X) will be …
a. equal to 6? b. less than 11? c. greater than 3? d. greater than 2 but less than 11?
2. You have collected the following data: 6 4 3 7 4 2 6 5 7 4
If you randomly select one of these 10 numbers, what is the probability the number (X) will be …
a. equal to 4? b. equal to 7? c. less than 5? d. greater than 2? e. greater than 4 but less than 7?
3. A lottery contains 500 tickets. In this lottery, there are 25 prizes of $1, 10 prizes of $5, and 5 prizes of $25. What is the probability of …
a. winning nothing ($0)? b. winning $25? c. winning more than $1?
4. (This example was introduced in Chapter 2.) A friend of yours asks 20 people to rate a movie using a 1- to 5-star rating: the higher the number of stars, the higher the recommendation. Their ratings are listed below:
Person # Stars
1 ***
2 *****
3 **
4 ****
5 ***
6 ***
7 ****
8 **
344
9 ***
10 *****
11 *
12 ****
13 ****
14 ***
15 ***
16 **
17 ***
18 **
19 ****
20 ***
a. What is the probability of this movie receiving three stars? b. What is the probability of four or more stars?
5. For the standard normal distribution, what is the probability of having a z-score … a. greater than 1.67 (z > 1.67)? b. greater than −1.67 (z > −1.67)? c. less than .75 (z < .75)? d. less than −2.75 (z < −2.75)? e. between 1.00 and 1.25 (1.00 < z < 1.25)? f. between –.25 and .25 (–.25 < z < .25)?
6. According to the test's publishers (www.act.org), scores on the ACT college entrance examination for students graduating in 2001 were normally distributed, with μ = 21 and σ = 5 (scores can range from 1-36).
a. What is the probability of having a score greater than 30? b. What is the probability of having a score greater than 20? c. What is the probability of having a score less than 17? d. What is the probability of having a score less than 23? e. What is the probability of having a score between 18 and 25?
7. A bottling company uses a machine to fill 16-ounce bottles with orange juice. The company finds that the standard deviation of the amount of juice in these bottles is equal to 1 4 ounce (σ = .25). What is the probability that a single bottle will contain …
a. less than 15.60 ounces? b. more than 16.30 ounces? c. between 16.10 and 16.60 ounces?
8. The length of time, in days, of pregnancy in healthy women is approximately normally distributed, with μ = 280 days and σ = 10 days. What is the probability a
345
woman will … a. give birth more than 1 week after her expected due date? b. give birth more than 2 weeks before her expected due date? c. give birth between 3 days before and 5 days after her expected due date?
9. (This example was introduced in Chapter 5.) How many base hits can a baseball team expect to get in a game? Frohlich (1994) recorded the number of hits by the 28 major league baseball teams for all of the games played in 1989 to 1993 (each team plays 162 games a year). He found that the number of hits the teams made in the games was normally distributed, with a mean of 8.72 and a standard deviation of 1.10.
a. What is the probability of a team getting more than 7 hits? b. What is the probability of a team getting less than 10 hits? c. What is the probability of a team getting less than 6 hits? d. What is the probability of a team getting between 8 and 11 hits?
NOTE: Exercises 10 to 15 use the binomial probabilities table.
10. Assuming a coin is fair, in a sample of 6 coin flips, what is the probability of getting …
a. 3 heads? b. less than 4 heads? c. more than 2 heads and less than 5 heads?
11. For a family with 5 children, assuming the probability of having a boy and having a girl are both .50 (50%), what is the probability of having …
a. 0 boys? b. 2 boys? c. more than 3 boys?
12. Assume an equal number of people prefer the two most popular brands of cola. Under this assumption, what is the probability of the following claim being true: “9 out of 10 people prefer Brand X over Brand Y”? 13. In each of four political races, Democrats are believed to have a 60% chance of winning. If so, what is the probability that Democrats will win …
a. none of the elections? b. at least one election? c. the majority of the elections?
14. A store advertises that there is a 90% chance their equipment will be trouble-free for a year. If you buy 6 of their products, what is the probability that …
a. all of the products will be trouble-free? b. half of the products will be trouble-free?
15. Assuming that 35% of all marriages end in divorce, if you encountered 8 adult men, what is the probability that all of them are still married? 16. For each of the following situations, state the two competing hypotheses to be
346
tested using words rather than mathematical symbols or formulas. a. A company designs a program aimed at helping people stop smoking. They
design a study aimed at testing the program. b. A study wished to examine whether using seat belts affected the severity of
injuries sustained by children in automobile accidents (Osberg & Di Scala, 1992).
c. A researcher hypothesizes that the more drivers use cellular phones, the greater the likelihood of getting into a traffic accident.
d. “It was hypothesized that … infants who spent greater amounts of time in center-based care would demonstrate more advanced exploratory behaviors than infants who did not spend as much time in center-based care” (Schuetze et al., 1999, p. 269).
e. “It is expected that achievement motivation will be a positive predictor of academic success” (Busato et al., 2000, p. 1060).
f. “The purpose of our study was to gain a better understanding of the relationship between social functioning and problem drinking…. We predicted that problem drinkers would endorse more social deficits than nonproblem drinkers” (Lewis & O'Neill, 2000, pp. 295–296).
17. For each of the following situations, state the two competing hypotheses to be tested using words rather than mathematical symbols or formulas.
a. Students who are taught effective learning skills will perform better on tests than students offered incentives to do well.
b. Increased use of Internet bulletin boards is associated with lower levels of television viewing.
c. Violent behavior in children may be reduced by teaching them conflict resolution skills.
d. The higher a person scores on the Graduate Management Admissions Test, the more likely the person is to succeed in graduate business school.
e. The more time children spend watching television, the more they will express a preference for unhealthy foods.
18. For each of the following situations, state a null hypothesis (H0) and a non- directional (two-tailed) alternative hypothesis (H1).
a. A bank advertises that customers never have to wait more than 5 minutes in line at its branches. A researcher standing in line decides to test this advertisement.
b. The average life expectancy in this country is 78 years. A researcher wishes to see whether students, when asked how long they expect to live, give estimates different from the actual life expectancy.
c. The American Statistical Association publishes a monthly magazine sent to all of its members. Not too long ago, an article was written asking, “How many chocolate chips are there in a bag of Chips Ahoy cookies?” (Warner & Rutledge, 1999). Nabisco, the makers of these cookies, claim there are 1,000
347
chocolate chips in each bag. The authors of this article wish to test this claim. d. An automobile company claims that their truck gets an average of 20 miles per
gallon. A disgruntled group of truck buyers disputes this claim. e. A company that makes batteries claims its batteries last an average of 25 hours
of continuous use. A consumer group believes they have data that suggest the batteries last less than 25 hours.
19. Returning to the Super Bowl example discussed earlier in this chapter, imagine that you found 13 of 18 teams that won the coin flip went on to win the game.
a. State the null and alternative hypotheses (H0 and H1). b. Identify the critical values for α = .05 and state a decision rule. c. Make a decision about the null hypothesis. d. Does the result of your analysis support or not support the research hypothesis
that winning the coin flip increases a team's chances of winning the game? 20. Two researchers test the same research hypothesis using the same instruments. One researcher rejects the null hypothesis but the other does not.
a. Which researcher is more likely to have had a larger sample size? Why? b. Which researcher is more likely to have had a smaller level of alpha? Why? c. Which researcher is more likely to have had a directional alternative
hypothesis? Why? 21. A company makes a test they say can detect whether or not someone is guilty of a crime. In this situation, what are the two types of errors that could be made?
348
Answers to Learning Checks
Learning Check 1
2. a. p = .13 b. p = .25 c. p = .50 d. p = .88 e. p = .38
3. a. p = .11 b. p = .66 c. p = .73 d. p = .09 e. p = .24
4. a. p = .16 b. p = .07 c. p = .55 d. p = .17 e. p = .03
Learning Check 2
2. a. H0: μ = .1370; H1: μ ≠ .1370 b. H0: μ = 164; H1: μ ≠ 164 c. H0: μ = 450; H1: μ ≠ 450 d. H0: μ = 900; H1: μ ≠ 900 e. H0: μ = .50; H1: μ ≠ .50
3. a. H0: μ = 10; H1: μ ≠ 10 b. For α = .05, if the number of wins is < 6 or > 14, reject H0; otherwise, do not
reject H0. c. Sixteen wins is > 14; therefore, reject H0 (p < .05). d. Thirteen of the 20 teams would not support the research hypothesis because
H0 would not be rejected (13 wins is not < 6 or > 14). 4.
349
a. H0: μ = 8; H1: μ ≠ 8 b. For α = .05, if the number of correct guesses is < 4 or > 12, reject H0;
otherwise, do not reject H0. c. Sixteen wins is > 15; therefore, reject H0 (p < .05). d. The result of this analysis does not support your friend's claim to have ESP.
Learning Check 3
2. a. critical values = 3 wins and 11 wins; 12/14 = 86% b. critical values = 4 wins and 12 wins; 13/16 = 81% c. critical values = 6 wins and 16 wins; 17/22 = 77%
3.
a. For α = .05, critical values = 2 wins and 8 wins (9/10 = 90%);
for α = .01, critical values = 1 wins and 9 wins (10/10 = 100%)
b. For α = .05, critical values = 4 wins and 12 wins (13/16 = 81%);
for α = .01, critical values = 3 wins and 13 wins (14/16 = 88%)
c. For α = .05, critical values = 7 wins and 17 wins (18/24 = 75%);
for α = .01, critical values = 6 wins and 18 wins (19/24 = 79%) 4.
a. For H1: μ ≠ 5, critical values = 2 wins and 8 wins;
for H1: μ > 5, critical value = 8 wins
b. For H1: μ ≠ 9, critical values = 5 wins and 13 wins;
for H1: μ < 9, critical value = 6 wins
c. For H1: μ ≠ 13, critical values = 8 wins and 18 wins;
for H1: μ > 13, critical value = 17 wins
350
Answers to Odd-Numbered Exercises
1. a. p = .20 b. p = .80 c. p = .60 d. p = .60
3. a. p = .92 b. p = .01 c. p = .03
5. a. p = .0475 b. p = .9525 c. p = .7734 d. p = .0030 e. p = .0531 f. p = .1974
7. a. p = .0548 (z = −1.60) b. p = .1151 (z = 1.20) c. p = .3364 (z = .40 and z = 2.40)
9. a. p = .9406 (z = −1.56) b. p = .8770 (z = 1.16) c. p = .0068 (z = −2.47) d. p = .7229 (z = –.65 and z = 2.07)
11. a. p = .0313 b. p = .3125 c. p = .1876
13. a. p = .0256 b. p = .9744 c. p = .4752
15. p = .0319 17.
a. H0: Students taught effective learning skills will perform the same on tests as students offered incentives to do well.
351
H1: Students taught effective learning skills will perform differently on tests than students offered incentives to do well.
b. H0: Use of Internet bulletin boards is unrelated to television viewing.
H1: Use of Internet bulletin boards is related to television viewing.
c. H0: Children taught conflict resolution skills do not differ in violent behavior compared to children not taught conflict resolution skills.
H1: Children taught conflict resolution skills differ in violent behavior compared to children not taught conflict resolution skills.
d. H0: Scores on the Graduate Management Admissions Test are not related to the likelihood of success in graduate business school.
H1: Scores on the Graduate Management Admissions Test are related to the likelihood of success in graduate business school.
e. H0: The amount of time children spend watching television is not related to their preference for unhealthy foods.
H1: The amount of time children spend watching television is related to their preference for unhealthy foods.
19. a. H0: μ = 9; H1: μ ≠ 9 b. For α = .05, if the number of wins is < 5 or > 13, reject H0; otherwise, do not
reject H0. c. Thirteen wins is not < 5 or > 13; therefore, do not reject H0 (p > .05). d. The result of this analysis does not support the research hypothesis that
winning the coin flip increases a team's chances of winning the game. 21. One type of error is to reject the null hypothesis when we should not. Assuming innocence is the null hypothesis, in this case, the company concluded that the person was guilty of the crime when in fact the individual was innocent. The second type of error is to not reject the null hypothesis when we should. In this case, the company concluded that the person was innocent when in fact the person was guilty.
352
Sharpen your skills with SAGE edge at edge.sagepub.com/tokunaga
SAGE edge for students provides a personalized approach to help you accomplish your coursework goals in an easy-to-use learning environment. Log on to access:
eFlashcards Web Quizzes Chapter Outlines Learning Objectives Media Links
353
Chapter 7 Testing One Sample Mean
354
Chapter Outline 7.1 An Example From the Research: Do You Read Me? 7.2 The Sampling Distribution of the Mean
Characteristics of the sampling distribution of the mean 7.3 Inferential Statistics: Testing One Sample Mean (σ Known)
State the null and alternative hypotheses (H0 and H1) Make a decision about the null hypothesis
Set alpha (α), identify the critical values, and state a decision rule Calculate a statistic: z-test for one mean Make a decision whether to reject the null hypothesis Determine the level of significance
Draw a conclusion from the analysis Relate the result of the analysis to the research hypothesis Assumptions of the z-test for one mean Summary
7.4 A Second Example From the Research: Unique Invulnerability 7.5 Introduction to the t-Distribution
Characteristics of the Student t-distribution 7.6 Inferential Statistics: Testing One Sample Mean (σ Not Known)
State the null and alternative hypotheses (H0 and H1) Make a decision about the null hypothesis
Calculate the degrees of freedom (df) Set alpha (α), identify the critical values, and state a decision rule Calculate a statistic: t-test for one mean Make a decision whether to reject the null hypothesis Determine the level of significance
Draw a conclusion from the analysis Relate the result of the analysis to the research hypothesis Assumptions of the t-test for one mean Summary
7.7 Factors Influencing the Decision About the Null Hypothesis Sample size Alpha (α) Directionality of the alternative hypothesis
7.8 Looking Ahead 7.9 Summary 7.10 Important Terms 7.11 Formulas Introduced in This Chapter 7.12 Using SPSS 7.13 Exercises
Chapter 6 introduced hypothesis testing, which is the process of conducting a statistical analysis to test a research hypothesis about a population. To introduce hypothesis testing, Chapter 6 used the relatively simple example of winning the pregame coin flip at the Super Bowl. The Super Bowl example used the binomial distribution, a distribution that allows us to state with exact precision the probability of every possible outcome within the population (e.g., getting 0 heads in 2 coin flips). In contrast, the statistical procedures
355
discussed in this chapter, as well as the chapters that follow, involve populations that cannot be fully defined or understood. This chapter will examine two of these procedures, the z-test for one mean and the t-test for one mean, which test the difference between the mean of a sample and a hypothesized population mean. Our discussion will center on the findings from actual, published research studies.
356
7.1 An Example from the Research: Do You Read Me?
What are the most critical goals of elementary school education? A team of researchers at Eastern Washington University headed by graduate student Jaclyn Reed was emphatic: “Without a doubt, reading is the most important skill that students can acquire in school…. Reading at high levels is associated with continued academic success, significantly reduced risk for school dropout, and higher rates of entering college and finding successful employment” (Reed, Marchand-Martella, Martella, & Kolts, 2007, p. 45).
Two fundamental aspects of reading are vocabulary and comprehension. Vocabulary pertains to the knowledge and understanding of individual words, whereas comprehension is the ability to understand why and how a passage of text is structured as well as to summarize a passage in terms of its meaning. In reviewing the relevant literature, the researchers wrote that “instruction that focuses on vocabulary building and text comprehension is critical for student success” (Reed et al., 2007, p. 46). Consequently, the goal of their study was to evaluate the effectiveness of a program designed to teach reading skills to fourth-grade students. On the basis of an evaluation of the literature, they hypothesized that students who went through the program would demonstrate higher vocabulary and comprehension skills than the general population of fourth graders.
The researchers collected data from students in four classrooms at one elementary school; these students received the yearlong Reading Success Level A program, which consisted of lesson plans, exercises, workbooks, quizzes, and exams. The study, which we will refer to as the reading skills study, is an example of quasi-experimental research, described in Chapter 1 as research methods using naturally formed or preexisting groups rather than employing random assignment to conditions. As the researchers noted, “Because this research project was conducted to evaluate Reading Success Level A for the elementary school and a control group was not available, it was not possible to assign students randomly to a control and an experimental group” (Reed et al., 2007, p. 65).
To measure students on their vocabulary and comprehension skills, the researchers used a software program known as Read Naturally. This program had each student read a passage of text out loud, after which the number of words that were spoken correctly and the amount of time taken to read the passage was recorded. The number of words correctly spoken was then divided by the amount of time needed to read the passage; this variable was called “number of words correct per minute” (WCPM). For example, correctly speaking 280 words in 2 minutes resulted in a WCPM score of 280 ÷ 2, or 140.
The sample in the study consisted of 93 students. However, to save space and time, the example in this chapter will use a smaller sample of 20 students (N = 20) designed to resemble the data from the original study. The words correct per minute (WCPM) for the 20 students are listed in Table 7.1.
357
To organize and illustrate the data, a grouped frequency distribution table and frequency polygon for the 20 students' WCPM is provided in Figure 7.1. Examining this figure, we see that the distribution of WCPM is roughly symmetrical, with the center of the distribution near the 141–150 and 151–160 intervals.
Table 7.1 Words Correct Per Minute (WCPM) for 20 Students Table 7.1 Words
Correct Per Minute (WCPM) for 20
Students
Student WCPM
1 159
2 140
3 148
4 156
5 141
6 178
7 169
8 126
9 150
10 136
11 168
12 153
13 134
14 161
15 146
16 174
17 144
18 119
19 160
20 145
One step in analyzing data that have been collected is to calculate descriptive statistics. The most typical descriptive statistics are a measure of central tendency, such as the mean, and a measure of variability, such as the standard deviation. The calculation of the mean ( X ¯ )
358
WCPM for the 20 students in the reading skills study is presented below (please refer to Chapter 3 for a review of this formula): X ¯ = ∑ X N = 159 + 140 + 148 + … + 119 + 160 + 145 20 = 3007 20 = 150.35
To describe the variability in this set of data, the standard deviation (s) is calculated using the definitional formula introduced in Chapter 4: s = ∑ ( X − X ¯ ) 2 N − 1 = ( 159 − 150.35 ) 2 + ( 140 − 150.35 ) 2 + ⋯ + ( 160 − 150.35 ) 2 + ( 145 − 150.35 ) 2 20 − 1 = 74.82 + 107.12 + ⋯ + 93.12 + 28.62 19 = 4640.55 19 = 224.24 = 15.63
Figure 7.1 Grouped Frequency Distribution Table and Frequency Polygon of Words Correct per Minute (WCPM) for 20 Students
Looking at the descriptive statistics X ¯ = 150.35 , s = 15.63 = 150.35, s = 15.63), we find that on average, students correctly read approximately 150 words per minute, with the majority of the students having WCPM scores between 135 and 165.
Preliminary conclusions about these students' reading skills may be drawn from examining the table, figure, and descriptive statistics. However, testing the study's research hypothesis, that students completing the reading program would demonstrate higher levels of
359
vocabulary and comprehension skills than the population of fourth graders, requires calculating a second type of statistic known as an inferential statistic.
For the reading skills study, comparing the study's sample with the population of fourth graders requires the calculation of an inferential statistic that evaluates a sample mean in terms of its difference from a hypothesized population mean. More specifically, we need to transform the sample mean into a statistic to determine the probability of obtaining the value of the sample mean. Although this may sound new to you, it's actually very similar to what we did in Chapter 5. In that chapter, we evaluated a score for a variable (X) by transforming it into a z-score (z); this transformation allowed us to determine the probability of the score. The critical difference between this chapter and Chapter 5 is that we'll use what we learned in Chapter 6 to determine whether the probability of the statistic is low enough for us to decide that the difference between the sample mean and the hypothesized population mean is statistically significant. Put another way, we'll use the probability of the statistic to make a decision whether to reject what is known as the null hypothesis, which is a statistical hypothesis that in this situation states a value for the population mean μ.
In the Chapter 6 Super Bowl example, to determine whether the probability of 5 wins in 12 games was low enough to reject the null hypothesis of 6 wins in 12 games (H0: μ = 6), we created a distribution of all of the possible wins that could occur in 12 games. That is, before we could determine the probability of our outcome, we needed to determine all of the possible outcomes. For the reading skills study, determining the probability of obtaining our sample mean of 150.35 also requires a distribution—a distribution of all possible sample means. This distribution, known as the sampling distribution of the mean, is discussed in the next section.
360
7.2 The Sampling Distribution of the Mean
The sampling distribution of the mean is the distribution of values of the sample mean for an infinite number of samples of size N that are randomly selected from the population. This distribution is an example of a sampling distribution, which is a distribution of statistics for samples randomly drawn from populations.
Let's use the example of age to illustrate both the sampling distribution of the mean and the concept of sampling distributions. We'll start by assuming the age of people who obtain their PhD degrees is normally distributed in the population with a mean of 30 years and a standard deviation of 7 years; in other words, μ = 30 and σ = 7. This is illustrated in Figure 7.2(a). Now, say we draw two random samples of three people (N = 3) from this population and calculate the mean of the ages for each sample. Would we expect the means of the two samples to be the same? Would we expect either or both of these sample means to be equal to the population mean of 30?
The answer to both of the above questions is “perhaps, but probably not.” Because the two samples are not only smaller than the population but also contain different people, we would expect the samples to differ from each other as well as from the population because of random, chance factors. This would result in differences between a statistic calculated from the samples (i.e., the sample mean) and the corresponding population parameter (i.e., the population mean); these differences represent the concept of sampling error introduced in Chapter 6.
Returning to the age example, imagine we continue to draw random samples of three people from this population and calculate the mean age of each sample until we have calculated an infinite number of sample means—all of the sample means that could occur within the population for a sample size of N = 3. The distribution of these sample means is an example of the sampling distribution of the mean, defined earlier as the distribution of sample means for an infinite number of samples of size N randomly drawn from the population. Figure 7.2(b) displays the sampling distribution of the mean for the age example.
361
Characteristics of the Sampling Distribution of the Mean
Like any other distribution, the sampling distribution of the mean such as the one in Figure 7.2(b) may be described in terms of its modality, symmetry, and variability. First, in terms of modality, because all of the samples are drawn from a population that has a mean equal to μ, the mean of the sample means is also expected to be equal to μ. That is, even though the samples may differ from each other, we expect the mean of the sample means to be equal to the mean of the population. In the age example, the mean of the sampling distribution of the mean is μ = 30.
Next, in terms of symmetry, the sampling distribution of the mean is an approximate normal (bell-shaped) distribution, assuming the samples are sufficiently large, which is typically defined as a sample size of at least N = 30. For the sample means to be normally distributed implies that, even though there is variability among the sample means, we expect the majority of the sample means to be relatively close to the population mean, especially if the samples are of adequate size.
In introducing measures of variability, Chapter 3 discussed the standard deviation, which is the average deviation of a score from the mean. However, because we are now working with sample means rather than individual scores, the variability of the sampling distribution of the mean is measured by the standard error of the mean, defined as the average deviation of a sample mean from the population mean. In essence, the standard error of the mean is the standard deviation of the sampling distribution of the mean.
You may wonder why the word error is used to describe variability in the sampling distribution of the mean. Theoretically, because all of the samples are drawn from the same population, the mean of every sample should be the same as the population mean μ. Any variability among these sample means is seen as being the result of random factors referred to as error. So rather than use the term standard deviation (which measures the variability of scores for a variable), we use standard error to represent the variability of statistics calculated from samples.
The modality, symmetry, and variability of the sampling distribution of the mean are defined by a statistical principle known as the central limit theorem. This theorem states that when an infinite number of random samples are drawn from a population, the sample means are approximately normally distributed with a mean equal to the population mean μ and a standard deviation equal to the standard error of the mean, assuming the samples are of sufficient size (N ≥ 30).
Figure 7.2 Distribution of Age in a Population and the Sampling Distribution of the Mean
362
That the sampling distribution of the mean is approximately normally distributed serves an important function to researchers. As we learned in Chapters 5 and 6, we can apply the principles of the standard normal distribution to normal distributions to determine the probability of any particular score in the distribution. Because the sampling distribution of the mean is approximately normally distributed, researchers are able to determine the probability of any particular sample mean. The next section describes how we can calculate an inferential statistic that uses this probability to test a hypothesis regarding the difference between a sample mean and a population mean.
363
Learning Check 1: Reviewing what you've Learned So Far
1. Review questions a. How could you create the sampling distribution of the mean for a variable? b. What are the main characteristics (modality, symmetry variability) of the sampling
distribution of the mean? c. What is measured by the standard error of the mean? d. What is the difference between a standard deviation and a standard error? e. Why is it important that the sampling distribution of the mean is normally distributed?
364
7.3 Inferential Statistics: Testing One Sample Mean (σ Known)
The research hypothesis in the reading skills study was that students completing the reading program would demonstrate higher levels of vocabulary and comprehension skills than the population of fourth graders. The process of hypothesis testing introduced in Chapter 6 will be used to test this research hypothesis. This process consists of four steps:
state the null and alternative hypotheses (H0 and H1), make a decision about the null hypothesis, draw a conclusion from the analysis, and relate the result of the analysis to the research hypothesis.
Each of these steps is described below, using the reading skills study to illustrate relevant concepts and calculations. For this example, hypothesis testing centers on the calculation of an inferential statistic that evaluates the difference between a sample mean and a hypothesized population mean when the population standard deviation (σ) for the variable is known and can be stated.
365
State the Null and Alternative Hypotheses (H0 and H1)
The process of hypothesis testing begins by stating two statistical hypotheses: the null hypothesis and the alternative hypothesis; ultimately, we will make a decision regarding which hypothesis is supported by the data. The null hypothesis (H0) implies no change, difference, or relationship exists among groups or variables in the population. In the reading skills study, if the program does not affect students' reading skills, the appropriate null hypothesis is that the mean words correct per minute (WCPM) of the sample is the same as the mean of the population.
To state the null hypothesis, we need a value of a hypothesized population mean (μ). In the reading skills study the researchers noted that the Read Naturally software program had been administered to thousands of students and that the programs publisher reported a population mean WCPM of 124.81. Consequently the null hypothesis for this analysis may be stated as H 0 : μ = 124.81
A mutually exclusive alternative to the null hypothesis is the alternative hypothesis (H1), which implies that a change, difference, or relationship does exist among groups or variables in the population. If the program in the reading skills study does in fact affect students' reading skills, one appropriate alternative hypothesis would reflect the belief that the mean WCPM of the sample is not the same as the population mean. This difference is represented by the following alternative hypothesis: H 1 : μ ≠ 124.81
As we learned in Chapter 6, an alternative hypothesis that includes the ≠ symbol is referred to as non-directional or two-tailed, which implies the null hypothesis will be rejected if the sample mean is either greater than or less than the population mean.
The relationship between research hypotheses and statistical hypotheses is sometimes confusing to students. For example, the research hypothesis in the reading skills study implies the WCPM for the sample should be greater than the population mean. As such, the alternative hypothesis could have been stated as H1: μ > 124.81, meaning that the null hypothesis will only be rejected if the sample mean is greater than the population mean. An alternative hypothesis that includes the less than (<) or greater than (>) symbol is referred to as directional, or one-tailed. However, non-directional alternative hypotheses such as H1: μ ≠ 124.81 are often used in research to allow for the possibility of unexpected findings, particularly when there is not a large body of relevant theory or research to enable researchers to be more specific in their predictions. Using a non-directional alternative hypothesis in the reading skills study allows for the possibility that the Reading Success Level A program somehow impairs students' reading skills, resulting in the sample having a
366
WCPM that is less than the mean of the population.
367
Make a Decision about the Null Hypothesis
The primary purpose of this step is to calculate the value of an inferential statistic and to make a decision about the null hypothesis, which in this example involves deciding whether the difference between the sample mean and the population mean is statistically significant. This section describes the steps involved in this decision-making process:
set alpha (α), identify the critical values, and state a decision rule; calculate a statistic: z-test for one mean; make a decision whether to reject the null hypothesis; and determine the level of significance.
Although the majority of these steps were introduced in Chapter 6, the last one is new and is discussed below.
Set Alpha (α), Identify the Critical Values, and State a Decision Rule
The first step in making the decision about the null hypothesis is to set alpha (α), which is the probability of the statistic needed to reject the null hypothesis. As we discussed in Chapter 6, alpha is traditionally set at .05, which implies the null hypothesis is rejected when the probability of obtaining the calculated value of a statistic is less than .05 (p < .05). In stating the value for alpha, it is useful to also indicate whether the alternative hypothesis is directional (one-tailed) or non-directional (two-tailed). For the reading skills study, because the alternative hypothesis (H1: μ ≠ 124.81) is non-directional, alpha may be stated as “α = .05 (two-tailed).”
To identify the critical values of the statistic, we must first determine which particular inferential statistic will be calculated. In the reading skills study, the publishers of the Read Naturally software program reported not only a value of the population mean (μ) but also a population standard deviation (σ) of 43.26 WCPM. When both the population mean and standard deviation of a variable are known, the appropriate distribution to evaluate the difference between a sample mean and a population mean is the standard normal distribution (normal curve) introduced in Chapter 5. Scores for the standard normal distribution are referred to as z-scores. However, because we'll use this distribution to calculate a statistic that tests the difference between a sample mean and a population mean, we'll refer to this statistic as a z-statistic.
Using the standard normal distribution, we need to determine the values of the z-statistic that divide the distribution into two regions: the region of rejection (containing the values of the statistic whose probability is low enough to lead to the decision to reject the null
368
hypothesis) and the region of non-rejection (containing the values of the statistic whose probability is high enough to lead to the decision to not reject the null hypothesis). Given that we have defined a “low” probability as .05 (α = .05), we must identify the values of the standard normal distribution that fall in a 5% region of rejection.
Because the alternative hypothesis for the reading skills study is non-directional, the 5% region of rejection is split into two halves, one at each tail of the distribution. Consequently, we divide the α = .05 probability by 2 (.05 ÷ 2 = .0250) and find the values of the z-statistic associated with .0250 in the “Area Beyond z” column of the normal curve table (Table 1). We first move down the “Area Beyond z” column until we reach the value .0250 and then move to the left until we are under the “z” column. Here, we find a z- statistic value of 1.96. Therefore, the 5% region of rejection is in the area beyond the z- statistics −1.96 and 1.96; this is illustrated by the shaded areas of the distribution in Figure 7.3.
To summarize, we have thus far stated a value of alpha and identified the critical values of the z-statistic. In other words, we have defined a “low” probability and identified the values of the statistic that have this low probability. For the reading skills study, this may be stated as follows: For α = . 05 two − tailed , critical values = ± 1.96
In stating the critical values, the ± symbol (i.e., ±1.96) represents “plus or minus.”
Once the critical values have been identified, a decision rule can be stated that explicitly specifies the values of the statistic that result in the rejection of the null hypothesis. For the reading skills study: If z < − 1.96 or>1.96, reject H 0 ; otherwise, do not reject H 0
Figure 7.3 Critical Values for z-Statistic, α = .05, Two-Tailed Alternative Hypothesis
This decision rule implies that the null hypothesis will be rejected if the value of the z- statistic calculated from the sample data is either less than −1.96 or greater than 1.96. In other words, the null hypothesis is rejected when the statistic is located in the region of rejection, which implies the statistic has a low (<.05) probability of occurring. On the other hand, if the value of the z-statistic is neither less than −1.96 or greater than 1.96, the null hypothesis will not be rejected. The decision to not reject the null hypothesis is made when
369
the statistic is located in the region of non-rejection, which implies the statistic does not have a low probability of occurring—that is, its probability is greater than α (p > .05).
Calculate a Statistic: Z-Test for One Mean
In this research situation, we want to evaluate a sample mean ( X ¯ ) in terms of its difference from a population mean (μ) to determine the probability of obtaining our value of the sample mean under the assumption that the null hypothesis is true. For the reading skills study, assuming the population mean WCPM is 124.81, is the probability of obtaining our sample mean of 150.35 low enough to reject the null hypothesis and conclude the difference between the sample mean and the population mean is statistically significant? We will answer this question by transforming the sample mean into z-statistic.
Evaluating a sample mean by transforming it may seem new to you, but its actually very similar to what we did in Chapter 5, where we evaluated scores for normally distributed variables by transforming them into z-scores using Formula 5-1: z = X − μ σ
Once this formula is used to transform a score into a z-score, we can evaluate the score by determining its probability. For example, in Chapter 5, we transformed an SAT score of 660 into a z-score of 1.60, where we found a .0548 probability of an SAT score being greater than 660.
For the reading skills study, to evaluate our sample mean ( X ¯ ) of 150.35 within its distribution of sample means, it must be transformed in a manner very similar to transforming scores. One statistical procedure used to evaluate sample means, the z-test for one mean, tests the difference between a sample mean and a population mean when σ is known:
(7-1) z = X ¯ − μ σ X ¯
where X ¯ is the sample mean, μ is the population mean, and σ X ¯ is the population standard error of the mean. As we can see, Formula 7-1 closely resembles Formula 5-1; however, the numerator in Formula 7-1 involves the difference between two means X ¯ − μ rather than the difference between a score and a mean (X – μ), and the denominator reflects the variability among sample means σ X ¯ instead of scores (σ).
The first step in calculating the z-statistic for the 2-test for one mean is to calculate the population standard error of the mean σ X ¯ , which is the standard deviation of the sampling distribution of the mean when the population standard deviation (σ) for the variable is known. (Later in this chapter, we will discuss a standard error of the mean that's used when the population standard deviation is not known.)
370
The formula for the population standard error of the mean σ X ¯ is presented in Formula 7-2:
(7-2) σ X ¯ = σ N
where σ is the population standard deviation and N is the size of the sample. You may wonder why the calculated value of the standard error of the mean, which is to say the amount of variability of sample means, is a function of the variability of individual scores (σ) and the size of the sample (N). First, the larger the amount of variability in scores in the population, the larger the amount of variability in sample means generated from this population. Second, the larger the samples used to calculate the sample means, the smaller the effect of random, chance factors that create variability among sample means; as a result, there should be less variability among sample means than when smaller samples are used.
In calculating the population standard error of the mean for the reading skills study, earlier we mentioned that the publishers of the Read Naturally program reported a population standard deviation (σ) of 43.26 WCPM. Inserting this value of σ and the sample size of 20 into Formula 7-2, σ X ¯ is calculated as follows: σ X ¯ = σ N = 43.26 20 = 43.26 4.47 = 9.68
It is important to remember that the standard deviation (σ) measures the variability of individual scores, whereas the standard error of the mean σ X ¯ estimates the variability of sample means.
Once the population standard error of the mean σ X ¯ has been calculated, we can calculate a value of the z-statistic. For the reading skills study, the values for the sample mean ( X ¯ = 150.35), population mean (μ = 124.81), and population standard error of the mean σ X ¯ = 9.68 are inserted into Formula 7-1 below: z = X ¯ − μ σ X ¯ = 150.35 − 124.81 9.68 = 25.54 9.68 = 2.64
Make a Decision Whether to Reject the Null Hypothesis
Now that a value for the z-statistic has been calculated, the next step in testing the difference between the sample mean and the population mean is to make a decision whether to reject the null hypothesis. Using the decision rule stated earlier, this decision is made by comparing the value of the z-statistic calculated from the sample with the critical values. For the reading skills study, we could state this decision the following way:
z = 2.64 is greater than the critical value 1.96; therefore, the decision is made to reject the null hypothesis because the probability of z = 2.64 is less than .05.
371
Learning Check 2: Reviewing what you've Learned So Far
1. Review questions a. In testing one mean, the value of what parameter is stated in the null and alternative
hypotheses? b. In testing one mean, what is implied by the null and alternative hypotheses? c. Why are non-directional (two-tailed) rather than directional (one-tailed) alternative
hypotheses predominantly used in research studies? d. What are the similarities and differences between the formula for the z-score (Chapter 5)
and the z-test for one mean? e. Under what conditions would one calculate the population standard error of the mean
rather than the standard error of the mean? f. What two factors influence the amount of variability in a distribution of sample means?
2. For each of the following situations, calculate the population standard error of the mean σ X ¯ . a. σ = 12.00; N = 16 b. σ = 1.00; N = 9 c. σ = 4.50; N = 26 d. σ = 39.76; N = 50
3. For each of the following situations, calculate the z-statistic (z). a. X ¯ = 10.00 = 10.00; μ = 8; σ = 3; N = 9 b. X ¯ = 4.00 ; μ = 7; σ = 6; N = 16 c. X ¯ = 3.52 ; μ = 3.29; σ = 1.18; N = 21 d. X ¯ = 13.25 ; μ = 11.87; σ = 3.42; N = 32
Another, more concise, way of stating this decision is as follows: z = 2.64 > 1.96 ∴ reject H 0 p < . 05
The ∴ symbol is the mathematical symbol for therefore.
Statistical Significance versus Nonsignificance
When the null hypothesis is rejected, it is common to say that the result of the analysis is “statistically significant”; in the reading skills study, we could say that “the difference between the sample mean of 150.35 and the population mean of 124.81 is statistically significant.” If, on the other hand, the decision had been made to not reject the null hypothesis, the result of the analysis may be referred to as “nonsignificant.” Please keep in mind that “nonsignificant” is a statistical concept. You may see researchers refer to a statistic as “insignificant”; however, the word insignificant is not a statistical concept but is rather a value judgment that implies the statistic is not meaningful or important.
The abbreviation “n.s.” may be used to indicate a nonsignificant result of a statistical analysis. Similarly, the notation “p > .05” informs the reader that the probability of the statistic was greater than .05, meaning the probability was not low enough to reject the null hypothesis. We have found students sometimes get confused over the difference between “p
372
< .05” and “p > .05.” Remember, the null hypothesis is rejected when the value of the statistic is large enough to have a low probability, p < .05, of occurring.
Determine the Level of Significance
In the reading skills study, the decision was made to reject the null hypothesis because the probability of the z-statistic of 2.64 was less than .05. When the decision is made to reject the null hypothesis, it is customary to go one step further and determine whether the probability of the statistic is not only less than .05 (p < .05) but is also less than .01 (p < .01). We do this in order to report the results of a statistical analysis as accurately and informatively as possible. This is similar to the difference between describing someone as “a 19-year-old” rather than as a “teenager” or saying ones yearly income is “$4, 500” rather than “less than $10,000.”
It is important for you to understand that when researchers calculate statistics using statistical software, the software provides the exact probability of a statistic rather than simply indicating whether the probability is less or greater than .05 or .01. As a result, in reading the results of statistical analyses in journal articles, you may see something such as “p = .021”; because .021 is less than .05, this implies the null hypothesis was rejected; “p = .004” would imply that the probability of the statistic is less than .01. However, in reporting the results of analyses in tables or figures rather than in the text of a journal article, it remains customary to use levels of significance such as .05, .01, and .001. It is for that reason that we include this step within the process of hypothesis testing.
To determine whether the probability of the z-statistic for the reading skills study is less than .01, we need to identify the values of the z-statistic that correspond to the outer 1% of the distribution because these values have a combined probability less than .01. First, for a two-tailed alternative hypothesis, we divide .01 by 2, which is .0050. Next, returning to the normal curve table (Table 1), we move down the “Area Beyond z” column until we reach the value .0050. Moving to the left to the “z” column, we find the z-statistic 2.58, which indicates there is less than a .01 probability of obtaining a z-statistic either less than −2.58 or greater than 2.58.
For the reading skill study, we determine whether the z-statistic meets the .01 level of statistical significance below: z = 2.64 > 2.58 ∴ p < . 01
Because a z-statistic value of 2.64 is greater than 2.58, we may conclude that its probability of occurrence is not only less than .05 but also less than .01. Figure 7.4 illustrates that a z- statistic value of 2.64 exceeds both the α =.05 and .01 cutoffs of 1.96 and 2.58, respectively. We have found it extremely useful for students to draw the distribution with its critical values in order to determine the appropriate level of significance.
373
Figure 7.4 Determining the Level of Significance for the Reading Skills Study
Why Determine Whether p <. 01?
As we mentioned earlier, researchers determine whether the probability of a statistic is less than .01 in order to present their results as precisely as possible. Unfortunately researchers sometimes attach inappropriate labels to different levels of significance. A common mistake is to say that a statistic with less than a .01 probability is not simply “statistically significant” but rather “highly significant.” This is actually inappropriate because statistical significance is a dichotomy (significant vs. nonsignificant) rather than a continuum. The purpose of reporting levels of significance such as p < .01 (or, when appropriate, p < .001 or p < .0001) is to provide accurate descriptions of statistical analyses, not to make judgmental declarations.
When Determine Whether p < .01?
It is only appropriate to determine whether the probability of a statistic is less than .01 when we have made the decision to reject the null hypothesis. When we do not reject the null hypothesis, this implies the probability of the statistic is greater than .05 (p > .05). If the probability is greater than .05, it obviously cannot be less than .01. Therefore, when we do not reject the null hypothesis, it is not necessary to determine whether p < .01.
What if p < .05 but not < .01?
In the reading skills study, what would we have concluded if we rejected the null hypothesis but the z-statistic had not been greater than 2.58? For example, what if z had been 2.32 rather than 2.64? Because z = 2.32 is greater than the α = .05 critical value of 1.96, we would reject the null hypothesis, implying that p < .05. But because 2.32 is less than the .01 cutoff of 2.58, the most accurate statement we could make would be that its probability is less than .05 (p < .05) but not less than .01. Consequently, we could state z = 2.32 < 2.58 ∴ p < . 05 but not <.01
This situation is illustrated in Figure 7.5. As we can see, when the calculated value of a statistic falls between the .05 and .01 critical values, its probability is less than .05 but not less than .01. When researchers report “p < .05” for an analysis, we assume they have
374
determined that the probability of their statistic is not less than .01.
375
Draw a Conclusion from the Analysis
What conclusion could we draw from the results of the analysis in the reading skills study? We could say that “the null hypothesis was rejected,” but this says nothing about the purpose or nature of the analysis. Concluding that “the difference was statistically significant” is slightly better, but it doesn't indicate the specific difference to which we are referring. Saying that “the sample mean (M = 150.35) is significantly different from the hypothesized population mean (μ = 124.81)” is better still, but perhaps there is a way to be even more specific and informative. One way we could state our conclusion regarding the reading skills study is the following:
The number of words correct per minute (WCPM) (M = 150.35) in a sample of 20 fourth-grade students was significantly greater than the national normative sample (μ = 124.81), z = 2.64, p < .01.
This brief sentence contains a great deal of information:
The variable that was analyzed: “the number of words correct per minute (WCPM)” The sample from whom data were collected: “a sample of 20 fourth-grade students”
Figure 7.5 Example of Situation where p < .05 but Not < .01
Descriptive statistics of the variables: “(M = 150.35) … (μ = 124.81)” The nature and direction of the findings: “was significantly greater than the national normative sample” (rather than simply saying the sample mean was significantly “different” from the hypothesized population mean) Information about the inferential statistic: “z = 2.64, p < .01” (which indicates the type of statistic calculated [z], the calculated value of the statistic [2.64], and the level of significance of the statistic [p < .01])
Its important to report the results of statistical analyses with as much relevant information as possible as this reduces the possibility that others will draw incomplete or inappropriate
376
conclusions about the analysis.
377
Relate the Result of the Analysis to the Research Hypothesis
The final step in hypothesis testing is to interpret the results of the statistical analysis in terms of the extent to which the analysis supports or does not support a study's research hypothesis. To recall, the research hypothesis in the reading skills study was that students who went through the Reading Success Level A program would demonstrate higher vocabulary and comprehension skills than the general population of fourth graders. Given that the researchers found the WCPM for their sample to be significantly greater than the population mean, here is how they communicated their conclusions regarding their research hypothesis:
This study revealed that teaching explicit, systematic reading comprehension strategies to fourth graders is likely to increase reading comprehension skills. (Reed et al., 2007, p. 64)
Notice that they did not say that their findings “proved” the effectiveness of the reading program. As we learned in Chapter 6, hypotheses regarding a population cannot be proven when data are collected from only a sample of the population.
378
Assumptions of the z-Test for One Mean
The goal of statistical procedures such as the z-test is to test hypotheses researchers have about populations by analyzing data collected from samples of these populations. These procedures are based on certain assumptions regarding such things as who the data have been collected from, how the variables in a research study are measured, and how scores for the variable are distributed in the population. To use these procedures appropriately, researchers must determine the extent to which their data meet these assumptions. This section discusses assumptions related to the z-test for one mean; these assumptions also apply to many of the statistical procedures discussed in later chapters of this book.
The first assumption related to the z-test for one mean is the assumption of random sampling, which is the assumption that the sample in a research study has been randomly selected from the population. The reason for this assumption is that probability distributions such as the sampling distribution of the mean and the standard normal distribution are based on random sampling. However, in conducting research, it is very difficult to meet this assumption. Imagine, for example, a researcher develops a hypothesis involving differences between American men and women. Given the millions of men and women who live in the United States, it would be extremely difficult to develop a completely random sample. As a result, researchers may use inferential statistics without random sampling but, in doing so, must assess their samples in terms of how representative they are of their relevant populations and describe any limitations in their ability to apply their findings beyond their samples.
The second assumption pertains to how the variables in a research study are measured. The assumption of interval or ratio scale of measurement implies that the statistical procedure is being applied to variables measured at the interval or ratio scale of measurement. In Chapter 1, we noted that the values of variables measured at these two scales, such as the number of WCPM, are equally spaced along a numeric continuum. This assumption is critical because the calculation of statistics such as the mean and standard deviation involves arithmetic operations such as addition and multiplication. These operations cannot be performed on variables measured at the nominal or ordinal scales of measurement, such as gender, whose values are qualitatively rather than quantitatively different from each other. One cannot, for example, add “male” and “female” to each other to calculate the “average gender” in a sample.
The third assumption is the assumption of normality, which is the assumption that scores for the variable are approximately normally distributed in the population. Because theoretical distributions of statistics are normally distributed, it is expected that they are based on variables that are normally distributed in the population. However, even when this assumption is not met, research has found that statistics such as the z-statistic are
379
robust, meaning they are able to withstand moderate violations of the assumption of normality. For example, an important aspect of the central limit theorem discussed earlier is that, assuming the samples are of sufficient size (N ≥ 30), the sampling distribution of the mean approximates a normal distribution even if the distribution of scores for the variable is not normally distributed in the population. This helps the z-test be robust to violations of the assumption of normality.
380
Summary
Table 7.2 Summary, Conducting the z-Test for One Mean (Reading Skills Example)
Table 7.2 Summary, Conducting the z-Test for One Mean (Reading Skills Example) State the null and alternative hypotheses (H0 and H1)
H σ : μ = 124.81 H 1 : μ ≠ 124.81 Make a decision about the null hypothesis
Set alpha (α), identify the critical values, and state a decision rule If z < −1.96 or > 1.96, reject H0; otherwise, do not reject H0 Calculate a statistic: z-test for one mean
Calculate the population standard error of the mean ( σ X ¯ ) σ X ¯ = σ N = 43.26 20 = 43.26 4.47 = 9.68 Calculate the z-statistic (z)
z = X ¯ − μ σ X ¯ = 150.35 − 124.81 9.68 = 25.54 9.68 = 2.64 Make a decision whether to reject the null hypothesis
z = 2.64 > 1.96 ∴ r e j e c t H 0 ( p < .05 ) Determine the level of significance
z = 2.64 > 2.58 ∴ p < .01
Draw a conclusion from the analysis
In a sample of 20 fourth-grade students, the number of words correct per minute (WCPM) (M = 150.35) was significantly greater than the national normative sample (μ = 124.81), z = 2.64, p < .01. Relate the result of the analysis to the research hypothesis
This study revealed that teaching explicit, systematic reading comprehension strategies to fourth graders is likely to increase reading comprehension skills (Reed et al., 2007, p. 64).
The process of testing the mean of a sample when the standard deviation (σ) is known by conducting the z-test for one mean is summarized in Table 7.2, using the reading skills study as an example. The z-test is used when the population standard deviation σ is known. However, as it is extremely rare for researchers to collect data from an entire population, the population standard deviation is typically not known, and researchers must estimate it using the standard deviation of the sample (s). Consequently, testing the difference between a sample mean and a hypothesized population mean when the population standard deviation is not known requires a different statistical procedure. Again using a published
381
research study, the next section introduces and discusses this procedure, known as the t-test for one mean.
382
Learning Check 3: Reviewing what you've Learned So Far
1. Review questions a. What is the difference between a statistic that is statistically “significant” and one that is
“nonsignificant”? b. What is the difference between a statistic referred to as “nonsignificant” and one referred to
as “insignificant”? c. Does “p < .05” imply the null hypothesis was rejected or not rejected? Why? d. When and why would you determine whether the probability of a statistic is less than .01? e. Which of these would you use when the value of a statistic falls in between the .05 and .01
critical values: p > .05, p < .05, or p < .01? f. What information about a statistical analysis is typically included when communicating the
results of the analysis? g. What are some assumptions applicable to the z-test for one mean?
2. For each of the following situations, calculate the z-statistic (z), make a decision about the null hypothesis (reject, do not reject), and indicate the level of significance (p > .05, p < .05, p < .01).
a. X ¯ = 12.00 ; μ = 6; σ X ¯ = 3.00 b. X ¯ = 3.00 ; μ = 4; σ X ¯ = 1.00 c. X ¯ = 14.92 ; μ = 11.76; σ X ¯ = 1.31
3. For each of the following situations, calculate the population standard error of the mean σ X ¯ and the z-statistic (z), make a decision about the null hypothesis, and indicate the level of significance.
a. X ¯ = 8.00 ; μ = 5; σ = 3; N = 9 b. X ¯ = 3.00 ; μ = 5; σ = 8; N = 16 c. X ¯ = 14.50 ; μ = 12.75; σ = 4.41; N = 23
383
7.4 A Second Example from the Research: Unique Invulnerability
How long do you expect to live? Each year, you may read about the number of traffic fatalities expected to occur on major holidays, or you may hear about someone who contracted a rare disease and died at a relatively young age. When you hear such forecasts or news events, you might think to yourself, “That happens to other people, not to me.” One study examined the tendency people have to “distort information so that negative human outcomes are less likely to happen to us than to other people” (Snyder, 1997, p. 197).
C. R. Snyder, a researcher at the University of Kansas, set out to demonstrate this bias, called “unique invulnerability,” in a classroom exercise. Dr. Snyder was interested in whether people display the unique invulnerability bias when it comes to predicting how long they expect to live. He hypothesized that, when asked to predict their age at the time of their death, people will provide estimates greater than the average life expectancy in the population.
For the sample in his study, Dr. Snyder used 17 graduate students (4 men and 13 women). The variable of interest in this study was the estimated age of death: how old (in years) each student expected to be when he or she died. To collect the data for this study, Dr. Snyder “began by informing the class that the actuarially predicted age of death for U.S. citizens (men and women together) is 75 years…. After delivering this information, I asked students to write down (anonymously) their estimated ages of death … on a blank slip of paper” (Snyder, 1997, p. 198). The estimated age of death for the 17 students is listed in Table 7.3.
Analyzing the data from the study is facilitated by constructing a grouped frequency distribution table and frequency polygon; these are provided in Figure 7.6. Looking at this figure, we see that 82% of the students' estimates of their age at death are greater than the stated population mean of 75 years, suggesting that they may indeed be displaying the unique invulnerability bias. In terms of the shape of the distribution, it appears the estimates are somewhat normally distributed, ranging from 66 years to 100 years, with the center of the distribution in the 81- to 85-year range.
Table 7.3 Estimated Age of Death (in Years) for 17 Students Table 7.3 Estimated Age of Death (in
Years) for 17 Students
Student Estimated Age of Death
1 85
2 92
384
3 75
4 80
5 84
6 73
7 86
8 90
9 84
10 83
11 68
12 95
13 90
14 100
15 80
16 83
17 80
Figure 7.6 Grouped Frequency Distribution Table and Frequency Polygon of Estimated Age of Death for 17 Students
385
Following is a numerical summary of the data for the study, achieved by calculating the sample mean X ¯ and standard deviation (s): X ¯ = ∑ X N = 85 + 92 + 75 + ⋯ + 80 + 83 + 80 17 = 1428 17 = 84.00 s = ∑ ( X − X ¯ ) 2 N − 1 = ( 85 − 84.00 ) 2 + ( 92 − 84.00 ) 2 + … + ( 83 − 84.00 ) 2 + ( 80 − 84.00 ) 2 17 − 1 = 1.00 + 64.00 + … + 1.00 + 16.00 16 = 1026.00 16 = 64.13 = 8.01
Looking at the descriptive statistics, we see that the average estimated age of death of the 17 students in this sample X ¯ = 84.00 is higher than the hypothesized population average of 75. The value for the standard deviation (s = 8.01) shows there is variability in the students' estimates, with the majority of the estimates between 76 and 92.
In the reading skills study discussed earlier in this chapter, the sample mean was transformed into a z-statistic that was evaluated using the standard normal distribution. The z-statistic is used to test the difference between a sample mean and a population mean when the population standard deviation for the variable is known. However, for the estimated age of death variable in the unique invulnerability study, the population standard deviation is not presumed to be known. Therefore, to test the study's research hypothesis, the sample mean must be transformed into a different type of statistic that is based on a different type of distribution. This distribution, known as the t-distribution, is introduced in the next section.
386
7.5 Introduction to the t-Distribution
The t-statistic is used to test the difference between a sample mean and a population mean when an unknown population standard deviation is estimated with the standard deviation of a sample. The distribution of t-statistics is known as the Student t-distribution, defined as the distribution of values of the t-statistic based on an infinite number of samples of size N randomly drawn from the population. The Student t-distribution was developed by W. S. Gosset, who wrote under the pen name “Student.” Figure 7.7 illustrates the t- distribution, using a sample size of N = 17 as an example.
387
Characteristics of the Student t-Distribution
The Student t-distribution is a theoretical distribution that is created in a similar manner to the sampling distribution of the mean. Using the unique vulnerability study to illustrate this distribution, imagine that the average estimated age of death in the population is 75. We draw a random sample from the population, calculate the mean estimated age of death for the sample, and transform the sample mean into a t-statistic. What should be the value of the t-statistic calculated from a sample? Similar to the z-statistic discussed earlier, because there should be no difference between the sample mean and the population mean, the t- statistic should be zero (0). But because of the random, chance factors associated with sampling error, we know there will be variability in the values of t-statistics such that there is a distribution of t-statistics.
The Student t-distribution shares many of the characteristics of the standard normal distribution. For example, in terms of modality, looking at the center of the distribution in Figure 7.7, we see that the mean of the Student t-distribution is equal to 0. Another similarity to the standard normal distribution is that the Student t-distribution is symmetrical, with an infinite number of values in both tails of the distribution. However, unlike the standard normal distribution, there is not just one t-distribution but instead a family of t-distributions: different distribution of t-statistics for each sample size. Because t- statistics involve estimating the population from a sample, and this estimate is a function of the size of the sample, it makes sense that there will be a different distribution of t-statistics for different sample sizes. This is similar to what we saw with the binomial distribution in Chapter 6, where, for example, the distribution of the number of heads in 12 coin flips was not the same as the distribution of the number of heads in 30 coin flips.
Figure 7.7 Example of the Student t-Distribution
Figure 7.8 Distribution of t for Different Sample Sizes
388
The three dashed lines in Figure 7.8 represent the t-distribution for three sample sizes: N = 2, N = 10, and N = 25; the solid line in this figure is the standard normal distribution. As we see, as the sample size increases, the more the t-distribution approximates the standard normal distribution. This is because the larger the sample, the more the sample resembles the population; the more the sample resembles the population, the more a distribution of statistics based on the sample (i.e., the t-distribution) resembles a distribution of statistics based on a population (i.e., the standard normal distribution).
389
Learning Check 4: Reviewing what you've Learned So Far
1. Review questions a. What are the main characteristics of the Student t-distribution? b. In what ways is the Student t-distribution similar to the standard normal distribution? c. Why is there a different distribution of t-statistics for different sample sizes? d. What happens to the shape of the t-distribution as the sample size grows larger?
390
7.6 Inferential Statistics: Testing One Sample Mean (σ Not Known)
As in the reading skills study, the unique invulnerability study hypothesizes a difference between a sample mean and a population mean. This difference is tested by calculating and testing an inferential statistic using the same four steps as before: state the null and alternative hypotheses (H0 and H1), make a decision about the null hypothesis, draw a conclusion from the analysis, and relate the result of the analysis to the research hypothesis. However, several of these steps contain differences due to the fact that, in this situation, the population standard deviation for the variable is not known. As we work through the steps, these differences will be highlighted and explained.
391
State the Null and Alternative Hypotheses (H0 and H1)
In testing one sample mean, the null hypothesis (H0) states a value of the population mean μ. In this example, if the unique invulnerability bias does not exist, the estimated age of death for this sample should be the same as the actual average age of death in the population. As mentioned earlier, the average age of death in the population of U.S. citizens was claimed to be 75. Thus, the null hypothesis may be stated as H 0 : μ = 75
The alternative hypothesis (H1) states a mutually exclusive alternative to the null hypothesis. In the unique invulnerability study, one way to state the alternative hypothesis (H1) reflects the belief that the estimated age of death in this sample is not the same as the actual average age of death in the population. This difference is represented by the following alternative hypothesis: H 1 : μ ≠ 75
As the study's research hypothesis suggests that the estimated age of death in the sample should be greater than the population mean, the alternative hypothesis could have been stated as the directional H1: μ > 75. But as we mentioned earlier in this chapter, a non- directional (two-tailed) alternative hypothesis was used to allow for the possibility that the sample mean may be significantly less than the population mean.
392
Make a Decision about the Null Hypothesis
The next step in hypothesis testing is to decide whether the difference between the sample mean and the population mean is statistically significant. The steps involved in making this decision when the population standard deviation is not known are as follows:
calculate the degrees of freedom (df) set alpha, identify the critical values, and state a decision rule; calculate a statistic: t-test for one mean; make a decision whether to reject the null hypothesis; and determine the level of significance.
The last four steps are the same as for the 2-test for one mean. However, when the population standard deviation is unknown, making the decision about the null hypothesis begins by calculating what are known as degrees of freedom. The next section defines and illustrates this concept.
Calculate the Degrees of Freedom (df)
The degrees of freedom (df) is defined as the number of values or quantities that are free to vary when a statistic is used to estimate a parameter. This step is included in this example because the standard deviation (s) is being used to estimate the population standard deviation (σ). The concept of degrees of freedom will appear throughout many of the remaining chapters of this book as different research situations and statistical procedures are discussed.
Lets illustrate the concept of degrees of freedom with a simple demonstration. Imagine a friend makes the following request of you: “Tell me any three numbers,” to which you respond, “2, 35, and —”. But your friend interrupts you before you can complete your response. “Stop!” he shouts. “I forgot to tell you this: the three numbers must add up to 99.” Once this has been said, you determine that the third number can only be 62. In this example, the first two numbers you chose were free to vary and could, in fact, have had any value. However, your friends constraint, in this case the sum of 99, meant the value of the third number was no longer free to vary but was instead fixed at 62. In this example, you had two “degrees of freedom.”
The concept of degrees of freedom is relevant to inferential statistics because these statistics involve estimating population parameters from data collected from samples. In the unique invulnerability example, from a sample of 17 participants (N = 17), the standard deviation of 8.01 is an estimate of an unknown population standard deviation. As a reminder, below is the formula for the standard deviation (s):
393
s = ∑ X − X ¯ 2 N − 1
We can see that the numerator of this formula involves summing deviations of scores from the mean X − X ¯ . In Chapter 4, we saw that, for any set of data, the sum of deviations ∑ X − X ¯ is equal to 0. In a sample of 17 participants, the first 16 deviations can be any value. However, once they are determined, the deviation of the last participant must enable the sum of the deviations to be equal to zero. Consequently, although the first 16 scores are free to vary, the 17th is not; it is predetermined by the first 16 scores.
Formula 7–3 calculates the degrees of freedom (df) for one sample:
(7-3) df = N − 1
where N is the size of the sample. In the unique invulnerability example, d f = N − 1 = 17 − 1 = 16
Therefore, when the sample size N is equal to 17, there are 16 degrees of freedom. Different formulas for the degrees of freedom will be presented in later chapters of this book as different statistical procedures are introduced.
Set Alpha (α), Identify the Critical Values, and State a Decision Rule
As is tradition, alpha (α) will be set to .05, meaning that the null hypothesis will be rejected if the value of the calculated statistic has less than a .05 probability of occurring under the assumption that the null hypothesis is true. Furthermore, because the alternative hypothesis in this example is non-directional (H1: μ ≠ 75), alpha can be referred to as “α = .05 (two- tailed).”
The next step is to identify the values of the t-statistic whose probability is low enough to lead to the decision to reject the null hypothesis. But because there are an infinite number of possible values of t and an infinite number of possible sample sizes, it would be impractical to have a table that lists the probability for every value of t for every sample size. Instead, because we are only interested in values of t that affect the decision about the null hypothesis, the only values of concern are the critical values, defined as the values of a statistic that separate the regions of rejection and nonrejection. A table of critical values for the t distribution is provided in Table 3 in the back of this book.
To determine the critical values of the t-statistic, three pieces of information are needed: the degrees of freedom (df), alpha (α), and the directionality of the alternative hypothesis (directional [one-tailed] or non-directional [two-tailed]). The first column of Table 7.4, labeled df, lists different degrees of freedom. Because there is a different t-distribution for
394
each sample size, each degrees of freedom has its own critical values. The bottom row of this table includes the infinity (∞) symbol; this does not imply that a sample may actually contain an infinite number of degrees of freedom; rather, it provides the critical values for a population assumed to be normally distributed.
The eight columns to the right of the df column list different values of alpha (.10, .05, .01, .001) corresponding to different levels of significance. The first four of these columns are used when the alternative hypothesis (H1) is directional (one-tailed); the last four columns are used for a non-directional (two-tailed) alternative hypothesis.
For the unique invulnerability study, the critical values may be identified by moving down the df column until we reach the appropriate row for this example: df = 16. Next, because the alternative hypothesis is non-directional (H1: μ ≠ 75), we move to the right until we are under the heading, “Level of Significance for Two-Tailed Test.” Within this set of columns, assuming α = .05, we find a critical t-value of 2.120. Therefore, for the unique invulnerability study, the critical values may be stated as follows: For α = .05 two − tailed and df = 16 , critical values = ± 2.120
The critical values and the regions of rejection and non-rejection for the unique invulnerability study are illustrated in Figure 7.9.
The next step in hypothesis testing is to state a decision rule that uses the critical values to specify the conditions under which the null hypothesis will be rejected. For the unique invulnerability study: If t < − 2.120 or > 2.120, reject H 0 ; otherwise, do not reject H 0
Table 7.4 Portions of the Table of Critical Values of the t-Statistic Table 7.4 Portions of the Table of Critical Values of the t-Statistic
Level of Significance for One-Tailed Test
Level of Significance for Two-Tailed Test
df .10 .05 .01 .001. 10 .05 .01 .001
1 3.078 6.314 31.821 318.310 6.314 12.706 63.657 636.619
2 1.886 2.920 6.965 22.326 2.920 4.303 9.925 31.598
3 1.638 2.353 4.541 10.213 2.353 3.182 5.841 12.941
4 1.533 2.132 3.747 7.173 2.132 2.776 4.604 8.610
5 1.476 2.015 3.365 5.893 2.015 2.571 4.032 6.859
11 1.363 1.796 2.7 18 4.025 1.796 2.201 3.106 4.437
12 1.356 1.782 2.681 3.930 1.782 2.179 3.055 4.318
13 1.350 1.771 2.650 3.852 1.771 2.160 3.012 4.221
395
14 1.345 1.761 2.624 3.787 1.761 2.145 2.977 4.140
15 1.341 1.753 2.602 3.733 1.753 2.131 2.947 4.073
16 1.337 1.746 2.583 3.686 1.746 2.120 2.921 4.015
17 1.333 1.740 2.567 3.646 1.740 2.110 2.898 3.965
18 1.330 1.734 2.552 3.610 1.734 2.101 2.878 3.922
19 1.328 1.729 2.639 3.579 1.729 2.093 2.861 3.883
20 1.325 1.725 2.528 3.552 1.725 2.086 2.845 3.850
90 1.291 1.662 2.368 3.183 1.662 1.987 2.632 3.403
120 1.289 1.658 2.358 3.232 1.658 1.980 2.617 3.373
∞ 1.282 1.645 2.326 3.183 1.645 1.960 2.576 3.291
This decision rule implies that if the value of t calculated for the sample exceeds either of the critical values, it falls in the region of rejection and the null hypothesis will be rejected.
Calculate a Statistic: T-Test for One Mean
The next step is to calculate a value of the t-statistic using a statistical procedure known as the t-test for one mean, which tests the difference between a sample mean and a hypothesized population mean when the population standard deviation is not known:
(7-4) t = X − μ ¯ S X ¯
Figure 7.9 Critical Values for the Unique Invulnerability Study (α = .05 [Two- Tailed] and df = 16)
where X ¯ is the sample mean, μ is the population mean, and S X ¯ is the standard error of the mean.
Formula 7–4 closely resembles the formula for the z-test for one mean (Formula 7–2) with
396
one critical difference: The denominator of this formula is the standard error of the mean S X ¯ , which is the standard deviation of the sampling distribution of the mean when the standard deviation (s) is used to estimate an unknown population standard deviation (σ). Formula 7–5 provides the formula for S X ¯ :
(7-5) S X ¯ = S N
where s is the standard deviation and N is the size of the sample. This formula differs from the formula for the population standard error of the mean (Formula 7–2) by having the standard deviation (s) rather than the population standard deviation (σ) in the numerator.
For the unique invulnerability study, for the sample of 17 students, the standard deviation of s = 8.01 is used to calculate the standard error of the mean as follows: S X ¯ = S N = 8.01 17 = 8.01 4.12 = 1.94
Once the standard error of the mean S X ¯ has been calculated, it is inserted into Formula 7–4, along with the sample X ¯ and population (μ) means to calculate a value of the t- statistic. For the unique invulnerability study: t = X ¯ − μ s X ¯ = 84.00 − 75 1.94 = 4.63 = 9.00 1.94
Now that the sample mean has been transformed into a t-statistic, it can be evaluated to make a decision about the null hypothesis; this is discussed in the next section.
397
Learning Check 5: Reviewing what you've Learned So Far
1. Review questions a. What is the generic definition of degrees of freedom (df)? b. Under what situations are degrees of freedom (df) calculated? c. Why is there a table of critical values for the t-statistic rather than a single critical value? d. What information is needed to identify the critical value for the t-statistic? e. Under what situations would you conduct the z-test for one mean or the t-test for one
mean? f. What are the differences between the formulas for the population standard error of the
mean and the standard error of the mean? g. What are the similarities and differences between the formulas for the z-test and the t-test?
2. For each of the following situations, calculate the degrees of freedom and identify the critical values (assume α = .05 [two-tailed]).
a. N = 9 b. N = 13 c. N = 18 d. N = 21
3. For each of the following situations, calculate the standard error of the mean S X ¯ a. s = 3.00; N = 16 b. s = 7.50; N = 25 c. s = 20.00; N = 50 d. s = 3.41; N = 38
4. For each of the following situations, calculate the t-statistic (t). a. X ¯ = 14.00 ; μ = 9; s = 12.00; N = 36 b. X ¯ = 3.00 ; μ = 5; s = 6.00; N = 16 c. X ¯ = .56 ; μ = .65; s = .18; N = 17
Make a Decision Whether to Reject the Null Hypothesis
Now that a value of the t-statistic has been calculated, we can make a decision whether to reject the null hypothesis. For the unique invulnerability study comparing the calculated value of the t-statistic with the critical values leads to the following decision: t = 4.63 > 2.120 ∴ reject H 0 p < . 05
Because the t-statistic value of 4.63 calculated from the sample exceeds the critical value 2.120, its probability is sufficiently low (p < .05) to lead to the decision to reject the null hypothesis.
Determine the Level of Significance
To be as precise as possible, when the null hypothesis has been rejected, we determine whether the probability of the t-statistic is not only less than .05 but also less than .01. For the unique invulnerability study, the .01 critical value is determined by returning to the
398
table of critical values in Table 3. After moving down to the df = 16 row, move to the right until we are under the .01 critical values of t for a two-tailed alternative hypothesis; here we find the critical value 2.921. Comparing the calculated value of t of 4.63 with the .01 critical value, t = 4.63 > 2.921 ∴ p < . 01
Figure 7.10 illustrates the location of t for this example relative to the .05 and .01 critical values. As can be seen from this figure, because 4.63 is greater than both 2.120 and 2.921, the probability of a t-statistic of 4.63 for df = 16 and α = .05 (two-tailed) is not only less than .05 but also less than .01.
Figure 7.10 Determining the Level of Significance for the Unique Invulnerability Study
399
Draw a Conclusion from the Analysis
We can state the following conclusion regarding the results of the statistical analysis for the unique invulnerability study:
The average estimated age of death (M = 84.00 years) for the 17 class members in this sample was significantly greater than the actual population average of 75 years, t(16) = 4.63, p < .01.
The above sentence describes the variable that was analyzed (“The average estimated age of death”), who was included in the analysis (“the 17 class members in this sample”), descriptive statistics of the variables (“(M = 84.00 years) … the actual population average of 75 years”), the nature and direction of the findings (“was significantly greater than the actual population average”), and information about the inferential statistic: t(16) = 4.63, p < .01, which indicates the type of statistic calculated (t), degrees of freedom (16), value of the statistic (4.63), and level of significance (p < .01).
400
Relate the Result of the Analysis to the Research Hypothesis
By finding that the mean estimate of age of death in this sample was significantly greater than the actual population average, we can relate the results of this analysis to the study's research hypothesis the following way:
The result of this analysis supports the research hypothesis that people will provide estimates of age of death greater than the average life expectancy in the population.
Dr. Snyder provided the following description of his conclusions from the study:
These results replicate other recent classroom-demonstration findings in which students maintain self-serving positive illusions in spite of knowing about the relevant research … instructors should help students to find ways of abandoning biases, such as unique invulnerability, especially when such biases increase the potential for harm in students' lives. (Snyder, 1997, p. 199)
401
Assumptions of the t-Test for One Mean
Given the similarities between the two procedures, it's not surprising that the assumptions of the z-test for one mean discussed earlier in this chapter (random sampling, interval or ratio scale of measurement, and normality) also apply to the t-test for one mean. One implication of the assumption of normality, which is the assumption that scores for the variable are approximately normally distributed in the population, is that distributions of scores for samples drawn from these populations are also expected to be approximately normal. Given we are estimating the population standard deviation from a sample, it's particularly important that the sample approximates a normal distribution. As we mentioned earlier, this is more likely to occur when the sample is sufficiently large (N ≥ 30).
Table 7.5 Summary, Conducting the t-Test for One Mean (Unique Invulnerability Example)
Table 7.5 Summary, Conducting the t-Test for One Mean (Unique Invulnerability Example)
State the null and alternative hypotheses (H0 and H1)
H 0 = μ = 75 H 1 : μ ≠ 75 Make a decision about the null hypothesis
Calculate the degrees of freedom (df) d f = N − 1 = 17 − 1 = 16
Set alpha (α), identify the critical values, and state a decision rule I f t < − 2.120 o r > 2.120 , r e j e c t H 0 ; o t h e r w i s e , d o n o t r e j e c t H 0 Calculate a statistic: t-test for one mean Calculate the standard error of the mean ( s X ¯ ) S X ¯ = s N = 8.01 17 = 8.01 4.12 = 1.94
Calculate the t-statistic (t) t = X ¯ − μ S X ¯ = 84.00 − 75 1.94 = 9.00 1.94 = 4.63 Make a decision whether to reject the null hypothesis t = 4.63 > 2.120 ∴ r e j e c t H 0 ( p < .05 ) Determine the level of significance t = 4.63 > 2.921 ∴ p < .01
Draw a conclusion from the analysis
The average estimated age of death (M = 84.00 years) for the 17 class members in
402
this sample was significantly greater than the actual population average of 75 years, t(16) = 4.63, p < .01.
Relate the result of the analysis to the research hypothesis
The result of this analysis supports the research hypothesis that people will provide estimates of age of death greater than the average life expectancy in the population.
403
Summary
The process of testing the mean of a sample when σ is unknown by conducting the t-test for one mean is summarized in Table 7.5, using the unique invulnerability study as an example. This chapter has presented several research studies that used the process of hypothesis testing first introduced in Chapter 6. The purpose of the next section is to further our discussion of another topic introduced in this earlier chapter: factors that directly affect the decision made about the null hypothesis.
404
Learning Check 6: Reviewing what you've Learned So Far
1. For each of the following situations, calculate the degrees of freedom (df), identify the critical values (assume α = .05 [two-tailed]), calculate the t-statistic (t), make a decision about the null hypothesis (reject, do not reject), and indicate the level of significance (p > .05, p < .05, p < .01).
a. X ¯ = 7.00 ; μ = 4 ; s X ¯ = 1.25 ; N = 25 b. X ¯ = 12.00 ; μ = 10 ; s X ¯ = 1.45 ; N = 16 c. X ¯ = 1.69 ; μ = 2.08 ; s X ¯ = 11 ; N = 19
2. For each of the following situations, calculate the degrees of freedom (df), identify the critical values (assume α = .05 [two-tailed]), calculate the standard error of the mean s X ¯ , calculate the t-statistic (t), make a decision about the null hypothesis, and indicate the level of significance.
a. X ¯ = 16.75 ; μ = 12.75 ; s = 4.00 ; N = 11 b. X ¯ = 3.27 ; μ = 2.98 ; s = 1.73 ; N = 27 c. X ¯ = 7.82 ; μ = 10.56 ; s = 5.21 ; N = 19
405
7.7 Factors Affecting the Decision about the Null Hypothesis
Within hypothesis testing, one of two decisions is ultimately made: Reject the null hypothesis or do not reject the null hypothesis. In Chapter 6, we introduced three factors that directly affect which of these two decisions are eventually made:
sample size, alpha, and the directionality of the alternative hypothesis.
This section will use the research situations covered in this chapter to continue and expand on our earlier discussion of these factors. It is critical to understand these factors as they are, to some degree, under the control of researchers.
406
Sample Size
Given its repeated appearance throughout the data collection and data analysis stages of the research process, sample size clearly plays a critical role in statistical analyses designed to test research hypotheses. As was mentioned in Chapter 6, the relationship between sample size and the decision regarding the null hypothesis is as follows:
The larger the sample size, the greater the likelihood of rejecting the null hypothesis.
In conducting an inferential statistic such as the t-test for one mean, sample size influences the decision about the null hypothesis in two ways. First, keeping all other aspects of the analysis constant, increasing the size of the sample increases the numeric value of the statistic, which in turn increases the likelihood of rejecting the null hypothesis. Second, increasing the size of the sample decreases the critical values of the statistic, which also increases the likelihood of rejecting the null hypothesis.
To illustrate the relationship between sample size and the numeric value of the statistic, imagine we conduct two studies of unique invulnerability. In both studies, the sample means are the same X ¯ = 80.00 , the standard deviations are the same (s = 7.00), and the hypothesized population mean is the same (μ = 75). However, the sample sizes in the two studies are different: N = 20 and N = 10, respectively. In Table 7.6, the standard error of the mean s X ¯ and the t-test for one mean are calculated for the two studies.
Looking at Table 7.6, using the larger sample of N = 20 results in a smaller value of the standard error of the mean ( s X ¯ = 1.57 vs. 2.22), which leads to a larger value of the t- statistic (t = 3.19 vs. 2.25). Even though we have not changed the difference between the sample mean and the population mean, increasing the sample size has increased the value of the t-statistic, which in turn increases the likelihood of rejecting the null hypothesis.
Next, in terms of the relationship between sample size and the critical values, turn to the table of critical values for the Student t-distribution in Table 7.4. Moving down the df column, we see that as the number of degrees of freedom increases, the critical value becomes smaller. This implies that as the sample gets larger, the value of the t-statistic needed to reject the null hypothesis gets smaller, thereby increasing the likelihood of rejecting the null hypothesis. Why do the critical values change as a function of sample size? As we mentioned earlier in this chapter, increasing the size of a sample decreases the effect of random factors that create variability among sample means; this results in less variability among sample means and less variability among statistics such as the t-statistic (see Figure 7.8). Because extreme values of a statistic are unlikely to occur in larger samples, a less
407
extreme value is needed to reject the null hypothesis.
Table 7.6 The Effect of Sample Size on the Numeric Value of the t-test for One Mean
Table 7.6 The Effect of Sample Size on the Numeric Value of the t-test for One Mean
N = 20 N = 10
S X ¯ = 7.00 20 = 7.00 4.47 = 1.57 S X ¯ = 7.00 10 = 7.00 3.16 = 2.22
t = 80.00 − 75 1.57 = 5.00 1.57 = 3.19 t = 80.00 − 75 2.22 = 5.00 2.22 = 2.25
408
Alpha (α) The relationship between alpha (α), the probability of a statistic needed to reject the null hypothesis, and the decision about the null hypothesis may be summarized as follows:
The larger the value of alpha, the greater the likelihood of rejecting the null hypothesis.
This is because an increase in alpha results in smaller critical values, which increases the likelihood of rejecting the null hypothesis.
Figure 7.11 illustrates the critical values for a sample size of N = 10 (df = 9) for two values of alpha: .05 and .10. Setting alpha to .10 implies that the null hypothesis will be rejected if the probability of the statistic is less than .10; this definition of a “low” probability is less stringent than α = .05. In this figure, we see that the critical values are smaller for α = .10 (±1.833) than for α = .05 (±2.262). Because the region of rejection is larger for α = .10, a smaller value of the statistic is needed to reject the null hypothesis, making it more likely the null hypothesis will be rejected.
Although it is tradition in academic research to set alpha to .05, researchers sometimes choose different definitions of a “low” probability. For example, researchers may believe that using an alpha level of .05 is too strict, particularly in research situations where it may not be possible or feasible to have large samples. In situations such as these, the likelihood of rejecting the null hypothesis may be increased by setting alpha to .10 rather than .05. On the other hand, researchers may believe that α = .05 is too lenient, choosing instead to reject the null hypothesis only when the statistic has a very small probability of occurring (i.e., .01 or even .001), thereby lowering the chances of rejecting the null hypothesis. Ultimately, it is each researcher's responsibility to set alpha at a level appropriate for the research situation at hand.
409
Directionality of the Alternative Hypothesis
The relationship between the directionality of the alternative hypothesis (H1) and the decision about the null hypothesis may be stated the following way:
There is a greater likelihood of rejecting the null hypothesis when the alternative hypothesis is directional (one-tailed) rather than non-directional (two-tailed).
Using a directional alternative hypothesis results in a smaller critical value, which increases the likelihood of rejecting the null hypothesis when the difference between the sample mean and the population mean is in the hypothesized direction.
Critical values for a non-directional and directional alternative hypothesis for N = 10 and α = .05 are illustrated in Figure 7.12. Notice that in this figure, because the 5% region of rejection is not split between the two ends of the distribution, the positive critical value is smaller for the directional alternative hypothesis (1.833) than a non-directional alternative hypothesis (2.262). Consequently, using a directional alternative hypothesis increases the likelihood of rejecting the null hypothesis when the result is in the predicted direction.
Figure 7.11 Relationship between Alpha (α) and Critical Values for t (df = 9)
Using a one-tailed rather than a two-tailed alternative hypothesis may be a source of
410
controversy, in that researchers could choose to use a one-tailed alternative hypothesis when the calculated value of their statistics is not large enough to reject the null hypothesis using a two-tailed alternative hypothesis, thereby leading to confusing or inappropriate conclusions regarding their statistical analyses. The decision to use a one-tailed or a two- tailed alternative hypothesis depends on several factors. The traditional and more conservative approach is to use a two-tailed alternative hypothesis, particularly if not enough is known about the topic or the population to make directional hypotheses. However, there are situations when using a one-tailed alternative hypothesis is appropriate. First, a large body of research may indicate the tendency for the predicted relationship between variables to be in only one direction. Also, it may not be possible for the relationship to exist in both directions. For example, if we conduct a study looking at the effects of a spray designed to greatly increase the length of one's hair, we may have no reason to expect the length of one's hair to decrease as a result of using the spray and consequently state an alternative hypothesis that only allows for an increase in length.
Figure 7.12 Relationship between Directionality of the Alternative Hypothesis and Critical Values for t (df = 9 and α = .05)
At first glance, one solution would be to use a one-tailed alternative hypothesis but switch to a two-tailed alternative hypothesis when the results are in the opposite, unexpected direction. As logical as this may sound, this solution presents problems of its own because choosing this strategy essentially means using a 7 1/2% region of rejection: 5% at one end of the distribution while still allowing for 2 1/2% at the other end. This is roughly analogous to the captain of a softball team calling “heads” while the coin is in the air and then changing her mind after the coin hits the ground. The directionality of the alternative hypothesis should be based on sound, defensible reasoning before statistical analyses are conducted, without vacillating once the choice has been made.
411
Learning Check 7: Reviewing what you've Learned So Far
1. Review questions a. What are the two ways in which sample size affects the decision about the null hypothesis in
testing the t-statistic? b. What is the relationship between alpha (α) and the decision about the null hypothesis? c. Why is it necessary to determine whether to use a one-tailed or two-tailed alternative
hypothesis before conducting a statistical analysis rather than after?
412
7.8 Looking Ahead
In this chapter, we have expanded our discussion of the research process in general and the process of hypothesis testing in particular. Using the z-test and t-test, we have tested the difference between a sample mean and a hypothesized population mean; later chapters will examine statistical procedures designed to test hypotheses in different research situations. In this chapter, we have also expanded on a discussion of issues relevant to hypothesis testing, such as factors that influence the decision about the null hypothesis. As we have seen, choices researchers make regarding sample size, alpha, and the directionality of the alternative hypothesis directly affect the conclusions drawn about research hypotheses. As such, hypothesis testing has been the source of concern and controversy among researchers. The next chapter discusses these concerns and presents an alternative to hypothesis testing.
413
7.9 Summary
To test hypotheses regarding the difference between a sample mean X ¯ and a population mean (μ), one of two statistical techniques may be used: the z-test for one mean or the t-test for one mean. The z-test is used when the population standard deviation (σ) for the variable being analyzed is known; the t-test is used when the population standard deviation is not known, and the standard deviation (s) is used to estimate σ.
Evaluating the difference between a sample mean and a population mean involves determining the probability of obtaining a particular value of the sample mean. To determine this probability, the sampling distribution of the mean, the distribution of values of the sample mean when an infinite number of samples of size N are randomly selected from the population, is used. The sampling distribution of the mean is an example of a sampling distribution, which is a distribution of statistics for samples randomly drawn from populations.
Three features of the sampling distribution of the mean are (1) the expected mean of the distribution is equal to the population mean (μ); (2) the distribution is an approximate normal distribution, even if the distribution of scores in the population is not normally distributed, assuming the samples are sufficiently large (typically defined as a sample size of at least N = 30); and (3) its variability is represented by the standard error of the mean, which is the standard deviation of the sampling distribution of the mean. These features are defined by a statistical principle known as the central limit theorem.
There are two types of standard error of the mean. The population standard error of the mean σ X ¯ is calculated when the population standard deviation (σ) is known; the standard error of the mean S X ¯ is calculated when the standard deviation (s) is used to estimate an unknown population standard deviation.
Making the decision about the null hypothesis for the z-test and the t-test involves many of the same steps; however, the z-test uses the standard normal distribution, whereas the t-test uses the Student t-distribution, which is the distribution of values of the t-statistic based on an infinite number of samples of size N randomly drawn from the population. Using the Student t-distribution involves calculating the degrees of freedom (df) for a sample, which are the number of values or quantities that are free to vary when a statistic is used to estimate a parameter.
Three factors affecting the decision about the null hypothesis are sample size, alpha, and the directionality of the alternative hypothesis. In terms of sample size, the larger the sample size, the greater the likelihood of rejecting the null hypothesis. In terms of alpha, the larger the value of alpha, the greater the likelihood of rejecting the null hypothesis. Finally, there is a greater likelihood of rejecting the null hypothesis when the alternative hypothesis is
414
directional (one-tailed) than when it is non-directional (two-tailed). It is important to understand these factors as they are, to some degree, under the control of researchers.
415
7.10 Important Terms
sampling distribution of the mean (p. 218) sampling distribution (p. 218) standard error of the mean (p. 219) central limit theorem (p. 219) z-test for one mean (p. 225) population standard error of the mean σ X ¯ (p. 225) assumption of random sampling (p. 231) assumption of interval or ratio scale of measurement (p. 231) assumption of normality (p. 231) robust (p. 231) Student t-distribution (p. 236) degrees of freedom (df) (p. 240) t-test for one mean (p. 242) standard error of the mean S X ¯ (p. 243)
416
7.11 Formulas Introduced in this Chapter
417
z-Test for One Mean
(7-1) z = X ¯ − μ σ X ¯
Population Standard Error of the Mean σ X ¯ (7-2) σ X ¯ = σ N
Degrees of Freedom (df) for One Sample
(7-3) df = N − 1
t-Test for One Mean
(7-4) t = X ¯ − μ S X ¯
Standard Error of the Mean S X ¯
(7-5) S X ¯ = S N
418
7.12 Using SPSS
419
Testing One Sample Mean (σ Not Known): The Unique Invulnerability Study (7.6)
1. Define variable (name, # decimal places, label for the variable) and enter data for the variable.
2. Select the t-test for one sample mean procedure within SPSS.
How? (1) Click Analyze menu, (2) click Compare Means, and (3) click One- Sample T Test.
3. Select the variable to be analyzed and set value of hypothesized population mean μ.
How? (1) Click variable and , (2) type value of hypothesized population mean (μ) in Test Value box, and (3) click OK.
420
4. Examine output.
421
7.13 Exercises
1. For each of the following situations, calculate the population standard error of the mean σ X ¯ .
a. σ = 8; N = 16 b. σ = 12; N = 64 c. σ = 2; N = 25 d. σ = 3.72; N = 18 e. σ = 32.86; N = 31
2. For each of the following situations, calculate the population standard error of the mean σ X ¯ .
a. σ = 18; N = 36 b. σ = 9.42; N = 49 c. σ = 1.87; N = 60 d. σ = .91; N = 22 e. σ = 21.43; N = 106
3. For each of the following situations, calculate the z-statistic (z). a. X ¯ = 14.00 ; μ = 11; σ = 6; N = 36 b. X ¯ = 7.00 ; μ = 8; σ = 3; N = 9 c. X ¯ = 2.86 ; μ = 2.69; σ = .40;N = 29 d. X ¯ = 10.12 ; μ = 7.98; σ = 4.59;N = 26 e. X ¯ = 92.87 ; μ = 101.55;σ = 20.65; N = 43
4. For each of the following situations, calculate the z-statistic (z). a. X ¯ = 8.00 ; μ = 5; σ = 6; N = 16 b. X ¯ = 4.00 ; μ = 2; σ = 8; N = 25 c. X ¯ = 11.50 ; μ = 9.25; σ = 5.75; N = 38 d. X ¯ = .95 ; μ = .82; σ = .31; N = 15 e. X ¯ = 74.59 ; μ = 81.29; σ = 13.54; N = 26
5. For each of the following situations, calculate the z-statistic (z), make a decision about the null hypothesis (reject, do not reject), and indicate the level of significance (p > .05, p < .05, p < .01).
a. X ¯ = 40.00 ; μ = 30; σ X ¯ = 10.00 b. X ¯ = 3.70 ; μ = 3.46; σ X ¯ = .14 c. X ¯ = 19.23 ; μ = 15.01; σ X ¯ = 1.47 d. X ¯ = 132.65 ; μ = 154.90; σ X ¯ = 11.79
6. For each of the following situations, calculate the z-statistic (z), make a decision about the null hypothesis (reject, do not reject), and indicate the level of significance (p > .05, p < .05, p < .01).
a. X ¯ = 8.00 ; μ = 16; σ X ¯ = 4.00 b. X ¯ = 39.54 ; μ = 34.22; σ X ¯ = 2.18 c. X ¯ = 1.19 ; μ = .92; σ X ¯ = .17
422
d. X ¯ = 56.81 ; μ = 64.45; σ X ¯ = 2.68 7. For each of the following situations, calculate the population standard error of the
mean σ X ¯ and the z-statistic (z), make a decision about the null hypothesis, and indicate the level of significance.
a. X ¯ = 4.00 ; μ = 2; σ = 6; N = 36 b. X ¯ = 24.52 ; μ = 19.76; σ = 10.93; N = 23 c. X ¯ = 8.11 ; μ = 10.12; σ = 3.28; N = 19 d. X ¯ = 4.54 ; μ = 3.89; σ = 2.32; N = 33
8. For each of the following situations, calculate the population standard error of the mean σ X ¯ and the z-statistic (z), make a decision about the null hypothesis, and indicate the level of significance.
a. X ¯ = 12.00 ; μ = 13; σ = 5; N = 25 b. X ¯ = 1.82 ; μ = 1.53; σ = .67; N = 40 c. X ¯ = 6.64 ; μ = 5.94; σ = 1.72; N = 34 d. X ¯ = 76.29 ; μ = 87.71; σ = 30.76; N = 26
9. Students applying to graduate schools in many disciplines are required to take the Graduate Record Examination (GRE); essentially, it is the graduate school equivalent of the SAT. Let's say an enterprising student develops a course she believes will increase students' GRE scores. She develops a sample of 25 students believed to be representative of the larger college student population and puts them through this course. Next, they take the GRE and receive a mean score of 1075.00. Assuming the GRE has a population mean (μ) of 1,000 and a standard deviation (σ) of 200, does the course appear to significantly increase GRE scores?
a. State the null and alternative hypotheses (H0 and H1) (allow for the possibility that the average GRE score may be less than 1,000).
b. Make a decision about the null hypothesis. 1. Set alpha (α), identify the critical values, and state a decision rule. 2. Calculate a statistic: z-test for one mean. 3. Make a decision whether to reject the null hypothesis. 4. Determine the level of significance.
c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
10. In the GRE test example (Exercise 9), what if it was believed that the only possible alternative to the null hypothesis is one in which the students' GRE scores increase (i.e., they cannot decrease).
a. State the null and alternative hypotheses (H0 and H1). b. Make a decision about the null hypothesis.
1. Set alpha (α), identify the critical values, and state a decision rule. Why is the decision rule different in this situation?
2. Calculate a statistic: z-test for one mean. Is it a different value in this situation? Why or why not?
3. Make a decision whether to reject the null hypothesis.
423
4. Determine the level of significance. c. Draw a conclusion from the analysis. d. What are the implications of this analysis for the GRE class?
11. One study examined the personal values of 116 students studying mortuary science with the intention of becoming funeral directors (Shaw & Duys, 2005). The students completed a well-established survey measuring different work-related values, one of which was the extent to which the student valued social interaction. The researchers tested the difference between the mean social interaction score of these students X ¯ = 12.91 with an estimated population mean (μ) of 14.50 and a standard deviation (σ) of 2.87.
a. State the null and alternative hypotheses (H0 and H1). b. Make a decision about the null hypothesis.
1. Set alpha (α), identify the critical values, and state a decision rule. 2. Calculate a statistic: z-test for one mean. 3. Make a decision whether to reject the null hypothesis. 4. Determine the level of significance.
c. Draw a conclusion from the analysis. d. What might you say about the extent to which mortuary science students value
social interaction as part of their jobs? 12. For each of the following situations, calculate the degrees of freedom (df) and
determine the critical values of t. a. N = 10; α = .05; H1: μ ≠ 5 b. N = 20; α = .05; H1: μ ≠ 5 c. N = 10; α = .01; H1: μ ≠ 5 d. N = 20; α = .01; H1: μ ≠ 5 e. N = 10; α = .05; H1: μ > 5 f. N = 20; α = .05; H1: μ > 5 g. N = 28; α = .05; H1: μ ≠ 5
13. For each of the following situations, calculate the standard error of the mean s X ¯ . a. s = 7.00; N = 49 b. s = 2.50; N = 14 c. s = 8.90; N = 23 d. s = 25.61; N = 54
14. For each of the following situations, calculate the standard error of the mean s X ¯ . a. s = 5.00; N = 16 b. s = 17.82; N = 10 c. s = 2.31; N = 37 d. s = 51.32; N = 21
15. For each of the following sets of numbers, calculate the sample size (N), the mean X ¯ , the standard deviation (s), and the standard error of the mean s X ¯ :
a. 2, 3, 3, 5, 7
424
b. 6, 7, 7, 9, 10, 12, 13, 16 c. 3, 3, 5, 6, 7, 7, 8, 9, 9, 10, 11 d. 2, 5, 7, 8, 9, 10, 13, 13, 14, 16, 18, 20 e. 82, 69, 51, 95, 78, 65, 62, 87, 47, 80, 73, 82, 55, 61, 96
16. For each of the following sets of numbers, calculate the sample size (N), the mean X ¯ , the standard deviation (s), and the standard error of the mean s X ¯ :
a. 4, 5, 8, 9, 11, 11 b. 27, 19, 14, 25, 19, 22, 26, 23 c. 2.7, 3.1, 3.4, 2.1, 2.8, 3.3, 3.0, 2.9, 3.8, 3.1 d. 12, 10, 9, 11, 9, 15, 10, 9, 8, 13, 9, 10, 7 e. 11, 9, 20, 12, 10, 19, 7, 8, 23, 14, 34, 9, 17, 6, 14, 19
17. For each of the following situations, calculate the t-statistic (t): a. X ¯ = 12.00 ; μ = 10; s X ¯ = 2.00 b. X ¯ = 6.00 ; μ = 9; s X ¯ = 1.50 c. X ¯ = 4.25 ; μ = 4.25; s X ¯ = .75 d. X ¯ = 5.62 ; μ = 5.25; s X ¯ = .20 e. X ¯ = 9.75 ; μ = 10.50; s X ¯ = 1.08
18. For each of the following situations, calculate the t-statistic (t): a. X ¯ = 11.00 ; μ = 5; s X ¯ = 3.00 b. X ¯ = 26.00 ; μ = 31; s X ¯ = 2.00 c. X ¯ = 19.60 ; μ = 22; s X ¯ = 3.25 d. X ¯ = 3.27 ; μ = 3; s X ¯ = .13 e. X ¯ = 74.92 ; μ = 65.50; s X ¯ = 5.16
19. For each of the following situations, calculate the t-statistic (t): a. X ¯ = 4.00 ; μ = 5; s = 3.00; N = 36 b. X ¯ = 1.50 ; μ = 1.25; s = .75; N = 25 c. X ¯ = 16.21 ; μ = 15; s = 3.23; N = 46 d. X ¯ = 8.26 ; μ = 6.31; s = 3.71; N = 12 e. X ¯ = 93.70 ; μ = 99.80; s = 8.16; N = 21
20. For each of the following situations, calculate the t-statistic (t): a. X ¯ = .45 ; μ =.52; s = .17; N = 56 b. X ¯ = 7.75 ; μ = 6; s = 3.98; N = 40 c. X ¯ = 3.31 ; μ = 4; s = 1.33; N = 35 d. X ¯ = 37.83 ; μ = 40.95; s = 7.74; N = 21 e. X ¯ = 127.85 ; μ = 125; s = 12.62; N = 18
21. For each of the following situations, calculate the t-statistic (t): a. X ¯ = 20.00 ; μ = 18; s X ¯ = 1.00 b. X ¯ = 20.00 ; μ = 13; s X ¯ = 1.00 c. X ¯ = 12.00 ; μ = 20; s X ¯ = 1.00 d. Looking at your answers to (a-c), how is the t-statistic affected by the size of the
difference between the sample mean and the population mean X ¯ − μ ? 22. For each of the following situations, calculate the degrees of freedom (df), identify the
425
critical values (assume α =.05 [two-tailed]), calculate the t-statistic (t), make a decision about the null hypothesis, and indicate the level of significance (p > .05, p < .05, p < .01).
a. X ¯ = 2.50 ; μ = 3.50; s X ¯ = . 67 ; N = 9 b. X ¯ = 25.00 ; μ = 20; s X ¯ = 2.00 ; N = 16 c. X ¯ = .78 ; μ = .59; s X ¯ = . 09 ; N = 14 d. X ¯ = 4.91 ; μ = 2.84; s X ¯ = . 60 ; N = 21
23. For each of the following situations, calculate the degrees of freedom (df), identify the critical values (assume α =.05 [two-tailed]), calculate the t-statistic (t), make a decision about the null hypothesis, and indicate the level of significance (p > .05, p < .05, p < .01).
a. X ¯ = 12.71 ; μ = 16.49; s X ¯ = 1.91 ; N = 9 b. X ¯ = 3.85 ; μ = 3; s X ¯ = .42 ; N = 16 c. X ¯ = 24.76 ; μ = 21.55; s X ¯ = 1.27 ; N = 12 d. X ¯ = 167.32 ; μ = 187.03; s X ¯ = 6.25 ; N = 26
24. For each of the following situations, calculate the degrees of freedom (df), identify the critical values (assume α =.05 [two-tailed]), calculate the standard error of the mean s X ¯ , calculate the t-statistic (t), make a decision about the null hypothesis (reject, do not reject), and indicate the level of significance (p > .05, p < .05, p < .01).
a. X ¯ = 75.00 ; μ = 65; s = 18.00; N = 9 b. X ¯ = 64.00 ; μ = 69; s = 6.00; N = 16 c. X ¯ = 75 ; μ = 63; s = .23; N = 29 d. X ¯ = 78.51 ; μ = 67.92; s = 18.53; N = 23
25. For each of the following situations, calculate the degrees of freedom (df), identify the critical values (assume α =.05 [two-tailed]), calculate the standard error of the mean s X ¯ , calculate the t-statistic (t), make a decision about the null hypothesis (reject, do not reject), and indicate the level of significance (p > .05, p < .05, p < .01).
a. X ¯ = 14.01 ; μ = 18.34; s = 8.82; N = 17 b. X ¯ = 6.41 ; μ = 5.72; s = 1.21; N = 15 c. X ¯ = 78.89 ; μ = 70; s = 23.54; N = 13 d. X ¯ = 471.76 ; μ = 500; s = 48.51; N = 25
26. For each set of data below, calculate the sample size (N), the mean X ¯ , the standard deviation (s), the standard error of the mean s X ¯ , and the t-statistic (t). Note: assume μ = 12.
a. 9, 14, 17, 23 b. 11, 9, 5, 13, 6, 8, 10 c. 4, 12, 16, 8, 10, 10, 9, 13, 8 d. 5, 16, 14, 8, 15, 11, 13, 9, 12, 19, 13, 10, 17
27. For each of the following situations, draw a normal distribution (a.k.a. “bell-shaped curve”), mark the critical value(s), shade in the region of rejection, compare the stated t-statistic to the critical value(s), and make a decision whether to reject the null hypothesis (H0).
426
a. Critical value = ±2.262; t = 1.25 b. Critical value = ±2.074; t = −2.50 c. Critical value = 1.734; t = −2.00 d. Critical value = −1.701; t = −1.79
28. The manager of a department store disputes the company's claim that the average age of customers who buy a particular brand of clothes is 16; she believes the average age is older than 16. A clerk for the store stops the next 25 people buying these clothes and asks for their age. She calculates a mean age of 18.50 years, with a standard deviation of 2.25 years. Use the data from this sample to test the manager's claim.
a. State the null and alternative hypotheses (H0 and H1) (allow for the possibility that the average age may be younger than 16).
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values (draw the distribution), and state
a decision rule. 3. Calculate a statistic: t-test for one mean. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
29. In the department store example (Exercise 28), imagine that the sample had been N = 4 rather than N = 25 (for the sake of this example, keep the sample mean and standard deviation at 18.50 and 2.25, respectively).
a. State the null and alternative hypotheses (H0 and H1) (allow for the possibility that the average age may be younger than 16).
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values, and state a decision rule. Why is
the decision rule different than when N = 25? 3. Calculate a statistic: t-test for one mean. Why it is a different value than
when N = 25? 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
c. Draw a conclusion from the analysis. d. What are the implications of this analysis for the store manager's claim?
30. One of the examples in this chapter discussed unique invulnerability, the belief that people will provide estimates of age at time of death that are greater than the average life expectancy in the population. In a follow-up to this article, the author again demonstrated this phenomenon, this time recognizing the fact that the actuarial age of death for people with higher levels of education is actually older than 75 and may in fact be 83 (Snyder, 1997). The estimated age of death provided by these students is presented below:
427
Student Estimated Age of Death
1 80
2 94
3 82
4 84
5 88
6 85
7 96
8 88
9 106
10 82
11 85
12 102
13 102
14 80
15 85
16 90
17 89
Use the data from this class to test whether the unique invulnerability bias occurs for people with higher levels of education.
a. Calculate the mean and standard deviation of the estimates of age of death. b. State the null and alternative hypotheses (H0 and H1) (allow for the possibility
that the estimated age of death may be less than the actuarial age in the population).
c. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values (draw the distribution), and state
a decision rule. 3. Calculate a statistic: t-test for one mean. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
d. Draw a conclusion from the analysis. e. Relate the result of the analysis to the research hypothesis.
31. The American Statistical Association publishes a monthly magazine sent to all of its members. This magazine once included an article that asked, “How many chocolate
428
chips are there in a bag of Chips Ahoy cookies?” (Warner & Rutledge, 1999). Nabisco, the makers of these cookies, claims there are (at least) 1,000 chocolate chips in each 18-ounce bag of cookies. The authors tested this claim by obtaining 42 bags of cookies, dissolving the cookies in water to separate the chocolate chips from the dough, and counting the number of chocolate chips in each bag. For the sake of simplicity, the number of chips in 18 of their bags is listed below (these 18 bags are representative of the larger sample of 42 bags):
Bag # Chips
1 1,103
2 1,219
3 1,345
4 1,258
5 1,307
6 1,419
7 1,121
8 1,185
9 1,325
10 1,269
11 1,440
12 1,132
13 1,219
14 1,191
15 1,166
16 1,270
17 1,215
18 1,514
Use these data to test Nabisco's “Chips Ahoy! 1,000 Chips Challenge.” a. Calculate the mean X ¯ and standard deviation (s) of the sample of bags of
cookies. b. State the null and alternative hypotheses (H0 and H1) (allow for the possibility
that the number of chips may either be less or greater than that claimed by Nabisco).
c. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values (draw the distribution), and state
429
a decision rule. 3. Calculate a statistic: t-test for one mean. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
d. Draw a conclusion from the analysis. e. What are the implications of this analysis for those who disbelieve Nabisco's
claim? 32. A sixth-grade teacher uses a new method of teaching mathematics to her students,
one she believes will increase their level of mathematical ability. To assess their mathematical ability, she administers a standardized test (the Test of Computational Knowledge [TOCK]). Scores on this test can range from a minimum of 20 to a possible maximum of 80. The TOCK scores for the 20 students are listed below:
Student TOCK Score
1 58
2 67
3 54
4 45
5 59
6 55
7 69
8 36
9 48
10 77
11 68
12 43
13 65
14 56
15 61
16 39
17 75
18 34
19 49
20 52
430
To interpret their level of performance on the test, she wishes to compare their mean with the hypothesized population mean (μ) of 50 for sixth graders in her state.
a. Calculate the mean X ¯ and standard deviation (s) of the TOCK scores. b. State the null and alternative hypotheses (H0 and H1) (allow for the possibility
that the new teaching method may, for some reason, lower students' mathematical ability).
c. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values (draw the distribution), and state
a decision rule. 3. Calculate a statistic: t-test for one mean. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
d. Draw a conclusion from the analysis. e. Relate the result of the analysis to the research hypothesis.
431
Answers to Learning Checks
Learning Check 2
2. a. σ X ¯ = 3.00 b. σ X ¯ = .33 c. σ X ¯ = .88 d. σ X ¯ = 5.62
3. a. z = 2.00 b. z = −2.00 c. z = .89 d. z = 2.28
Learning Check 3
2. a. z = 2.00; Reject H0; p < .05 b. z = −1.00; Do not reject H0; p > .05 c. z = 2.41; Reject H0; p < .05
3. a. σ X ¯ = 1.00 ; z = 3.00; Reject H0; p < .01 b. σ X ¯ = 2.00 ; z = −1.00; Do not reject H0; p > .05 c. σ X ¯ = .92 ; z = 1.90; Do not reject H0; p > .05
Learning Check 5
2. a. df = 8; critical value = ±2.306 b. df = 12; critical value = ±2.179 c. df = 17; critical value = ±2.110 d. df = 20; critical value = ±2.086
3. a. s X ¯ = . 75 b. s X ¯ = 1.50 c. s X ¯ = 2.83 d. s X ¯ = . 55
4.
432
a. t = 2.50 b. t = −1.33 c. t = −2.06
Learning Check 6
1. a. df = 24; ±2.064; t = 2.40; Reject H0; p < .05 b. df = 15; ±2.131; t = 1.38; Do not reject H0; p > .05 c. df = 18; ±2.101; t = −3.55; Reject H0; p < .01
2. a. df = 10; ±2.228; s X ¯ = 1.21 ; t = 3.32; Reject H0; p < .01 b. df = 26; ±2.060; s X ¯ = . 33 ; t = .87; Do not reject H0; p > .05 c. df = 18; ±2.101; s X ¯ = 1.20 ; t = −2.29; Reject H0; p < .05
433
Answers to Odd-Numbered Exercises
1. a. σ X ¯ = 2.00 b. σ X ¯ = 1.50 c. σ X ¯ = .40 d. σ X ¯ = .88 e. σ X ¯ = 5.90
3. a. z = 3.00 b. z = −1.00 c. z = 2.43 d. z = 2.38 e. z = −2.76
5. a. z = 1.00; Do not reject H0 p > .05 b. z = 1.71; Do not reject H0 p > .05 c. z = 2.87; Reject H0 p < .01 d. z = −1.89; Do not reject H0 p > .05
7. a. σ X ¯ = 1.00 ; z = 2.00; Reject H0 p < .05 b. σ X ¯ = 2.28 ; z = 2.09; Reject H0 p < .05 c. σ X ¯ = .75 ; z = −2.68; Reject H0 p < .01 d. σ X ¯ = .40 ; z = 1.63; Do not reject H0 p > .05
9. a. H0: μ = 1,000; H1: μ ≠ 1,000 b.
1. If z < −1.96 or > 1.96, reject H0; otherwise, do not reject H0 2. σ X ¯ = 40.00 ; z = 1.88 3. z = 1.88 is not < −1.96 or > 1.96 ∴ do not reject H0 (p > .05) 4. Not applicable (H0 not rejected)
c. The mean GRE score of the 25 students taking the course (M = 1,075.00) was not significantly different from the population mean of 1,000, z = 1.88, p > .05.
d. The result of this analysis does not support the students belief that the course will significantly increase students' GRE scores.
11. a. H0: μ = 14.50; H1: μ ≠ 14.50 b.
434
1. If z < −1.96 or > 1.96, reject H0; otherwise, do not reject H0 2. σ X ¯ = .27 ; z = −5.89 3. z = −5.89 < −1.96 ∴ reject H0 (p < .05) 4. z = −5.89 < −2.58 ∴ p < .01
c. The mean social interaction value of the 116 mortuary students (M = 12.91) was significantly lower than the population mean of 14.50, z = −5.89, p < .01.
d. The result of this analysis suggests that mortuary science students do not appear to value social interaction as part of their jobs.
13. a. s X ¯ = 1.00 b. s X ¯ = .67 c. s X ¯ = 1.86 d. s X ¯ = 3.49
15. a. N = 5, X ¯ = 4.00 , s = 2.00, s X ¯ = .89 b. N = 8, X ¯ = 10.00 , s = 3.46, s X ¯ = 1.22 c. N = 11, X ¯ = 7.09 , s = 2.66, s X ¯ = .80 d. N = 12, X ¯ = 11.25 , s = 5.38, s X ¯ = 1.55 e. N = 15, X ¯ = 72.20 , s = 15.27, s X ¯ = 3.94
17. a. t = 1.00 b. t = −2.00 c. t = .00 d. t = 1.85 e. t = –.69
19. a. t = −2.00 b. t = 1.67 c. t = 2.52 d. t = 1.82 e. t = −3.43
21. a. t = 2.00 b. t = 7.00 c. t = −8.00 d. The greater the difference between the sample mean and the population mean,
the greater the absolute value of the t-statistic. 23.
a. df = 8; ±2.306; t = −1.98; Do not reject H0; p > .05 b. df = 15; ±2.131; t = 2.02; Do not reject H0; p > .05 c. df = 11; ±2.201; t = 2.53; Reject H0; p < .05
435
d. df = 25; ±2.060; t = −3.15; Reject H0; p < .01 25.
a. df = 16; ±2.120; s X ¯ = 2.14 ; t = −2.02; Do not reject H0; p > .05 b. df = 14; ±2.145; s X ¯ = .31 ; t = 2.23; Reject H0; p < .05 c. df = 12; ±2.179; s X ¯ = 6.53 ; t = 1.36; Do not reject H0; p > .05 d. df = 24; ±2.064; s X ¯ = 9.70 ; t = −2.94; Reject H0; p < .01
27.
a.
b.
c.
d.
29. a. H0: μ = 16; H1: μ ≠ 16 b.
1. df = 3 2. If t < −3.182 or > 3.182, reject H0; otherwise, do not reject H0. The
decision rule is different because the critical values are dependent on sample size—the smaller the sample, the larger the critical values.
3. s X ¯ = 1.13 ; t = 2.21. The t-statistic is different because its value is dependent on sample size—the smaller the sample size, the smaller the value of t.
436
4. t = 2.21 < 3.182 ∴ do not reject H0 (p > .05) 5. Not applicable (H not rejected)
c. The mean age of the four shoppers in the sample (M = 18.50) is not significantly different from the company's hypothesized population mean of 16 years, t(3) = 2.21, p > .05.
d. The result of this analysis does not support the managers claim that the average age of customers is older than 16.
31. a. X ¯ = 1 , 261.00 ; s = 114.01 b. H0: μ = 1,000; H1: μ ≠ 1,000 c.
1. df = 17
2. If t < −2.110 or > 2.110, reject H0; otherwise, do not reject H0 3. s X ¯ = 26.87 ; t = 9.71 4. t = 9.71 > 2.110 ∴ reject H0 (p < .05) 5. t = 9.71 > 2.898 \ p < .01
d. In the sample of 18 bags of cookies, the average number of chocolate chips (M = 1,261.00) was significantly greater than the claimed amount of 1,000, t(17) = 9.71, p < .01.
e. The results of this analysis support Nabisco's claim that there are at least 1,000 chocolate chips in each bag.
437
Sharpen your skills with SAGE edge at edge.sagepub.com/tokunaga
SAGE edge for students provides a personalized approach to help you accomplish your coursework goals in an easy-to-use learning environment. Log on to access:
eFlashcards Web Quizzes Chapter Outlines Learning Objectives Media Links SPSS Data Files
438
Chapter 8 Estimating the Mean of a Population
439
Chapter Outline 8.1 An Example From the Research: Salary Survey 8.2 Introduction to the Confidence Interval for the Mean
Using the sampling distribution of the mean to create the confidence interval for the mean 8.3 The Confidence Interval for the Mean (σ Not Known)
State the desired level of confidence Calculate the confidence interval and confidence limits
Calculate the standard error of the mean ( s X ¯ ) Calculate the degrees of freedom (df) and identify the critical value of t Calculate the confidence interval and confidence limits
Draw a conclusion about the confidence interval Probability of an interval or the probability of a population mean?
Summary 8.4 The Confidence Interval for the Mean (σ Known)
State the desired level of confidence Calculate the confidence interval and confidence limits
Calculate the population standard error of the mean ( σ X ¯ ) Identify the critical value of z Calculate the confidence interval and confidence limits
Draw a Conclusion About the Confidence Interval Summary
8.5 Factors Affecting the Size of the Confidence Interval for the Mean Sample size
Estimating the sample size needed for a desired confidence interval Level of confidence
Choosing between different levels of confidence: probability versus precision 8.6 Interval Estimation and Hypothesis Testing
Concerns about hypothesis testing The interpretation of statistical significance and nonsignificance The influence of sample size on the decision about the null hypothesis Interval estimation as an alternative to hypothesis testing
8.7 Looking Ahead 8.8 Summary 8.9 Important Terms 8.10 Formulas Introduced in This Chapter 8.11 Using SPSS 8.12 Exercises
Chapter 7 discussed how to test a research hypothesis regarding the difference between the mean of a sample and a hypothesized population mean (μ). However, it is important to understand there are situations in which the population mean for a variable is not known. In such cases, the goal of a research study may be to use a sample mean to develop, with a desired level of confidence, an estimate of an unknown population mean. The purpose of this chapter is to introduce interval estimation, defined as the estimation of a population parameter (such as μ) by specifying a range, or interval, of values within which one has a certain degree of confidence the population parameter falls. This chapter will discuss how these intervals, known as confidence intervals, are constructed and interpreted. In our
440
discussion of the relationship between confidence intervals and hypothesis testing, we will examine a variety of important issues faced by researchers.
441
8.1 An Example from the Research: Salary Survey
Psychology is a diverse discipline, with areas of specialization that include clinical, developmental, social, personality, and biological psychology. Another, less well-known, area of psychology is industrial and organizational (I-O) psychology. The main goal of I-O psychology is to apply psychological theory and methods to organizations; I-O psychologists study topics such as leadership, work motivation, and employee selection.
The Society for Industrial and Organizational Psychology (SIOP), the professional organization for I-O psychologists, publishes a newsletter for its members. In a 2007 article, Charu Khanna and Gina Medsker reported the results of a survey in which they collected employment and salary information from SIOP members (Khanna & Medsker, 2007). One of the goals of the study, which will be referred to as the salary survey study, was to learn about the salaries of those who responded to the survey. In this chapter, we will use some of the data in the study to estimate the annual salary of a population of I-O psychologists.
To collect information from a large and representative sample, a survey was e-mailed to every member of SIOP. In this example, only a small portion of this sample, the income of 36 SIOP members who had recently (within the previous 4 years) received their master's degrees, will be used. Respondents reported their income in response to a question that asked for “total income from your primary job (in thousands of U.S. dollars).” This variable, which will be called income, is measured at the ratio level of measurement. The income for the 36 participants is listed in Table 8.1.
To gain an initial sense of the data from the salary survey study, the incomes for the sample of 36 SIOP members are organized into the grouped frequency distribution table and frequency polygon in Figure 8.1. Looking at this figure, the distribution of income for the 36 participants is somewhat normally distributed, with the center of the distribution in the $50,001 to $60,000 interval.
The next step in analyzing data is to calculate descriptive statistics: the sample mean ( X ¯ ) and standard deviation (s): X ¯ = ∑ X N = 70000 + 41000 + … + 47000 + 42000 36 = 2032329 36 = 56453.58 s = ∑ ( X − X ¯ ) 2 N − 1 = ( 70000 − 56453.58 ) 2 + … + ( 42000 − 56453.58 ) 2 36 − 1 = 183505404.51 + … + 208906071.17 35 = 5148748012.75 35 = 147107086.10 = 12128.77
Looking at the descriptive statistics, we see that the mean income of 36 SIOP members who recently received their master's degrees was $56,453.58. As the mean is in the middle of the modal income interval ($50,001–$60,000), the presence of a similar mean and mode provides corroborating evidence that the incomes in this sample are normally distributed.
442
The standard deviation ($12,128.77) reflects variability in income around the $56,453.58 mean.
Table 8.1 Annual Income of 36 SIOP Members with Master's Degrees Table 8.1 Annual Income of 36 SIOP Members with
Master's Degrees
SIOP Member Income
1 70,000
2 41,000
3 52,500
4 44,200
5 61,000
6 92,000
7 37,200
8 60,000
9 70,000
10 62,723
11 45,000
12 77,353
13 49,000
14 65,500
15 64,896
16 51,500
17 63,000
18 41,416
19 55,000
20 56,000
21 70,040
22 55,000
23 50,000
24 34,000
25 56,000
26 57,614
443
27 54,000
28 62,500
29 54,000
30 49,000
31 52,500
32 76,887
33 60,000
34 52,500
35 47,000
36 42,000
444
8.2 Introduction to the Confidence Interval for the Mean
At this point in discussing the salary survey study, we might expect to encounter the steps in hypothesis testing (stating the null and alternative hypotheses, etc.) described in earlier chapters. A different approach is required, however, when the population mean is not presumed to be known. Rather than test the difference between the sample mean and a known, stated population mean, in this chapter we'll use the data from the study to estimate the value of an unknown population mean.
For the salary survey study, a sample mean of $56,453.58 has been calculated. This sample mean may be thought of as an estimate of the mean income of the population of SIOP members who recently received their master's degrees. As such, the sample mean is a point estimate, defined as a single value used to estimate an unknown population parameter. Figure 8.2(a) illustrates the concept of the point estimate.
The sample mean of $56,453.58 serves as a useful starting point for estimating the unknown average salary of SIOP members with master's degrees. However, from the discussion of the sampling distribution of the mean (Chapter 7), we know that if we were to draw additional samples of 36 SIOP members from the population, we would calculate a variety of values for the sample mean. So, rather than rely on a sample mean as the sole estimate of the population mean, the features of the sampling distribution of the mean can be used to surround the sample mean with a range of values that, with a stated degree of confidence, may include the population mean. This range of values is an example of a confidence interval, defined as a range or interval of values that has a stated probability of containing an unknown population parameter.
Figure 8.1 Grouped Frequency Distribution Table and Frequency Polygon of Income for 36 SIOP Members
445
Figure 8.2 Example of a Point Estimate and Confidence Interval
In this chapter, we will calculate the confidence interval for the mean, which is an interval of values for a variable that has a stated probability of containing an unknown population mean. The concept of the confidence interval for the mean is illustrated by the gray band in Figure 8.2(b). Constructing a confidence interval for the mean implies we have a stated level of confidence that the gray band in this figure contains an unknown population mean.
446
Using the Sampling Distribution of the Mean to Create the Confidence Interval for the Mean
Given that we are using the mean of a sample to estimate a population mean, we will rely on the sampling distribution of the mean to help construct the confidence interval for the mean. In Chapter 7, the sampling distribution of the mean was defined as the distribution of values of the sample mean for an infinite number of random samples of size N drawn from the population. Furthermore, because the confidence interval for the mean consists of a range of values around a sample mean, we are particularly interested in measuring the variability of sample means in the sampling distribution of the mean.
The variability of sample means within the sampling distribution of the mean is measured by a statistic known as the standard error of the mean. There are two versions of the standard error of the mean: the population standard error of the mean ( σ X ¯ ), which is calculated when the standard deviation for the variable in the population is known, and the standard error of the mean ( s X ¯ ), which is calculated when the standard deviation for the variable in the population is unknown and is estimated with a standard deviation (s). In this chapter, we will create confidence intervals for the mean using both versions of the standard error of the mean.
In Chapter 7, to evaluate a sample mean, the sampling distribution of the mean and the standard error of the mean were used to transform the sample mean into one of two statistics. First, the sample mean was transformed into a z-statistic in the situation where the population standard deviation for the variable (σ) was known; the z-statistic was then evaluated using the standard normal distribution. Next, the sample mean was transformed into a t-statistic when the population standard deviation was not known; the t-statistic was evaluated using the Student t-distribution, with a different distribution of t-statistics for each degrees of freedom (df), defined as the number of values or quantities that are free to vary when a statistic is used to estimate a parameter.
Similar to the test of one mean discussed in Chapter 7, the confidence interval for the mean may be calculated in one of two ways, with the choice being a function of whether the population standard deviation (σ) for the variable is known or not known. In the next section, we discuss how the confidence interval for the mean may be determined for a variable with an unknown population standard deviation. Later in this chapter, we'll construct a confidence interval for the mean for a variable with a known population standard deviation; we're discussing the two confidence intervals in this order because it is extremely rare to collect enough data from a population to be able to state a population standard deviation with any degree of certainty.
447
8.3 The Confidence Interval for the Mean (σ Not Known) This section discusses how to calculate and interpret the confidence interval for the mean when the population standard deviation is not known. There are three steps in calculating this confidence interval:
state the desired level of confidence, calculate the confidence interval and confidence limits, and draw conclusions about the confidence interval.
Each of these steps is explained and completed below, using the income variable in the salary survey study.
448
State the Desired Level of Confidence
The size or width of the confidence interval for the mean is a function of the desired level of confidence—that is, how confident one wants to be that the interval contains the population mean. As an extreme example, how big would the confidence interval have to be if we wanted to be absolutely (which is to say 100%) confident that the interval contains the population mean? To be absolutely confident, the interval would have to include all possible values of the population mean; in other words, a 100% confidence interval would have to be infinitely wide. In our current example, it would be akin to saying, “The average income in the population of SIOP members is somewhere between 0 and infinity” Although this is true, a 100% confidence interval is uninformative due to its lack of precision.
To be of practical use, a confidence interval must have a desired level of confidence less than 100%. One traditional practice is to set the desired level of confidence at 95%, such that researchers construct a 95% confidence interval around the mean: Desired level of confidence = 95%
A 95% confidence interval for the mean implies that one can say with 95% confidence that the interval contains the unknown population mean. Another way of stating this is that there is a .95 probability that the confidence interval contains the population mean and, conversely, a .05 probability the confidence interval does not contain the population mean. Later in this chapter, intervals of different levels of confidence will be calculated and compared with the 95% confidence interval.
449
Calculate the Confidence Interval and Confidence Limits
Once we have determined the desired level of confidence, the next step is to calculate the confidence interval. The formula for the confidence interval (CI) for the mean when the population standard deviation is unknown is presented in Formula 8–1:
(8-1) CI = X ¯ ± t ( s X ¯ )
where X ¯ is the point estimate (i.e., the sample mean), t is the critical value of the t-statistic for the desired level of confidence, and s X ¯ is the standard error of the mean. When σ is not known and is estimated using the standard deviation, we must rely on the t-statistic and the Student t-distribution to help construct the confidence interval. The inclusion of the ± symbol in Formula 8–1 indicates that the confidence interval extends from the sample mean in both the left and right directions.
Calculating the confidence interval for the mean using Formula 8–1 requires the completion of three steps:
calculate the standard error of the mean ( s X ¯ ), calculate the degrees of freedom (df) and identify the critical value of t, and calculate the confidence interval and confidence limits.
Each of these steps will be introduced and illustrated using the income variable in the salary survey study.
Calculate the Standard Error of the Mean ( s X ¯ )
The confidence interval for the mean is based on a sample mean, which serves as the point estimate of the population mean. To build an interval of values around the sample mean, one piece of information needed to construct the confidence interval is a measure of variability of sample means. When the population standard deviation for a variable is not known, the variability of sample means is represented by the standard error of the mean ( s X ¯ ).
To calculate the standard error of the mean for the data from the salary survey study the standard deviation (s) of 12,128.77 calculated earlier and the sample size (N = 36) are inserted into the following formula (Formula 7–5 in Chapter 7): s X ¯ = s N = 12128.77 36 = 12128.77 6 = 2021.46
The width of a confidence interval for the mean is a function of the amount of variability of sample means—the greater the variability of sample means, the wider an interval must be
450
to have the desired level of confidence that the interval contains the population mean.
Calculate the Degrees of Freedom (df) and Identify the Critical Value of t
Because we are using the standard deviation as an estimate of the population standard deviation, the Student t-distribution is used to help construct the confidence interval for the mean. As we saw in Chapter 7, using this distribution requires the calculation of the degrees of freedom (df) for the sample. Using Formula 7–5, the degrees of freedom for the sample in the salary survey study is calculated below: d f = N − 1 = 36 − 1 = 35
In a sample of 36 scores, the first 35 scores are free to vary; that is, there are 35 degrees of freedom.
Once the degrees of freedom have been determined, the next step is to identify the appropriate critical value of the t-statistic for the desired level of confidence. For the salary survey study, what value of t is associated with a 95% confidence interval? As a 95% confidence interval implies there is a .05 probability that the interval does not contain the population mean, we can identify the critical value for the confidence interval by identifying the values of t that have less than a .05 probability of occurring. Consequently, stating a desired level of confidence is similar to stating a significance level of alpha (α); within the context of hypothesis testing, alpha is the probability of a statistic needed to reject a null hypothesis.
Stating a 95% level of confidence implies we must identify the critical value for the α = .05 (two-tailed) region of rejection. For the salary survey study, we need to find the critical value of t for df = 35 and α = .05 (two-tailed). Turning to the table of critical t values in Table 3, we start at the df column. Moving down this column, we don't find a row for df = 35; instead, we find rows for df = 30 and df = 40. To find the critical value for the salary survey example, the df = 30 row is chosen; it would be inappropriate to use the df = 40 row because this would imply the sample is larger than it really is. Within the df = 30 row, the critical value of t for α = .05 (two-tailed) is identified by moving to the right until we are under the .05 column under the “Level of significance for two-tailed test” label; here, we find the critical value 2.042. Therefore, for the data from the salary survey study: For a 95% confidence interval and d f = 35 , critical value of t = 2.042
Calculate the Confidence Interval and Confidence Limits
Inserting the information obtained from the earlier steps (desired level of confidence [95%], point estimate [ X ¯ = 56453.58 ], standard error of the mean [ s X ¯ = 2021.46 ], and critical value of t [2.042]) into Formula 8–1, the 95% CI for the mean for the salary
451
survey study is calculated as follows: 95 % CI = X ¯ ± t ( s X ¯ ) = 56453.58 ± 2.042 ( 2021.46 ) = 56453.58 ± 4127.82
By subtracting and adding (±) 4127.82 from the point estimate of 56453.58, confidence limits, defined as the lower and upper boundaries of the confidence interval, may be calculated. The lower and upper confidence limits for the income variable are as follows: Lower limit = 56453 . 58 − 4127 . 82 Upper limit = 56453 . 58 + 4127 . 82 =52325 . 76 =60581 . 40
Therefore, the 95% confidence interval for the income variable ranges from 52325.76 to 60581.40, with the point estimate (the sample mean of 56453.58) located in the center of the interval. The confidence interval for the mean may be represented the following way: 95 % CI = ( 52325.76 , 60581.40 )
where the lower and upper confidence limits are placed within parentheses and separated using a comma sign.
452
Draw a Conclusion about the Confidence Interval
The third and final step is to draw a conclusion about the confidence interval that has been calculated. This conclusion for the salary survey study may be stated as follows:
There is a .95 probability that the interval of $52,325.76 to $60,581.40 contains the mean income for the population of recent masters degrees recipients in I-O psychology.
In other words, one can say with 95% confidence that this interval contains the unknown population mean.
Income and employment surveys of the type conducted in the salary survey study are useful for a variety of reasons. For example, an estimate of the mean income for the population may assist students in developing realistic expectations regarding their future employment. In one study, students were asked to estimate the starting salaries of those completing doctoral degrees in clinical psychology (Gallucci, 1997). Compared with the results of a salary survey, students overestimated starting salaries by as much as $20,000. Developing an estimate of the mean income in the population using a confidence interval may help students make more informed choices regarding their choice of academic specialization and their vocational goals.
Probability of an Interval or the Probability of a Population Mean?
For the salary survey study, we concluded that there is a .95 probability that the interval of $52,325.76 to $60,581.40 contains the mean income for the population of recent master's degrees recipients in I-O psychology. Let's use the sampling distribution of the mean to help explain what is inferred by this probability. One feature of this distribution is that, assuming the samples are sufficiently large (typically defined as a sample size of N ≥ 30), 95% of the sample means fall within ±1.96 standard errors of the population mean μ. (This is analogous to the standard normal distribution, in which 95% of z-scores are between ±1.96.) This feature of the sampling distribution of the mean is illustrated for a hypothetical variable in Figure 8.3(a), where the unshaded area in the middle of the distribution represents the 95% of sample means located in the area μ ± 1.96 σ X ¯ .
Still working with the hypothetical variable in Figure 8.3(a), suppose we were to draw a large number of random samples of sufficient size, and for each sample we calculate the confidence interval for the mean. The confidence intervals for these samples are graphically represented by the bars in Figure 8.3(b). Because 95% of the sample means are located
453
within ±1.96 standard errors of the population mean μ, it stands to reason that 95% of the confidence intervals will contain the population mean μ. Consequently, for any single sample, there is a .95 probability the confidence interval calculated for that sample will contain the population mean.
It is useful to make a distinction at this point between the probability of a confidence interval and the probability of a population mean. When interpreting a confidence interval, one may be tempted to draw a conclusion regarding the probability that the population mean falls in the interval. In the salary survey study, this would be represented by stating, “There is a .95 probability that the mean income for the population is between $52,325.76 and $60,581.40.” This is actually an incorrect statement, however. The .95 probability refers to the probability that the confidence interval contains an unknown population mean —it does not refer to the probability that the population mean is in the interval.
Let's look at Figure 8.3(b) to explain the distinction between the probability of a confidence interval and the probability of a population mean. In this figure, we see that the population mean μ, represented by the vertical line below the μ symbol, remains the same across all of the samples. In other words, the population mean does not change—it is either located within an interval or it is not. As such, it does not make sense to talk about its “probability” of being in an interval. However, the intervals do change and the probability that an interval contains the population mean varies from sample to sample. Again, we describe a confidence interval for the mean as the probability the interval contains the population mean, not the probability that the population mean is in the interval.
Figure 8.3 Understanding Conclusions Drawn regarding the Confidence Interval for the Mean
454
455
Summary
Table 8.2 summarizes the steps in calculating the confidence interval for the mean when the population standard deviation σ is not known, using the salary survey study as the example. In the next section, we discuss the relatively rare situation in which the confidence interval for the mean is calculated for a variable with a known population standard deviation (σ).
Table 8.2 Summary, Calculating the Confidence Interval for the Mean (σ Not Known) (Salary Survey Study Example) Table 8.2 Summary, Calculating the Confidence Interval for the Mean (σ Not
Known) (Salary Survey Study Example) State the desired level of confidence
Desired level of confidence = 95% Calculate the confidence interval and confidence limits
Calculate the standard error of the mean ( s X ¯ ) s X ¯ = s N = 12128.77 36 = 121288.77 6 = 2012.46
Calculate the degrees of freedom (df) and identify the critical value of t
Calculate the degrees of freedom (df) d f = N − 1 = 36 − 1 = 35
Identify the critical value For a 95% confidence interval and df = 35, critical value of t = 2.042
Calculate the confidence interval and confidence limits
Calculate the confidence interval 95 % C I = X ¯ ± t ( s X ¯ ) = 56453.58 ± 2.042 ( 2021.46 ) = 56453.58 ± 4127.82
Calculate the confidence limits
L o w e r l i m i t = X ¯ − t ( s X ¯ ) = 56453.58 − 4127.82 = 52325.76 U p p e r l i m i t = X ¯ + t ( s X ¯ ) = 56453.58 + 4127.82 = 60581.40
Draw a conclusion about the confidence interval
There is a .95 probability that the interval of $52,325.76 to $60,581.40 contains the mean income for the population of recent master's degrees recipients in I-O psychology.
456
Learning Check 1: Reviewing what you've Learned So Far
1. Review questions a. For what types of situations might one calculate the confidence interval for the mean rather
than the t-test for one mean? b. What is the difference between a point estimate and the confidence interval for the mean? c. Why is a 100% confidence interval of relatively little use? d. “There is a .95 probability that the population mean is between 3 and 6.” Why is this
conclusion about a confidence interval inappropriate? 2. For each of the following situations, calculate a 95% confidence interval for the mean (σ not
known), beginning with the step, “Calculate the degrees of freedom (df) and identify the critical value of t.”
a. N = 11, X ¯ = 3.00 , s X ¯ = .50 b. N = 22, X ¯ = 7.50 , s X ¯ = 1.25 c. N = 26, X ¯ = 16.42 , s X ¯ = 2.27
3. For each of the following sets of numbers, calculate a 95% confidence interval for the mean (σ not known); before calculating the confidence interval, the sample mean ( X ¯ ) and standard deviation (s) must be calculated.
a. 3, 6, 4, 2, 5 b. 9, 13, 19, 6, 20, 15, 11 c. 59, 54, 61, 72, 50, 66, 48, 70, 53, 63
4. A team of researchers interested in reducing alcohol-related problems in college fraternity members asked 159 members to report the number of drinks consumed on a typical occasion (Larimer et al., 2001). The mean number of drinks in their sample was 4.93, with a standard deviation of 2.73. Calculate a 95% confidence interval for this set of data.
a. State the desired level of confidence.
b. Calculate the confidence interval and confidence limits. 1. Calculate the standard error of the mean ( s X ¯ ). 2. Calculate the degrees of freedom (df) and identify the critical value of t. 3. Calculate the confidence interval and confidence limits.
c. Draw a conclusion about the confidence interval.
457
8.4 The Confidence Interval for the Mean (σ Known) This section illustrates how to calculate the confidence interval for the mean for the uncommon situation where the population standard deviation (σ) for a variable is known. Let's return to the example of SAT scores introduced in Chapter 5, where the math section of the SAT has a stated population standard deviation (σ) of 100. Imagine that a high school counselor wants to develop an estimate of the mean SAT score for the population of students attending her high school. She obtains the SAT scores from a sample of 12 seniors from her school and calculates a sample mean of 550.00 ( X ¯ = 550.00 ). (We don't need to calculate the standard deviation (s) for the sample because its not needed to calculate the confidence interval mean when σ is known.)
Calculating the confidence interval for the mean when σ is known involves the same three steps as when σ is not known:
state the desired level of confidence, calculate the confidence interval and confidence limits, and draw a conclusion about the confidence interval.
However, even though the steps are the same, there are a few differences in completing these steps when σ is known; these differences are noted below.
458
State the Desired Level of Confidence
The first step requires stating the desired level of confidence that the confidence interval contains the unknown population mean. In this example, the desired level of confidence will once again be set at the traditional 95%. Desired level of confidence = 95 %
A desired level of confidence of 95% implies that there is a .95 probability that the calculated confidence interval contains an unknown population mean.
459
Calculate the Confidence Interval and Confidence Limits
The formula for the confidence interval (CI) for the mean when σ is known is provided in Formula 8–2:
(8-2) CI = X ¯ ± z ( σ X ¯ )
where X ¯ is the point estimate (i.e., the sample mean), z is the critical value of the z- statistic for the desired level of confidence, and σ X ¯ is the population standard error of the mean. There are two critical differences between Formula 8–2 and the formula for the confidence interval for the mean when σ is not known (Formula 8–1). First, the population standard error of the mean ( σ X ¯ ) is used to estimate the variability of sample means rather than the standard error of the mean ( s X ¯ ). Second, the critical value is based on the standard normal distribution of z-statistics rather than the Student t-distribution.
Using Formula 8–2 to calculate the confidence interval for the mean involves three steps:
calculate the population standard error of the mean ( σ X ¯ ), identify the critical value of z, and calculate the confidence interval and confidence limits.
Each of these three steps is described below, using the SAT example to illustrate the necessary calculations.
Calculate the Population Standard Error of the Mean ( σ X ¯ )
When σ is known, the variability of sample means is measured by the population standard error of the mean ( σ X ¯ ). Using the SATs stated population standard deviation of σ = 100 and the sample size for this example (N = 12), σ X ¯ is calculated for the SAT example using Formula 7–2 from Chapter 7: σ X ¯ = σ N = 100 12 = 100 3.46 = 28.90
Identify the Critical Value of z
The next step is to identify the critical value of a distribution associated with the desired level of confidence. When σ is known, the appropriate distribution is the standard normal distribution—this is the same distribution used in Chapter 7 to test a single mean when σ is known.
460
In the SAT example, a stated level of confidence of 95% requires us to identify the values of the z-statistic associated with the middle 95% of the standard normal distribution. In Chapter 7, these values were found to be ±1.96. Therefore, the critical value of z may be stated as follows: For a 95% confidence interval, critical value of z = 1.96
Note that, unlike the income variable in the salary survey study, we did not calculate the degrees of freedom (df) for the sample; this is because we aren't estimating the population standard deviation from the samples standard deviation.
Calculate the Confidence Interval and Confidence Limits
Using the desired level of confidence (95%), point estimate ( X ¯ = 550.00 ), population standard error of the mean ( σ X ¯ = 28.90 ), and critical value of z (1.96), we are ready to calculate the confidence interval for the mean for the SAT example using Formula 8–2: 95 % CI = X ¯ ± z ( σ X ¯ ) = 550.00 ± 1.96 ( 28.90 ) = 550.00 ± 56.64
Next, the lower and upper confidence limits for this 95% confidence interval may be calculated: Lower limit = 550.00 − 56.64 Upper limit = 550.00 + 56.64 = 493.36 = 606.64
The 95% confidence interval for the SAT example may be represented by the following: 95 % CI = ( 493.36 , 606.64 )
461
Draw a Conclusion about the Confidence Interval
A conclusion regarding the confidence interval for the hypothetical SAT example may be described in the following way:
There is a .95 probability that the interval of 493.36 to 606.64 contains the mean SAT score for the population of students attending the counselors high school.
In drawing conclusions about the confidence interval for the mean, its important to explicitly identify the level of confidence (“.95 probability”), the variable being measured (“SAT score”), and the population whose mean is being estimated (“the population of students attending the counselors high school”).
462
Summary
Table 8.3 summarizes the steps in calculating the confidence interval for the mean when the population standard deviation σ is known, using the SAT variable as the example. The main purpose of the confidence interval for the mean is to construct a range of values for a variable that has a stated probability of containing an unknown population mean. As such, the smaller the interval, the more precisely the population mean can be estimated. In the next section, we discuss two factors under the control of researchers that influence the width or size of a confidence interval for the mean.
463
8.5 Factors Affecting the Width of the Confidence Interval for the Mean
In Chapter 7, we discussed factors that affect the decision to reject the null hypothesis. Similarly, there are factors that influence the width or size of a confidence interval for the mean, which in turn influences a researchers ability to estimate an unknown population mean with a desired degree of precision. In this section, we will use the income variable from the SIOP salary survey study to illustrate two of these factors:
sample size, and the desired level of confidence.
As with hypothesis testing, it is important to understand the role these factors play in the development of a confidence interval as they are to a certain degree under the control and discretion of researchers.
Table 8.3 Summary, Calculating the Confidence Interval for the Mean (σ Known) (SAT Example)
Table 8.3 Summary, Calculating the Confidence Interval for the Mean (σ Known) (SAT Example)
State the desired level of confidence
Desired level of confidence = 95% Calculate the confidence interval and confidence limits
Calculate the population standard error of the mean ( σ X ¯ ) σ X ¯ = σ N = 100 12 = 100 3.46 = 28.90 Identify the critical value of z For a 95% confidence interval, critical value of z = 1.96 Calculate the confidence interval and confidence limits
Calculate the confidence interval 95 % C I = X ¯ ± z ( σ X ¯ ) = 55000 ± 1.96 ( 28.90 ) = 550.00 ± 56.64 Calculate the confidence limits L o w e r l i m i t = X ¯ − z ( σ X ¯ ) = 550.00 − 56.64 = 493.36 U p p e r l i m i t = X ¯ + z ( σ X ¯ ) = 550.00 + 56.64 = 606.64 Draw a conclusion about the confidence interval
There is a .95 probability that the interval of 493.36 to 606.64 contains the mean SAT math score for the population of students attending the counselor's high school.
464
Learning Check 2: Reviewing what you've Learned So Far
1. Review questions a. How does the calculation of the confidence interval for the mean change depending on
whether the population standard deviation σ is known?
2. For each of the following situations, calculate a 95% confidence interval for the mean (σ known), beginning with the step, “Identify the critical value of z.”
a. X ¯ = 4.00 , σ X ¯ = 1.25 b. X ¯ = 14.50 , σ X ¯ = 2.16 c. X ¯ = 97.34 , σ X ¯ = 3.61
3. For each of the following sets of numbers, calculate a 95% confidence interval for the mean (σ known); before calculating the confidence interval, the sample mean ( X ¯ ) and standard deviation (s) must be calculated.
a. 7, 5, 8, 7, 4, 6 (assume σ = 1.00) b. 23, 19, 21, 22, 20, 18, 22, 21 (assume σ = 2.75) c. 3.05, 2.72, 3.69, 3.11, 2.38, 2.90, 3.37, 2.56, 3.73, 3.02, 2.66 (assume σ = .50)
4. Although the link between obesity and ones physical condition is well established, a research team stated that less was known about the relationship between obesity and ones mental state of mind (Fontaine, Cheskin, & Barofsky 1996). These researchers asked 312 obese people to complete a survey measuring social functioning. The researchers reported a sample mean ( X ¯ ) on social functioning of 77.10 and a population standard deviation (σ) of 22.30. Calculate a 95% confidence interval for the mean for the social functioning variable.
a. State the desired level of confidence.
b. Calculate the confidence interval and confidence limits. a. Calculate the population standard error of the mean ( σ X ¯ ). b. Identify the critical value of z. c. Calculate the confidence interval and confidence limits.
c. Draw a conclusion about the confidence interval.
465
Sample Size
The relationship between the size of a sample and the width of the confidence interval for the mean may be summarized as follows: The larger the sample size, the narrower the confidence interval for the mean .
In other words, the larger the sample size, the more precisely one is able to estimate a population mean.
To illustrate the impact of sample size on confidence intervals, suppose the sample from the salary survey study had consisted of 100 master's degree recipients rather than 36. Table 8.4 shows the calculation of the confidence interval for the mean for the two sample sizes; to focus on the effect of sample size, we'll assume both samples have the same standard deviation (s = 12128.77). Looking at the bottom row of this table, we find that the confidence interval for N = 100 ($54,043.59, $58,863.57) is narrower than the confidence interval for N = 36 ($52,325.76, $60,581.40), which means we can more precisely estimate the population mean with the larger sample.
To further illustrate the effect of sample size on the width of a confidence interval, Figure 8.4 illustrates confidence intervals for the SIOP income variable for five different sample sizes (N = 10, 20, 36, 100, and 300). Comparing the width of the gray bands in this figure, we find that the bands become narrower as the sample size increases. As the sample size increases, the population mean can be estimated with greater precision.
Table 8.4 The Effect of Sample Size on the Confidence Interval for the Mean
Table 8.4 The Effect of Sample Size on the Confidence Interval for the Mean
Step N = 36 N = 100
Calculate s X ¯ s X ¯ = 12128.77 36 = 2021.46 s X ¯ = 12128.77 100 = 1212.88
Calculate df d f = 36 − 1 = 35 d f = 100 − 1 = 99
Identify critical value of t
2.042 1.987
Calculate confidence interval
CI = 56453.58 ± 2.042(2021.46) = 56453.58 ± 4127.82
CI = 56453.58 ± 1.987(1212.88) = 56453.58 ± 2409.99
Calculate confidence limits
($52,325.76, $60,581.40) ($54,043.59, $58,863.57)
466
The relationship between sample size and the width of a confidence interval may be explained by the effect of sampling error on the relationship between samples and populations. In Chapter 6, we defined sampling error as the difference between a statistic calculated from a sample (i.e., the sample mean) and the corresponding population parameter (i.e., the population mean) due to random, chance factors. Because larger samples more closely approximate the population than smaller samples, larger samples have less sampling error. This implies that the distribution of sample means for larger samples is more closely clustered around the population mean, with less variability among sample means. Looking at the calculations in Table 8.4, we see that the larger (N = 100) sample has a smaller value for the standard error of the mean ( s X ¯ ), which in turn narrows the width of the confidence interval. In summary, the larger the sample, the more confidence a researcher has that the sample mean estimates the unknown population mean μ and the smaller the confidence interval around the sample mean.
Estimating the Sample Size Needed for a Desired Confidence Interval
The previous section illustrated the impact of sample size on confidence intervals. As researchers may wish to estimate a population mean with a stated degree of precision, it is possible to determine the minimum sample size needed to attain a confidence interval of a desired width. The relationship between a desired interval width and necessary sample size may be described the following way:
Figure 8.4 Confidence Interval for the Mean for Different Sample Sizes, Salary Survey Study Example
The smaller the desired width of a confidence interval, the larger the necessary sample size.
To illustrate the determination of sample size for a confidence interval, imagine it is now
467
several years since the last SIOP income survey was conducted. The researchers again set out to estimate the mean income of masters degree recipients; however, they now have the goal of creating a confidence interval with a stated degree of precision: a width of $4,000. In other words, they want to surround their point estimate (the sample mean X ¯ ) with an interval that stretches from $2,000 below the sample mean to $2,000 above the sample mean. How large of a sample should they have to achieve this result?
The estimated sample size for the confidence interval for the mean of a desired width may be calculated using Formula 8–3:
(8-3) N ′ = ( z ( σ ^ ) CI ÷ 2 ) 2
where N' is the necessary sample size, z is the critical value of the z-statistic for a desired level of confidence, σ ^ is an estimate of the population standard deviation, and CI is the desired width of confidence interval.
To calculate the necessary sample size for the income variable, the first piece of information needed is the z-statistic associated with a stated confidence level. Assuming the researchers wish to have the traditional 95% confidence that the interval contains the population mean, the critical value of z is 1.96. Next, we need an estimate of the population standard deviation for the variable ( σ ^ ). As this estimate is often obtained from the results of existing or previous research, the researchers could use as their estimate the standard deviation (s) from the 2007 income survey (12128.77). Finally, in terms of the denominator of Formula 8–3, it was proposed earlier that they wanted the confidence interval to have a width of $4,000; therefore, CI = 4000.
Based on these three pieces of information (z = 1.96, σ ^ = 12128.77 , CI = 4000), the necessary sample size (N') for the new salary survey study is calculated as follows: N ′ = ( z ( σ ^ ) CI ÷ 2 ) 2 = ( 1.96 ( 12128.77 ) 4000 ÷ 2 ) 2 = ( 23772.39 2000 ) 2 = ( 11.89 ) 2 = 141.28
Rounding the result of these calculations, the researchers need a sample size of at least 142 participants to construct a 95% confidence interval with a width of $4,000.
To illustrate the relationship between desired confidence interval width and necessary sample size, Table 8.5 lists the necessary sample size for 95% confidence intervals of different widths for the SIOP income variable, using z = 1.96 and σ ^ = 12128.77 . Looking at this table, you see that the smaller the desired width of a confidence interval, the larger the necessary sample size. In other words, the greater the desired precision of the estimate of the population mean, the larger the sample must be to attain this precision.
It must be emphasized that calculating a necessary sample size does not guarantee that the confidence interval will contain the population mean; it simply determines the sample size
468
necessary to achieve the desired width of the interval. Nonetheless, the ability to determine minimum sample sizes can be of great use to researchers before they begin their research. For example, as a result of these calculations, a researcher may decide that he or she lacks the resources to collect the amount of data needed to attain their desired level of precision and may consequently change the goals of the study.
469
Level of Confidence
In addition to sample size, the width of the confidence interval is also affected by a researchers desired level of confidence that the interval contains the population mean. The relationship between level of confidence and the width of the confidence interval can be stated as follows:
The higher the desired level of confidence, the wider the confidence interval.
470
Learning Check 3: Reviewing what you've Learned So Far
1. Review questions a. What happens to the width of a confidence interval for the mean as sample size increases? b. What are the two ways sample size affects the width of the confidence interval for the mean? c. Why might a researcher want to estimate the sample size needed for a confidence interval of
a desired width?
2. Counselors at two colleges (College Blue and College Gold) want to estimate the average grade point average (GPA) of students attending their respective colleges. The College Blue counselor collects GPAs from 10 students (N = 10); however, the College Gold counselor collects GPAs from 30 students (N = 30). In analyzing their data, by an amazing coincidence, they calculate the same mean and standard deviation ( X ¯ = 2.81 , s = .90).
a. Calculate the 95% confidence interval for the mean (σ not known) for the two colleges. b. Which college has the narrower confidence interval? Why?
3. Earlier in this chapter, a 95% confidence interval for the mean (σ known) was calculated for the SAT example (N = 12). Assuming the sample mean ( X ¯ = 550.00 ) and population standard deviation (σ = 100) stay the same …
a. What is the confidence interval for the mean for N = 25 students rather than 12? b. What is the confidence interval for the mean for N = 50 students rather than 12? c. What is the confidence interval for the mean for N = 100 students rather than 12? d. What is the effect of increasing sample size on the width of the confidence interval?
4. In the SAT example, for N = 12, the 95% confidence interval had a width of about 115 points (the difference between the lower limit of 493.36 and the higher limit of 606.64). If we again use 100 as the estimate of the population standard deviation ( σ ^ ) …
a. How large should the sample be for a desired width of 80 points rather than 115? b. How large should the sample be for a desired width of 40 points rather than 115? c. How large should the sample be for a desired width of 20 points rather than 115?
In other words, the greater the desired probability that the confidence interval contains the population mean, the wider the interval must be to attain this probability.
Table 8.5 Estimated Sample Size, 95% Confidence Intervals of Different Widths, Salary Survey Study Example
Table 8.5 Estimated Sample Size, 95% Confidence Intervals of Different Widths, Salary Survey Study Example
Desired Width of Interval Estimated Sample Size (N')
12000 15.70
8000 35.32
4000 141.28
2000 565.13
1000 2260.51
471
All of our examples thus far have calculated a 95% confidence interval for the mean. This is the traditional approach, as it corresponds to the traditional value of .05 for alpha (α), the level of significance used to test the null hypothesis within hypothesis testing. However, researchers can also determine confidence intervals that are either more or less than 95%. For example, one researcher may want the probability that the interval contains the population mean to be greater than 95%, even if this wider interval implies a loss of precision in estimating the population mean. However, another researcher may choose a level of confidence less than the traditional 95%, even if a narrower interval lowers the researcher's confidence that the interval contains the population mean. In this section, we will calculate and interpret confidence intervals of different levels of confidence.
One alternative to a 95% confidence interval is a 99% confidence interval, which is an interval that has a .99 probability of containing an unknown population mean. Consequently, a 99% confidence interval has a higher probability of containing the population mean than does a 95% confidence interval. Another common alternative is a 90% confidence interval, which is an interval with a .90 probability of containing the unknown population mean. A 90% confidence interval has a lower probability of containing the population mean than does a 95% or 99% confidence interval.
In Table 8.6, 99% and 90% confidence intervals are constructed for the income variable in the SIOP salary survey study. Looking at the bottom row of this table, we see that the 99% confidence interval ($50,894.56, $62,012.60) is wider than the 90% confidence interval ($53,023.16, $59,884.00). Looking at the steps in the calculations, the 99% confidence interval is wider because the critical value of t is larger for the 99% interval (2.750) than for the 90% interval (1.697). A larger critical value is needed for a 99% confidence interval in order to encompass a larger percentage of the possible values of the population mean.
Choosing between Different Levels of Confidence: Probability versus Precision
The three bars in Figure 8.5 illustrate the 90%, 95% and 99% confidence intervals for the SIOP income variable. Comparing the three confidence intervals, we see that the 99% confidence interval is wider than the 95% interval, which in turn is wider than the 90% interval. In other words, the higher the desired level of confidence, the larger the confidence interval.
Table 8.6 The Effect of Level of Confidence on the Confidence Interval for the Mean
Table 8.6 The Effect of Level of Confidence on the Confidence Interval for the Mean
Step 99% Confidence 90% Confidence
Calculate s X ¯ s X ¯ = 12128.77 36 = 2021.46 s X ¯ = 12128.77 36 = 2021.46
472
Calculate df d f = 36 − 1 = 35 d f = 36 − 1 = 35
Identify critical value of t
2.750 1.697
Calculate confidence interval
CI = 56453.58 ± 2.750 ( 2021.46 ) = 56453.56 ± 5559.02
CI = 56453.58 ± 1.697 ( 2021.46 ) = 56453.56 ± 3430.42
Calculate confidence limits
($50,894.56, $62,012.60) ($53,023.16, $59,884.00)
Figure 8.5 Confidence Interval for the Mean for Different Levels of Confidence, Salary Survey Study Example
The choice of level of confidence is ultimately a trade-off between probability and precision. The higher the level of confidence, the greater the probability the interval contains the population mean. However, this heightened probability results in a corresponding loss of precision. For example, comparing the 99% and 90% confidence intervals for the income variable, although the 99% confidence interval has a greater probability of containing the population mean, it is almost twice as wide as the 90% confidence interval, meaning that the 99% confidence interval estimates the population mean with less precision.
The choice as to which level of confidence one might choose may depend on the other factor discussed earlier: sample size. The relationship between level of confidence and sample size may be summarized as follows:
The larger the sample size, the smaller the effect of level of confidence on the width of the confidence interval.
Table 8.7 displays 90%, 95%, and 99% confidence intervals for the income variable for three sample sizes: N = 20, 36, and 100 (some of these confidence intervals have been
473
calculated earlier in this chapter). Keeping the sample mean and standard deviation the same for the different combinations of level of confidence and sample size, we see that as the sample size increases, the level of confidence has less of an impact on the width of the confidence interval. For the largest sample size (N = 100), there is little difference between the lower and upper limits for the 90%, 95%, and 99% confidence intervals. On the other hand, a much bigger difference exists between the three confidence intervals for the smallest sample size (N = 20). Put another way, the smaller the sample, the greater the trade-off between probability and precision.
474
8.6 Interval Estimation and Hypothesis Testing
Throughout this chapter, we have identified many similarities between the confidence interval for the mean and the test of one mean discussed in Chapter 7. Both involve many of the same pieces of information, such as hypothesized population means, sample means, sampling distributions, and probability. However, they differ in their primary purpose. The main purpose of hypothesis testing is to make a decision about the null hypothesis; in testing one mean, we decide whether or not a sample mean is significantly different from a stated value of a population mean μ. In contrast, the purpose of the confidence interval for the mean is to use a sample mean to construct an estimate of an unknown population mean. In this section, we will introduce and discuss concerns about hypothesis testing and the relationship between interval estimation and hypothesis testing.
Table 8.7 Combined Effect of Sample Size and Level of Confidence on Width of the Confidence Interval, Salary Survey Study Example
Table 8.7 Combined Effect of Sample Size and Level of Confidence on Width of the Confidence Interval, Salary Survey Study Example
Level of Confidence
N = 20 N = 36 N = 100
Lower Upper Lower Upper Lower Upper
90% 51764.39 61142.77 53023.16 59884.00 54437.77 58469.39
95% 50777.20 62129.96 52325.76 60581.40 54043.59 58863.57
99% 48694.32 64212.84 50894.56 62012.60 53261.28 59645.88
475
Learning Check 4: Reviewing what you've Learned So Far
1. Review questions a. What happens to the size of a confidence interval if the level of confidence is changed from
95% to 99%? From 95% to 90%? b. What are relative strengths and weaknesses of confidence intervals of different levels of
confidence?
2. Using our earlier example of SAT (N = 12, X ¯ = 550.00 , σ = 100) … a. Calculate the 90% and 99% confidence interval for the mean (σ known). b. Which confidence interval is wider, and why?
3. As part of an examination of baseballs used in major league baseball games (Rist, 2001), a research team weighs a sample of 80 baseballs manufactured in 2000. They find the mean weight to be 5.11 ounces, with a standard deviation of .06 ounces.
a. Calculate the 90%, 95%, and 99% confidence interval for the mean (σ not known). b. Would you say the three confidence intervals are very similar or very different? Why?
476
Concerns about Hypothesis Testing
Since the 1950s, researchers have expressed concerns about the uses, and potential misuses, of hypothesis testing. For example, in 1996, one psychologist, Frank Schmidt, wrote that “my conclusion is that we must abandon the statistical significance test” (Schmidt, 1996, p. 116). Another writer, Dr. Geoffrey Loftus, wrote an article entitled, “Psychology Will Be a Much Better Science When We Change the Way We Analyze Data” (Loftus, 1996). More recently, Geoff Cumming (2014) wrote, “I conclude from the arguments and evidence I have reviewed that best research practice is not to use NHST [null hypothesis significance testing] at all” (p. 26). The following sections focus on two concerns about hypothesis testing raised by researchers such as Drs. Schmidt, Loftus, and Cumming that highlight differences between hypothesis testing and interval estimation:
the interpretation of statistical significance and nonsignificance, and the influence of sample size on the decision about the null hypothesis.
Keep in mind that these concerns are not inherent in hypothesis testing but are rather due to how researchers interpret the results of statistical analyses.
The Interpretation of Statistical Significance and Nonsignificance
Hypothesis testing centers on two mutually exclusive statistical hypotheses: the null hypothesis and the alternative hypothesis. Based on the probability of a calculated value of a statistic such as the t-statistic, one of two decisions is made: reject or do not reject the null hypothesis. When the null hypothesis is rejected, we say that the result of the analysis is “statistically significant”; if the null hypothesis is not rejected, the result is referred to as “nonsignificant.”
Dichotomizing statistical analyses as “significant” or “nonsignificant” leads to the possibility that researchers may misinterpret the results of their analyses. For example, there is the tendency to diminish or ignore nonsignificant findings, due at least in part to the tendency of academic journals to publish studies with significant rather than nonsignificant results. Consequently, readers may mistakenly believe “that if a difference or relationship is not statistically significant, then it is zero” (Schmidt, 1996, p. 126). For example, a nonsignificant statistic with a probability of .06 (just slightly greater than the α = .05 cutoff) may be regarded in a similar manner to a statistic with a probability of .99.
The Influence of Sample Size on Statistical Analyses
An overreliance on the significant versus nonsignificant dichotomy to evaluate the results of
477
a statistical analysis is also problematic because researchers may fail to take into account factors that affect the decision about the null hypothesis, one of which is sample size. As we mentioned in Chapters 6 and 7, the larger the sample size, the greater the likelihood of rejecting the null hypothesis. Uncertainty regarding the existence of a hypothesized effect or relationship can be magnified when researchers independently studying the same effect make different decisions regarding the null hypothesis simply because of differences in sample size.
Joseph Rossi, a researcher at the University of Rhode Island, was interested in the concept of “spontaneous recovery” (Rossi, 1997). To illustrate this concept, imagine you're given a set of material to learn and are tested on your recall of this material. As you might expect, you remember some, but not all, of the material. Next, you're given a second set of material to learn that's similar to the first set. Given that the two sets of material are similar, spontaneous recovery refers to the possibility that you suddenly remember information in the first set that you did not remember earlier.
Dr. Rossi analyzed 40 published research studies that tested for the existence of spontaneous recovery, where he found that half of the studies found a statistically significant spontaneous recovery effect but half did not. As you might imagine, there was confusion and disagreement among researchers regarding whether the spontaneous recovery effect did in fact exist. In analyzing these studies, Dr. Rossi found that a critical difference between the studies that found a significant effect and those that did not was the size of the sample—the studies that found a significant spontaneous recovery effect had larger samples than the studies that did not. As he concluded, “These results suggest that the inconsistency among spontaneous recovery studies may have been due to the emphasis reviewers and researchers placed on the level of [statistical] significance” (Rossi, 1997, p. 183).
478
Interval Estimation as an Alternative to Hypothesis Testing
We have thus far discussed two concerns related to the decision to reject or not reject the null hypothesis within hypothesis testing. The first concern is that, because this decision categorizes the results of a statistical analysis as either “significant” or “nonsignificant,” researchers may incorrectly conclude that a hypothesized effect either completely exists or completely does not exist. The second concern is that researchers may fail to take into consideration the critical role of sample size in making this decision and this categorization.
The decision-making process that is an integral part of hypothesis testing has the potential to lead to confusion and misunderstanding. As a result, researchers such as Drs. Schmidt and Loftus have suggested that hypothesis testing be replaced with or supplemented by interval estimation. For example, rather than make a decision as to whether the difference between a sample mean and a hypothesized population mean is statistically significant, researchers could instead use the sample mean to develop a confidence interval for the mean to estimate the population mean. Interval estimation is seen as being of particular value when researchers are not certain of the value of the population mean.
The concerns about hypothesis testing discussed in this chapter are valid and important. Nevertheless, hypothesis testing will be used as the primary method for testing research hypotheses throughout this book. As psychologist Raymond Nickerson (2000) pointed out, there are several reasons for retaining hypothesis testing. First, in relatively complex research situations, it may be possible to do hypothesis testing but not interval estimation. Second, like hypothesis testing, the width of confidence intervals is affected by such things as sample size and alpha; as such, the evaluation and interpretation of confidence intervals are subject to some of the same concerns as hypothesis testing.
One strategy employed by researchers to address the concerns related to hypothesis testing is to calculate and report statistics that estimate the size or magnitude of the hypothesized effect being tested in a statistical analysis. These statistics, known as measures of effect size, are used to supplement the decision regarding the null hypothesis. Measures of effect size, which we will introduce in Chapter 10, are valuable in that they do not involve decision making and are not affected by sample size.
Ultimately, we believe many of the perceived flaws and weaknesses of hypothesis testing are not inherent to the statistical procedures themselves but are rather a function of how researchers use and interpret these procedures. It is up to researchers to understand how to interpret and communicate the results of their statistical analyses in an appropriate and informed manner. For example, rather than focusing solely on the decision about the null hypothesis, researchers must understand and discuss how factors such as sample size and alpha may have affected this decision.
479
8.7 Looking Ahead
The purpose of this chapter was to introduce interval estimation and confidence intervals, which are statistical procedures designed to estimate unknown population parameters rather than test research hypotheses. Although there are many similarities between interval estimation and hypothesis testing, there remain several fundamental differences. We ended the chapter with a discussion of concerns about hypothesis testing, concerns that will be relevant for the remainder of this book. In the next chapter, we again introduce a statistical procedure designed to test research hypotheses. However, unlike Chapter 7, in which we were testing one mean, in the next chapter, we will test hypotheses regarding the difference between two means.
480
8.8 Summary
Interval estimation is the estimation of a population parameter by specifying a range, or interval, of values within which one has a certain degree of confidence the population parameter falls.
A point estimate is a single value used to estimate an unknown population parameter. A confidence interval is a range or interval of values that has a stated probability of containing an unknown population parameter. The confidence interval for the mean is a range of values for a variable that has a stated probability of containing an unknown population mean. Confidence limits are the lower and upper boundaries of the confidence interval.
Confidence intervals for the mean may be calculated when the population standard deviation (σ) is not known or it is known. When the population standard deviation is not known, the confidence interval for the mean is developed using the t-statistic and the Student t-distribution. When the population standard deviation is known, the confidence interval for the mean is developed using the z-statistic and the standard normal distribution.
Two factors influence the width or size of a confidence interval for the mean. The first factor is sample size, such that the larger the sample size, the narrower the width of the confidence interval. It is possible to determine the minimum sample size needed to attain a confidence interval of a desired width; the smaller the desired width of a confidence interval, the larger the necessary sample size. The second factor is the desired level of confidence, such that the higher the desired level of confidence, the wider the confidence interval.
Although there are many similarities between confidence intervals and the test of one mean (Chapter 7), there are also some critical differences between hypothesis testing and interval estimation. Researchers have expressed concerns regarding two aspects of hypothesis testing: The first is the interpretation of statistical significance and nonsignificance, such that dichotomizing statistical analyses as either “significant” or “nonsignificant” may lead to inappropriate interpretations of the results of statistical analyses. The second concern pertains to the influence of sample size on the decision about the null hypothesis. Because interval estimation does not require researchers to make a decision about the null hypothesis, researchers have suggested that hypothesis testing be replaced with or supplemented by interval estimation, particularly when researchers are not certain of the value of the population mean.
481
8.9 Important Terms
interval estimation (p. 271) point estimate (p. 273) confidence interval (p. 275) confidence interval for the mean (p. 275) confidence limits (p. 279)
482
8.10 Formulas Introduced in this Chapter
Confidence Interval for the Mean (σ Not Known)
(8-1) CI = X ¯ ± t ( s X ¯ )
Confidence Interval for the Mean (σ Known)
(8-2) CI = X ¯ ± z ( σ X ¯ )
Estimated Sample Size for the Confidence Interval for the Mean
(8-3) N ′ = ( z ( σ ^ ) CI ÷ 2 ) 2
483
8.11 Using Spss
Calculating the Confidence Interval for the Mean: The Salary Survey Study (8.1)
1. Define variable (name, # decimal places, label for the variable) and enter data for the variable.
2. Select the t-test for one sample mean procedure within SPSS.
How? (1) Click Analyze menu, (2) click Compare Means, and (3) click One- Sample T Test
3. Select the variable to be analyzed.
How? (1) Click variable and , (2) click .
484
4. Examine output.
485
8.12 Exercises
1. For each of the following situations, calculate a 95% confidence interval for the mean (σ not known), beginning with the step, “Calculate the degrees of freedom (df) and identify the critical value of t.“
a. N = 15, X ¯ = 6.00 , s X ¯ = 1.50
b. N = 24, X ¯ = 23.40 , s X ¯ = 1.73
c. N = 13, X ¯ = 8.50 , s X ¯ = .73
d. N = 19, X ¯ = 3.37 , s X ¯ = .11
2. For each of the following situations, calculate a 95% confidence interval for the mean (σ not known), beginning with the step, “Calculate the degrees of freedom (df) and identify the critical value of t.“
a. N = 11, X ¯ = 3.00 , s X ¯ = .13
b. N = 20, X ¯ = 17.83 , s X ¯ = 2.37
c. N = 17, X ¯ = 1.56 , s X ¯ = .14
d. N = 30, X ¯ = 69.71 , s X ¯ = 4.76
3. For each of the following sets of numbers, calculate a 95% confidence interval for the mean (σ not known); before going through the steps in calculating the confidence interval, the sample mean ( X ¯ ) and standard deviation (s) must first be calculated.
a. 7, 2, 5, 9, 6, 6 b. 4, 8, 13, 6, 7, 11, 15, 10 c. 6, 25, 20, 7, 10, 9, 21, 14, 11, 15 d. 89, 92, 87, 84, 90, 88, 91, 80, 87, 93, 85
4. For each of the following sets of numbers, calculate a 95% confidence interval for the mean (σ not known); before going through the steps in calculating the confidence interval, the sample mean ( X ¯ ) and standard deviation (s) must first be calculated.
a. 12, 11, 16, 9, 14, 10 b. 3.50, 2.25, 3.30, 4.75, 2.60, 4.00, 3.80 c. 24, 31, 28, 26, 19, 33, 22, 17, 25 d. 7, 4, 13, 8, 6, 10, 9, 5, 17, 10, 7, 12
5. For each of the following sets of numbers, calculate a 95% confidence interval for the mean (σ not known); before going through the steps in calculating the confidence interval, the sample mean ( X ¯ ) and standard deviation (s) must first be calculated. Next, draw a conclusion about each confidence interval.
486
a. 22, 29, 30, 25, 21, 19, 17 b. 10, 1, 7, 4, 5, 5, 6, 2 c. 2.25, 1.48, 3.31, 1.90, 2.82, 3.07, 1.98, 1.54, 2.09, 2.56, 3.81 d. 15, 29, 10, 14, 23, 9, 19, 12, 21, 34, 11, 5, 17, 13, 25
6. The exercises in Chapters 7 included a study designed to measure the number of chocolate chips in a bag of Chips Ahoy cookies (Warner & Rutledge, 1999). From this example, for the sample of 18 bags of cookies, the mean ( X ¯ ) and standard deviation (s) of the number of chips were found to be 1,261.00 and 114.01, respectively. Imagine you want to estimate the average number of chocolate chips in the population of Chips Ahoy bags of cookies. Calculate a 95% confidence interval for the mean for this set of data.
a. State the desired level of confidence.
b. Calculate the confidence interval and confidence limits. 1. Calculate the standard error of the mean ( s X ¯ ). 2. Calculate the degrees of freedom (df) and identify the critical value of t. 3. Calculate the confidence interval and confidence limits.
c. Draw a conclusion about the confidence interval.
7. In an earlier SIOP salary survey (Katkowski & Medsker, 2001), 73 of the respondents (33 males, 40 females) received their master's degrees in the 10-year period from 1991 to 2000. Listed below are descriptive statistics of the income of the two genders:
Gender N Mean ( X ¯ ) s.d. (s)
Male 33 $70,727 $41,845
Female 40 $63,200 $30,182
Calculate two 95% confidence intervals, one for males and one for females. For each gender …
a. State the desired level of confidence.
b. Calculate the confidence interval and confidence limits. 1. Calculate the standard error of the mean ( s X ¯ ). 2. Calculate the degrees of freedom (df) and identify the critical value of t. 3. Calculate the confidence interval and confidence limits.
c. Draw a conclusion about the confidence interval.
8. How long does it take an ambulance to respond to a request for emergency medical aid? One of the goals of one study was to estimate the response time of ambulances using warning lights (Ho & Lindquist, 2001). They timed a total of 67 runs in a
487
small rural county in Minnesota. They calculated the mean response time to be 8.51 minutes, with a standard deviation of 6.64 minutes. Calculate a 95% confidence interval for the mean for this set of data.
a. State the desired level of confidence.
b. Calculate the confidence interval and confidence limits. 1. Calculate the standard error of the mean ( s X ¯ ). 2. Calculate the degrees of freedom (df) and identify the critical value of t. 3. Calculate the confidence interval and confidence limits.
c. Draw a conclusion about the confidence interval.
9. For each of the following situations, calculate a 95% confidence interval for the mean (σ known), beginning with the step, “Identify the critical value of z.”
X ¯ = 7.00 , σ X ¯ = 1.00 X ¯ = 65.50 , σ X ¯ = 2.18 X ¯ = 26.40 , σ X ¯ = 1.05 X ¯ = 112.00 , σ X ¯ = 10.13
10. For each of the following situations, calculate a 95% confidence interval for the mean (σ known), beginning with the step, “Identify the critical value of z.”
X ¯ = 50.00 , σ X ¯ = 3.00 X ¯ = 1.58 , σ X ¯ = .12 X ¯ = 7.34 , σ X ¯ = 1.87 X ¯ = 32.56 , σ X ¯ = 3.70
11. For each of the following sets of numbers, calculate a 95% confidence interval for the mean (σ known); before going through the steps in calculating the confidence interval, the sample mean ( X ¯ ) and standard deviation (s) must first be calculated. Next, draw a conclusion about each confidence interval.
a. 9, 2, 8, 16, 7, 4, 12 (assume σ = 5.00) b. 1.50, .75, 3.22, 1.25, .78 (assume σ = .30) c. 116, 123, 97, 108, 112, 101, 119, 104 (assume σ = 15)
12. For each of the following sets of numbers, calculate a 95% confidence interval for the mean (σ known); before going through the steps in calculating the confidence interval, the sample mean ( X ¯ ) and standard deviation (s) must first be calculated. Next, draw a conclusion about each confidence interval.
a. 4, 7, 3, 6, 2, 5, 2, 4, 3 (assume σ = 1.50) b. 43, 34, 48, 31, 39, 32, 45, 40, 46, 37, 42 (assume σ = 6.25)
488
c. 18, .15, .20, .16, .14, .18, .22, .17, .26, .13, .15, .09, .23, .13 (assume σ = .06)
13. An exercise in Chapter 7 referred to a program designed to improve scores on the Graduate Record Examination (GRE). The class had 25 students and they scored a mean of 1075.00 on the GRE. Assuming the population standard deviation for the GRE (σ) is 200, estimate the mean GRE score for the population of students who take this program by calculating the 95% confidence interval for the mean.
a. State the desired level of confidence.
b. Calculate the confidence interval and confidence limits. 1. Calculate the population standard error of the mean ( σ X ¯ ). 2. Identify the critical value of z. 3. Calculate the confidence interval and confidence limits.
c. Draw a conclusion about the confidence interval.
14. William and Meagan are working on their senior projects, both of which involve estimating the average alcohol consumption of students on their college campuses. Imagine they achieve the same mean and standard deviation ( X ¯ = 2.60 drinks per week, s = 1.10), but the sample sizes of the two studies differ (William: N = 25, Meagan: N = 100).
a. Calculate the 95% confidence interval for the mean (σ not known) for each of the two studies.
b. Compare the width of the confidence intervals and explain why they differ.
15. In the Chips Ahoy cookie example (Exercise 6), a 95% confidence interval of (1204.30, 1317.70) was constructed for N = 18 bags of cookies. Assuming the mean and standard deviation remained the same ( X ¯ = 1261.00 , s = 114.01) …
a. What is the confidence interval for the mean for N = 36 bags rather than 18? b. What is the confidence interval for the mean for N = 200 bags rather than 18? c. What happens to the confidence interval as the sample gets larger (18 to 36 to
200)?
16. In the ambulance response time example (Exercise 8), a 95% confidence interval was constructed for N = 67 ambulance runs. Assuming the mean and standard deviation remained the same ( X ¯ = 8.51 , s = 6.64) …
a. What is the confidence interval for the mean for N = 125 ambulance runs rather than 67?
b. What is the confidence interval for the mean for N = 33 ambulance runs rather than 67?
c. How do these two confidence intervals compare with the 95% confidence interval of (6.89, 10.13) calculated in Exercise 8?
17. In the Chips Ahoy cookie example (Exercise 6), a 95% confidence interval with a
489
width of approximately 100 (the difference between the lower limit of 1204.30 and the higher limit of 1317.70) was constructed when the sample consisted of 18 bags of cookies (N = 18). Using the standard deviation of 114.01 as the estimate of the population standard deviation ( σ ^ ) …
a. How large should the sample be for a desired 95% confidence interval width of 50 chocolate chips?
b. How large should the sample be for a desired 95% confidence interval width of 30 chocolate chips?
c. How large should the sample be for a desired 95% confidence interval width of 10 chocolate chips?
18. In the ambulance response time example (Exercise 8), a 95% confidence interval with a width of approximately 3 1/4 minutes was constructed based on a sample of 67 runs (N = 18). Using the standard deviation of 6.64 as the estimate of the population standard deviation ( σ ^ ) …
a. How large should the sample be for a desired 95% confidence interval width of 2 minutes?
b. How large should the sample be for a desired 95% confidence interval width of 1 minute?
c. How large should the sample be for a desired 95% confidence interval width of 1/2 minute (30 seconds)?
19. Calculate the 90% confidence interval for the first three situations in Exercise 1 (a– d).
20. Calculate the 99% confidence interval for the first three situations in Exercise 1 (a– d).
21. Calculate the 90% and 99% confidence interval for the Chips Ahoy example (Exercise 6).
22. Calculate the 90% and 99% confidence interval for the ambulance response time example (Exercise 8).
23. A pediatrician wanted to estimate the average temperature of children who come to her for treatment. She records the temperature of 23 children and calculates a mean of 98.83° and a standard deviation of 2.11°.
a. Calculate the 90%, 95%, and the 99% confidence interval for the mean (σ not known) for her data.
b. Which of the three confidence intervals is the widest? Which is the narrowest?
490
Answers to Learning Checks
Learning Check 1
2. a. df = 10; critical value = 2.228; 95% CI = 3.00 ± 1.11; confidence limits =
(1.89, 4.11) b. df = 21; critical value = 2.080; 95% CI = 7.50 ± 2.60; confidence limits =
(4.90, 10.10) c. df = 25; critical value = 2.060; 95% CI = 16.42 ± 4.68; confidence limits =
(11.74, 21.10) 3.
a. X ¯ = 4.00 ; s = 1.58; s X ¯ = .71 ; df = 4; critical value = 2.776; 95% CI = 4.00 ± 1.97; confidence limits = (2.03, 5.97)
b. X ¯ = 14.50 ; s = 5.12; s X ¯ = 1.93 ; df = 6; critical value = 2.447; 95% CI = 14.50 ± 4.72; confidence limits = (8.56, 18.01)
c. X ¯ = 97.34 ; s = 8.29; s X ¯ = 2.62 ; df = 9; critical value = 2.262; 95% CI = 97.34 ± 5.93; confidence limits = (53.67, 65.53)
4. a. Desired level of confidence = 95% b. s X ¯ = .22 ; df = 158; critical value = 1.980; 95% CI = 4.93 ± .44; confidence
limits = (4.49, 5.37) c. There is a .95 probability the interval of 4.49 to 5.38 contains the mean
number of drinks consumed in the population of fraternity members.
Learning Check 2
2.
a. Critical value = 1.96;
95% CI = 4.00 ± 2.45;
confidence limits = (1.55, 6.45)
b. Critical value = 1.96;
95% CI = 14.50 ± 4.23;
confidence limits = (10.27, 18.73)
c. Critical value = 1.96;
95% CI = 97.34 ± 7.08;
confidence limits = (90.26, 104.42)
3. a. X ¯ = 6.17 ; σ X ¯ = .41 ; critical value = 1.96; 95% CI = 6.17 ± .80;
confidence limits = (5.36, 6.97); we are 95% confident that the interval of 5.36 to 6.97 contains the population mean.
491
b. X ¯ = 20.75 ; σ X ¯ = .97 ; critical value = 1.96; 95% CI = 20.75 ± 1.90; confidence limits = (18.85, 22.65); we are 95% confident that the interval of 18.85 to 22.65 contains the population mean.
c. X ¯ = 3.02 ; σ X ¯ = .15 ;critical value = 1.96; 95% CI = 3.02 ± .29; confidence limits = (2.72, 3.31); we are 95% confident that the interval of 2.72 to 3.31 contains the population mean.
4. a. Desired level of confidence = 95% b. σ X ¯ = 1.26 ; critical value = 1.96; 95% CI = 77.10 ± 2.47; confidence limits =
(74.63, 79.57) c. There is a .95 probability the interval of 74.63 to 79.57 contains the mean
level of social functioning in the population of people seeking treatment for obesity.
Learning Check 3
2. a. College Blue:
1. Desired level of confidence = 95% 2. s X ¯ = .28 ; df = 9; critical value = 2.262; 95% CI = 2.81 ±
.63; confidence limits = (2.18, 3.44)
College Gold: 1. Desired level of confidence = 95% 2. s X ¯ = .16 ; df = 29; critical value = 1.045; 95% CI = 2.81
± .33; confidence limits = (2.48, 3.14) The confidence interval for College Gold is narrower and more precise because of its larger sample size, which decreases the standard error of the mean and lowers the critical value. 3.
a. N = 25 1. Desired level of confidence = 95% 2. σ X ¯ = 20.00 ; critical value = 1.960; 95% CI = 550.00 ± 39.20;
confidence limits = (510.80, 589.20) b. N = 50
1. Desired level of confidence = 95% 2. σ X ¯ = 14.14 ; critical value = 1.960; 95% CI = 550.00 ± 27.71;
confidence limits = (522.29, 577.71) c. N = 100
1. Desired level of confidence = 95% 2. σ X ¯ = 10.00 ; critical value = 1.960; 95% CI = 550.00 ± 19.60;
confidence limits = (530.40, 569.60)
492
d. Increasing the sample size narrows the width of the confidence interval. For example, increasing the sample size from N = 25 to N = 100 cuts the width of the confidence interval in half.
4.
a. z = 1.96;
σ ^ = 100.00 ;
CI = 80;
N' = 24.01, sample size of 25 is needed
b. z = 1.96;
σ ^ = 100.00 ;
CI = 40;
N' = 96.04, sample size of 97 is needed
c. z = 1.96;
σ ^ = 100.00 ;
CI = 20;
N' = 384.16,
sample size of 385 is needed
Learning Check 4
2. a. 90%:
1. Desired level of confidence = 90% 2. σ X ¯ = 28.90 ; critical value = 1.64; 90% CI = 550.00 ± 47.40;
confidence limits = (502.60, 597.40)
99%: 1. Desired level of confidence = 99% 2. σ X ¯ = 28.90 ; critical value = 2.58; 99% CI = 550.00 ± 74.56;
confidence limits = (475.44, 624.56) b. The 99%: confidence interval is wider than the 90% confidence interval—in
this example, it is 50 points wider. The 99% interval is wider to have a greater level of confidence that the interval contains the population mean.
3. a. 90%:
1. Desired level of confidence = 90% 2. s X ¯ = .007 ; df = 79; critical value = 1.671; 90% CI = 5.11 ± .012;
confidence limits = (5.098, 5.122)
95%: 1. Desired level of confidence = 95% 2. s X ¯ = .007 ; df = 79; critical value = 2.000; 95% CI = 5.11 ± .014;
confidence limits = (5.096, 5.124)
99%: 1. Desired level of confidence = 99% 2. s X ¯ = .007 ; df = 79; critical value = 2.660; 99% CI = 5.11 ± .019;
confidence limits = (5.091, 5.129)
493
b. The three confidence intervals appear to be very similar in width; this is in part due to the small standard deviation (s = .06), which indicates that there is very little variability in the weights of baseballs. The lack of variability increases the precision with which the population mean may be estimated.
494
Answers to Odd-Numbered Exercises
1. a. df = 14; critical value = 2.145; 95% CI = 6.00 ± 3.22; confidence limits =
(2.78, 9.22) b. df = 23; critical value = 2.069; 95% CI = 23.40 ± 3.58; confidence limits =
(19.82, 26.98) c. df = 12; critical value = 2.179; 95% CI = 8.50 ± 1.59; confidence limits =
(6.91, 10.09) d. df = 18; critical value = 2.101; 95% CI = 3.37 ± .23; confidence limits = (3.14,
3.60) 3.
a. X ¯ = 5.83 ; s = 2.32; s X ¯ = .95 ; df = 5; critical value = 2.571; 95% CI = 5.83 ± 2.44; confidence limits = (3.39, 8.28)
b. X ¯ = 9.25 ; s = 3.69; s X ¯ = 1.31 ; df = 7; critical value = 2.365; 95% CI = 9.25 ± 3.10; confidence limits = (6.15, 12.35)
c. X ¯ = 13.80 ; s = 6.41; s X ¯ = 2.03 ; df = 9; critical value = 2.262; 95% CI = 13.80 ± 4.49; confidence limits = (9.21, 18.39)
d. X ¯ = 87.82 ; s = 3.82; s X ¯ = 1.15 ; df = 10; critical value = 2.228; 95% CI = 87.82 ± 2.56; confidence limits = (85.26, 90.38)
5. a. X ¯ = 23.29 ; s = 4.92; s X ¯ = 1.86 ; df = 6; critical value = 2.447; 95% CI =
23.29 ± 4.55; confidence limits = (18.73, 27.84); we are 95% confident that the interval of 18.73 to 27.84 contains the population mean.
b. X ¯ = 5.00 ; s = 2.83; s X ¯ = 1.00 ; df = 7; critical value = 2.365; 95% CI = 5.00 ± 2.37; confidence limits = (2.64, 7.37); we are 95% confident that the interval of 2.64 to 7.37 contains the population mean.
c. X ¯ = 2.44 ; s = .75; s X ¯ = .23 ; df = 10; critical value = 2.228; 95% CI = 2.44 ± .51; confidence limits = (1.92, 2.95); we are 95% confident that the interval of 1.92 to 2.95 contains the population mean.
d. X ¯ = 17.13 ; s = 8.02; s X ¯ = 2.07 ; df = 14; critical value = 2.145; 95% CI = 17.13 ± 4.44; confidence limits = (12.69, 21.57); we are 95% confident that the interval of 12.69 to 21.57 contains the population mean.
7. Males: a. Desired level of confidence = 95% b. s X ¯ = 7290.07 ; df = 32; critical value = 2.042; 95% CI = 70727.00 ±
14886.32; confidence limits = (55840.68, 85613.32) c. There is a .95 probability the interval of $55,840.68 to $85613.32 contains the
mean income for the population of males who recently received their master's degrees in I/O psychology.
495
Females: a. Desired level of confidence = 95% b. s X ¯ = 4775.63 ; df = 39; critical value = 2.042; 95% CI = 63200.00 ±
9751.84; confidence limits = (53448.16, 72951.84) c. There is a .95 probability the interval of $53,448.16 to $72,951.84 contains
the mean income for the population of females who recently received their master's degrees in I/O psychology.
9. a. Critical value = 1.96; 95% CI = 7.00 ± 1.96; confidence limits =
(5.04, 8.96) b. Critical value = 1.96; 95% CI = 65.50 ± 4.27; confidence limits =
(61.23, 69.77) c. Critical value = 1.96; 95% CI = 26.40 ± 2.06; confidence limits =
(24.34, 28.46) d. Critical value = 1.96; 95% CI = 112.00 ± 19.85; confidence limits =
(92.15, 131.85) 11.
a. X ¯ = 8.29 ; σ X ¯ = 1.89 ; critical value = 1.96; 95% CI = 8.29 ± 3.70; confidence limits = (4.58, 11.99); we are 95% confident that the interval of 4.58 to 11.99 contains the population mean.
b. X ¯ = 1.50 ; σ X ¯ = .13 ; critical value = 1.96; 95% CI = 1.50 ± .25; confidence limits = (1.25, 1.75); we are 95% confident that the interval of 1.25 to 1.75 contains the population mean.
c. X ¯ = 110.00 ; σ X ¯ = 5.30 ; critical value = 1.96; 95% CI = 110.00 ± 10.39; confidence limits = (99.61, 120.39); we are 95% confident that the interval of 99.61 to 120.39 contains the population mean.
13. a. Desired level of confidence = 95% b. σ X ¯ = 40.00 ; critical value = 1.96; 95% CI = 1075.00 ± 78.40; confidence
limits = (996.60, 1153.40) c. There is a .95 probability the interval of 996.60 to 1153.40 contains the mean
GRE score for the population of students that take the program. 15.
a. N = 36 1. s X ¯ = 19.00 ; df = 35; critical value = 2.042; 95% CI = 1261.00 ±
38.80; confidence limits = (1222.20, 1299.80) b. N = 200
1. s X ¯ = 8.06 ; df = 199; critical value = 1.980; 95% CI = 1261.00 ± 15.96; confidence limits = (1245.04, 1276.96)
c. As the sample size increases, the width of the confidence gets smaller. We are able to estimate the population mean with greater precision.
17.
496
a. z = 1.96; σ ^ = 114.01 ; CI = 25; N' = 79.92, a sample size of 80 is needed b. z = 1.96; σ ^ = 114.01 ; CI = 15; N' = 222.01, a sample size of 223 is needed. c. z = 1.96; σ ^ = 114.01 ; CI = 5; N' = 1997.20, a sample size of 1998 is needed.
19. a. s X ¯ = 1.50 ; df = 14; critical value = 1.761; 90% CI = 6.00 ± 2.64; confidence
limits = (3.36, 8.64) b. s X ¯ = 1.73 ; df = 23; critical value = 1.714; 90% CI = 23.40 ± 2.97;
confidence limits = (20.43, 26.37) c. s X ¯ = .73 ; df = 12; critical value = 1.782; 90% CI = 8.50 ± 1.30; confidence
limits = (7.20, 9.80) d. s X ¯ = .11 ; df = 18; critical value = 1.734; 90% CI = 3.37 ± .19; confidence
limits = (3.18, 3.56) 21.
a. 90% confidence interval 1. s X ¯ = 26.89 ; df = 17; critical value = 1.740; 90% CI = 1261.00 ±
46.79; confidence limits = (1214.21, 1307.79) b. 99% confidence interval
1. s X ¯ = 26.89 ; df = 17; critical value = 2.898; 99% CI = 1261.00 ± 77.93; confidence limits = (1183.07, 1338.93)
23. a. 90% confidence interval
1. s X ¯ = .44 ; df = 22; critical value = 1.717; 90% CI = 98.83 ± .76; confidence limits = (98.07, 99.59)
95% confidence interval
2. s X ¯ = .44 ; df = 22; critical value = 2.074; 95% CI = 98.83 ± .91; confidence limits = (97.92, 99.74)
99% confidence interval
3. s X ¯ = .44 ; df = 22; critical value = 2.819; 99% CI = 98.83 ± 1.24; confidence limits = (97.59, 100.07)
b. The 99% confidence interval is the widest; the 90% confidence interval is the narrowest (the greater the level of desired confidence, the wider the interval has to be to contain the population mean).
497
Sharpen your skills with SAGE edge at edge.sagepub.com/tokunaga
SAGE edge for students provides a personalized approach to help you accomplish your coursework goals in an easy-to-use learning environment. Log on to access:
eFlashcards Web Quizzes Chapter Outlines Learning Objectives Media Links SPSS Data Files
498
Chapter 9 Testing the Difference between Two Means
499
Chapter Outline 9.1 An Example From the Research: You Can Just Wait 9.2 The Sampling Distribution of the Difference
Characteristics of the sampling distribution of the difference 9.3 Inferential Statistics: Testing the Difference Between Two Sample Means
State the null and alternative hypotheses (H0 and H1) Make a decision about the null hypothesis
Calculate the degrees of freedom (df) Set alpha (α), identify the critical values, and state a decision rule Calculate a statistic: t-test for independent means Make a decision whether to reject the null hypothesis Determine the level of significance Draw a conclusion from the analysis Relate the result of the analysis to the research hypothesis Assumptions of the t-test for independent means Summary
9.4 Inferential Statistics: Testing the Difference Between Two Sample Means (Unequal Sample Sizes)
An example from the research: kids' motor skills and fitness Inferential statistics: testing the difference between two sample means (unequal sample sizes)
State the null and alternative hypotheses (H0 and H1) Make a decision about the null hypothesis Draw a conclusion from the analysis Relate the result of the analysis to the research hypothesis Assumptions of the t-test for independent means (unequal sample sizes) Summary
9.5 Inferential Statistics: Testing the Difference Between Paired Means Research situations appropriate for within-subjects research designs An example from the research: web-based family interventions Inferential statistics: testing the difference between paired means
Calculate the difference between the paired scores State the null and alternative hypotheses (H0 and H1) Make a decision about the null hypothesis Draw a conclusion from the analysis Relate the result of the analysis to the research hypothesis Assumptions of the t-test for dependent means Summary
9.6 Looking Ahead 9.7 Summary 9.8 Important Terms 9.9 Formulas Introduced in This Chapter 9.10 Using SPSS 9.11 Exercises
This chapter returns to a discussion of the process of calculating inferential statistics to test research hypotheses. Chapter 7 introduced this process with the simplest example, the test of one mean, which tested a hypothesis about the mean of a population by evaluating the difference between a sample mean ( X ¯ ) and a hypothesized population mean (μ). This
500
chapter, on the other hand, will discuss inferential statistics that test hypotheses about two populations by evaluating the difference between the means of two samples drawn from two populations. Although the calculations in this chapter are slightly more complicated than those in Chapter 7, throughout this chapter, we will find many critical similarities between the test of one mean and the difference between two means.
501
9.1 An Example from the Research: You can Just Wait
Like most college towns, Berkeley has an inordinate number of coffee houses. The author of this book has spent a great deal of time in these businesses preparing lectures, grading exams, and, of course, drinking a lot of coffee. It is for this reason that this book's author became introduced to a sign next to a restroom door:
Remember, how long a minute is depends on which side of the door you're on.
Does it ever seem to you that people move slower when they know you're waiting for them than when they're unaware of your presence? Well, perhaps it isn't your imagination.
Two sociologists at Pennsylvania State University, Barry Ruback and Daniel Juieng, were interested in studying territorial behavior, defined as “marking, occupying, or defending a location in order to indicate presumed rights to the particular place” (Ruback & Juieng, 1997, p. 821). Although we may think of territorial behavior as protecting our homes from burglars, the researchers studied this behavior in public places. For example, imagine you're in a busy library working on a computer terminal when someone walks up and, without saying a word to you, makes it apparent he wishes to use the computer. Do you (a) speed up to finish your work more quickly, (b) continue to work in the same manner as if no one were waiting, or (c) deliberately slow down and actually take longer than if no one were there?
The theory of territorial behavior states that people sometimes select choice (c). In describing this behavior, Ruback and Juieng (1997) proposed that when a person possesses a limited resource that is desired by others, the person will maintain possession of the resource to defend it from “intruders” and “would be territorial even when they had completed their task at the location and the territory no longer served any function to them” (p. 823).
The researchers chose to test their beliefs in a setting that may be familiar to you: a shopping mall. Picture yourself driving in a crowded parking lot when you see someone leave the mall and walk to his car. You drive over and wait for him to leave. And you wait … and wait … and wait. Based on the theory of territorial behavior, the research hypothesis in the Ruback and Juieng (1997) study was that even if people no longer need a parking space, they will take longer to leave that parking space when another driver is waiting than when no such “intruder” is present.
To collect data for their study, the researchers watched drivers leave parking spaces at a large shopping mall. As such, this study is an example of observational research, defined in
502
Chapter 1 as the systematic and objective observation of naturally occurring events, with little or no intervention on the part of the researcher.
The researchers were interested in measuring drivers in terms of the amount of time taken to leave a parking space. Using a stopwatch, they “started timing the moment the departing shopper opened the driver's side car door and stopped timing when the car had completely left the parking space” (Ruback & Juieng, 1997, p. 823). They also recorded whether or not another driver was waiting to use the parking space. If there was another driver waiting, the departing driver was defined as having an “intruder”; if not, the departing driver was defined as having “no intruder” present. Therefore, this study, which we will refer to as the parking lot study, involved two variables. Driver group, the independent variable, consisted of two groups: Intruder and No intruder. Time, the dependent variable, was measured as the amount of time in seconds taken to leave the parking space.
The original study consisted of 200 drivers. To save space, the example in this chapter will use a smaller sample of 30 drivers, equally divided between the Intruder and No intruder groups. (Although the data in this example differ from the original study, the results of the analysis mirror those reached by the researchers.) The time in seconds taken by 15 drivers in each group to leave their parking spaces is presented in Table 9.1(a); the data for each group have been organized into grouped frequency distribution tables (Table 9.1(b)). An examination of these tables shows that the departure times of the Intruder group are generally longer than those in the No intruder group. For example, the modal interval for the Intruder group is 31 to 40 seconds as opposed to 21 to 30 seconds for the No intruder group. However, the shape of the distribution for both groups is somewhat normal, with each group having a range of about 40 seconds from the quickest driver to the slowest.
The next step in analyzing the collected data is to calculate descriptive statistics of the dependent variable for each level of the independent variable (Table 9.2). This table includes additions and changes to the notational system used in this book. For example, subscripts have been added to the mathematical symbols to distinguish each group's descriptive statistics—in the parking lot study, the sample sizes for the Intruder and No intruder groups are symbolized by N1 and N2 rather than simply N. Also, the subscript i is used instead of a number to represent a group without specifying any particular one. For example, the symbol X ¯ i represents the mean of either of the two groups.
Previous chapters have created figures to illustrate the distribution of scores for a variable. For example, bar charts were used in Chapter 2 to display the frequencies for the values of a variable measured at the nominal or ordinal level of measurement, and histograms and frequency polygons were used for interval or ratio variables. This chapter introduces figures used to illustrate descriptive statistics such as measures of central tendency and variability. Just as there are different types of figures for variables, there are different ways of displaying descriptive statistics, the choice being a function of the nature of the variables.
503
When the independent variable is measured at the nominal or ordinal level of measurement, descriptive statistics are typically displayed using a bar graph. A bar graph is a figure in which bars are used to represent the mean of the dependent variable for each level of the independent variable. For the parking lot study, the bar graph in Figure 9.1 displays the descriptive statistics for the dependent variable Time for the Intruder and No intruder groups.
The height of the bars in the bar graph in Figure 9.1 represents the mean of the dependent variable (Time) for each of the two groups. The sample mean serves as an estimate of the mean of the population from which the sample was drawn. However, from our discussion of sampling error in Chapters 6 and 7, we know that the means of samples drawn from the population may vary as the result of random, chance factors; this variability is estimated using a statistic known as the standard error of the mean. To illustrate the variability of sample means, the T-shaped lines extending above and below the mean in each bar in Figure 9.1 measure one standard error of the mean above and below the sample mean ( ± 1 s X ¯ ).
Table 9.1 The Time (in Seconds) for Drivers in the Intruder and No Intruder Groups
Using Formula 7–5 from Chapter 7, the standard error of the mean ( s X ¯ ) for the two groups are calculated as follows:
Table 9.2 Descriptive Statistics of Time for Drivers in the Intruder and No Intruder Groups
Table 9.2 Descriptive Statistics of Time for Drivers in the Intruder and No Intruder
504
Groups (a) Mean ( X ¯ i )
Intruder No Intruder
X ¯ 1 = Σ X N = 23 + 62 + … + 41 + 40 15 = 611 15 = 4073
X ¯ 2 = Σ X N = 54 + 19 + … + 29 + 30 15 = 475 15 = 31.67
(b) Standard Deviation (si)
s 1 = Σ ( X − X ¯ ) 2 N − 1 = ( 23 − 40.73 ) 2 + … + ( 40 − 40.73 ) 2 15 − 1 = 314.47 + … + .54 14 = 108.50 = 10.42
s 2 = Σ ( X − X ¯ ) 2 N − 1 = ( 54 − 31.67 ) 2 + … + ( 30 − 31.67 ) 2 15 − 1 = 498.78 + … + 2.78 14 = 101.52 = 10.08
Intruder No Intruder
s X ¯ = s 1 N 1 = 10.42 15 = 10.42 3.87 = 2.69
s X ¯ = s 2 N 2 = 10.08 15 = 10.08 3.87 = 2.60
In Figure 9.1, for the Intruder group, the area covered by the T-shaped lines extends from 40.73 ± 2.69. From our discussion of confidence intervals in Chapter 8, we know that the range represented by the T-shaped line represents an interval or range with a stated probability of containing the mean of the population on the dependent variable.
Looking at Table 9.2 and Figure 9.1, we see that the mean time for the Intruder group (M = 40.73) is 9.06 seconds longer than that for the No intruder group (M = 31.67). This difference in departure time provides initial support for the hypothesis that people will take longer to leave a parking space when an intruder is present than when there is no intruder. However, to formally test the study's research hypothesis, the next step is to calculate an inferential statistic to determine whether the difference between the two sample means is statistically significant.
Figure 9.1 Bar Graph of Time for Drivers in the Intruder and No Intruder Groups
505
In Chapter 7, testing the difference between a sample mean and a population mean ( X ¯ − μ ) involved determining the probability of obtaining a particular value of the sample mean. To do this, we relied on the sampling distribution of the mean, defined as the distribution of all possible values of the sample mean when an infinite number of samples of size N are randomly selected from the population. However, as the goal in this chapter is to test the difference between two sample means ( X ¯ 1 − X ¯ 2 ), we now need to determine the probability of obtaining our particular difference between the two sample means. In the parking lot study, what is the probability of obtaining our difference in departure time of 9.06 seconds between the Intruder and No intruder groups? To answer this question, we need to create a different type of sampling distribution: a distribution of differences between sample means. This distribution is introduced and illustrated in the next section.
506
9.2 The Sampling Distribution of the Difference
The sampling distribution of the difference is the distribution of all possible values of the difference between two sample means when an infinite number of pairs of samples of size N are randomly selected from two populations. It, like the sampling distribution of the mean (Chapter 7), is an example of sampling distribution, which is a distribution of statistics for samples randomly drawn from populations. The sampling distribution of the difference is used to determine the probability of obtaining any particular difference between two sample means. As such, we will rely on this distribution to test the difference between the departure times of the Intruder and No intruder groups in the parking lot study.
To illustrate the sampling distribution of the difference, imagine you have two populations that do not differ on some variable, such that the two population means μ1 and μ2 are equal to each other. You randomly draw samples of equal size from these two populations, calculate the mean for each of the two samples, and then calculate the difference between the two sample means ( X ¯ 1 − X ¯ 2 ). What should be the difference between the two sample means? Because we have stated that the two populations do not differ, we would expect the difference between the two sample means to be equal to zero. However, given that the concept of sampling error implies variability among sample means drawn from the populations, we must also assume there will be variability among differences between sample means. Consequently, the difference between the two sample means may not be equal to zero. Furthermore, if we were to repeat this process of randomly drawing samples and calculating the difference between sample means an infinite number of times, we could create a distribution of these differences. This distribution is known as the sampling distribution of the difference.
507
Characteristics of the Sampling Distribution of the Difference
Like any other distribution, the sampling distribution of the difference may be characterized in terms of its modality, symmetry, and variability. In terms of its modality, the mean of the sampling distribution of the difference is equal to 0. This is because we are working under the assumption that the two population means are equal, such that the average difference between sample means should be equal to zero. Next, in terms of its symmetry, the sampling distribution of the difference is approximately normal, assuming the two samples are of sufficient size (typically defined as N ≥ 30); like the t-distribution discussed in Chapter 7, the shape of the distribution changes as a function of the size of the samples. Finally, the variability of the sampling distribution of the difference is measured by the standard error of the difference ( s X ¯ 1 − X ¯ 2 ), defined as the standard deviation of the sampling distribution of the difference (the formula used to calculate the standard error of the difference will be presented later in this chapter).
The characteristics of the sampling distribution of the difference enable researchers to determine the probability of obtaining any particular difference between the means of two samples. For the parking lot study, determining the probability of obtaining our difference of 9.06 seconds between the Intruder (M = 40.73) and No intruder (M = 31.67) groups will allow us to test the study's research hypothesis that people will take longer to leave a parking space when another driver is waiting than when no such “intruder” is present. The next section describes the process of calculating and evaluating an inferential statistic that employs the sampling distribution of the difference to test differences between two sample means.
508
Learning Check 1: Reviewing what you've Learned So Far
1. Review questions a. What changes to this books notational system are introduced as a function of having two
groups rather than one? b. What are the differences between visual displays of variables such as bar charts and visual
displays of descriptive statistics such as bar graphs? c. What is the main purpose and characteristics of the sampling distribution of the difference? d. What is the difference between the sampling distribution of the mean (Chapter 7) and the
sampling distribution of the difference? e. What is the difference between the standard error of the mean (Chapter 7) and the standard
error of the difference? 2. Construct a bar graph for each of the following situations (assume the independent variable is
Gender and the dependent variable is Score): a. Males (N = 6, M = 3.00, s = 1.79); Females (N = 6, M = 5.00, s = 1.41) b. Males (N = 9, M = 13.00, s = 1.73); Females (N = 9, M = 11.00, s = 2.06) c. Males (N = 12, M = .64, s = .27); Females (N = 12, M = .45, s = .25)
3. For each of the following situations, (a) calculate the mean, standard deviation, and standard error of the mean for each group and (b) construct a bar graph (assume the independent variable is Type of pet and the dependent variable is Obedience).
a. Dog: 1, 7, 1, 0, 1
Cat: 6, 10 5, 1, 8
b. Dog: .76, .80, .67, .42, .56, .78, .49, .31, .24, .69
Cat: .92, .51, .48, .65, .32, .71, .93, .26, .37, .14
c. Dog: 13, 7, 9, 11, 12, 7, 8, 10, 6, 7, 5, 10, 9, 8
Cat: 5, 4, 8, 10, 7, 6, 8, 4, 6, 9, 3, 5, 6, 7
509
9.3 Inferential Statistics: Testing the Difference between Two Sample Means
This section describes the steps involved in calculating and evaluating an inferential statistic designed to test a research hypothesis regarding the difference between two sample means. In the parking lot study, do drivers take longer to leave a parking space when another driver is waiting than when there is no intruder? To test this hypothesis, we will follow the steps introduced in earlier chapters:
state the null and alternative hypotheses (H0 and H1), make a decision about the null hypothesis, draw a conclusion from the analysis, and relate the result of the analysis to the research hypothesis.
In discussing each of these steps, we will note both similarities and differences between the research situations discussed in this chapter versus those in Chapter 7, which involved one sample mean rather than the difference between two sample means.
510
State the Null and Alternative Hypotheses (H0 and H1)
The process of hypothesis testing begins by stating the null hypothesis (H0), which reflects the conclusion that no change, difference, or relationship exists among groups or variables in the population. When testing the difference between two sample means, the null hypothesis implies that the means of the two populations are equal to each other. For the parking lot study, the null hypothesis is that the mean amount of time taken to leave a parking space for drivers with or without an intruder is equal to each other. This absence of difference between the two populations is reflected in the following null hypothesis: H 0 : μ Intruder = μ No intruder
The alternative hypothesis (H1), which is mutually exclusive from the null hypothesis, states that in the population there does in fact exist a change, difference, or relationship between groups or variables. For the parking lot study, one way to state the alternative hypothesis is to reflect the conclusion that the mean departure time in the two populations is not equal to one another. This conclusion is represented by the following alternative hypothesis: H 1 : μ Intruder ≠ μ No intruder
The use of the “not equals” (≠) symbol in the above alternative hypothesis implies that the mean departure time of drivers in the No intruder group in the population may either be greater than or less than the mean departure time of drivers in the Intruder group in the population. As such, this alternative hypothesis is considered “non-directional” or “two- tailed.” From the study's research hypothesis, which predicts that people will take a greater amount of time to leave in the presence of an intruder than when there is no intruder, it may have been reasonable for the researchers to propose the directional (one-tailed) alternative hypothesis H1: μintruder > μNo intruder However, they chose the traditional approach of a non-directional alternative hypothesis to allow for the possibility of a statistically significant difference in the opposite direction of what was anticipated.
511
Make a Decision about the Null Hypothesis
Once the null and alternative hypotheses have been stated, the next step is to make the decision whether to reject the null hypothesis. This involves completing the steps introduced in Chapter 7:
calculate the degrees of freedom (df); set alpha (α), identify the critical values, and state a decision rule; calculate a statistic: t-test for independent means; make a decision whether to reject the null hypothesis; and determine the level of significance.
Below we will complete each of these steps for the parking lot study, highlighting changes new to this chapter.
Calculate the Degrees of Freedom (df)
The degrees of freedom (df) is defined as the number of values or quantities that are free to vary when a statistic is used to estimate a parameter. In testing the mean of one sample in Chapter 7, we calculated the degrees of freedom as equal to N – 1 (the number of scores minus 1). However, because samples have now been drawn from two populations rather than one, the number of degrees of freedom has changed. Formula 9–1 presents the formula for the degrees of freedom for the difference between two sample means:
(9-1) d f = ( N 1 − 1 ) + ( N 2 − 1 )
where N1 and N2 are the sample sizes of the two groups.
For the parking lot study, where both groups have a sample size of 15 (N1 = 15), the degrees of freedom are calculated as follows: d f = ( N 1 − 1 ) + ( N 2 − 1 ) = ( 15 − 1 ) + ( 15 − 1 ) = 14 + 14 = 28
Therefore, there are a total of 28 degrees of freedom (df = 28) for the data in the parking lot study.
Set Alpha (α), Identify the Critical Values, and State a Decision Rule
The second step in making the decision to reject the null hypothesis consists of three parts. The first part is to set alpha (α), which is the probability of the inferential statistic needed
512
to reject the null hypothesis. The second part is to use the degrees of freedom and stated value of alpha to identify the critical values of the statistic; the critical values determine the values of the statistic that result in the decision to reject the null hypothesis. The third part is to state a decision rule, which is a rule that explicitly states the logic to be followed in making the decision whether to reject the null hypothesis.
In this chapter, as in earlier chapters, alpha will be set to the traditional value of .05, meaning that the null hypothesis will be rejected if the probability of the calculated value of the inferential statistic is less than .05. Furthermore, because the alternative hypothesis in the parking lot study (H1: μintruder ≠ μNo intruder) is non-directional, alpha may be stated as “α = .05 (two-tailed).”
To identify the critical values, we must first determine which particular inferential statistic will be calculated. In Chapter 7, when the population standard deviation for a variable was unknown, the sample mean was transformed into a t-statistic, which was part of the Student t-distribution. In this chapter, we have a similar situation in that the population standard deviations for the two groups are unknown. As such, we will transform the difference between the two sample means into a t-statistic and again rely on the Student t- distribution.
To demonstrate how to identify the critical value of t, we return to the table of critical values for the t-distribution provided in Table 3. Three pieces of information are needed to identify the critical value: the degrees of freedom (df), alpha (α), and the directionality of the alternative hypothesis (one-tailed or two-tailed). For the parking lot study, we calculated the degrees of freedom to be equal to 28 (df = 28); therefore, we begin by moving down the df column of Table 3 until we reach the row associated with the number 28. Next, because our alternative hypothesis is non-directional, we move to the right until we reach the set of columns labeled “Level of significance for two-tailed test.” Within this set of columns, we move to the column associated with the stated value of .05 for alpha (α). Using these three pieces of information, we find a critical t-value of 2.048. Therefore, for the parking lot study, the critical values may be stated as the following: For α = .05 ( two-tailed ) and d f = 28 , critical values ± 2.048.
The critical values and regions of rejection and non-rejection for the parking lot study are illustrated in Figure 9.2.
Once the critical values have been identified, it is helpful to explicitly state a decision rule specifying the logic to be followed in making the decision to reject or not reject the null hypothesis. For the parking lot study, the following decision rule is stated: If t < − 2.048 or > 2.048 , reject H 0 ; otherwise, do not reject H 0 .
In this situation, the null hypothesis will be rejected if the value of t we calculate from our data is either less than −2.048 or greater than 2.048 because such a value is located in the
513
region of rejection, meaning it has a low (< .05) probability of occurring.
Calculate a Statistic: T-Test for Independent Means
The next step in hypothesis testing is to calculate a value of an inferential statistic. As mentioned earlier, in testing the difference between two sample means, we will calculate a value of a t-statistic. Formula 9-2 provides the formula for the t-test for independent means, defined as an inferential statistic that tests the difference between the means of two samples drawn from two populations:
(9-2) t = X ¯ 1 − X ¯ 2 s X ¯ 1 − X ¯ 2
where X ¯ 1 and X ¯ 2 are the means for the two groups and s X ¯ 1 − X ¯ 2 is the standard error of the difference. This statistic is called the t-test for “independent” means to indicate that the data were collected from two samples drawn from two different populations and that the scores in one sample are unrelated to the scores in the other sample. (Later in this chapter, we will discuss research situations in which the two scores are related to each other because both are collected from the same participant.) For the parking lot study, “independence” implies that the No intruder and Intruder drivers are drawn from two different populations of drivers.
Figure 9.2 Critical Values and Regions of Rejection and Non-Rejection for the Parking Lot Study
If you compare Formula 9-2 with Chapter 7's Formula 7–4 (the t-test for one mean), you will find the two formulas are very similar in that the numerator of both formulas involves the difference between means and the denominator contains a measure of variability. However, the numerator of Formula 9-2 consists of the difference between two sample means rather than the difference between a sample mean and a population mean, and the denominator of this formula contains the variability of differences between sample means rather than the variability of sample means.
The first step in calculating a value of the t-statistic is to calculate the standard error of the
514
difference ( s X ¯ 1 − X ¯ 2 ) using Formula 9-3:
(9-3) s X ¯ 1 − X ¯ 2 = s 1 2 N 1 + s 2 2 N 2
where s1 and s2 are the standard deviations for the two groups and N1 and N2 are the sample sizes of the two groups. The formula for the standard error of the difference is very similar to the formula for the standard error of the mean ( s X ¯ ) discussed in Chapter 7, the critical difference being that we must now take into account the variability in two samples rather than one.
Using the standard deviations calculated in Table 9.2, the standard error of the difference for the parking lot study is calculated as follows: s X ¯ 1 − X ¯ 2 = s 1 2 N 1 + s 2 2 N 2 = ( 10.42 ) 2 15 + ( 10.08 ) 2 15 = 7.27 + 6.77 = 14.01 = 3.74
Inserting our value of the standard error of the difference and the Intruder and No intruder sample means into Formula 9-2, a value of the t-statistic may now be calculated: t = X ¯ 1 − X ¯ 2 s X ¯ 1 − X ¯ 2 = 40.73 − 31.67 3.74 = 9.06 3.74 = 2.42
As you see, the calculated value for the t-statistic (t = 2.42) is a positive number. This is because the mean of the first group, X ¯ Intruder = 40.73 , is greater than the mean of the second group, X ¯ No intruder = 31.67 . Because the designation of first and second group is arbitrary, the t-statistic would have been −2.42 if the No intruder group had been the first group and the Intruder group the second.
515
Learning Check 2: Reviewing what you've Learned So Far
1. Review questions a. In testing the difference between two sample means, what conclusions are represented by
the null and alternative hypotheses (H0 and H1)? b. What is the difference between the degrees of freedom (df) for the test of one sample mean
(Chapter 7) and the difference between two sample means? c. What are the main differences between the formulas for the t-test for one sample mean
(Chapter 7) and the t-test for independent means? d. What is the main difference between the formulas for the standard error of the mean
(Chapter 7) and the standard error of the difference? 2. State the null and alternative hypotheses (H0 and H1) for the following research questions:
a. Do children from single-parent families have higher self-esteem than children from intact families?
b. Does tap water taste any different than bottled water? c. Who is more likely to intervene when a crime takes place: someone walking alone or
someone walking with other people?
3. For each of the following situations, calculate the degrees of freedom (df) and identify the critical value of t (assume α = .05).
a. N11 = 7, N2 = 7, H1: μ1 ≠ μ2 b. N1 = 11, N2 = 11, H1: μ1 ≠ μ2 c. N1 = 8, N2 = 8, H1: μ1 > μ2
4. For each of the following situations, calculate the standard error of the difference ( s X ¯ 1 − X ¯ 2 ). a. N1 = 9, s1 = 3.00, N2 = 9, s2 = 6.00 b. N1 = 16, s1 = 8.00, N2 = 16, s2 = 12.00 c. N1 = 24, s1 = 3.38, N2 = 24, s2 = 3.76
5. For each of the following situations, calculate the standard error of the difference ( s X ¯ 1 − X ¯ 2 ) and the t-test for independent means.
a. N1 = 6, X ¯ 1 = 9.00 , s1 = 3.00, N2 = 6, X ¯ 2 = 6.00 , s2 = 2.00 b. N1 = 15, X ¯ 1 = 53.62 , s1 = 13.54, N2 = 15, X ¯ 2 = 49.12 , s2 = 11.87 c. N1 = 36, X ¯ 1 = 7.63 , s1 = 2.90, N2 = 36, X ¯ 2 = 10.50 , s2 = 3.38
Make a Decision Whether to Reject the Null Hypothesis
The next step is to make a decision whether to reject the null hypothesis by comparing the value of the t-statistic calculated from the data with the identified critical values. If the t- statistic exceeds one of the critical values, it falls in the region of rejection, and the decision is made to reject the null hypothesis. This decision implies that the difference between the means of the two groups is statistically significant, meaning it is unlikely to occur as the result of chance factors.
516
For the parking lot study, the decision about the null hypothesis may be stated the following way: t = 2.42 > 2.048 ∴ reject H 0 ( p < .05 )
Because our calculated value of t of 2.42 is greater than the critical value 2.048, it falls in the region of rejection and the decision is made to reject the null hypothesis because its probability is less than the α = .05 cutoff; in other words, p < .05. As a result, we have made the decision that the difference between the means of the two groups ( X ¯ Intruder = 40.73 and X ¯ No intruder = 31.67 ) is statistically significant.
Determine the Level of Significance
If and when the null hypothesis is rejected, it is appropriate to determine whether the probability of the calculated value of the t-statistic is not only less than .05 but also less than .01. In Chapter 7, we discussed that when researchers calculate a statistic using statistical software, they often report the exact probability of the statistic (i.e., “p = .017”) rather than use cutoffs such as “p < .05.” However, in this textbook, we will include the step of determining whether the probability of a statistic meets the .01 level of significance because “p < .01” is commonly included in tables and figures of statistics in journal articles.
To determine whether the probability of the t-statistic for the parking lot study is less than .01, we return to the table of critical values in Table 3. Moving down to the df = 28 row, we move to the right until we are under the .01 column under “Level of significance for two-tailed test.” Here, we find the critical value 2.763. Comparing the calculated t-statistic of 2.42 with the α = .01 critical value of 2.763, we reach the following conclusion: t = 2.42 < 2.763 ∴ p < .05 ( but not < .01 )
Figure 9.3 illustrates the location of the value of t from this example relative to the .05 and .01 critical values. Because the calculated t-statistic of 2.42 is greater than the .05 critical value of 2.048, it falls in the region of rejection. However, because 2.42 is less than the .01 critical value of 2.763, its probability is less than .05 but not less than .01. Consequently, we would report “p < .05” as the level of significance.
517
Draw a Conclusion from the Analysis
Given that we have made the decision to reject the null hypothesis for the data in the parking lot study, what conclusion can we make about the difference between the departure times of the Intruder and No intruder groups? One way to report the results of the analysis is the following:
Figure 9.3 Determining the Level of Significance for the Parking Lot Study
The mean departure time for the 15 drivers in the Intruder group (M = 40.73 s) is significantly greater than the mean departure time for the 15 drivers in the No intruder group (M = 31.67 s), t(28) = 2.42, p < .05.
Note that this single sentence provides the following information:
the dependent variable (“The mean departure time”), the two samples (“15 drivers … 15 drivers”), the independent variable (“Intruder group … No intruder group”), descriptive statistics (“M = 40.73 s” … “M = 31.67 s”), the nature and direction of the findings (“significantly greater than”), and information about the inferential statistic (“t(28) = 2.42, p < .05”), which indicates the inferential statistic calculated (t), the degrees of freedom (28), the value of the statistic (2.42), and the level of significance (p < .05).
518
Relate the Result of the Analysis to the Research Hypothesis
It is critical to relate the result of the statistical analysis back to the research hypothesis it was designed to test. In the parking lot study, does the finding that drivers took a significantly longer amount of time to leave a parking space in the presence of an intruder support or not support the study's research hypothesis? Here is what the authors of the study had to say:
The present series of studies is consistent with prior findings that people display territorial defense in public territories…. What is new about the present research is that it suggests people sometimes display territorial behavior merely to keep others from possessing the space even when it no longer has any value to them. (Ruback & Juieng, 1997, p. 831)
It is important to note that the researchers did not say that their study “proved” their hypothesis but rather that it “suggests people sometimes display territorial behavior.” As was discussed in Chapter 6, research hypotheses cannot be proven because data are collected from a sample rather than the entire population.
The process of testing the difference between two sample means is summarized in Table 9.3, using the parking lot study as an example. In closing, you may remember an old saying: “A watched pot never boils.” Research suggests that this old saying may really have something to teach us about how people behave in social situations. Sometimes a watched person really does take longer to move than does someone who's not being observed!
519
Assumptions of the t-Test for Independent Means
The goal of statistical procedures such as the t-test is to test hypotheses researchers have about populations by analyzing data collected from samples of these populations. These procedures make certain mathematical assumptions about populations; more specifically, they make assumptions about the distribution of scores for variables in the population. To appropriately use these procedures, researchers must determine whether the data in their samples meet these assumptions. This section discusses assumptions related to the t-test for independent means and strategies available to researchers if their data do not meet these assumptions.
Table 9.3 Summary, Testing the Difference between Two Sample Means (Parking Lot Study Example)
The first assumption related to the t-test for independent means is the assumption of normality, which was introduced in Chapter 7. This is the assumption that scores for the dependent variable in each of the two populations are approximately normally distributed; that is, the shape of the two distributions resembles normal (bell-shaped) curves. The implication of this assumption is that distributions of scores for samples drawn from these populations are also expected to be approximately normal. The assumption of normality is common to many of the statistical procedures discussed in this book.
520
The second assumption applicable to the t-test for independent means is homogeneity of variance, which is the assumption that the variance of scores for a variable in the two populations is the same. Assuming this assumption is true, the variance in two samples drawn from the two populations should also have the same amount of variance. Figure 9.4(a) illustrates two distributions that meet the assumptions of normality and homogeneity of variance.
A violation of homogeneity of variance might take place if the variance in one sample is much less or much greater than the variance in the other sample; this is illustrated in Figure 9.4(b). When the homogeneity of variance assumption is violated, it's possible that the two samples are not representative of their populations, which in turn raises doubt about the results of statistical analyses conducted on the samples.
Figure 9.4 Illustration of the Homogeneity of Variance Assumption
Violating the assumptions of normality and homogeneity of variance increases the possibility of making the wrong decision regarding the null hypothesis for a statistic such as the t-test. For example, a researcher may decide to reject the null hypothesis and conclude the difference between the two sample means is significant when the groups do not actually differ in the populations from which the samples were drawn. (These types of errors in decision making are discussed in greater detail in the next chapter.)
Statistical tests have been developed to determine whether the assumptions of normality and homogeneity of variance have been met in a set of data. However, we will not discuss these tests for two reasons: They are beyond the scope of this textbook, and the need to conduct these tests has been called into question (i.e., Zimmerman, 2004). As we mentioned in Chapter 7, research has found that statistics such as the t-statistic are robust, meaning they are able to withstand moderate violations of their mathematical assumptions. To withstand a violation of an assumption implies that, even if the data from a sample violate an assumption, the decision made regarding the null hypothesis is the same decision that would have been made if the assumption had not been violated. For example, imagine
521
that two populations do not differ on a variable that is normally distributed in both populations. Even if data from samples of these populations happen to violate the assumptions, a robust statistic such as the t-test will lead to the correct decision, which in this case is to not reject the null hypothesis. This is particularly true when the two samples are of equal size and are sufficiently large (i.e., Ni ≥ 30).
Even when a statistic is considered robust, researchers may wish to address possible violations of assumptions by altering how they analyze their data. One way to do this is to use different formulas to calculate and evaluate the t-test (Welch, 1938); we will not present these formulas for the same reasons mentioned earlier. A second way to address extreme violations of assumptions is to employ an alternative statistical procedure, one that is not based on these assumptions. For example, rather than calculate the t-test for independent means, the data can be analyzed using a statistical procedure known as the Mann-Whitney U. Examples of these alternative types of statistical procedures are introduced in Chapter 14 of this textbook.
522
Learning Check 3: Reviewing what you've Learned So Far
1. Review questions a. What are the two mathematical assumptions related to the t-test for independent means? b. What does it mean to violate the assumption of heterogeneity of variance? c. What does it mean for a statistic to be robust? d. What can researchers do when they believe assumptions have been violated?
2. How much do you think the average adult woman weighs? Let's imagine that a researcher, in response to available evidence from both the media and society, entertains the assumption that men believe that women weigh less than they actually do. Consequently, the researcher hypothesizes that men will give lower estimates of the average woman's weight than will women. To test this hypothesis, the researcher asks samples of men and women, “How much (in pounds) do you believe the average adult woman weighs?” The researcher reports the following descriptive statistics regarding the mean estimate of weight: Men: N = 13, M = 135.62, s = 19.78, Women: N = 13, M = 137.69, s = 12.68.
a. State the null and alternative hypotheses (H0 and H1).
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. For α = .05 (two-tailed), identify the critical values and state a decision rule. 3. Calculate the t-test for independent means. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
523
Summary
The parking lot study compared the means of two groups calculated from samples of equal size (Ni = 15). Having sample sizes equal to each other is preferable because it helps ensure that the two samples are treated equally in analyzing the data. However, as it is possible that the two groups may not have the same sample size, the next section describes how to test the difference between two sample means that are based on unequal sample sizes.
524
9.4 Inferential Statistics: Testing the Difference between Two Sample Means (Unequal Sample Sizes)
The example in this section compares the means of two samples when the means are based on different sample sizes. With unequal Ns, the calculation of the t-test is slightly more complicated because the formula for the standard error of the difference ( s X ¯ 1 − X ¯ 2 ) must be modified.
525
An Example from the Research: Kids' Motor Skills and Fitness
Childhood obesity is a serious concern in this country as it has been linked to both psychological problems, such as lowered self-esteem and depression, and physical problems such as diabetes and high blood pressure. A variety of interventions have been developed to combat childhood obesity, some of which have the goal of changing the long-term lifestyles and eating habits of older children and adolescents. However, two researchers at the University of Northern Iowa, Oksana Matvienko and Iradge Ahrabi-Fard, wondered whether a brief intervention focusing on “the development of motor skills applied in popular sports and games is an effective approach for increasing physical activity particularly among young and preadolescent children” (Matvienko & Ahrabi-Fard, 2010, p. 299).
The researchers developed a short, 4-week intervention aimed at developing the motor skills of kindergarten and first-grade students. As part of a 90-minute after-school program, students received short classroom lessons on topics such as human anatomy and nutrition, played exercises and games designed to increase their physical strength and endurance, and learned different motor skills such as throwing and kicking balls and rope jumping.
The researchers obtained the permission of four elementary schools to participate in the study; students at two of the schools received the intervention, and students at the other two schools did not. This study is an example of quasi-experimental research, which involves comparing preexisting groups rather than randomly assigning participants to conditions. We will refer to the independent variable in this study as Group, which consists of two levels: Intervention and Control.
Students in both the Intervention and Control groups were measured on different physical activities. In this example, we will focus on the number of times the student was able to successfully jump over a jump rope in 30 seconds; the dependent variable in this example will be called Jumps. To investigate how long any effect of the intervention may last, the number of jumps for each student was measured 4 months after the end of the intervention.
Because the number of students attending each of the schools differed, the number of students in the Intervention and Control groups was not equal to each other. For this example, the number of students in the Intervention and Control groups will be 16 and 11, respectively. Data reflecting the results of this study for the two groups are presented in Table 9.4.
The descriptive statistics for the number of jumps variable are calculated in Table 9.5; the means of the two groups are displayed in Figure 9.5. Looking at this table and figure, we
526
see that students in the Intervention group averaged a higher number of jumps in 30 seconds (M = 27.31) than did students in the Control group (M = 11.91).
527
Inferential Statistics: Testing the Difference between Two Sample Means (Unequal Sample Sizes)
As in the parking lot study, the next step in analyzing the data in the motor skills study is to calculate an inferential statistic designed to test the study's research hypothesis. As you will see in this section, the steps followed in testing the difference between two sample means are—with one important difference—the same whether or not the two sample sizes are equal. The difference is that with unequal sample sizes, the calculation of the standard error of the difference ( s X ¯ 1 − X ¯ 2 ) is slightly more complicated.
State the Null and Alternative Hypotheses (H0 and H1)
In this example, the null hypothesis is that the average number of jumps for students who receive the intervention will be the same as for students who do not receive the intervention; this represents the conclusion that the intervention does not affect students' physical fitness. The null hypothesis for the motor skills study is stated as follows: H 0 : μ Intervention = μ Control
Table 9.4 Number of Jumps for Students in the Intervention and Control Groups
Table 9.5 Descriptive Statistics of Jumps for Students in the Intervention
528
and Control Groups Table 9.5 Descriptive Statistics of Jumps for Students in the Intervention and Control
Groups (a) Mean ( X ¯ i )
Intervention Control
X ¯ 1 = Σ X N = 37 + 16 + … + 38 + 29 16 = 437 16 = 27.31
X ¯ 2 = Σ X N = 16 + 6 + … + 3 + 6 11 = 131 11 = 11.91
(b) Standard Deviation (si)
s 1 = Σ ( X − X ¯ ) 2 N − 1 = ( 37 − 27.31 ) 2 + … + ( 29 − 27.31 ) 2 16 − 1 = 93.85 + … + 2.85 15 = 103.70 = 10.18
s 2 = Σ ( X − X ¯ ) 2 N − 1 = ( 16 − 11.91 ) 2 + … + ( 16 − 11.91 ) 2 11 − 1 = 16.74 + … + 34.92 10 = 94.89 = 9.74
In this example, a non-directional (two-tailed) alternative hypothesis will be used, stating that the mean number of jumps for the two groups of students is not equal: H 1 : μ Intervention ≠ μ Control
Make a Decision about the Null Hypothesis
The first step in making the decision whether to reject the null hypothesis is to calculate the degrees of freedom (df). Inserting the sample sizes N1 = 16 and N2 = 11 for the Intervention and Control groups into Formula 9–1, the degrees of freedom for the motor skills study are d f = ( N 1 − 1 ) + ( N 2 − 1 ) = ( 16 − 1 ) + ( 11 − 1 ) = 15 + 10 = 25
Figure 9.5 Bar Graph of Jumps for Students in the Intervention and Control Groups
The next step is to set alpha (α), identify the critical values, and state a decision rule. In this example, alpha will once again be defined as “α = .05 (two-tailed).” Using the table of critical values in Table 3, the critical values are identified by moving down the df column until we reach the df = 25 row. For α = .05 (two-tailed), we find a critical value of 2.060. Consequently,
529
For α = .05 ( two-tailed ) and d f = 25 , critical value = ± 2.060.
On the basis of our values of alpha and the critical values, we can state a decision rule used to make the decision about the null hypothesis. For the data in the motor skills study: If t < − 2.060 or > 2.060 , reject H 0 ; otherwise, do not reject H 0 .
The next step in making a decision about the null hypothesis is to calculate a statistic—in this case, the t-test for independent means. Looking back at Formula 9-2, the first part of calculating a value of the t-statistic is to calculate the standard error of the difference ( s X ¯ 1 − X ¯ 2 ). This is the one aspect of the analysis that differs depending on whether the sample sizes are equal. Formula 9–4 presents the formula for the standard error of the difference (unequal sample sizes):
(9-4) s X ¯ − X ¯ 2 = ( N 1 − 1 ) s 1 2 + ( N 2 − 1 ) s 2 2 N 1 + N 2 − 2 ( 1 N 1 + 1 N 2 )
where N1 and N2 are the sample sizes of the two groups, and s1 and s2 are the standard deviations of the two groups. Notice that this formula is more complicated algebraically than Formula 9-3 (the standard error of the difference with equal sample sizes) in that the sample size for each group (N1 and N2) must be represented multiple times. The standard error of the difference for the motor skills example is calculated as follows: s X ¯ − X ¯ 2 = ( N 1 − 1 ) s 1 2 + ( N 2 − 1 ) s 2 2 N 1 + N 2 − 2 ( 1 N 1 + 1 N 2 ) = ( 16 − 1 ) ( 10.18 ) 2 + ( 11 − 1 ) ( 9.74 ) 2 16 + 11 − 2 ( 1 16 + 1 11 ) = ( 15 ) ( 103.63 ) + ( 10 ) ( 94.87 ) 25 ( .06 + .09 ) = 2503.15 25 ( .15 ) = 15.02 = 3.88
Once the standard error of the difference has been calculated, we are ready to calculate a value of the t-statistic using Formula 9-2: t = X ¯ 1 − X ¯ 2 s X ¯ 1 − X ¯ 2 = 27.31 − 11.91 3.88 = 15.40 3.88 = 3.97
Next, we make a decision whether to reject the null hypothesis. In this example: t = 3.97 > 2.060 ∴ reject H 0 ( p < .05 )
Here, because the calculated t-statistic of 3.97 is greater than the critical value 2.060 identified earlier, the null hypothesis is rejected because 3.97 lies in the region of rejection at the right end of the t-distribution.
Because the decision has been made to reject the null hypothesis, it is appropriate to determine the level of significance. Returning to the table of critical values in Table 3, we need to locate the critical value for df = 25 and a .01 (two-tailed) probability. Comparing the obtained t-value with the identified .01 critical value of 2.787, we draw the following conclusion: t = 3.97 > 2.787 ∴ p < .01
Because the t-value of 3.97 is greater than the .01 critical value of 2.787, its probability is
530
not only less than .05 but also less than .01 (see Figure 9.6).
Draw a Conclusion from the Analysis
What conclusions could we make on the basis of this analysis? Following the format of the earlier examples, the following statement could be made:
The average number of rope jumps in 30 seconds is significantly greater for the 16 students who received the intervention (M = 27.31) than for the 11 students in the control group who did not receive the intervention (M = 11.91), t(25) = 3.97, p < .01.
Relate the Result of the Analysis to the Research Hypothesis
The purpose of the motor skills study was to evaluate the effectiveness of a brief intervention aimed at improving the physical activity levels of kindergarten and first-grade students. What are the implications of our statistical analysis for evaluating the effectiveness of this intervention? Here is what the authors of the study said:
Figure 9.6 Determining the Level of Significance for the Motor Skills Study
This finding suggests that programs emphasizing the enhancement of basic motor skills that children apply in a variety of games and sports may be an effective approach to increasing overall activity and fitness levels of young children. (Matvienko & Ahrabi-Fard, 2010, p. 303)
The process of testing the difference between two sample means with unequal samples is summarized in Table 9.6.
531
Assumptions of the t-Test for Independent Means (Unequal Sample Sizes)
Earlier in this chapter, we introduced two mathematical assumptions related to the t-test for independent means: the assumption of normality, which is the assumption that the distribution of scores for the dependent variable in the two populations is approximately normal, and the assumption of homogeneity of variance, in which the variance of scores in the two populations is assumed to be the same. The data collected by researchers should meet these assumptions to draw appropriate conclusions from the results of statistical analyses. This is of particular concern when the sample sizes of the two groups are not equal to each other, as unequal sample sizes exacerbate the effects of differences in the shape and variance of the distributions of the two samples.
Statistics such as the t-statistic have been found to be robust, meaning that even if the data from a sample violate an assumption, the decision made regarding the null hypothesis is the same that would have been made if the assumption had not been violated. However, researchers have found that the robustness of statistics such as the t-statistic is lessened when the sample sizes of the groups are not equal to each other, especially when the variances are also unequal. The combination of unequal sample sizes and unequal variances may result in the need to consider the alternatives to the t-test for independent means discussed earlier in this chapter.
532
Summary
The examples discussed thus far in this chapter have compared participants in different conditions or groups. For example, the parking lot study compared the departure times of drivers who either had or did not have an intruder waiting to take their space, and the motor skills study measured the jumping rope ability of students who either received or did not receive the after-school intervention. As such, these studies are examples of between- subjects research designs, defined as research designs in which each research participant appears in only one level or category of the independent variable. The word between indicates these designs involve testing differences between different groups of participants. The next section introduces a different type of research design, one in which each participant appears in all levels or categories of the independent variable.
Table 9.6 Summary, Testing the Difference between Two Sample Means (Unequal Sample Sizes) (Motor Skills Example)
533
Learning Check 4: Reviewing what you've Learned So Far
1. Review questions a. How does the calculation of the t-test for the difference between two sample means change
when the sample sizes of the two groups are unequal?
2. For each of the following situations, calculate the standard error of the difference ( s X ¯ 1 − X ¯ 2 ) using the formula for unequal sample sizes.
a. N1 = 4, s1 = 2.00, N2 = 5, s2 = 3.00 b. N1 = 20, s1 = 6.00, N2 = 10, s2 = 4.00 c. N1 = 11, s1 = 8.64, N2 = 17, s2 = 11.75
3. To assist parents whose children have been diagnosed as having the eating disorder anorexia nervosa, one study evaluated the impact of a program designed to provide parents access to information and support groups (Carlton & Pyle, 2007). At the end of the program, parents who participated in the program (N = 29) and a group of parents of children with anorexia who did not participate (N = 53) completed a survey asking them to what extent they know what to feed their child at home (the higher the score, the higher the level of knowledge). The researchers reported the following descriptive statistics: participated (M = 4.68, s = 1.57) and did not participate (M = 2.77, s = 2.02).
a. State the null and alternative hypotheses (H0 and H1).
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. For α = .05 (two-tailed), identify the critical values and state a decision rule. 3. Calculate a value for the t-test for independent means (when calculating the
standard error of the difference s X ¯ 1 − X ¯ 2 , be sure to use the formula for unequal sample sizes).
4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
c. Draw a conclusion from the analysis. d. What conclusions might the researchers draw regarding the impact of the program on
parents' knowledge of anorexia nervosa?
534
9.5 Inferential Statistics: Testing the Difference between Paired Means
This section will again use a research study to introduce a statistical procedure that tests the difference between two means. However, this study differs from the earlier examples in that each research participant appears in both levels of the independent variable rather than just one. It is an example of a within-subjects research design, which is a research design in which each participant appears in all levels or categories of the independent variable. Rather than testing the difference between different groups of participants, within-subjects research designs test differences within the same participant.
535
Research Situations Appropriate for Within-Subjects Research Designs
There are several types of research situations where within-subjects research designs are employed. First, within-subjects designs are used to examine differences within a person regarding different situations or stimuli. A simple example of this situation would be to have people taste two types of ice cream and rate both types in terms of their flavor. One research study that used a within-subjects design for this purpose looked at police officers' beliefs regarding eyewitnesses' ability to provide accurate information about a crime (Kebbell & Milne, 1998). A sample of police officers were asked how often they believed eyewitnesses provided accurate information regarding four different aspects of a crime: the person who committed the crime, the action taken by the perpetrator, the object or target of the crime, and the surroundings in which the crime took place. Comparing these officers' ratings of these four aspects, the researchers found that officers believed eyewitnesses were more likely to provide useful information regarding the action involved in the crime than about the person committing the crime, the object of the crime, or the surroundings. That is, they believed witnesses are better able to describe the actions of a criminal than they are able to describe the actual criminal.
A second use of within-subjects designs involves collecting data on the same variable across repeated administrations. Longitudinal research is a research design in which the same information is collected from a research participant over two or more administrations to assess changes occurring within the person over time. An example of a longitudinal study, conducted by the author of this textbook (Tokunaga, 1985), examined the relationship between experiencing the death of a close friend or relative and changes in one's own fears of death and dying during the initial bereavement period. People who had experienced this type of loss were contacted and asked to complete a survey measuring death-related fears and attitudes every 4 months over a 12-month period.
Another type of longitudinal research, the pretest-posttest research design, involves collecting data from participants before and after the introduction of an intervention or experimental manipulation to determine whether the intervention or manipulation is associated with changes in the dependent variable. The research study described below is an example of a pretest-posttest design.
536
An Example from the Research: Web-Based Family Interventions
Given the large number of families in which both parents work, concern has been expressed regarding parents' ability to be aware of their children's health and well-being. A team of researchers lead by Diane Deitz set out to develop and evaluate a program designed to increase parents' awareness and knowledge of problems such as childhood anxiety and depression (Deitz, Cook, Billings, & Hendrickson, 2009). As the researchers wrote, “The prevalence of mental disorders in youth is substantial; however, these disorders often go unrecognized by parents and those closest to them … [therefore] the purpose of the project was to test a web-based program providing working parents with the knowledge and skills necessary for prevention and early intervention of mental health problems in youth” (p. 488).
In the study, the researchers developed a web-based program in which parents could go online and work through a series of modules that covered such things as information about different mental disorders, treatment options for these disorders, and techniques used to build better relationships between parent and children. As part of their study, the researchers hypothesized that “parents receiving the web-based program would exhibit significant gains in … knowledge of mental health issues in youth” (Deitz et al., 2009, p. 489).
The participants in the study, which will be referred to as the web-based intervention study, were working parents who currently had at least one child living at home. Before starting the program, each parent answered a test of 32 true/false items designed to assess their knowledge of depression, anxiety, treatment options, and parenting. Approximately 3 weeks later, after completing the program, parents completed the test a second time. In the study, the dependent variable, Knowledge, is the number of items each parent answered correctly. The independent variable, Time, consists of two levels: the test score before starting the program (Pretest) and after completing the program (Posttest).
The findings of the study will be illustrated using a sample of 20 parents. The parents' scores on the knowledge variable both before (Pretest) and after (Posttest) the program are presented in Table 9.7(a). Notice that each of the research participants appears in both levels of the independent variable. For example, before the program was introduced, the first parent answered 16 of the 32 items correctly; after the program, this same parent received a score of 24.
Table 9.8 calculates the mean and standard deviation for the Knowledge variable for the Pretest and Posttest time periods. From examining the descriptive statistics and the bar graph in Figure 9.7, we see that the mean Knowledge scores among these parents increased
537
from the Pretest (M = 15.55) to Posttest (M = 21.15).
538
Inferential Statistics: Testing the Difference between Paired Means
To test the study's research hypothesis, the next step is to calculate an inferential statistic to determine whether the difference between the two means is statistically significant. As in the earlier examples, the inferential statistic calculated in this situation will be a t-test. However, having each research participant appear in both levels of the independent variable alters the steps used to test the difference between the two means.
Calculate the Difference between the Paired Scores
In a between-subjects design, scores in the different groups are independent of each other. For example, in the parking lot study, the departure time for the first driver in the Intruder group is completely unrelated to the departure time of the first driver in the No intruder group. However, in a pretest-posttest research design, the scores in the two groups are related to each other. For example, in the web-based intervention study, the first pretest and posttest knowledge scores of 16 and 24 in Table 9.7 were produced by the same person.
Table 9.7 Knowledge of Anxiety and Depression for Parents at Pretest and Posttest
539
Table 9.8 Descriptive Statistics of Knowledge for Parents at the Pretest and Posttest Time Periods
Table 9.8 Descriptive Statistics of Knowledge for Parents at the Pretest and Posttest Time Periods
(a) Mean ( X ¯ i )
Pretest Posttest
X ¯ 1 = Σ X N = 16 + 23 + … + 12 + 19 20 = 311 20 = 15.55
X ¯ 2 = Σ X N = 24 + 23 + … + 21 + 18 20 = 423 20 = 21.15
(b) Standard Deviation (si)
s 1 = Σ ( X − X ¯ ) 2 N − 1 = ( 16 − 15.55 ) 2 + … + ( 19 − 15.55 ) 2 20 − 1 = .20 + … + 11.90 19 = 16.47 = 4.06
s 2 = Σ ( X − X ¯ ) 2 N − 1 = ( 24 − 24.15 ) 2 + … + ( 18 − 11.91 ) 2 20 − 1 = 8.12 + … + 9.92 19 = 27.29 = 5.22
Having the same research participant appear in all levels of an independent variable has important consequences for how the data are analyzed. When we discussed sampling error in Chapters 6 and 7, we noted that samples may vary from each other as the result of random, chance factors. One of the factors that causes variability across samples is having different people in the different samples. Using a pretest-posttest design eliminates this source of error. Therefore, analyzing data collected using this type of design requires explicit recognition that the scores in the two groups are related to each other and may be paired together. To indicate this pairing, we calculate the difference between each participant's two scores.
For the web-based intervention study, the difference between the knowledge scores at the posttest and pretest for each of the 20 parents, represented by the symbol D, is calculated in the last column of Table 9.9(a). Notice that some of these differences are negative numbers. The presence of negative numbers simply indicates that a participant's posttest score is greater than his or her pretest score. For example, the first participant in Table 9.9 had a pretest score of 16 and a posttest score of 24, resulting in a difference of −8. Because they will be needed later in the analysis, Table 9.9(b) calculates the mean ( X ¯ D ) and standard deviation (sD) of the difference scores.
The consequence of pairing the two scores is that rather than having two scores for each participant, we now have just one score—the difference score (D). As a result, the steps in hypothesis testing for paired means are essentially identical to those presented in Chapter 7, in which we compared the mean of one sample against a hypothesized population mean. These steps are illustrated below.
Figure 9.7 Bar Graph of Knowledge for Parents at the Pretest and Posttest Times
540
State the Null and Alternative Hypotheses (H0 and H1)
As in the earlier examples in this chapter, the null hypothesis in a pretest-posttest research design reflects the belief that the two means in the population are equal to each other. In the web-based intervention study the null hypothesis is that there is no difference between pretest and posttest knowledge scores. This absence of difference is reflected in the following null hypothesis:
where the subscript D stands for “difference.” Stating that the mean difference between the two populations is equal to zero is simply another way of stating that the two population means are equal to each other (μPretest = μPosttest).
Table 9.9 Difference (D) Scores of Knowledge for Parents at Pretest and Posttest
541
As in the previous examples, the alternative hypothesis may be either directional or non- directional. In the web-based intervention study the following non-directional alternative hypothesis will be used: H 1 : μ D ≠ 0
Therefore, in this example, the null hypothesis will be rejected if the posttest knowledge scores are either significantly less or significantly greater than pretest knowledge scores.
Make a Decision about the Null Hypothesis
The first step in making the decision to reject the null hypothesis is to calculate the degrees of freedom (df). Formula 9–5 presents the formula for the degrees of freedom for paired means:
(9-5) d f = N D − 1
where ND is the number of difference scores. This degrees of freedom differs from the degrees of freedom for the difference between two sample means presented earlier in this chapter because we are now working with one set of scores (the difference [D] scores) rather than two. For the web-based intervention study, because there are 20 parents, the degrees of freedom are d f = N D − 1 = 20 − 1 = 19
Once the degrees of freedom have been calculated, the next step is to set alpha (α), identify
542
the critical values, and state a decision rule. Assuming alpha is set to .05, the non-directional alternative hypothesis in this example (H1: μD ≠ 0) leads us to state that “α = .05 (two- tailed).” To determine the critical values, using Table 3, we move down the df column until we reach df = 19 and then to the right until we are under the .05 heading within the “Level of significance for two-tailed test” columns. For the web-based intervention study: For α = .05 ( two-tailed ) and d f = 19 , critical value = ± 2 . 093
We can now state a decision rule regarding the conditions under which the null hypothesis will be rejected. Given the critical values we have just identified: If t < − 2.093 or > 2.093 , reject H 0 ; otherwise, do not reject H 0 .
The next step is to calculate a statistic. As in the other examples in this chapter, the inferential statistic we will calculate will be a t-test. However, in this situation, we will calculate the t-test for dependent means (t), which is an inferential statistic that tests the difference between two means that are based on the same participant or paired participants. The word dependent is used to indicate that the two means are based on one sample drawn from one population, as opposed to the two samples and two populations that are the basis of the t-test for independent means discussed earlier in this chapter.
The formula for the t-test for dependent means is presented in Formula 9–6:
(9-6) t = X ¯ D − μ D s D ¯
where X ¯ D is the mean of the difference scores, μD is the hypothesized population mean of the difference scores, and s D ¯ is the standard error of the difference scores. Except for the D in the subscripts, we see that Formula 9-6 is virtually identical to the formula for the t-test of one mean (Formula 7-4) presented in Chapter 7.
The three pieces of information needed to calculate a value for the t-test for dependent means are located in different places. First, the value for X ¯ D is found in the descriptive statistics (Table 9.9(b)). For the web-based intervention study, X ¯ D = − 5.60 . Second, the population mean μD is located in the null hypothesis. Because of how the null hypothesis in this example has been stated (H0: μD = 0), μD is equal to 0. The third piece of information, the standard error of the difference scores ( s D ¯ ), is calculated using Formula 9-7:
(9-7) s D ¯ = s D N
where sD is the standard deviation of the difference scores and N is the sample size. Obtaining the standard deviation of the difference scores (sD) from Table 9.9(b), the standard error of the difference scores ( s D ¯ ) for the 20 parents in the web-based
543
intervention study is as follows: s D ¯ = s D N = 4.35 20 = 4.35 4.47 = .97
Placing these three pieces of information into Formula 9.7, we may now calculate a value of the t-statistic: t = X ¯ D − μ D s D ¯ = − 5.60 − 0 .97 = − 5.60 .97 = − 5.77
The value of the t-statistic for the web-based intervention study is a negative number because the first mean ( X ¯ Pretest = 15.55 ) is less than the second mean ( X ¯ Posttest = 21.15 ).
We are now ready to make a decision whether to reject the null hypothesis. Comparing the t- statistic of −5.77 calculated from this example with the critical values, we make the following decision: t = − 5.77 < − 2.093 ∴ reject H 0 ( p < .05 )
Because the null hypothesis has been rejected, it is appropriate to determine the level of significance; more specifically, this means determining whether the probability of our value of the t-statistic is less than .01. In Table 3, for α = .01 (two-tailed) and df = 19, we find the critical value −2.861. Comparing the calculated t-value for this example with the .01 critical value: t = − 5.77 < − 2.861 ∴ p < .01
Because the t-value of −5.77 is less than the .01 critical value of −2.861, the level of significance for this set of data is p < .01 (see Figure 9.8).
Draw a Conclusion from the Analysis
For the web-based intervention study, we've found that the 20 parents' mean knowledge scores at the pretest and posttest are significantly different from each other. To present the results of this analysis in a more informative way, we could say the following:
The average knowledge scores for the 20 parents were significantly higher after completing the web-based intervention program (M = 21.15) than before beginning the program (M = 15.55), t(19) = −5.77, p < .01.
Relate the Result of the Analysis to the Research Hypothesis
One purpose of the web-based intervention study was to determine whether the program could provide working parents knowledge that may help them and their children with mental health-related problems. The researchers discuss the implication of finding a significant increase in knowledge from the pretest to the posttest below:
544
These findings indicate that the program can be an effective intervention for improving parents' knowledge of children's mental health problems and boost their confidence in handling such issues…. The study findings lend support to the growing literature on the utility of offering web-based programs to improve the health of the general population. (Deitzetal., 2009, p. 492)
Figure 9.8 Determining the Level of Significance for the Web-Based Intervention Study
545
Assumptions of the t-Test for Dependent Means
As with the t-test for independent means, the appropriate use of the t-test for dependent means requires that certain mathematical assumptions be met. One of the fundamental assumptions of statistics such as the t-test is the assumption of normality: that the data being analyzed are normally distributed. Given that the t-test for dependent means involves analyzing difference (D) scores, it is not surprising that this statistic is based on the assumption that difference scores in the population are normally distributed. The larger the sample size, the greater the likelihood of meeting this assumption.
546
Summary
The steps in testing the difference between paired means are summarized in Table 9.10. Looking at this table, we see a great amount of similarity between these steps and those used to test a single mean in Chapter 7. Although research studies may differ in their purpose, hypotheses, variables of interest, or method of data collection, we've seen that the analysis of the data collected in these studies follows a common sequencing and logic.
Table 9.10 Summary, Testing the Difference Between Paired Means (Web-Based Intervention Example)
547
Learning Check 5: Reviewing what you've Learned So Far
1. Review questions a. What is the main difference between between-subjects and within-subjects research designs? b. For what types of research situations might you use a within-subjects design? c. In analyzing a pretest-posttest research design, why do we calculate the difference between
the scores in the two groups? d. When would you calculate the t-test for dependent means rather than the t-test for
independent means?
2. You conduct a class project in which you divide classmates into seven pairs; each pair consists of a shy student and an outgoing student. You have each pair of students work on a puzzle but do not allow them to talk to each other. After completing the puzzle, you ask each student to indicate how much he or she enjoyed working on the puzzle:
Pair Shy Outgoing
1 8 4
2 4 6
3 8 2
4 6 6
5 3 7
6 6 3
7 5 4
a. Calculate the difference score (D) for each pair of students. b. Calculate the mean ( X ¯ D ), standard deviation (sD), and standard error ( s D ¯ ) of the
difference scores.
3. For each of the following situations, calculate the standard error of the difference scores (sD) and the t-test for dependent means.
a. ND = 9, X ¯ D = 2.00 , s D ¯ = 1.75 b. N = 12, X ¯ D = 1.50 , s D ¯ = 3.50 c. N = 21, X ¯ D = 12.63 , s D ¯ = 24.82
4. As adults get older, their lives may become increasingly sedentary Given concerns about physical and psychological problems associated with a lack of physical activity, one study investigated whether some office tasks such as typing, working on a computer, and reading may be performed at a satisfactory level while walking on a treadmill rather than sitting (John, Bassett, Thompson, Fairbrother, & Baldwin, 2009). One part of the study involved having 20 adults take a typing test two times: once while sitting at a desk and once while walking on a treadmill. The average number of words per minute (WPM) typed correctly for the two conditions was as follows: sitting (M = 40.20, s = 10.20) and treadmill (M = 36.90, s = 11.10). Furthermore, the researchers reported a mean difference score ( X ¯ D ) of 3.30 and a standard deviation of the difference scores (sD) of 4.70. Test the difference in typing speed for the two conditions.
a. State the null and alternative hypotheses (H0 and H1).
b. Make a decision about the null hypothesis.
548
1. Calculate the degrees of freedom (df). 2. For α = .05 (two-tailed), identify the critical values and state a decision rule. 3. Calculate a value for the t-test for dependent means. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
c. Draw a conclusion from the analysis. d. What conclusions might the researchers draw regarding whether the task of typing may be
performed at a satisfactory level while walking on a treadmill rather than sitting?
549
9.6 Looking Ahead
The main purpose of the present chapter was to further your understanding of the process of hypothesis testing using a situation slightly more complicated than those presented in earlier chapters. As you move further along in this book, you will encounter research situations of growing complexity. Although the statistical procedures we discuss may become more challenging, understand that the steps in hypothesis testing will remain essentially the same. The main purpose of hypothesis testing is to make one of two decisions about the null hypothesis: reject or not reject. This decision is based on the probability of obtaining a calculated value of an inferential statistic when the null hypothesis is true. Because we must rely on probability, there exists the possibility that the decision made about the null hypothesis may be in error. The next chapter discusses these errors in greater detail, as well as what researchers can do to minimize the possibility and impact of making these errors.
550
9.7 Summary
To test the difference between two sample means, we create the sampling distribution of the difference, which is the distribution of all possible values of the difference between two sample means when an infinite number of pairs of samples of size N are randomly selected from two populations. Three characteristics of the sampling distribution of the difference are that its mean is equal to zero (0), the distribution is approximately normal in shape, and the variability of this distribution is measured by the standard error of the difference ( s X ¯ 1 − X ¯ 2 ), defined as the standard deviation of the sampling distribution of the difference. These characteristics of the sampling distribution of the difference enable researchers to determine the probability of obtaining any particular difference between the means of two samples.
The t-test for independent means, an inferential statistic that tests the difference between the means of two samples drawn from two populations, is used to test a hypothesis about the difference between two population means. The t-test may be calculated when the sample sizes for the two groups are either equal or unequal; with unequal sample sizes, the calculation of the t-test is slightly more complicated because the formula for the standard error of the difference must be modified.
The t-test for independent means is used in between-subjects research designs, in which each research participant appears in only one level or category of the independent variable, and these designs involve testing differences between different groups of participants. The t-test for dependent means is used in within-subjects designs, in which each participant appears in all levels or categories of the independent variable involve collecting information from each participant more than once. An example of within-subjects designs is longitudinal research, which involves collecting the same information from a research participant over two or more administrations to assess changes occurring within the person over time. One example of longitudinal research is the pretest-posttest research design, which involves collecting data from participants twice—before and after the introduction of an intervention or experimental manipulation—to determine whether the intervention or manipulation is associated with changes in the dependent variable.
551
9.8 Important Terms
bar graph (p. 314) sampling distribution of the difference (p. 317) standard error of the difference ( s X ¯ 1 − X ¯ 2 ) (p. 318) t-test for independent means (p. 322) homogeneity of variance (p. 329) between-subjects research designs (p. 338) within-subjects research designs (p. 341) longitudinal research (p. 341) pretest-posttest research design (p. 341) t-test for dependent means (t) (p. 347)
552
9.9 Formulas Introduced in this Chapter
Degrees of Freedom (df), Difference Between Two Sample Means
(9-1) d f = ( N 1 − 1 ) + ( N 2 − 1 )
t-Test for Independent Means
(9-2) t = X ¯ 1 − X ¯ 2 s X ¯ 1 − X ¯ 2
Standard Error of the Difference ( s X ¯ 1 − X ¯ 2 )
(9-3) s X ¯ 1 − X ¯ 2 = s 1 2 N 1 + s 2 2 N 2
Standard Error of the Difference (Unequal Sample Sizes)
(9-4) s X 1 ¯ − X 2 ¯ = ( N 1 − 1 ) s 1 2 + ( N 2 − 1 ) s 2 2 N 1 + N 2 − 2 ( 1 N 1 + 1 N 2 )
Degrees of Freedom (df), Difference Between Paired Means
(9-5) d f = N D − 1
t-Test for Dependent Means
(9-6) t = X ¯ D − μ D s D ¯
Standard Error of Difference Scores (sD)
(9-7) s D ¯ = s D N
553
9.10 Using SPSS
554
Testing the Difference between Two Sample Means: The Parking Lot Study (9.1)
1. Define independent and dependent variables (name, # decimals, labels for the variables, labels for values of the independent variable) and enter data for the variables.
NOTE: Numerically code values of the independent variable (i.e., 1 = Intruder, 2 = No intruder) and provide labels for these values in the Values box within Variable View.
2. Select the t-test for independent means procedure within SPSS.
How? (1) Click Analyze menu, (2) click Compare Means, and (3) click Independent-Samples T Test.
3. Identify the dependent variable, the independent variable, and the values of the independent variable.
How? (1) Click dependent variable and → Test Variable, (2) click independent variable and → Grouping Variable, (3) click Define Groups and type the values for the independent variable, (4) click Continue, and (5) click OK.
555
4. Examine output.
Testing the Difference between Paired Means: The Web- Based Intervention Study (9.5)
1. Define the two levels of the independent variable (name, # decimals, labels for the variables) and enter data for each level.
NOTE: Each participant has scores on two variables—each variable represents one of the two levels of the independent variable (i.e., pretest, posttest).
556
2. Begin the t-test for dependent means procedure within SPSS.
How? (1) Click Analyze menu, (2) click Compare Means, and (3) click Paired- Samples T Test.
3. Identify the paired variables.
How? (1) Click the two variables and → Paired Variables and (2) click OK.
4. Examine output.
557
558
9.11 Exercises
1. Construct a bar graph for each of the following (assume the independent variable is group and the dependent variable is time):
a. Group A (N = 5, M = 4.00, s = 1.58); Group B (N = 5, M = 6.00, s = 2.12) b. Group A (N = 8, M = 26.00, s = 2.56); Group B (N = 8, M = 23.00, s = 2.33) c. Group A (N = 11, M = 1.77, s = .29); Group B (N = 11, M = 1.53, s = .22) d. Group A (N = 18, M = 63.59, s = 23.70); Group B (N = 18, M = 71.42, s =
20.91) 2. Construct a bar graph for each of the following (assume the independent variable is
Group and the dependent variable is time):
a. Group A (N = 21, M = 14.05, s = 3.63);
Group B (N = 21, M = 12.33, s = 3.26)
b. Group A (N = 28, M = 6.79, s = 3.11);
Group B (N = 28, M = 7.93, s = 2.36)
c. Group A (N = 16, M = 52.56, s = 23.77);
Group B (N = 16, M = 60.38, s = 21.91)
d. Group A (N = 25, M = 5.76, s = 2.14);
Group B (N = 25, M = 4.43, s = 2.27)
3. For each of the following, (a) calculate the mean, standard deviation, and standard error of the mean for each Group and (b) construct a bar graph (assume the independent variable is condition and the dependent variable is test score).
a. Experimental: 3, 7, 4, 1, 10, 4, 6
Control: 10, 7, 12, 4, 8, 9
b. Experimental: 15, 12, 10, 14, 17, 15, 18, 19
Control: 22, 12, 17, 19, 20, 21, 16, 17
c. Experimental: 3.89, 3.04, 3.95, 2.91, 2.72, 3.70, 3.16, 3.21, 2.86
Control: 2.96, 3.38, 2.82, 2.07, 3.22, 2.56, 2.44, 3.11, 2.68
559
d. Experimental: 2, 4, 3, 5, 1, 4, 3, 5, 4, 1, 4, 5, 3
Control: 5, 2, 5, 1, 4, 3, 5, 2, 1, 4, 5, 1, 4 4. State the null and alternative hypotheses (H0 and H1) for each of the following
research questions: a. Are the average starting salaries for clinical psychologists in private practice the
same as or different from that of psychological researchers in business or the government?
b. In Chapter 3, we looked at psychology majors' and non– psychology majors' belief in the myth that we only use 10% of our brains. Do the two groups differ in this belief?
c. Are men paid more than women for doing the same job? d. Do Republicans and Democrats similarly support a national health insurance
program or does one group favor this more than the other?
5. For each of the following, calculate the degrees of freedom (df) and determine the critical values of t (assume α = .05).
a. N1 = 21, N2 = 21, H1: μ1 ≠ μ2 b. N1 = 14, N2 = 14, H1: μ1 ≠ μ2 c. N1 = 4, N2 = 4, H1: μ1 > μ2 d. N1 = 32, N2 = 32, H1: μ1 < μ2
6. For each of the following, calculate the degrees of freedom (df) and determine the critical values of t (assume α = .05).
a. N1 = 5, N2 = 5, H1: μ1 ≠ μ2 b. N1 = 12, N2 = 12, H1: μ1 < μ2 c. N1 = 9 N2 = 9, H1: μ1 ≠ μ2 d. N1 = 16, N2 = 16, H1: μ1 > μ2
7. For each of the following, calculate the standard error of the difference ( s X ¯ 1 − X ¯ 2 ).
a. N1 = 35, s1 = 1.50, N2 = 35, s2 = 3.25 b. N1 = 4, s1 = 4.30, N2 = 4, s2 = 2.10 c. N1 = 23, s1 = 8.10, N2 = 23, s2 = 7.50 d. N1 = 20, s1 = 1.20, N2 = 20, s2 = 1.75
8. For each of the following, calculate the standard error of the difference ( s X ¯ 1 − X ¯
560
2 ). a. N1 = 10, s1 = 2.00, N2 = 10, s2 = 3.00 b. N1 = 19, s1 = 1.73, N2 = 19, s2 = 1.48 c. N1 = 27, s1 = 24.91, N2 = 27, s2 = 27.02 d. N1 = 50, s1 = 12.29, N2 = 50, s2 = 10.63
9. For each of the following, calculate the t-test for independent means. a. X ¯ 1 = 3.49 , X ¯ 2 = 3.14 , s X ¯ 1 − X ¯ 2 = .31 b. X ¯ 1 = 13.27 , X ¯ 2 = 16.45 , s X ¯ 1 − X ¯ 2 = 1.52 c. X ¯ 1 = .76 , X ¯ 2 = .91 , s X ¯ 1 − X ¯ 2 = .09 d. X ¯ 1 = 1.52 , X ¯ 2 = 1.36 , s X ¯ 1 − X ¯ 2 = .05
10. For each of the following, calculate the t-test for independent means.
a. X ¯ 1 = 7.00 , X ¯ 2 = 11.00 ,
s X ¯ 1 − X ¯ 2 = 1.17
b. X ¯ 1 = 65.56 , X ¯ 2 = 60.92 ,
s X ¯ 1 − X ¯ 2 = 2.88
c. X ¯ 1 = 137.73 , X ¯ 2 = 114.09 ,
s X ¯ 1 − X ¯ 2 = 10.71
d. X ¯ 1 = 73.24 , X ¯ 2 = 81.53 ,
s X ¯ 1 − X ¯ 2 = 4.39
11. For each of the following, calculate the standard error of the difference ( s X ¯ 1 − X ¯ 2 ) and the t-test for independent means.
a. N1 = 6, X ¯ 1 = 18.50 , s1 = 2.00,
N2 = 6, X ¯ 2 = 19.00 , s2 = 2.50
b. N1 = 13, X ¯ 1 = 36.23 , s1 = 4.17,
N2 = 13, X ¯ 2 = 29.59 , s2 = 6.01
c. N1 = 29, X ¯ 1 = 7.80 , s1 = 1.25,
N2 = 29, X ¯ 2 = 7.17 , s2 = 1.63
d. N1 = 38, X ¯ 1 = 21.14 , s1 = 4.38
N2 = 38, X ¯ 2 = 23.29 , s2 = 3.91
12. For each of the following, calculate the standard error of the difference ( s X ¯ 1 − X ¯ 2 ) and the t-test for independent means.
561
a. N1 = 11, X ¯ 1 = 12.00 , s1 = 2.00,
N2 = 11, X ¯ 2 = 10.00 , s2 = 3.00
b. N1 = 21, X ¯ 1 = 39.85 , s1 = 5.23,
N2 = 21, X ¯ 2 = 44.16 , s2 = 4.60
c. N1 = 33, X ¯ 1 = 4.37 , s1 = 1.07,
N2 = 33, X ¯ 2 = 4.92 , s2 = .94
d. N1 = 57, X ¯ 1 = 53.98 , s1 = 6.96,
N2 = 57, X ¯ 2 = 50.74 , s2 = 6.03
13. When you're interviewing for a job, is your behavior influenced by beliefs the interviewer has about you? Researchers Ridge and Reber (2002) observed the interactions of 54 men, each of whom was interviewing a woman for a job. Right before conducting their interviews, half of the men were told the woman whom they were interviewing was attracted to them. The researchers hypothesized that “these same men would elicit relatively more flirtatious behavior from women than would men holding no such belief” (p. 2). The interviews were videotaped, and the number of flirtatious behaviors performed by each woman was counted. The researchers reported the following descriptive statistics regarding the mean number of flirtatious behaviors: attraction condition (N = 27, M = 37.44, s = 5.21) and no attraction condition (N = 27, M = 34.59, s = 4.54).
a. State the null and alternative hypotheses (H0 and H1).
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values, and state a decision rule. 3. Calculate a value for the t-test for independent means. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
14. An advertising agency is interested in learning how to fit its commercials to the interests and needs of the viewing audience. It asked samples of 41 men and 41 women to report the average amount of television watched daily. The men reported a mean television time of 1.70 hours per day with a standard deviation of .70. The women reported a mean of 2.05 hours per day with a standard deviation of .80. Use these data to test the manager's claim that there is a significant gender difference in television viewing.
a. State the null and alternative hypotheses (H0 and H1).
562
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values (draw the distribution), and state
a decision rule. 3. Calculate a value for the t-test for independent means. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
c. Draw a conclusion from the analysis. d. What are the implications of this analysis for the advertising agency?
15. A pair of researchers was interested in studying men's preferences for a potential mate and how these preferences may change as men get older (Alterovitz & Mendelsohn, 2009). The researchers believed that men prefer women who are younger than themselves; furthermore, they hypothesized that the preferred difference in age is greater for older men than for younger men. As part of a larger study, they examined personal advertisements placed by men who were either 20 to 34 years old or 40 to 54 years old; in looking at these ads, the researchers calculated the difference between the man's age and the man's preferred age of a potential partner. The descriptive statistics for the difference in age variable for the two age groups are as follows:
20 to 34 years old: N1 = 25, X ¯ 1 = 1.04 , s1 = 2.72
40 to 54 years old: N1 = 25, X ¯ 1 = 4.98 , s1 = 3.80 a. State the null and alternative hypotheses (H0 and H1).
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values (draw the distribution), and state
a decision rule. 3. Calculate a value for the t-test for independent means. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
16. A third-grade teacher is interested in comparing the effectiveness of two styles of instruction in language comprehension: imagery, in which the students are asked to picture a situation involving the word, and repetition, in which the students repeat the definition of the word multiple times. After 6 weeks of instruction, she gives the students a language comprehension test. The following scores are the number of correct answers for each student. Determine whether the two styles of instruction differ in their effectiveness.
Imagery: 12, 13, 11, 11, 13, 13, 15, 12, 9, 12
563
Repetition: 6, 11, 10, 12, 9, 10, 11, 12, 10, 8 a. For each group, calculate the sample size (N), mean ( X ¯ ), and standard
deviation (s). b. State the null and alternative hypotheses (H0 and H1).
c. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values (draw the distribution), and state
a decision rule. 3. Calculate a value for the t-test for independent means. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
d. Draw a conclusion from the analysis. e. What are the implications of this analysis for the teacher?
17. Imagine you believe that men and women have different beliefs regarding the average cost of a wedding. More specifically, because they are traditionally more involved in the planning of weddings and therefore have a more complete understanding of their costs, you hypothesize that young women will provide a higher estimate of the cost of the average wedding than will men. To test your hypothesis, you ask a sample of college students (13 women and 13 men), “How much (in thousands) do you believe the average wedding costs?” The estimates (in thousands of dollars) of these students are below:
Women: 20, 9, 14, 11, 25, 18, 12, 33, 24, 10, 30, 14, 22 Men: 15, 6, 19, 16, 33, 4, 21, 13, 10, 7, 24, 5, 20
a. For each group, calculate the sample size (N), mean ( X ¯ ), and standard deviation (s).
b. State the null and alternative hypotheses (H0 and H1).
c. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values (draw the distribution), and state
a decision rule. 3. Calculate a value for the t-test for independent means. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
d. Draw a conclusion from the analysis. e. Relate the result of the analysis to the research hypothesis.
18. For each of the following, calculate the standard error of the difference ( s X ¯ 1 − X ¯ 2 ) (be sure to use the formula for unequal sample sizes).
a. N1 = 8, s1 = 2.30, N2 = 12, s2 = 2.00
564
b. N1 = 25, s1 = 5.15, N2 = 22, s2 = 7.80 c. N1 = 13, s1 = 4.50, N2 = 19, s2 = 5.89 d. N1 = 50, s1 = 21.00, N2 = 45, s2 = 17.50
19. For each of the following, calculate the standard error of the difference ( s X ¯ 1 − X ¯ 2 ) (be sure to use the formula for unequal sample sizes).
a. N1 = 9, s1 = 4.00, N2 = 7, s2 = 2.00 b. N1 = 10, s1 = 8.50, N2 = 20, s2 = 7.80 c. N1 = 33, s1 = 1.87, N2 = 41, s2 = 5.89 d. N1 = 21, s1 = 16.21, N2 = 19, s2 = 3.56
20. A team of researchers stated that effective methods of therapy help people access and process their emotions (Watson & Bedard, 2006). They looked at the level of emotional processing brought about by two different types of therapy for clients diagnosed with depression. The first group (N = 17) took part in cognitive-behavioral therapy (CBT). The second group (N = 21) took part in process-experiential therapy (PET). The Experiencing Scale measured their level of emotional processing. The CBT group scored a mean of 2.73 on the scale, with a standard deviation of .46. The PET group scored a mean of 3.04, with a standard deviation of .42. Use the steps of hypothesis testing to determine whether one therapeutic technique brings about more emotional processing than the other.
a. State the null and alternative hypotheses (H0 and H1).
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values (draw the distribution), and state
a decision rule. 3. Calculate a value for the t-test for independent means (when calculating
the standard error of the difference [ s X ¯ 1 − X ¯ 2 ], be sure to use the formula for unequal sample sizes).
4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
c. Draw a conclusion from the analysis. d. Does one of the therapy methods appear to be more effective than the other?
21. A team of audiologists was interested in examining whether their patients' satisfaction with their hearing aids was related to how long they had used hearing aids (Williams, Johnson, & Danhauer, 2009). They divided their patients into two categories, new users (N = 30) and experienced users (N = 34), and asked them to indicate how satisfied they were with their hearing aids; the higher the score, the greater the satisfaction. The new users reported a mean satisfaction of 26.90 on the scale (standard deviation = 3.96), and the experienced users reported a mean satisfaction of
565
28.03 (standard deviation = 5.04). Use the steps of hypothesis testing to test the difference in satisfaction between new and experienced users.
a. State the null and alternative hypotheses (H0 and H1).
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values (draw the distribution), and state
a decision rule. 3. Calculate a value for the t-test for independent means (when calculating
the standard error of the difference [ s X ¯ 1 − X ¯ 2 ], be sure to use the formula for unequal sample sizes).
4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
c. Draw a conclusion from the analysis. d. Is patients' satisfaction with hearing aids related to how long they have used
hearing aids?
22. Imagine a researcher asks a sample of five people to drive two types of cars and rate each of them on a 1 to 20 scale. Listed below are the data she collected:
Person Type A Type B
1 20 10
2 10 11
3 11 4
4 23 12
5 7 0
a. Calculate the difference score (D) for each of the five participants. b. Calculate the mean ( X ¯ D ), standard deviation ( s D ¯ ), and standard error (
s D ¯ ) of the difference scores.
23. For each of the following, calculate the standard error of the difference scores ( s D ¯ ) and the t-test for dependent means.
a. ND = 25, X ¯ D = 10.00 , s D ¯ = 8.00
b. ND = 5, X ¯ D = 6.80 , s D ¯ = 4.71
c. ND = 16, X ¯ D = 2.12 , s D ¯ = 2.17
d. ND = 20, X ¯ D = .80 , s D ¯ = 3.10
24. For each of the following, calculate the standard error of the difference scores ( s D ¯ )
566
and the t-test for dependent means.
a. ND = 10, X ¯ D = 12.00 , s D ¯ = 9.00
b. ND = 19, X ¯ D = 3.52 , s D ¯ = 6.14
c. ND = 8, X ¯ D = 2.89 , s D ¯ = 4.86
d. ND = 30, X ¯ D = 1.44 , s D ¯ = 8.32
25. As people in our society live longer, there is a growing need for older adults to maintain a healthy lifestyle, part of which is their physical fitness. One study evaluated a program designed to increase older adults' level of fitness using weight and strength training (Doll, 2009). In the study, eight adults (mean age = 75.60 years) received training on a weight machine over an 8-week period, doing exercises such as bench presses and abdominal crunches. Before and after beginning the training, each adult was measured on a number of functional tasks, one of which was the number of times they could lift a 5-pound weight over their heads. The descriptive statistics for the number of overhead lifts were the following: pretest (M = 45.60, s = 15.00) and posttest (M = 53.90, s = 18.60). Furthermore, the mean difference score ( X ¯ D ) was −8.30; the standard deviation of the difference scores ( s D ¯ ) was 9.84. Test the difference in overhead lifts between the pretest and posttest.
a. State the null and alternative hypotheses (H0 and H1).
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values (draw the distribution), and state
a decision rule. 3. Calculate a value for the t-test for dependent means. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
c. Draw a conclusion from the analysis. d. What conclusions might the researchers draw regarding the effectiveness of a
weight and strength training program designed for older adults?
26. Another part of the study discussed in Exercise 17 had a second sample of nine older adults take part in an 8-week calisthenics program in which they did squats, hamstring curls, and lunges. As the program was designed to help these adults perform everyday activities more easily, before and after the program, each adult was asked how difficult it was for them to get in and out of a bathtub (the higher the number, the greater the difficulty). The descriptive statistics for this bathtub-related difficulty variable were the following: pretest (M = 4.30, s = 2.00) and posttest (M = 2.90, s = .90). Furthermore, the mean difference score ( X ¯ D ) was 1.40; the standard deviation of the difference scores ( s D ¯ ) was 1.86. Test the difference in
567
bathtub-related difficulty between the pretest and posttest. a. State the null and alternative hypotheses (H0 and H1).
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values (draw the distribution), and state
a decision rule. 3. Calculate a value for the t-test for dependent means. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
c. Draw a conclusion from the analysis. d. What conclusions might the researchers draw regarding the effectiveness of a
calisthenics program designed for older adults?
27. A group of friends decided to work together to decrease their cigarette smoking. They hypothesize that they can decrease their level of smoking by doing such things as giving one another encouragement, using carrots and candy to replace the physical action of smoking, and using deep breathing and counting through cravings. Each friend keeps track of how many cigarettes he or she smokes each day. The first set of data is each person's baseline (before beginning the plan), and the second set of data is after 4 weeks of using the tools. Complete the steps below to test their hypothesis.
Person Before After
1 8 0
2 25 13
3 17 0
4 5 1
5 12 14
6 20 10
a. For each of the two time periods, calculate the sample size (Ni), the mean ( X ¯ i ), and the standard deviation (si).
b. Calculate the mean ( X ¯ D ) and the standard deviation of the difference scores (sD).
c. State the null and alternative hypotheses (H0 and H1).
d. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values (draw the distribution), and state
a decision rule.
568
3. Calculate a value for the t-test for dependent means. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
e. Draw a conclusion from the analysis. f. Relate the result of the analysis to the research hypothesis.
28. A researcher hypothesizes that an afternoon dose of caffeine lessens the amount of time a person sleeps at night. The participants are instructed to report their average hours of sleep per night for a week without any caffeine in the afternoons and their average amount of sleep for a week in which they drink two caffeinated beverages between 2 and 3 p.m. Using their reported data, run through the steps of hypothesis testing to test the research hypothesis.
Person Without Caffeine With Caffeine
1 9.00 6.00
2 6.00 5.50
3 8.00 8.00
4 7.00 6.25
5 7.50 5.00
6 7.00 5.00
7 6.00 7.00
8 9.00 8.00
9 8.00 7.00
10 7.00 6.00
a. For each of the two caffeine conditions, calculate the sample size (Ni), the mean ( X ¯ i ), and the standard deviation (si).
b. Calculate the mean ( X ¯ D ) and the standard deviation of the difference scores (sD).
c. State the null and alternative hypotheses (H0 and H1).
d. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values (draw the distribution), and state
a decision rule. 3. Calculate a value for the t-test for dependent means. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
e. Draw a conclusion from the analysis.
569
f. Relate the result of the analysis to the research hypothesis.
570
Answers to Learning Checks
Learning Check 1
2.
a.
b.
c.
3.
a. Dog: X ¯ = 2.00 , s = 2.83, S X ¯ = 1.27 ;
Cat: X ¯ = 6.00 ; s = 3.39, s X ¯ = 1.52
b. Dog: X ¯ = .57 , s = .20, s X ¯ = .06 ;
Cat: X ¯ = .53 ; s = .27, s X ¯ = .09
571
c. Dog: X ¯ = 8.71 , s = 2.30, s X ¯ = .62 ;
Cat: X ¯ = 6.29 ; s = 2.02, s X ¯ = .54
Learning Check 2
2.
a. H0: μSingle parent = μintact; H1: μSingle parent ≠ μintact b. H0: μTap water = μBottled water; H1: μTap water ≠ μBottled water c. H0: μAlone = μWith others; H1: μAlone ≠ μWith others
3.
a. df = 12; critical value = ±2.179
b. df = 20; critical value = ±2.086
c. df = 14; critical value = 1.761
4. a. s X ¯ 1 − X ¯ 2 = 2.24 b. s X ¯ 1 − X ¯ 2 = 3.61 c. s X ¯ 1 − X ¯ 2 = 1.03
5.
a. s X ¯ 1 − X ¯ 2 = 1.47 ; t = 2.04
572
b. s X ¯ 1 − X ¯ 2 = 4.65 ; t = .97
c. s X ¯ 1 − X ¯ 2 = .74 ; t = −3.88
Learning Check 3
2. a. H0: μMen = μWomen H1: μMen ≠ μWomen b.
1. df = 24 2. If t < −2.064 or > 2.064, reject H0; otherwise, do not reject H0 3. s X ¯ 1 − X ¯ 2 = 6.52 ; t = –.32 4. t = –.32 is not < −2.064 or > 2.064 ∴ do not reject H0 (p > .05) 5. Not applicable (H0 not rejected)
c. The average estimate of a woman's weight in the sample of 13 men (M = 135.62) and 13 women (M = 137.69) was not significantly different, t(24) = –.32, p > .05.
d. The results of this analysis do not support the research hypothesis that men will give lower estimates of the average woman's weight than will women.
Learning Check 4
2. a. s X ¯ 1 − X ¯ 2 = 1.76 b. s X ¯ 1 − X ¯ 2 = 2.11 c. s X ¯ 1 − X ¯ 2 = 4.13
3. a. H0: μParticipate = μNot participate H1: μParticipate ≠ μNot participate b.
1. df = 80 2. If t < −2.000 or > 2.000, reject H0; otherwise, do not reject H0 3. s X ¯ 1 − X ¯ 2 = .42 ; t = 4.55 4. t = 4.55 > 2.000 ∴ reject H0 (p < .05) 5. t = 4.55 > 2.660 ∴ p < .01
c. The average level of knowledge regarding what to feed their children at home was significantly higher for the sample of 29 parents who participated in the program (M = 4.68) than the sample of 53 parents who did not participate (M = 2.77), t(80) = 4.55, p < .01.
d. The results of this analysis suggest that the program may provide information for parents of children with anorexia nervosa.
573
Learning Check 5
2. a.
Pair Shy Outgoing Difference (D)
1 8 4 4
2 4 6 –2
3 8 2 6
4 6 6 0
5 3 7 –4
6 6 3 3
7 5 4 1
b. X ¯ D = 1.14 ; s D ¯ = 3.48 ; s D ¯ = 1.32 3.
a. s D ¯ = .50 ; t = 6.00 b. s D ¯ = .38 , t = 3.95 c. s D ¯ = 5.42 , t = 2.33
4. a. H0: μD = 0; H1: μD ≠ 0 b.
1. df = 19 2. If t < −2.093 or > 2.093, reject H0; otherwise, do not reject H0 3. s D ¯ = 1.05 ; t = 3.14 4. t = 3.14 > 2.093 ∴ reject H0 (p < .05) 5. t = 3.14 > 2.861 ∴ p < .01
c. The average number of words per minute (WPM) typed correctly for a sample of 20 adults was significantly greater when tested while sitting (M = 40.20) than while they were walking on a treadmill (M = 36.90), t(19) = 3.14, p < .01.
d. The results of this analysis suggest that typing is not a task that may be performed at a satisfactory level while walking on a treadmill rather than sitting.
574
Answers to Odd-Numbered Exercises
1.
a.
b.
c.
d.
3.
a. Experimental: X ¯ = 5.00 , s = 2.94, s X ¯ = 1.11 ;
Control: X ¯ = 9.00 ; s = 3.06, s X ¯ = 1.16
b. Experimental: X ¯ = 15.00 , s = 3.02, s X ¯ = 1.07
Control: X ¯ = 18.00 ; s = 3.21, s X ¯ = 1.13
575
c. Experimental: X ¯ = 3.27 , s = .46, s X ¯ = .15 ;
Control: X ¯ = 2.80 ; s = .41, s X ¯ = .14
d. Experimental: X ¯ = 3.38 , s = 1.39, s X ¯ = .39
Control: X ¯ = 3.23 ; s = 1.64, s X ¯ = .46
5.
df = 40; critical value = ±2.021
df = 26; critical value = ±2.056
df = 6; critical value = 1.943
df = 62; critical value = −1.671
7. a. s X ¯ 1 − X ¯ 2 = .60 b. s X ¯ 1 − X ¯ 2 = 2.39 c. s X ¯ 1 − X ¯ 2 = 2.30 d. s X ¯ 1 − X ¯ 2 = .47
9. a. t = 1.13
576
b. t = −2.09 c. t = −1.67 d. t = 3.20
11.
s X ¯ 1 − X ¯ 2 = 1.31 ; t = –.38
s X ¯ 1 − X ¯ 2 = 2.03 ; t = 3.27
s X ¯ 1 − X ¯ 2 = .37 ; t = 1.70
s X ¯ 1 − X ¯ 2 = .95 ; t = −2.26
13. a. H0: μAttraction = μNo attraction H1: μAttraction ≠ μNo attraction b.
1. df = 52 2. If t < −2.009 or > 2.009, reject H0; otherwise, do not reject H0 3. s X ¯ − X ¯ 2 = 1.33 ; t = 2.14 4. t = 2.14 > 2.009 ∴ reject H0 (p < .05) 5. t = 2.14 < 2.678 ∴ p < .05 (but not < .01)
c. The average number of women's flirtatious behaviors was significantly higher in the sample of 13 interviewers in the attraction group (M = 37.44) than the 13 interviewers in the no attraction group (M = 34.59), t(24) = 2.14, p < .05.
d. The results of this analysis support the research hypothesis that interviewers who believe that the women they interview are attracted to them elicit more flirtatious behavior from these women than interviewers who do not hold this belief.
15. a. H0: μ20-34 = μ40-54; H1: μ20-34 ≠ μ40-54 b.
1. df = 48 2. If t < −2.021 or > 2.021, reject H0; otherwise, do not reject H0
3. s X ¯ 1 − X ¯ 2 = .94 ; t = −4.22 4. t = −4.19 < −2.021 ∴ reject H0 (p < .05) 5. t = −4.19 < −2.704 ∴ p < .01
c. The average difference in age between themselves and their preferred female
577
partners was significantly higher for the sample of 25 men aged 40 to 54 years (M = 4.98) than the sample of 25 men aged 20 to 34 years (M = 1.04), t(48) = −4.19, p < .01.
d. The results of this analysis support the research hypothesis that the preferred difference in age between themselves and their preferred partners is greater for older men than for younger men.
17.
a. Women: N = 13, X ¯ = 18.62 , s = 7.81
Men: N = 13, X ¯ = 14.85 , s = 8.55 b. H0: μWomen = μMen; H1: μWomen ≠ μMen c.
1. df = 24 2. If t < −2.064 or > 2.064, reject H0; otherwise, do not reject H0
3. s X ¯ 1 − X ¯ 2 = 3.21 ; t = 1.17 4. t = 1.17 is not < −2.064 or > 2.064 ∴ do not reject H0 (p > .05) 5. Not applicable (H0 not rejected)
d. The average estimates of the cost of a wedding (in thousands of dollars) for a sample of 13 women (M = 18.62) and a sample of 13 men (M = 14.85) were not significantly different, t(24) = 1.17, p > .05.
e. The result of this analysis does not support the research hypothesis that young women will provide a higher estimate of the cost of the average wedding than will men.
19. a. s X ¯ 1 − X ¯ 2 = 1.65 b. s X ¯ 1 − X ¯ 2 = 3.11 c. s X ¯ 1 − X ¯ 2 = 1.02 d. s X ¯ 1 − X ¯ 2 = 3.80
21. a. H0: μNew user = μExperienced user; H1: μNew user ≠ μExperienced user b.
1. df = 62 2. If t < −2.000 or > 2.000, reject H0; otherwise, do not reject H0
578
3. s X ¯ 1 − X ¯ 2 = 1.12 ; t = −1.01 4. t = −1.01 is not < −2.000 or > 2.000 ∴ do not reject H0 (p > .05) 5. Not applicable (H0 not rejected)
c. The average level of satisfaction with their hearing aids of the sample of 30 new users (M = 26.90) and a sample of 34 experienced users (M = 28.03) was not significantly different, t(62) = −1.01, p > .05.
d. The result of this analysis indicates that these patients' satisfaction with their hearing aids is not related to how long they've used their hearing aids.
23. a. s D ¯ = 1.60 ; t = 6.25 b. s D ¯ = 2.11 ; t = 3.22 c. s D ¯ = .54 ; t = 3.93 d. s D ¯ = .69 ; t = 1.16
25. a. H0: μD = 0; H1: μD ≠ 0 b.
1. df = 7
2. If t < −2.365 or > 2.365, reject H0; otherwise, do not reject H0 3. s D ¯ = 3.48 ; t = −2.39 4. t = −2.39 < −2.365 ∴ reject H0 (p < .05) 5. t = −2.39 > −3.499 ∴ p < .05 (but not < .01)
27.
a. Before: N1 = 6, X ¯ 1 = 14.50 , s1 = 7.56
After: N2 = 6, X ¯ 2 = 6.33 , s2 = 6.71 b. X ¯ D = 8.17 ; s D ¯ = 6.59 c. H0: μD = 0; H1: μD ≠ 0
579
d. 1. df = 5 2. If t < −2.571 or > 2.571, reject H0; otherwise, do not reject H0
3. s D ¯ = 2.69 ; t = 3.04 4. t = 3.04 > 2.571 ∴ reject H0 (p < .05) 5. t = 3.04 < 4.032 ∴ p < .05 (but not < .01)
e. The average number of cigarettes smoked per day for a sample of six smokers significantly decreased from before (M = 14.50) to 4 weeks after (M = 6.33) using the tools, t(5) = 3.04, p < .05.
f. Using tools such as group support, substituting food for cigarettes, and deep breathing may be effective means of reducing cigarette smoking.
580
Sharpen your skills with SAGE edge at edge.sagepub.com/tokunaga
SAGE edge for students provides a personalized approach to help you accomplish your coursework goals in an easy-to-use learning environment. Log on to access:
eFlashcards Web Quizzes Chapter Outlines Learning Objectives Media Links SPSS Data Files
581
Chapter 10 Errors in Hypothesis Testing, Statistical Power, and Effect Size
582
Chapter Outline 10.1 Hypothesis Testing vs. Criminal Trials 10.2 An Example From the Research: Truth or Consequences
Correct decisions and mistakes in criminal trials 10.3 Two Errors in Hypothesis Testing: Type I and Type II Error
Type I error: the risk in rejecting the null hypothesis What is the probability of making a Type I error? What is the probability of not making a Type I error? Why is Type I error a concern? Why does Type I error occur?
Type II error: the risk in not rejecting the null hypothesis What is the probability of making a Type II error? What is the probability of not making a Type II error? Why is Type II error a concern? Why does Type II error occur?
Type I and Type II error: summary
10.4 Controlling Type I and Type II Error Controlling Type I error
Concerns about controlling Type I error Controlling Type II error
Increasing sample size Raising alpha (α) Using a directional alternative hypothesis (H1) Increasing between-group variability Decreasing within-group variability
Controlling Type I and Type II error: summary 10.5 Measures of Effect Size
Interpreting inferential statistics: the influence of sample size The inappropriateness of “highly significant” Effect size as variance accounted for
Measure of effect size for the difference between two sample means: r2
Interpreting measures of effect size Presenting measures of effect size Reasons for calculating measures of effect size A second measure of effect size for the difference between two sample means: Cohen's d
10.6 Looking Ahead 10.7 Summary 10.8 Important Terms 10.9 Formulas Introduced in This Chapter 10.10 Exercises
In the last few chapters, we've discussed several statistical procedures used to test research hypotheses. Hypothesis testing is centered on making one of two decisions, reject or do not reject the null hypothesis, based on the probability of a statistic that has been calculated. Making the decision whether to reject the null hypothesis implies either that a hypothesized difference or relationship exists in the population or it does not. However, because this
583
decision is based on probability rather than certainty, the possibility remains that whatever decision is made about the null hypothesis may either be correct or incorrect. In this chapter, two types of errors in hypothesis testing will be defined and described. More important, we will discuss strategies used by researchers to minimize the occurrence and consequences of these errors. To facilitate this discussion, we will use what may at first glance seem to be an unlikely analogy of hypothesis testing: the American justice system.
584
10.1 Hypothesis Testing vs. Criminal Trials
As part of conducting research, the process of hypothesis testing follows a number of well- defined steps:
1. A researcher, on the basis of an evaluation of a literature, states a research hypothesis regarding the relationship between variables and collects data to be analyzed.
2. At the start of a statistical analysis, a null hypothesis, one that states a hypothesized relationship does not exist, is presumed to be true.
3. The researcher statistically analyzes the collected data. 4. On the basis of the analysis of the data, the researcher makes one of two decisions:
reject or do not reject the null hypothesis. If the probability of the value of the statistic calculated from the data is sufficiently low, the null hypothesis is rejected; otherwise, the null hypothesis is not rejected.
Note that a decision to “not reject the null hypothesis” is not the same as “accept the null hypothesis.” As discussed in Chapter 6, a statistical analysis cannot “prove” a hypothesized relationship does not exist in a population because data have been collected only from a limited sample of the population. The purpose of an analysis, therefore, isn't to prove the null hypothesis is true but rather to determine whether there is sufficient evidence to reject the null hypothesis.
The procedure for conducting criminal trials within the American justice system in many ways resembles the steps taken in hypothesis testing, even though the two processes occur in completely different settings and have different objectives:
1. A law enforcement agency, on the basis of its evaluation of a crime that has taken place, collects evidence regarding an identified suspect, who may eventually be brought to trial.
2. At the start of a trial, the accused person is presumed to be innocent. 3. The attorneys present their evidence to the jury, which evaluates it. 4. On the basis of its evaluation of the evidence, the jury makes one of two decisions:
guilty or not guilty. In order to reject the presumption of innocence, the jury must decide that the evidence is convincing beyond a reasonable doubt. If the jury decides the reasonable doubt standard has been met, the presumption of innocence is rejected, and the person is found to be guilty; otherwise, the person is found to be not guilty.
It is important to note that a verdict of “not guilty” is not the same as deciding the person is “innocent,” in that the purpose of a criminal trial isn't to prove the person's innocence but rather to determine whether there is sufficient evidence to conclude the person is guilty.
585
As we can see, the goal of both hypothesis testing and a criminal trial is to analyze and evaluate collected evidence to make one of two decisions regarding a stated initial assumption: The assumption is true or the assumption is not true. In both situations, the decision is based on probability rather than certainty; as such, it's possible to make the wrong decision. The next section discusses a research study illustrating the two types of incorrect decisions that may occur in the justice system; this will lead to a discussion of two types of errors that can be made in hypothesis testing.
586
10.2 An Example from the Research: Truth or Consequences
One piece of evidence used to determine whether a suspect in a criminal investigation may be guilty is the result of a psychophysiological detection device. The most well-known example of this type of device is the polygraph examination, sometimes referred to as the lie detector test. A polygraph examination involves attaching tubes and electrodes to a suspect's body to record respiration, skin response, and cardiovascular activity. Next, a trained test examiner asks the suspect a series of questions and records and evaluates physiological responses to these questions. For example, the examiner may ask the suspect where he or she was the day the crime was committed, during which any changes in such things as the suspect's blood pressure or heart rate are noted. Ultimately, the examiner may make a decision whether or not the suspect should be considered guilty of the crime. It is important to understand that polygraph examinations cannot objectively prove someone's guilt; instead, the test examiner's decision is based on the likelihood an innocent person would have reacted the way the suspect did to the questions.
Can polygraph examinations accurately identify criminals? This has been the subject of debate for decades. One questioning technique used in polygraph examinations is the guilty knowledge test (GKT). The GKT is a multiple-choice examination measuring physiological responses to facts known to be relevant to a crime. For example, a suspect may be asked, “What kind of gun was used to shoot Mr. Doe? Was it a .22-caliber rifle, a 12-gauge shotgun, a .38-caliber revolver, an M16 rifle, or a 9-mm handgun?” The suspect's physiological response to hearing the correct choice is compared to his or her responses to the incorrect choices. As noted by Eitan Elaad and his colleagues, “Innocent subjects, who are unable to distinguish relevant from irrelevant alternatives, are not expected to respond differentially to the relevant and irrelevant items … [but] if a subject's physiological responses are consistently greater to the relevant item than to the irrelevant ones, knowledge about a crime is inferred” (Elaad, Ginton, & Jungman, 1992, p. 757).
The goal of the GKT is to identify guilty and innocent suspects as accurately as possible. Consequently, Mr. Elaad and his research team conducted a study examining the accuracy of the GKT in actual crime investigations (Elaad et al., 1992). Part of this study consisted of giving GKT examinations to 68 people interrogated by police during criminal investigations. What is particularly unique about this study is that, after taking these examinations, 38 of the 68 suspects confessed to having actually committed the crimes. As a result, this study was able to assess the accuracy of the GKT by comparing the actual innocence or guilt of the suspect with the decision (not guilty vs. guilty) made by the GKT.
587
Correct Decisions and Mistakes in Criminal Trials
In the GKT study, two pieces of information were recorded for each suspect: (1) the suspect's actual innocence or guilt (this study assumed that a suspect who did not confess to committing the crime and was eventually found not guilty during the trial was in fact innocent) and (2) the GKT's decision of whether the suspect should be considered not guilty or guilty. The four possible outcomes that could occur as a result of combining these two things are illustrated in Figure 10.1.
Figure 10.1 Correct and Incorrect Decisions Using the Guilty Knowledge Test (GKT)
The top row of Figure 10.1 corresponds to suspects who were actually innocent of the crime. Within this row, Cell 1 represents innocent subjects whom the GKT classified as not guilty; this cell represents a correct GKT decision. Cell 2, on the other hand, comprises people innocent of the crime that the GKT defined as guilty. This cell represents an incorrect GKT decision, a mistake made by the justice system in which an innocent person could be sent to jail.
The bottom row of Figure 10.1 contains suspects actually guilty of the crime. On the right side of this row (Cell 4) are guilty suspects identified by the GKT as guilty; this cell represents a correct GKT decision. However, Cell 3 represents people who confessed to committing the crime but were categorized by the GKT as not guilty. This is an incorrect decision, a mistake made by the justice system in which a guilty person may go free rather than be punished.
In summary, Figure 10.1 represents four possible outcomes that can occur using the GKT to make decisions regarding a person's innocence or guilt. Two of these outcomes represent correct decisions: classifying an innocent person as not guilty (Cell 1) and classifying a guilty person as guilty (Cell 4). As the goal of the justice system is to protect the innocent while convicting the guilty, we would want the proportion of people in these two cells to be as large as possible.
588
The other two outcomes, deciding an innocent person is guilty (Cell 2) and deciding a guilty person is not guilty (Cell 3), are incorrect decisions. We would certainly want to minimize the possibility of making either of these mistakes; however, do we consider them equally undesirable? Simply put, which do we as a society believe is worse: sending an innocent person to jail or letting a guilty person go free? Historically, a large number of laws and judicial rulings have been based on the belief that a greater miscarriage of justice has occurred if an innocent person is incarcerated than if a criminal escapes punishment. As a result, although we want the proportion of decisions in Cells 2 and 3 to be as small as possible, it is preferable to label a guilty person as not guilty (Cell 3) than to decide an innocent person is guilty (Cell 2).
In terms of the GKT study's findings regarding the accuracy of the GKT in criminal investigations, Figure 10.2 provides the number of people in each of the four possible outcomes. Looking first at the innocent subjects (Cells 1 and 2), we see that 37 of the 38 innocent subjects (97%) were correctly categorized by the GKT as not guilty, with only 1 of the 38 (3%) mistakenly labeled as guilty. Therefore, the GKT was found to be highly accurate in identifying innocent suspects. However, looking at the bottom row of Figure 10.2, the GKT was not as precise in identifying guilty suspects, with the 30 guilty suspects almost equally likely to be labeled not guilty (14/30 or 47%) as guilty (16/30 or 53%). Comparing Cells 2 and 3, the GKT was much more likely to lead to the incorrect decision of letting a guilty person go free (Cell 3) than of sending an innocent person to jail (Cell 2).
589
10.3 Two Errors in Hypothesis Testing: Type I and Type II Error
The GKT study illustrates correct and incorrect decisions that can be made in the American justice system. Just as mistakes can be made regarding a suspect's innocence or guilt, in hypothesis testing, it is possible to commit errors in concluding whether a hypothesized effect exists in the population. Two types of errors that can be made in hypothesis testing, Type I and Type II error, are defined and described below.
590
Type I Error: The Risk in Rejecting the Null Hypothesis
Hypothesis testing starts with the assumption that the null hypothesis is true, which implies that the hypothesized effect or relationship does not exist in the population. This is similar to the process of criminal trials, which start with the assumption that the accused is innocent and has not committed a crime. The top row in Figure 10.3 portrays the situation where the null hypothesis is in fact true; within this row, Cell 1 represents the situation where the null hypothesis is true and, in analyzing data, the decision is made to not reject the null hypothesis. In other words, the hypothesized effect does not exist in the population, and from our analysis, we conclude that it does not exist. This represents a correct decision, similar to the GKT deciding that an innocent suspect is not guilty.
Figure 10.2 Accuracy of GKT Test (Elaad et al., 1992)
Cell 2, on the other hand, represents a Type I error, which is an error that occurs when the null hypothesis is true but the decision is made to reject the null hypothesis. That is, the hypothesized effect does not exist in the population, but on the basis of the analysis of the data from our sample, we conclude that it does exist. Type I error may also be expressed as “rejecting a true null hypothesis,” “rejecting the null hypothesis when we shouldn't,” or “concluding an effect exists when it actually does not.”
To illustrate Type I error, let's return to our Super Bowl example in Chapter 6, in which we tested a hypothesis regarding whether winning the pregame coin flip affects a team's chances of winning the Super Bowl. In this study, a Type I error would occur if winning the pregame coin flip does not actually affect a team's chances of winning the game, but from our analysis of a set of data, we decide that it does. In the parking lot study from Chapter 9, a Type I error would happen if the two groups of drivers in the population, Intruder and No intruder, do not differ in the time taken to leave a parking space, yet in our study we conclude that they do differ. Within the GKT study, the equivalent of a Type I error occurred when the GKT concluded a person was guilty of the crime when he or she was in fact innocent.
What is the Probability of Making a Type I Error?
591
Figure 10.3 Correct Decisions and Errors in Hypothesis Testing
In analyzing a set of data, what is the chance we will decide to reject the null hypothesis when in fact we shouldn't? Given that a Type I error occurs only when the decision is made to reject the null hypothesis, the probability of making this type of error is the same as the probability of rejecting the null hypothesis. That is, the probability of making a Type I error is equal to alpha (α), defined in Chapter 6 as the probability of a statistic used to make the decision whether to reject the null hypothesis:
(10-1) p ( Type I error ) = α
Because alpha is traditionally set to .05, the probability of making a Type I error is typically .05. Consequently, in making the decision whether to reject the null hypothesis, there is a .05 probability of concluding an effect exists when it in fact does not.
What is the Probability of Not Making a Type I Error?
If the probability of making a Type I error is equal to the probability of rejecting the null hypothesis, it stands to reason that the probability of not making a Type I error is equal to the probability of not rejecting the null hypothesis. If the probability of making a Type I error is equal to alpha (α), the probability of not making this error is equal to 1 – α: p ( not making Type I error ) = 1 − α
Assuming α = .05, the probability of not making a Type I error is equal to 1 – .05 or .95. Therefore, in making the decision whether to reject the null hypothesis, there is a .95 probability of correctly concluding an effect does not exist.
Note that there is a much lower probability of making a Type I error (.05) than not making a Type I error (.95). In hypothesis testing, researchers assume the null hypothesis is true unless the data provide convincing evidence to the contrary. The assumption that an effect does not exist is rejected only when the evidence is very strong. Again, this closely parallels the situation in criminal trials, in which the presumption of innocence is rejected only when the evidence is convincing beyond a reasonable doubt.
592
Why is Type I Error a Concern?
It is important for researchers to avoid making Type I errors because these errors may result in the communication of incorrect information. Science is built on the accumulation of research such that researchers base their studies on the results of others' findings. Consequently, stating that a relationship or effect exists when it does not may adversely affect others' thinking, time, energy, and resources.
Imagine, for example, you are a researcher for a pharmaceutical company trying to refine a drug used to fight lung cancer. You most likely base your research on the results of published research studies regarding the effectiveness of the drug. If it turned out that the findings of some of these studies were incorrect in that the drug was in fact not effective, not only might you have lost a great deal of time and effort, but you would also need to alter and adjust how you think about and study the topic. Furthermore, Type I error not only affects researchers but also may have negative consequences for a broader audience. Returning to the drug example, learning that a drug is not as effective as previously believed may dramatically affect the lives of people relying on the drug to combat their ailments.
Why does Type I Error Occur?
Type I errors occur as the result of random, chance factors. To illustrate how and why these errors occur, let's return to the example in Chapter 6 of flipping a coin 30 times (Table 6.6 and Figure 6.6). In this example, we developed a distribution of the number of heads that can occur in 30 coin flips under the assumption that coins are in fact fair and there is an equal probability of getting heads or tails. Using this distribution, we developed the decision rule that the initial assumption that coins are fair would be rejected if we were to get either less than 10 heads or more than 20 heads in 30 coin flips. Therefore, although it's possible to get, say, 24 heads in 30 coin flips even when coins are fair, the probability of this event is so low that we'll make the decision to reject the null hypothesis in favor of the alternative hypothesis.
Given the possible undesirable effects of Type I error, you may wonder why researchers allow this mistake to happen. Wouldn't it be more reasonable to create a situation where the null hypothesis can never be incorrectly rejected? Eliminating the possibility of Type I error is untenable for one basic reason: The only way to never make a Type I error is to never reject the null hypothesis, no matter how low the probability of the value of a calculated statistic. A similar dilemma exists in the justice system, where the only way to never convict an innocent person is to never convict anyone, no matter how convincing the evidence.
Creating a situation in which the null hypothesis could never be rejected (or suspects could never be convicted) would render hypothesis testing (and the justice system) ineffective and
593
obsolete.
594
Learning Check 1: Reviewing what you've Learned So Far
1. Review questions a. How is making the decision whether to reject the null hypothesis similar to making the
decision regarding guilt in a criminal trial? b. What are the different ways of describing a Type I error? c. If you conducted a study comparing two groups, what would be a Type I error? d. Why is the probability of Type I error equal to alpha (α)? e. Why is the probability of not making a Type I error equal to 1 – alpha (1 – α)? f. Why is Type I error a concern? g. Why do researchers allow for the possibility of making Type I errors?
2. For each of the situations below, describe what would be considered a Type I error. a. A restaurant owner wants to see if his customers can tell if they are drinking alcoholic or
nonalcoholic beer. b. A dentist tests two types of anesthesia on her patients to see which type is better at
alleviating pain. c. A consumer protection agency wants to know if computer users are less likely to get infected
if they pay for virus protection software rather than download free software.
Thus, the process of hypothesis testing must allow for the possibility of making a Type I error. Later in this chapter, we will discuss methods used by researchers to minimize the possibility of Type I error.
595
Type II Error: The Risk in Not Rejecting the Null Hypothesis
In addition to rejecting the null hypothesis when it is in fact true, a second type of error can be made in hypothesis testing: not rejecting the null hypothesis when the hypothesized effect does exist in the population. Returning to Figure 10.3, the bottom row in this figure pertains to the situation where the alternative hypothesis (H1) is true, meaning that the hypothesized effect does in fact exist in the population. Starting this time with the cell on the right half of this row, Cell 4 represents the situation where the alternative hypothesis is true and the decision is made to reject the null hypothesis. Deciding an effect exists when it does in fact exist is a correct decision, similar to the GKT test correctly identifying a guilty suspect in a criminal investigation.
Unfortunately, it is possible to make a wrong decision regarding the alternative hypothesis. Cell 3 represents a Type II error, which is an error that occurs when the decision is made to not reject the null hypothesis even though the alternative hypothesis is true. That is, the hypothesized effect exists in the population, but on the basis of the data from our sample, we conclude that it does not exist. Other ways of describing a Type II error include “failing to reject a false null hypothesis,” “not rejecting the null hypothesis when we should,” and “concluding an effect does not exist when it actually does.”
Returning to the Super Bowl example in Chapter 6, a Type II error would occur if winning the pregame coin flip does in fact affect a team's chances of winning the game, but from our analysis of a set of data, we decide that it does not. In the Chapter 9 parking lot study, a Type II error would be if Intruder and No intruder drivers in the population differ in the time taken to leave a parking space, yet in our study, we conclude that they do not differ. In the GKT study, the equivalent of a Type II error occurred when the GKT concluded a person was not guilty, but the person later confessed to committing the crime.
What is the Probability of Making a Type II Error?
The probability of making a Type II error, of not rejecting the null hypothesis when we should, is represented by the Greek symbol beta (β):
(10-2) p ( Type II error ) = β
Mathematically, what is the probability of making a Type II error? If the probability of making a Type I error is typically .05 (α = .05), it might seem reasonable to assume that the probability of a Type II error (β) is equal to .95. However, β is not equal to 1 – α; this is because Type I and Type II errors take place under different situations.
596
Looking at Figure 10.3, we see that Type I error occurs in situations where the null hypothesis (H0) is true, whereas Type II error occurs in situations where it is the alternative hypothesis (H1) that is true. Using the GKT study (Figure 10.2) as an illustration, if we were to randomly pick one of the 38 innocent people, the probability of making what would be considered a Type I error, labeling an innocent suspect as guilty, is 1/38 or .03. However, looking at the bottom left cell of this figure, the probability of making a Type II error, labeling a guilty suspect as innocent, is not .97 (37/38) but rather .47 (14/30).
Unlike the GKT study, in which the suspects could be definitively identified as being either innocent or guilty in a statistical analysis it is extremely difficult to determine the exact probability of making a Type II error. This difficulty is due to the inherent nature of how the alternative hypothesis is stated. One way to explain this is by providing an example that compares the null hypothesis (where the precise value of the population mean is specified) and the alternative hypothesis (where the precise value of the population mean is not specified).
In the reading skills study example in Chapter 7, the null hypothesis was stated as H0: μ = 124.81. Because the null hypothesis defines a specific value for the population mean (μ = 124.81), we can create one sampling distribution of the mean with μ = 124.81 in the center of this distribution and use this distribution to identify the values of the sample mean that have a probability less than alpha (α). As a result, we can determine the probability of making a Type I error, of rejecting the null hypothesis when it is true.
Turning to the probability of making a Type II error, the alternative hypothesis for the reading skills study was stated as H1: μ ≠ 124.81. By definition, the alternative hypothesis does not indicate a specific value for the population mean. Consequently, a true alternative hypothesis in the reading skills study simply implies μ is equal to some value other than 124.81. Given the many possible values of μ, it becomes extremely difficult to determine the probability of either accepting or failing to accept the alternative hypothesis for any value of the sample mean. In short, the probability of making a Type II error in a statistical analysis is extremely difficult to determine because the value of a parameter such as the population mean μ is unknown when the alternative hypothesis is true.
What is the Probability of Not Making a Type II Error?
If the probability of making a Type II error is represented symbolically by β, the probability of not making a Type II error is represented by 1 – β: p ( not making Type II error ) = 1 − β
Given our earlier discussion of the difficulty in determining the exact probability of making a Type II error, it's not surprising that it's equally difficult to determine the probability of not making this error.
597
The probability of not making a Type II error has been given a special name, statistical power, defined as the probability of rejecting the null hypothesis when the alternative hypothesis is true. In other words, the statistical power of an analysis is the probability of detecting an effect when it does in fact exist in the population. Because rejecting the null hypothesis and detecting effects typically supports a study's research hypothesis, it is in a researchers best interests to maximize the statistical power of an analysis. In Section 10.4 of this chapter, we discuss strategies used by researchers to maximize statistical power.
Why is Type II Error a Concern?
Type II error, concluding an effect doesn't exist when in fact it does, is as important a concern to researchers as Type I error. However, it's important for a very different reason: A Type II error may result in the non-communication of correct information. A Type II error occurs when the null hypothesis isn't rejected and a researcher concludes a difference or relationship doesn't exist; this conclusion implies a study's research hypothesis has not been supported. Because research studies whose hypotheses haven't been supported do not typically get published in academic journals or presented at conferences, research findings where the null hypothesis has not been rejected might not be communicated to other people. Returning to the drug example introduced earlier, a Type II error would occur if a drug is in fact effective but the researcher concludes it isn't effective. In such a situation, people's lives would be adversely affected in that they might not get access to a drug that could help them. However, they may not become aware that this error has occurred because the drug will not be produced or sold.
Why does Type II Error Occur?
Type II errors occur when research studies do not find differences or relationships that actually exist in the population. As with Type I error, this can take place as the result of random, chance factors. Also, the effect being studied may be very small and difficult to detect. However, Type II error may also result from how researchers conduct their studies. For example, the sample in the study may have been small or unrepresentative of the population, or the methods used to measure or manipulate variables may have been inappropriate, incomplete, or flawed in some way.
As you can see, it is difficult to identify the precise reasons why the null hypothesis is not rejected. However, in conducting their studies, researchers have control over factors that increase the likelihood of rejecting the null hypothesis. Later in this chapter, methods used by researchers to decrease the probability of making a Type II error, and thereby increase the statistical power of a statistical analysis, will be described and evaluated.
598
Type I and Type II Error: Summary
To summarize, statistical analyses designed to test research hypotheses begin with two statistical hypotheses: the null hypothesis (H0) and the alternative hypothesis (H1). In the population, one of two actualities may be true: The null hypothesis is true or the alternative hypothesis is true. In conducting a statistical analysis, one of two decisions is made: Reject or do not reject the null hypothesis.
By combining the two actualities with the two decisions, we are able to demonstrate that one of four possible outcomes may occur whenever a decision is made regarding the null hypothesis. The top row of Figure 10.4 represents the two outcomes that may occur when the null hypothesis is true. Here, deciding to not reject a true null hypothesis (Cell 1) is a correct decision, but rejecting a true null hypothesis (Cell 2) is an incorrect decision known as a Type I error. The probability (p) of making a Type I error is equal to alpha (α), typically set at .05; the probability of not making this error, 1 – α, is typically equal to .95. Type I error is a concern because the communication of incorrect information negatively affects others' thinking, time, energy, and resources. Type I errors occur as the result of random, chance factors, but researchers must allow for the possibility of making this error because the only way it may be eliminated is to never reject the null hypothesis regardless of the evidence.
The two outcomes depicted in the bottom row of Figure 10.4 take place when the alternative hypothesis is true. It is possible to commit a Type II error (Cell 3), which is failing to reject a false null hypothesis. The exact probability of committing a Type II error, represented by beta (β), is extremely difficult to determine because of the inherently imprecise nature of the alternative hypothesis. Type II errors are a concern because they may result in the non-communication of correct information that is potentially useful to others. Type II errors occur as the result of random, chance factors as well as how researchers conduct their studies, and the probability of making this error is affected by a number of factors discussed later in this chapter. Finally, Cell 4 represents a correct decision, in which the null hypothesis is rejected when the alternative hypothesis is true. The probability of this outcome (1 – β), which is the probability of not making a Type II error, is the statistical power of an analysis.
Figure 10.4 Summary of Type I and Type II Errors
599
600
Learning Check 2: Reviewing what you've Learned So Far
1. Review questions a. What are the different ways of describing a Type II error? b. If you conducted a study comparing two groups, what would be a Type II error? c. What is the difficulty in calculating the probability of Type II error? d. What is the concept of statistical power? What is its relationship to Type II error? e. Why is Type II error a concern? f. Why do Type II errors occur?
2. For each of the situations below, identify an example of a Type II error. a. A snowboarder tries two different kinds of wax to see which one lasts longer. b. A high school counselor wants to see whether students who pay for test preparation services
are more likely to get into college than students who take free practice exams. c. A driver trying to save money uses both regular and premium gasoline and wants to know
whether the type of gasoline affects her cars performance.
3. For each of the following hypotheses, create a chart that shows the four possible outcomes that could occur as a result of hypothesis testing. There should be two correct decisions and two incorrect decisions. Label the type of error for each of the incorrect decisions.
a. People who drive electric vehicles live closer to their workplaces than people who drive gasoline-driven vehicles.
b. People who drive after playing driving-related video games are more likely to get into automobile accidents than people who drive after playing stationary video games.
601
10.4 Controlling Type I and Type II Error
This section discusses methods used by researchers to minimize the occurrence of Type I and Type II error, as well as concerns and drawbacks related to these methods. To demonstrate these methods, we will revisit the parking lot study discussed in Chapter 9. Table 10.1 lists the steps used to test the difference between the departure times of the Intruder and No intruder groups. As we will see, these steps change depending on the particular method used to control for Type I or Type II error.
602
Controlling Type I Error
Given that a Type I error can occur when the decision is made to reject the null hypothesis, the probability of making this error can be reduced by making it more difficult to reject the null hypothesis. More specifically, Type I error may be controlled by lowering the value of alpha (α), the probability of a statistic needed to reject the null hypothesis. For example, rather than using the traditional α = .05 probability cutoff, a researcher could choose to reject the null hypothesis only when the probability of a calculated statistic is less than .01 (p < .01).
Table 10.2 demonstrates how to reduce Type I error by lowering alpha, using the parking lot study as an example. In this table, the difference between the means of the Intruder and No intruder groups is tested under two conditions: α = .05 and α = .01. Comparing the steps in the two conditions, we see that changing alpha does not alter how the null and alternative hypotheses (H0 and H1) are stated or how the inferential statistic (t) is calculated. However, lowering alpha from .05 to .01 increases the critical value of the statistic, in this example from 2.048 to 2.763, which makes the region of rejection smaller. Consequently, when α is equal to .01, a more extreme value of the statistic is needed to reject the null hypothesis than when α = .05.
Table 10.1 Summary of Hypothesis Testing, Parking Lot Study (Chapter 9)
Looking at Table 10.2, we see that the null hypothesis in the parking lot study is not rejected when α is equal to .01, even though it was rejected when α = .05. Reducing the probability of making a Type I error by lowering alpha means we are adopting a more conservative approach, only rejecting the null hypothesis when the calculated value of a
603
statistic has an extremely low probability of occurring.
Concerns about Controlling Type I Error
Reducing Type I error by lowering the value of alpha is a relatively simple matter. However, there are several concerns about this method. First, statistical significance is typically and traditionally defined using the α = .05 cutoff. Researchers may be reluctant to set α to .01 because doing so may lead to confusion regarding the interpretation of the term statistical significance.
A second concern about controlling Type I error is that reducing the probability of Type I error increases the probability of Type II error. That is, the more difficult we make it to conclude an effect exists, the less likely we are to detect the effect when it does in fact exist. Using the justice system analogy, the harder we make it to decide a person is guilty, the less likely we are to detect those who are in fact guilty.
Table 10.2 Controlling for Type I Error by Lowering Alpha (α = .05 vs. .01)
Figure 10.5 illustrates the trade-off between Type I and Type II error. The figure in the upper-left corner, Figure 10.5(a), represents a sampling distribution of a hypothetical statistic under the assumption the null hypothesis (H0) is true; the shaded area in this figure represents the probability of Type I error when α = .05. Moving down, we find Figure 10.5(b), which represents a distribution of the statistic under the assumption that the alternative hypothesis (H1) is true; this figure is a symbolic rather than actual representation of this distribution because, as mentioned earlier, the mean of this distribution cannot be defined. The shaded area in Figure 10.5(b) represents the probability of making a Type II error using the same α = .05 cutoff as in Figure 10.5(a).
604
What is the consequence of reducing Type I error by lowering α? In the distribution in the upper-right hand corner (Figure 10.5(c)), the shaded area represents the probability of making a Type I error when α = .01 rather than .05. As we see by comparing Figures 10.5(a) and 10.5(c), lowering α from .05 to .01 reduces the probability of Type I error. However, comparing the shaded areas in the two bottom distributions, Figure 10.5(b) and 10.5(d), we see that lowering alpha from .05 to .01 increases the probability of Type II error.
Figure 10.5 Trade-off between Type I and Type II Error (α =.05 vs. α =.01)
605
Controlling Type II Error
A Type II error occurs when the decision is made to not reject the null hypothesis when the alternative hypothesis is true. Reducing Type II error, thereby increasing the statistical power of an analysis, involves increasing the likelihood of rejecting the null hypothesis. As discussed in several chapters of this book, a number of factors influence the likelihood of rejecting the null hypothesis: the size of the sample, the value of alpha (α), and the directionality of the alternative hypothesis. This section illustrates how these factors may be changed to increase the statistical power of an analysis; furthermore, we will introduce two additional factors that may be altered to increase the likelihood of rejecting the null hypothesis: between-group and within-group variability.
Increasing Sample Size
Perhaps the most commonly used method to increase the likelihood of rejecting the null hypothesis (i.e., increase the statistical power of an analysis) is to increase the size of a sample. In Table 10.3, the parking lot study is altered by doubling the sample size from N = 15 to N = 30. Keeping all other aspects of the study the same, increasing the sample size alters a number of aspects of the analysis.For example, a larger sample size results in smaller critical values (±2.009 vs. ±2.048); the smaller the critical values, the greater the likelihood of rejecting the null hypothesis. Also, the larger sample results in a bigger calculated value of the t-statistic (3.39 vs. 2.42); the larger the calculated value of a statistic, the greater the likelihood of rejecting the null hypothesis.
Table 10.3 Controlling for Type II Error by Increasing Sample Size
Although it may be simple to say, “If you want to reject the null hypothesis, use a big sample,” the statistical advantages of a large sample must be balanced against pragmatic
606
concerns. Increasing the sample size increases the amount of time, energy, and resources needed to collect data. As such, the answer to the question, “How large should my sample be?” may depend on the size of the hypothesized effect being studied. The smaller the effect, the larger the sample must be to detect it.
Raising Alpha (α) The probability of making a Type II error may also be decreased by raising alpha (α). Table 10.4 alters the parking lot study by raising alpha from .05 to .10, implying that the null hypothesis will be rejected when the probability of the calculated statistic is only less than .10 (rather than the more stringent .05). Looking at this table, we see that raising alpha lowers the critical value used to make the decision to reject the null hypothesis (±1.701 vs. ±2.048), which in turn increases the size of the region of rejection. Because a smaller value of the statistic is needed to reject the null hypothesis when α = .10 than when α = .05, the likelihood the null hypothesis will be rejected has been increased.
Table 10.4 Controlling for Type II Error by Raising Alpha (α = .10 vs. .05)
The drawbacks of raising alpha from .05 to .10 to increase the power of an analysis are similar to the concerns about controlling Type I error by lowering alpha from .05 to .01. First, using a value of alpha other than the traditional value of .05 may create uncertainty over the interpretation of the term statistical significance. Second, raising alpha from .05 to .10 to lower the probability of Type II error increases the probability of making a Type I error. In other words, increasing the likelihood of deciding an effect exists also increases the possibility of concluding the effect exists when it does not. Given our concerns about Type I error, researchers do not typically increase the power of an analysis by raising alpha.
607
Using a Directional Alternative Hypothesis (H1)
Another, somewhat controversial, way to increase the probability of rejecting the null hypothesis, and therefore increase the statistical power of an analysis, is to use a directional (one-tailed) rather than a non-directional (two-tailed) alternative hypothesis. In the parking lot study, the non-directional alternative hypothesis of H1: μIntruder ≠ μNo intruder was used to allow for the possibility that the departure time for the Intruder group may be longer or shorter than the No intruder group. However, given the hypothesis that the Intruder group should take longer to leave than the No intruder group, what if we were to use a directional hypothesis such as H1: μIntruder > μNo intruder?
Table 10.5 analyzes the parking lot study using both a directional and non-directional alternative hypothesis. The critical value is smaller for a directional hypothesis than for a non-directional hypothesis (1.701 vs. 2.048), with the .05 region of rejection at one end of the t-distribution rather than split between the two ends, making it more likely to reject the null hypothesis.
Using a directional alternative hypothesis to decrease the probability of making a Type II error is similar to changing a two-tailed alpha from .05 to .10. However, there is one important distinction between the two approaches: A one-tailed test does not allow for results in the direction opposite from that specified by the alternative hypothesis. For example, stating the alternative hypothesis as H1: μIntruder > μNo intruder requires the Intruder group mean to be greater than the No intruder mean to reject the null hypothesis. In this situation, we would not reject the null hypothesis if, for some reason, the Intruder mean was much less than the No intruder mean.
Table 10.5 Controlling for Type II Error by Using a Directional Alternative Hypothesis
608
Using a directional alternative hypothesis may be justified when it is reasonable to believe the hypothesized effect may only occur in one direction. When this is not the case, however, using a directional alternative hypothesis simply to increase the likelihood of rejecting the null hypothesis is a questionable practice.
Increasing Between-Group Variability
The purpose of this section is to introduce another factor that may be altered to increase the likelihood of rejecting the null hypothesis: between-group variability. This factor is introduced below in the formula for the t-test for independent means (Formula 9-2): t = X ¯ 1 − X ¯ 2 s X 1 ¯ − X 2 ¯
We've seen that the larger the value of the t-statistic calculated from a set of data, the more likely we are to reject the null hypothesis. One factor affecting the value of the t-statistic is the value of the numerator of this formula: the difference between means ( X ¯ 1 − X ¯ 2 ). The greater the difference between means, the larger the calculated value of t and therefore the greater the likelihood of rejecting the null hypothesis.
As the numerator of the formula for the t-test involves the difference between group means, it may be referred to as between-group variability, defined as differences among the means of the different groups that comprise an independent variable. The greater the amount of between-group variability, the greater the likelihood of rejecting the null hypothesis.
Table 10.6 illustrates the effect of increasing between-group variability on the decision to reject the null hypothesis. In the original parking lot study, the means of the Intruder and No intruder groups were 40.73 and 31.67, respectively. In Table 10.6, this difference has been increased by adding 5 seconds to the Intruder group mean and subtracting 5 seconds from the No intruder mean. Keeping all other aspects of the analysis the same, increasing between-group variability increases the calculated value of the t-statistic (5.09 vs. 2.42), which increases the likelihood the null hypothesis will be rejected.
How can the amount of between-group variability in a study be increased? This may be accomplished by using research methods and manipulations that are effective and are sensitive to differences between groups. For example, in the parking lot study, we could have one of the researchers act as the intruding driver rather than using actual drivers. The researcher could behave in a way that enhanced the level of intrusion, such as honking the horn or talking to the departing driver, which may increase the difference between the departing times of the Intruder and No intruder groups. Returning to the drug example introduced earlier in this chapter, in testing the effectiveness of the drug, you could decide to administer it only to people who have no other medical conditions that might interfere with or reduce the drugs ability to fight lung cancer.
609
Decreasing Within-Group Variability
Another factor affecting the likelihood of rejecting the null hypothesis is within-group variability. This factor is illustrated in the denominator of the formula for the t-test for two means: the standard error of the difference ( s X ¯ 1 − X ¯ 2 ). Below is the formula for the standard error of the difference (Formula 9-3):
Table 10.6 Controlling for Type II Error by Increasing Between-Group Variability
s X ¯ 1 − X ¯ 2 = s 1 2 N 1 + s 2 2 N 2
Looking at this formula, the calculated value of s X ¯ 1 − X ¯ 2 is a function of the standard deviations of the two groups (s1 and s2). The smaller the standard deviations, the smaller the standard error of the difference; the smaller the value of the standard error of the difference, the larger the value of the t-test and the greater the likelihood of rejecting the null hypothesis.
The standard deviations represent the variability of scores within each of the groups. Therefore, the denominator of the formula for a statistic such as the t-test may be referred to as within-group variability, defined as the variability of scores within the groups that comprise an independent variable. The smaller the amount of within-group variability in a study, the greater the likelihood of rejecting the null hypothesis.
Table 10.7 modifies the parking lot study, this time decreasing within-group variability by dividing each of the two standard deviations by 2. Again keeping all other aspects of the analysis the same, decreasing within-group variability ultimately increases the value of the t- statistic (4.84 vs. 2.42), thereby increasing the likelihood of rejecting the null hypothesis.
To answer the question regarding how researchers can reduce the amount of within-group
610
variability in a study, we must first consider what within-group variability represents. How much variability should there be among the scores for a group? Theoretically, because all of the members of a group are presumed to be the same, there should be zero or no variability. For example, all of the drivers who are intruded upon should take exactly the same amount of time to leave their parking spaces. Any variability in the departure times of these drivers cannot be explained or accounted for and therefore is seen as the result of factors, referred to as error. Within-group variability is sometimes referred to as error variance, which is variability that cannot be explained or accounted for.
There are several ways to reduce the amount of within-group variability in a study. Perhaps the simplest way is to increase the size of the sample. As a sample grows in size and more closely approximates the population, the amount of variability among scores decreases. But error may also be reduced by collecting data in a systematic, standardized manner. For example, in the parking lot study, error variance was reduced by having all of the data collectors measure the variable “departure time” the same way: the difference between when the departing drivers opened their door and when the car had completely left the parking space.
Table 10.7 Controlling for Type II Error by Decreasing Within-Group Variability
611
Controlling Type I and Type II Error: Summary
Table 10.8 summarizes the methods available to researchers to control for Type I and Type II error. It is important to note that these methods differ in both the concerns they raise and the demands they place on researchers. For example, raising alpha or using a directional alternative hypothesis is a relatively simple method to decrease Type II error; however, researchers may be reluctant to use these methods because they are contrary to traditional practices. Consequently, researchers might instead focus on how data are collected rather than how they are analyzed, perhaps by increasing the size of the sample or determining ways to maximize between-group variability or minimize within-group variability, even though this may add to the time and resources necessary to conduct a study.
Thus far in this chapter, we've defined errors that can occur in making the decision to reject the null hypothesis, the probability of making these errors, why these errors occur, and methods used to reduce the occurrence of these errors. The next section of this chapter introduces a strategy used by researchers to appropriately interpret the results of their statistical analyses regardless of the decision about the null hypothesis.
Table 10.8 Summary, Methods Controlling for Type I and Type II Error Table 10.8 Summary, Methods Controlling for Type I and Type II Error
Type of Error Probability of Making Error
Method for Controlling
Concern About Controlling
Type I (rejecting H0 when H0 is true)
α (typically .05)
Lower alpha (i.e., .05 to .01)
Traditional use of α = .05
Increases p(Type II error)
Type II (not rejecting H0 when H1 is true)
Difficult to calculate
Increase sample size
Pragmatic concerns (time, energy, cost)
Raise alpha (α) (i.e., .05 to .10)
Traditional use of α = .05
Increases p (Type I error)
Use directional (one-tailed) H1
Ability to hypothesize direction of relationship
Does not allow for result in opposite
612
direction
Increase between group variability
Determining method to collect data
Decrease within- group variability
Determining method to collect data
613
Learning Check 3: Reviewing what you've Learned So Far
1. Review questions a. What is the main method used to control Type I error? What are concerns with this
method? b. What are some methods for controlling Type II error? What are concerns with each of these
methods? c. What is the relationship between the probability of making Type I and Type II errors? d. What happens to the critical values and the region of rejection when you control Type I
error by lowering alpha or control Type II error by raising alpha? e. What happens to the critical values and the region of rejection when you control Type II
error by using a directional alternative hypothesis? f. What methods for increasing the statistical power of an analysis require the least amount of
effort from a researcher? What methods require the greatest amount of effort? g. What is the difference between between-group variability and within-group variability?
2. Imagine two studies are independently conducted regarding the same topic. Somehow, both studies find the same means and standard deviations for two groups: X ¯ 1 = 7.00 , s1, = 2.00, and X ¯ 2 = 5.00 , s2 = 1.00. However, the two studies differ in the sample sizes of the two groups: Ni = 4 in the first study and Ni = 9 in the second study.
a. Calculate the standard error of the difference ( s X ¯ 1 − X ¯ 2 ) and the t-test for independent means for each study.
b. How does increasing the sample size from Ni = 4 to Ni = 9 affect the calculated value of the t-statistic, as well as the probability of making a Type II error?
3. Two researchers set out to test the same research hypothesis. Both studies are based on the same sample size (Ni = 23), and somehow both studies end up with the same standard deviations for the two groups (s1 = 5.00 and s2 = 6.00). However, the experimental manipulation is more effective for the second researcher than for the first researcher, which results in a bigger difference between the means of the two groups: X ¯ 1 = 10.00 and X ¯ 2 = 12.00 for the first researcher and X ¯ 2 = 9.00 and X ¯ 2 = 13.00 for the second researcher.
a. Calculate the standard error of the difference ( s X ¯ 1 − X ¯ 2 ) and the t-test for independent means for each study.
b. How does the increased between-group variability in the second study affect the calculated values of the t-statistic, as well as the probability of making a Type II error?
614
10.5 Measures of Effect Size
Within the research process, a research hypothesis states a hypothesized relationship between variables. Drawing a conclusion regarding whether a study's research hypothesis has been supported is typically based on whether the null hypothesis in a statistical analysis has been rejected. The previous section discussed a number of factors that directly influence the decision to reject the null hypothesis. Because these factors affect the decision about the null hypothesis, they also impact how researchers interpret the results of their statistical analyses. This section first illustrates a situation that may develop when researchers do not take into account the impact of sample size on the decision to reject the null hypothesis. Next, strategies used by researchers to address this situation are presented and illustrated.
615
Interpreting Inferential Statistics: The Influence of Sample Size
Several times in this book, we have discussed the critical role of sample size in research studies. Because sample size directly influences the decision regarding the null hypothesis, it also affects conclusions drawn from statistical analyses. In this section, the Chapter 9 parking lot study will be used to illustrate the role of sample size in interpreting and communicating the results of statistical analyses. When we discussed this study in Chapter 9, the decision was made to reject the null hypothesis and conclude that the departure time of the Intruder group (M = 40.73s) was significantly greater than that of the No intruder group (M = 31.67s), t(28) = 2.42, p < .05.
Table 10.9 presents a modified version of the data from the parking lot study. In this new version, the original 15 scores in each of the two groups have been duplicated. For example, the last 15 scores for the Intruder group are the same as the first 15 scores. This example is similar to what we did earlier in showing how Type II error can be reduced by increasing sample size; however, this time we will discuss how sample size influences both the calculations and interpretation of inferential statistics.
Table 10.9(b) provides the mean and standard deviation of the departure times for the two groups. It is important to notice that the two means (M = 40.73 and M = 31.67) are exactly the same as in the original example; even though the size of the sample has increased, the difference between the groups has not changed. On the other hand, the standard deviations of the two groups (10.24 and 9.90) for Ni = 30 are smaller than when Ni = 15 (10.42 and 10.08), reinforcing our earlier statement that within-group variance may be reduced by increasing sample size.
Table 10.10 tests the difference between the two means for the modified data. Comparing this table with Table 10.1, we see that moving from Ni = 15 to Ni = 30 results in a number of differences in the conducting and results of the analysis. For example, the degrees of freedom have increased from 28 to 58, the critical value has decreased from 2.048 to 2.009, and the standard error of the difference ( s X ¯ 1 − X ¯ 2 ) has decreased from 3.74 to 2.60. For the purpose of this discussion, let's focus on two other differences: (1) The larger sample size results in a larger value of the t-statistic (t = 3.49 vs. t = 2.42), and (2) the level of significance for Ni = 30 is p < .01 as opposed to p < .05 for the original study. So even though the difference between the two groups did not change, increasing the sample size changed the level of significance from p < .05 to p < .01. This highlights an important limitation of inferential statistics such as the t-test:
The statistical significance of inferential statistics is a direct function of sample
616
size: The larger the sample size, the more likely an inferential statistic will be statistically significant.
Table 10.9 Parking Lot Study Data, Sample Size Doubled
617
The Inappropriateness of “Highly Significant”
The difference in level of significance in the two versions of the parking lot study may lead to differences in how the two analyses are interpreted and communicated. Researchers sometimes distinguish between different levels of significance by referring to p < .05 findings as “significant” and p < .01 or p < .001 findings as “highly significant.” Furthermore, the phrase highly significant may be used to suggest that analyses at the .01 or .001 level of significance are somehow more noteworthy meaningful, or better than findings at the .05 level.
Table 10.10 Analysis, Modified Parking Lot Study Data
In Chapter 7, we mentioned that the phrase highly significant is inappropriate because statistical significance is a dichotomy (significant vs. nonsignificant) rather than a continuum. In hypothesis testing, we either reject or do not reject the null hypothesis; describing a result as “highly significant” incorrectly implies the null hypothesis wasn't just rejected but somehow was really rejected. Although different values of a statistic may have different levels of significance, one cannot say or imply that one value is “more” significant than another.
Furthermore, the analysis of the modified parking lot study data provides another reason why the phrase “highly significant” is improper: The level of significance of a statistical analysis is a direct function of sample size. Two researchers studying the same effect may find the same difference between groups; however, they may reach a different level of significance simply as a result of having different sample sizes. Therefore, when a statistic has a probability less than .01, this doesn't necessarily mean the effect or relationship is
618
greater or more pronounced than another statistic that has a probability less than .05. Although significance levels of .01 or .001 may provide precise descriptions of the result of a statistical analysis, they may not reflect the size or magnitude of the effect that is being studied.
At this point, it's understandable if you feel like you've received two opposing messages. On one hand, larger samples are desirable because they increase the probability of rejecting the null hypothesis. On the other hand, larger samples may be problematic if they lead researchers to misinterpret the results of their analyses. There are two points to remember in resolving this apparent contradiction. First, samples should in fact be as large as possible because the larger the sample, the greater the confidence in any decision made regarding the null hypothesis. That is, not only does a large sample increase the probability of rejecting the null hypothesis, but it also increases confidence about a decision to not reject the null. In other words, if an analysis has a high likelihood of rejecting the null hypothesis and the null hypothesis is not rejected, we may begin to conclude that the hypothesized effect may not exist.
Second, levels of significance such as .01 and .001 should be included to describe the results of an analysis as precisely as possible; it is up to researchers to avoid drawing inappropriate conclusions regarding the size of the effect being studied. Ultimately, to draw conclusions about the relative magnitude of a hypothesized effect, what is needed is a statistic that estimates the size of the effect without being affected by sample size. The next section introduces these statistics, which are known as measures of effect size.
619
Effect Size as Variance Accounted for
To create a statistic that provides information regarding the size of an effect yet is not influenced by sample size, let's return to the concept of variance first introduced in Chapter 4. Statistically, variance measures the amount of variability in a distribution of scores for a variable. Conceptually, variance represents differences in the phenomena studied by researchers. People differ in such things as behavior, personality, attitudes, and beliefs. The primary goals of a discipline such as psychology are to describe, understand, explain, and predict these differences. Put another way, the goal of psychology is to account for variance.
The notion of “accounting for variance” may be introduced with a simple example. Imagine for a moment there is a perfect relationship between people's heights and their weights, such that knowing a person's height enables us to predict with perfect accuracy how much the person weighs. For example, a perfect relationship between height and weight might imply that everyone who is 5′ 7″ tall weighs exactly the same amount, such as 155 pounds, and everyone who is 6′ 2″ tall weighs exactly 210 pounds. Put another way, this perfect relationship allows us to completely explain differences—that is to say, variance —in people's weights by their differences in height.
Given the presumed perfect relationship between the two variables, how much of the variance in people's weights can be explained or accounted for by differences in height? The answer is all, or 100%, of the variance. On the other hand, what if there was absolutely no relationship between the two variables? A complete lack of relationship implies that none, or 0%, of the variance in weight can be explained by differences in height. In between these two extremes of 0% and 100% lie relationships of varying strengths.
620
Measure of Effect Size for the Difference between Two Sample Means: R2
A measure of effect size is a statistic that measures the magnitude of the relationship between variables. In this section, we introduce several measures of effect size that are appropriate for the first research situation described in Chapter 9: testing the difference between two sample means. First, we discuss a measure of effect size that represents the percentage of variance in a variable that is accounted for by another variable, ranging from .00 to 1.00 (0% to 100%). Next, we will discuss a second measure of effect size that measures the effect in terms of the difference between the means of the two groups.
When testing the difference between two sample means, one measure of effect size that may be calculated is the r squared (r2) statistic, defined as the percentage of variance in one variable that is accounted for by another variable. The formula for r2 is presented below:
(10-3) r 2 = t 2 t 2 + d f
where t is the calculated value of the t-test for two sample means and df is the degrees of freedom for two sample means. Using Table 10.1, lets calculate r2 for the original Chapter 9 parking lot study: r 2 = t 2 t 2 + d f = ( 2.42 ) 2 ( 2.42 ) 2 + 28 = 5.86 33.86 = .17
621
Interpreting Measures of Effect Size
How can the r2 value of .17 for the parking lot study be interpreted? Because this measure of effect size is the percentage of the variance in one variable accounted for by another variable, we could conclude the following:
Seventeen percent of the variance in drivers' departure times may be accounted for by whether or not they are intruded upon by another driver.
But what does it mean to account for 17% of the variance in a variable? Jacob Cohen, a psychologist and statistician at New York University, devoted much of his 30-year career to developing estimates of statistical power for different types of research studies. In 1997, his important contributions to this area of research were honored by the American Psychological Association, which presented him with a lifetime achievement award.
Categorizing effect sizes as “small,” “medium,” and “large,” Cohen (1988) provided the following guidelines for the r2 statistic for the difference between two sample means:
A “small” effect produces an r2 of .01.
A “medium” effect produces an r2 of .06.
A “large” effect produces an r2 of .15 or greater.
Using Cohen's guidelines, the r2 of .17 for the parking lot study would be considered an effect of large magnitude. It is important to distinguish words such as large, which describe the magnitude of the effect, from words such as meaningful and important, which are subjective value judgments.
You may be surprised and perhaps discouraged by these guidelines regarding effect size. For example, even an effect considered large (r2 = .17) fails to account for the wide majority (1 – .17 = 83%) of the variance in a variable. The relatively small amount of accounted for variance is a reflection of the complexity of phenomena studied by social scientists. People vary in such things as intelligence, personality, and cognition for a wide variety of reasons; it is not surprising that any one variable can explain only a small amount of these differences.
622
Presenting Measures of Effect Size
One purpose of measures of effect size is to supplement the results of inferential statistical analyses such as the t-test. Therefore, these measures are typically provided at the end of the reporting of an inferential statistic. For example, for the Chapter 9 parking lot study, we could write the following:
The mean departure time for the 15 drivers in the Intruder group (M = 40.73 s) is significantly greater than the mean departure time for the 15 drivers in the No intruder group (M = 31.67 s), t(28) = 2.42, p < .05, r2 = .17. The difference between the two groups of drivers, using Cohen's (1988) guidelines, comprises a large effect.
Many of the statistical techniques discussed in the remainder of this book will illustrate how to report and communicate different measures of effect size.
623
Reasons for Calculating Measures of Effect Size
There are a number of reasons for calculating and reporting measures of effect size. First, measures of effect size are not affected by sample size. To illustrate this point, we return to the modified parking lot study data in Tables 10.9 and 10.10, where we found that increasing the sample size from Ni = 15 to Ni = 30 changed the level of significance from p < .05 to p < .01 even though the difference between the means of the two groups were the same. Below we calculate r2 for the modified (Ni = 30) data: r 2 = t 2 t 2 + d f = ( 3.49 ) 2 ( 3.49 ) 2 + 58 = 12.18 70.18 = .17
Comparing this value for r2 with the one calculated earlier for Ni = 15, we see that the two r2 values are identical (r2 = .17). As opposed to inferential statistics such as the t-test, measures of effect size are not affected by sample size.
A second reason for calculating measures of effect size such as r2 is that, because they are measured on the same .00 to 1.00 scale and are unaffected by sample size, they are comparable across different samples and different studies. Because inferential statistics such as the t-test are affected by sample size, the results of analyses conducted in two different studies cannot be compared to each other, even when they are from the same population. For example, the same t-value of 2.02 may be statistically significant in a study involving 250 participants but nonsignificant in a sample of 25.
A third reason for using measures of effect size is that they aid in the interpretation of statistically significant findings. It is possible for an analysis to be statistically significant yet have a very small value for a measure of effect size. For example, we may read the following: t(158) = 2.76, p < .01, r2 = .05. Rather than focusing solely on the level of significance (p < .01), the small value for r2 informs us that only a small percentage of the variance in the variable is accounted for. The relationship between an inferential statistic and its corresponding measure of effect size is sometimes referred to as the difference between “statistical significance” and “practical significance.” Whereas statistical significance focuses on the decision to reject the null hypothesis, practical significance addresses the degree to which the analysis is able to explain or account for the variance in the phenomenon being examined.
Finally, measures of effect size not only help interpret statistically significant findings but also assist in the understanding of analyses in which the null hypothesis is not rejected. For example, how might we interpret the following analysis: t(18) = 2.05, p > .05, r2 = .19? Although the null hypothesis was not rejected (p > .05), the relatively large r2 value of .19 suggests the effect may in fact be large in nature. Just as a statistically significant finding may have a small effect size, a nonsignificant finding may have a relatively large effect size.
624
The large value for r2 in this example suggests the researcher may wish to reexamine the effect, perhaps collecting data from a larger sample to increase the likelihood of rejecting the null hypothesis.
625
A Second Measure of Effect Size for the Difference between Two Sample Means: Cohen's d
As a measure of effect size, the r2 statistic represents the percentage (%) of variance in a variable that is accounted for by another variable. As we will see in later chapters of this book, the r2 statistic is particularly useful as it can be used to measure the size of an effect in a variety of research situations. However, a variety of different measures of effect size may be used for a specific research situation. In this section, we introduce a measure of effect size that measures an effect in terms of the difference between the means of two groups.
The Cohen's d statistic is an estimate of the magnitude of the difference between the means of two groups measured in standard deviation units. In other words, it measures the size of the treatment effect in terms of the number of standard deviations the means of the two groups differ from each other. As such, the Cohen's d is similar to the z-score, introduced in Chapter 5, which measures a score in standard deviation units (the number of standard deviations a score is different from the mean).
There exist a number of different formulas for the Cohen's d; these formulas differ depending on whether the sample sizes or standard deviations of the two groups are equal or unequal. Formula 10-4 presents a formula for the Cohen's d that may be used when the sample sizes of the two groups are equal (Rosenthal & Rosnow, 1991):
(10-4) d = 2 t d f
where t is the calculated value of the t-test for two sample means and df is the degrees of freedom for two sample means.
Below we calculate Cohen's d for the original Chapter 9 parking lot study (Table 10.1): d = 2 t d f = 2 ( 2.42 ) 28 = 4.84 5.29 = .91
The Cohen's d value of .91 for the parking lot study may be interpreted as the difference between the means of the two groups in standard deviations. More specifically, it indicates that the mean departure times of the Intruder and No intruder groups are .91 standard deviations different from each other. (Note that if t is a negative number, the calculated value of Cohen's d will also be negative; however, it is always reported as a positive value because it represents the number of standard deviations the groups differ from each other.)
As with the r2 statistic, Cohen (1988) provided guidelines for interpreting values of the Cohen's d:
626
A “small” effect produces a Cohen's d of .20.
A “medium” effect produces a Cohen's d of .50.
A “large” effect produces a Cohen's d of .80 or greater.
Using these guidelines, the Cohen's d of .91 for the parking lot study would be considered an effect of large magnitude. Note that this is the same conclusion reached from the r2
value of .17 calculated from the same data.
627
Learning Check 4: Reviewing what you've Learned So Far
1. Review questions a. Why is it inappropriate to use terms such as highly significant to describe analyses significant at the p < .01 or .001 level? b. Why is it useful to include a measure of effect size when reporting the results of a statistical analysis such as the t-test? c. Conceptually, what does it mean for one variable to account for the variance of another variable? d. What are Cohen's (1988) guidelines for interpreting different effect sizes? e. Why are measures of effect size comparable across different research studies? e. What is the difference between statistical significance and practical significance? f. A researcher conducts an analysis in which the null hypothesis is not rejected, yet a large measure of effect size is calculated. What advice might you give the researcher, and why?
g. What is the difference between the r2 statistic and Cohen's d measure of effect size?
2. For each of the situations below, calculate the r2 and Cohen's d measures of effect size and specify whether each is a “large,” “medium,” or “small” effect as outlined in the chapter.
a. df = 28; t = 2.08
b. df = 10; t = 1.47
c. df = 80; t = 2.14
3. Below are sets of data for two studies. Note that the data for the second study double the data from the first study.
Original data: Boy: 6, 7, 4, 5, 9, 6
Girl: 3, 5, 1, 7, 4, 6
Doubled data: Boy: 6, 7, 4, 5, 9, 6, 6, 7, 4, 5, 9, 6
Girl: 3, 5, 1, 7, 4, 6, 3, 5, 1, 7, 4, 6
Conduct the following analyses for the two sets of data: a. For each group, calculate the sample size (Ni), mean ( X ¯ i ), and standard deviation (si). b. State the null and alternative hypothesis (H0 and H1).
c. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. For α = .05 (two-tailed), identify the critical values and state a decision rule. 3. Calculate a value for the t-test for independent means. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
6. Calculate r2 and Cohen's d and describe size of effect (small, medium, large). d. Compare the results of the analysis of the original and doubled data in terms of level of
significance and measure of effect size.
628
10.6 Looking Ahead
This chapter began by discussing errors that may occur in making the decision to reject the null hypothesis. Next, methods used by researchers to minimize the occurrence of these errors were described and evaluated. For conceptual and pragmatic reasons, completely eliminating errors in hypothesis testing is neither possible nor desirable. Instead, it is important to understand the consequences of making these errors. Both the research process in general and hypothesis testing in particular are influenced by a researcher's beliefs, motivations, and skills. Much of the concerns about hypothesis testing are not directed at the statistical procedures themselves but rather how these procedures are used and perhaps misused by researchers. Consequently, it is important to learn to draw appropriate conclusions from statistical analyses. This becomes particularly true when researchers conduct studies requiring relatively complex statistical analyses. The next chapter, for example, discusses a statistical procedure used to compare differences between the means of three or more groups that comprise one independent variable.
629
10.7 Summary
Within hypothesis testing, the decision to reject or not reject the null hypothesis is based on probability rather than certainty. Consequently, the possibility remains that whatever decision is made about the null hypothesis may either be correct or incorrect. There are two types of errors that can be committed in making the decision to reject the null hypothesis: A Type I error occurs when the null hypothesis is true but the decision is made to reject the null hypothesis; a Type II error occurs when the decision is made to not reject the null hypothesis even though the alternative hypothesis is true.
In terms of Type I error, the probability of making this error is equal to alpha (α), typically .05; the probability of not making a Type I error is equal to 1 – α, typically .95. Type I error is a concern because the communication of incorrect information negatively affects others' thinking, time, energy, and resources. Type I errors occur as the result of random, chance factors, and the probability of making this error may be reduced by making the null hypothesis more difficult to reject (e.g., setting alpha to .01 rather than .05). However, decreasing the likelihood of making a Type I error increases the likelihood of making a Type II error. Ultimately, researchers must allow for the possibility of making this error because the only way it may be eliminated is to never reject the null hypothesis regardless of the evidence.
In terms of Type II error, the probability of committing a Type II error is represented by beta (β), which is difficult to calculate because of the imprecise nature of the alternative hypothesis. Type II errors are a concern because they may result in the non-communication of correct information that is potentially useful to others. Type II errors occur as the result of how researchers conduct their studies, and the probability of not making a Type II error is what is known as statistical power. Reducing Type II error, thereby increasing the statistical power of an analysis, involves increasing the likelihood of rejecting the null hypothesis. Factors such as the sample size, alpha, directionality of the alternative hypothesis, between-group variability, and within-group variability may be altered to increase statistical power.
One limitation of inferential statistics such as the t-test is that the statistical significance of these statistics is a function of sample size. To draw conclusions about the relative magnitude of a hypothesized effect, researchers can calculate a measure of effect size, which is a statistic that measures the magnitude of the relationship between variables. Two measures of effect size related to the difference between two sample means are r2, which is the percentage of variance in one variable that is accounted for by another variable, and Cohen's d, which is an estimate of the magnitude of the difference between the means of two groups measured in standard deviation units. Some reasons for calculating measures of effect size is that they are not affected by sample size, they are comparable across different
630
samples and different studies, and they aid in the interpretation of statistically significant and nonsignificant findings.
631
10.8 Important Terms
Type I error (p. 382) Type II error (p. 385) statistical power (p. 386) between-group variability (p. 396) within-group variability (p. 397) error variance (p. 398) measure of effect size (p. 405) r squared (r2) (p. 405) Cohen's d (p. 407)
632
10.9 Formulas Introduced in this Chapter
Probability of Type I Error
(10-1) p ( Type I error ) = α
Probability of Type II error
(10-2) p ( Type II error ) = β
r Squared (r2)
(10-3) r 2 = t 2 t 2 + d f
Cohen's r
(10-4) d = 2 t d f
633
10.10 Exercises
1.For each of the following hypotheses, create a chart that shows the four possible outcomes that could occur as a result of hypothesis testing. There should be two correct decisions and two incorrect decisions. Label the type of error for each of the incorrect decisions.
a. A new reading program in the elementary schools is hypothesized to improve reading comprehension.
b. People who take at least two vacations each year are happier, on average, than people who do not take at least two vacations per year.
c. A college student cramming for finals hypothesizes that caffeine improves his ability to concentrate.
2.If we reject the null hypothesis and we were wrong to do so, what type of error have we made? Give an example of a research situation in which this could occur. 3.If we fail to reject the null hypothesis but should have rejected it, what type of error have we made? Provide an example of a research situation in which this could occur. 4.For each of the situations below, describe what would be considered a Type I and Type II error.
a. A consumer group wants to see if people can tell whether they are drinking tap water or bottled water.
b. A college instructor wishes to see whether his students prefer to work on assignments individually or in groups.
c. A teacher evaluating a program designed to improve math skills gives her students a test before and after the program.
d. A college student conducts searches on two different Internet search engines to see if they differ in terms of the number of relevant results provided.
5.For each of the situations below, describe what would be considered a Type I and Type II error.
a. A political scientist compares older and younger adults' attitudes toward the death penalty.
b. A movie critic wants to know if people prefer to see movies in 3-D or stereoscopic (not 3-D).
c. An owner of a grocery store tests shoppers' ability to tell the difference between name brand cereal and generic cereal.
d. A teacher wants to know if boys and girls differ in their career goals. 6.
Below are two hypothetical situations:
Situation 1: Ni = 30, α = .05 (two-tailed test)
634
Situation 2: Ni = 30, α = .01 (two-tailed test) a. Find the critical values for each of the two situations. b. In which situation is there less of a chance of making a Type I error? Why? c. What is the effect of changing alpha from .05 to .01 on the probability of
making a Type II error? 7.
Below are two hypothetical situations:
Study A: N1 = 20, X ¯ 1 = 14.50 , s1 = 2.50
N2 = 20, X ¯ 2 = 12.75 , s2 = 3.25
Study B: N1 = 65, X ¯ 1 = 14.50 , s1 = 2.50
N2 = 65, X ¯ 2 = 12.75 , s2 = 3.25 a. Calculate the standard error of the difference ( s X ¯ 1 − X ¯ 2 ) and the t-test
for independent means for each study. b. How does increasing the sample size from Ni = 20 to Ni = 65 affect the
calculated values of the standard error of the difference and the t-statistic? c. How does increasing the sample size affect the probability of making a Type II
error? 8.Imagine two studies examining the same research question find the same means and standard deviations for two groups: X ¯ 1 = 10.00 , s1 = 3.00, and X ¯ 2 = 12.00 , s2 = 2.00. However, the two studies differ in sample sizes: Ni = 9 in the first study and Ni = 16 in the second study.
a. Calculate the standard error of the difference ( s X ¯ 1 − Z ¯ 2 ) and the t-test for independent means for each study.
b. How does increasing the sample size from Ni = 9 to Ni = 16 affect the calculated values of the standard error of the difference and the t-statistic, as well as the probability of making a Type II error?
9.Two researchers separately testing the same research hypothesis end up with the same sample size (Ni = 12) and standard deviations for the two groups (s1 = .50 and s2 = .75). However, the experimental manipulation is more effective for the second researcher, which results in a bigger difference between the means of the two groups: X ¯ 1 = 2.50 and X ¯ 2 = 2.15 (first researcher), as well as X ¯ 1 = 2.70 and X ¯ 2 = 1.95 (second researcher).
a. Calculate the standard error of the difference ( s X ¯ 1 − X ¯ 2 ) and the t-test for independent means for each study.
b. How does the increased between-group variability in the second study affect the calculated value of the t-statistic, as well as the probability of making a Type II error?
635
10.
A study compares the effectiveness of a new (Drug A) and an old medicine (Drug B) and finds no difference in effectiveness: t(68) = 1.34, p > .05. The researchers run the study a second time: Both studies are based on the same sample size (Ni= 35), and somehow both studies end up with the same standard deviations for the two groups. However, in the second study, the size of the dosage of Drug A was increased, which increased the difference between the two groups (between-group variability):
Drug A: X ¯ 1 = 21.00 ,
Drug B: X ¯ 1 = 10.50 ,
s X ¯ 1 − X ¯ 2 = 3.60 a. Calculate the t-test for independent means. b. What is the effect of increasing between-group variability on the likelihood of
rejecting the null hypothesis? c. What is the effect of increasing between-group variability on the likelihood of
making a Type II error? 11.Two research studies compare two groups on the same dependent variable. Both studies are based on the same sample size (Ni = 17), and somehow both studies end up with the same means for the two groups: X ¯ 1 = 77.62 and X ¯ 2 = 68.05 . However, in collecting her data, the second researcher provides clearer instructions to the research participants; as a result, the scores in the second study are less influenced by random, chance factors. The standard deviations for the two groups are s1 = 11.23 and s2 = 10.86 for the first study and s1 = 9.47 and s2 = 8.59 for the second study.
a. Calculate the standard error of the difference ( s X ¯ 1 − X ¯ 2 ) and the t-test for independent means for each study.
b. How does the decreased within-group variability in the second study affect the calculated value of the t-statistic, as well as the probability of making a Type II error?
12.
A researcher wishes to test a drug that claims to improve memory. The researcher sets up an experiment where a control group receives a placebo (i.e., a sugar pill) and an experimental group receives the memory pill. The following data are the number of sentences each participant correctly recalls after listening to a two-paragraph story. Determine if there is a significant difference between the two groups.
Placebo: 10, 3, 4, 6, 9, 1, 7
Experimental: 12, 5, 6, 10, 4, 7, 8 a. For each group, calculate the sample size (Ni), mean ( X ¯ i ), and standard
636
deviation (si). b. State the null and alternative hypothesis (H0 and H1). c. Make a decision about the null hypothesis.
1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values, and state a decision rule. 3. Calculate a statistic: the t-test for independent means. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
d. Draw a conclusion from the analysis. 13.
The researcher in Exercise 12 notices the large amount of variability within each of the two groups. Investigating, he finds the research assistants counting the number of correctly remembered sentences differed in what they counted as a correct sentence. Some experimenters counted them as correct if the participant remembered the general idea of the sentence, whereas others only counted it as correct if every single word was recalled in the correct order. To decrease this source of error, he handed out instructions on what should be counted as “correct.” Using the following new data and the steps of hypothesis testing, determine if there is a significant difference between the groups.
Placebo: 7, 4, 5, 6, 5, 7, 6
Experimental: 9, 6, 7, 8, 5, 8, 9 a. For each group, calculate the sample size (Ni), mean ( X ¯ i ), and standard
deviation (si). b. State the null and alternative hypothesis (H0 and H1).
c. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values, and state a decision rule. 3. Calculate a statistic: the t-test for independent means. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
d. Draw a conclusion from the analysis. e. What type of error (Type I or Type II) was committed the first time the study
was conducted? 14.For each of the situations below, calculate both the r2 and Cohen's d measure of effect size and specify whether each is a “small,” “medium,” or “large” effect as outlined in the chapter.
a. df = 18, t = 2.23 b. df = 30, t = 1.90
637
c. df = 20, t = 2.09 d. df = 80, t = 2.09 e. df = 74, t = 3.49
15.For each of the situations below, calculate both the r2 and the Cohen's d measure of effect size and specify whether each is a “small,” “medium,” or “large” effect as outlined in the chapter.
a. df = 12, t = 2.07 b. df = 24, t = 1.90 c. df = 24, t = 2.51 d. df = 62, t = 2.96 e. df = 32, t = 1.31
15.For each of the situations below, calculate both the r2 and the Cohen's d measure of effect size and specify whether each is a “small,” “medium,” or “large” effect as outlined in the chapter.
e. df = 12, t = .98 16.
A researcher is investigating the birth-order theory that middle siblings are greater risk takers than are first-born siblings. The following statistics are obtained for the study as a measure of risk taking:
Middle siblings: N 1 = 50 , X ¯ = 23.75 , s 1 = 3.40
First born: N 2 = 50 , X ¯ = 21.80 , s 2 = 3.56
a. Calculate the standard error of the difference ( s X ¯ 1 − X ¯ 2 ) and the t-test for independent means.
b. Is this difference statistically significant? At what level of significance? c. Calculate r2 and Cohen's d and specify whether it is a “small,” “medium,” or
“large” effect. d. Report the result of this analysis in APA format.
17.
Below are sets of data for two studies: The data for the second study double the data from the first study.
Original data: Group A: 20, 14, 18, 12, 24, 14, 22
Group B: 15, 18, 16, 13, 10, 17, 8
Doubled data: Group A: 20, 14, 18, 12, 24, 14, 22, 20, 14, 18, 12, 24, 14, 22
638
Group B: 15, 18, 16, 13, 10, 17, 8, 15, 18, 16, 13, 10, 17, 8
Conduct the following analyses for the two sets of data. a. For each group, calculate the sample size (Ni), mean ( X ¯ i ), and standard
deviation (si). b. State the null and alternative hypothesis (H0 and H1).
c. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values, and state a decision rule. 3. Calculate a value for the t-test for independent means. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate r2 and Cohen's d and describe size of effect (small, medium,
large). d. Compare the results of the analysis of the original and doubled data in terms of
level of significance and measure of effect size.
639
Answers to Learning Checks
Learning Check 1
2. a. His customers think they are drinking alcoholic beer when they are actually
drinking nonalcoholic beer. b. The dentist decides the two types of anesthesia differ in their ability to alleviate
pain when they actually don't. c. The consumer protection agency concludes computer users are less likely to get
infected if they pay for virus protection software when in fact they are not.
Learning Check 2
2. a. The snowboarder decides the two different kinds of wax last the same amount
of time when in fact they don't. b. The high school counselor decides that whether or not one pays for test
preparation services isn't related to the likelihood of getting into college when in fact it is.
c. The driver decides the type of gasoline doesn't affect her car's performance when in fact it does.
3. a.
Do not reject H0 Reject H0
H0 is true
Correct decision: Conclude the two types of drivers live the same distance from work and they really do.
Type I error: Conclude the two types of drivers do not live the same distance from work but they really do.
H1 is true
Type II error: Conclude the two types of drivers live the same distance from work and they really do not.
Correct decision: Conclude the two types of drivers do not live the same distance from work and they really do not.
b.
Do not reject H0 Reject H0 Type I error: Conclude the
640
H0 is true
two types of video games are not related to automobile accidents and they really are not.
two types of video games are related to automobile accidents and they really are not.
H1 is true
Type II error: Conclude the two types of video games are not related to automobile accidents and they really are.
Correct decision: Conclude the two types of video games are related to automobile accidents and they really are.
Learning Check 3
2. a.
First study: s X ¯ 1 − X ¯ 2 = 1.12 ; t = 1.79
Second study: s X ¯ 1 − X ¯ 2 = .74 ; t = 2.70
b. Increasing the sample size increases the value of the t-statistic and reduces the probability of making a Type II error.
3. a.
First study: s X ¯ 1 − X ¯ 2 = 1.63 ; t = −1.23
Second study: s X ¯ 1 − X ¯ 2 = 1.63 ; t = −2.45
b. Increasing between-group variability increases the value of the t-statistic and reduces the probability of making a Type II error.
Learning Check 4
2. a. r2 = .13; Cohen's d = .79; medium effect b. r2 = .18; Cohen's d = .93; large effect c. r2 = .05; Cohen's d = .48; small effect
3. Original data a.
Boy: N1 = 6, X ¯ 1 = 6.17 , s1 = 1.72
Girl: N2 = 6, X ¯ 2 = 4.33 , s2 = 2.16
641
b. H0: μBoy = μGirl; H1: μBoy ≠ μGirl c.
1. df = 10 2. If t < −2.228 or > 2.228, reject H0; otherwise, do not reject H0 3. s X ¯ 1 − X ¯ 2 = 1.13 ; t = 1.63 4. t = 1.63 < 2.228 ∴ do not reject H0 (p > .05) 5. Not applicable (H0 not rejected) 6. r2 = .21, d = 1.03; large effect
Doubled data a.
Boy: N1 = 12, X ¯ 1 = 6.17 , s1 = 1.64
Girl: N2 = 12, X ¯ 2 = 4.33 , s2 = 2.06
b. H0: μBoy = μGirl; H1: μBoy ≠ μGirl c.
1. df = 22 2. If t < −2.074 or > 2.074, reject H0; otherwise, do not reject H0 3. s X ¯ 1 − X ¯ 2 = .75 ; t = 2.45 4. t = 2.45 > 2.074 ∴ reject H0 (p < .05) 5. t = 2.45 < 2.819 ∴ p < .05 (but not < .01) 6. r2 = .21, d = 1.03; large effect
d. The null hypothesis was not rejected for the original data (t(10) = 1.63, p > .05) but was rejected for the doubled data (t(22) = 2.45, p < .05). However, the measures of effect size for the two sets of data were the same (r2 = .15, d = 1.03), showing that effect size is not influenced by sample size.
642
Answers to Odd-Numbered Exercises
1. a. Decision re: null hypothesis
Do not reject H0 Reject H0
H0 is true
Correct decision: Conclude the program does not improve comprehension and it really does not
Type I error: Conclude the program improves comprehension but it really does not
H1 is true
Type II error: Conclude the program does not improve comprehension but it really does
Correct decision: Conclude the program improves comprehension and it really does
b. Decision re: null hypothesis
Do not reject H0 Reject H0
H0 is true
Correct decision: Conclude vacationing people are not happier than nonvacationing people and they really are not
Type I error: Conclude vacationing people are happier than nonvacationing people but they really are not
H1 is true
Type II error: Conclude vacationing people are not happier than nonvacationing people but they really are
Correct decision: Conclude vacationing people are happier than nonvacationing people and they really are
c. Decision re: null hypothesis
Do not reject H0 Reject H0 H0 is true
Correct decision: Conclude caffeine does not improve attention and it really does not
Type I error: Conclude caffeine improves attention but it really does not
H1 is true
Type II error: Conclude caffeine does not improve attention but it really does
Correct decision: Conclude caffeine improves attention and it really does
3. Type II error 5.
643
a. Type I error: deciding older and younger adults' attitudes differ when they actually don't.
Type II error: deciding older and younger adults' attitudes don't differ when they actually do.
b. Type I error: deciding people's movie-watching preferences differ when they actually don't.
Type II error: deciding people's movie-watching preferences don't differ when they actually do.
c. Type I error: deciding shoppers can tell the difference between the two types of cereals when they actually can't.
Type II error: deciding shoppers can't tell the difference between the two types of cereals when they actually can.
d. Type I error: deciding boys and girls differ in their career goals when they actually don't.
Type II error: deciding boys and girls don't differ in their career goals when they actually do.
7. a.
Study A: s X ¯ 1 − X ¯ 2 = .92 ; t = 1.90
Study B: s X ¯ 1 − X ¯ 2 = .51 ; t = 3.43
b. Increasing the sample size reduces the standard error of the difference and increases the value of the t-statistic.
c. Increasing the sample size reduces the probability of making a Type II error. 9.
a.
First study: s X ¯ 1 − X ¯ 2 = .26 ; t = 1.25
Second study: s X ¯ 1 − X ¯ 2 = .26 ; t = 2.88
b. Increasing between-group variability increases the value of the t-statistic and reduces the probability of making a Type II error.
11. a.
644
First study: s X ¯ 1 − X ¯ 2 = 3.79 ; t = 2.53
Second study: s X ¯ 1 − X ¯ 2 = 3.10 ; t = 3.09
b. Decreasing within-group variability increases the value of the t-statistic and reduces the probability of making a Type II error.
13.
a. Placebo: N1 = 7, X ¯ 1 = 5.71 , s1 = 1.11
Experimental: N2 = 7, X ¯ 2 = 7.43 , s2 = 1.51
b. H0: μPlacebo = μExperimental;
H1: μPlacebo ≠ μExperimental c.
1. df = 12 2. If t < −2.179 or > 2.179, reject H0; otherwise, do not reject H0 3. s X ¯ 1 − X ¯ 2 = .71 ; t = −2.42 4. t = −2.42 < −2.179 ∴ reject H0 (p < .05) 5. t = −2.42 > −3.055 ∴ p < .05 (but not < .01)
d. The number of sentences recalled by the Experimental group (M = 7.43) was significantly greater than the Placebo group (M = 5.71), t(12) = −2.42, p < .05.
e. The first study is most likely to have made a Type II error, failing to spot an effect that is there. By decreasing within-group variability, the probability of making this error was decreased.
15. a. r2 = .26; Cohen's d = 1.20; large effect b. r2 = .13; Cohen's d = .78; medium effect c. r2 = .21; Cohen's d = 1.02; large effect d. r2 = .12; Cohen's d = .75; medium effect e. r2 = .05; Cohen's d = .46; small effect
17. Original data
a. Group A: N1 = 7, X ¯ 1 = 17.71 , s1 = 4.54
Group B: N2 = 7, X ¯ 2 = 13.86 , s2 = 3.72 b. H0: μGroup A = μGroup B; H1: μGroup A ≠ μGroup B c.
1. df = 12 2. If t < −2.179 or > 2.179, reject H0; otherwise, do not reject H0 3. s X ¯ 1 − X ¯ 2 = 2.22 ; t = 1.73
645
4. t = 1.73 < 2.179 ∴ do not reject H0 (p > .05) 5. Not applicable (H0 not rejected) 6. r2 = .20, d = 1.00; large effect
Doubled data
a. Group A: N1 = 14, X ¯ 1 = 17.71 , s1 = 4.36
Group B: N2 = 14, X ¯ 2 = 13.86 , s2 = 3.57 b. H0: μGroup A = μGroup B; H1: μGroup A ≠ μGroup B c.
1. df = 26 2. If t < −2.056 or > 2.056, reject H0; otherwise, do not reject H0 3. s X ¯ 1 − X ¯ 2 ; t = 2.55 4. t = 2.55 > 2.056 ∴ reject H0 (p < .05) 5. t = 2.55 < 2.779 ∴ p < .05 (but not < .01) 6. r2 = .20, d = 1.00; large effect
d. The null hypothesis was not rejected for the original data (t(14) = 1.73, p > .05) but was rejected for the doubled data (t(26) = 2.55, p < .05). However, the measures of effect size for the two sets of data were the same (r2 = .20, d = 1.00), showing that effect size is not influenced by sample size.
646
Sharpen your skills with SAGE edge at edge.sagepub.com/tokunaga
SAGE edge for students provides a personalized approach to help you accomplish your coursework goals in an easy-to-use learning environment. Log on to access:
1. eFlashcards 2. Web Quizzes 3. Chapter Outlines 4. Learning Objectives 5. Media Links
647
Chapter 11 One-Way Analysis of Variance (ANOVA)
648
Chapter Outline 11.1 An Example From the Research: It's Your Move 11.2 Introduction to Analysis of Variance (ANOVA)
Review: testing the difference between the means of two samples Understanding between-group and within-group variability Relating between-group and within-group variability
Introduction to analysis of variance (ANOVA) and the F-ratio Characteristics of the F-ratio distribution
11.3 Inferential Statistics: One-Way Analysis of Variance (ANOVA) State the null and alternative hypotheses (H0 and H1) Make a decision about the null hypothesis
Calculate the degrees of freedom (df) Set alpha (α), identify the critical value, and state a decision rule Calculate a statistic: F-ratio for one-way ANOVA Make a decision whether to reject the null hypothesis Determine the level of significance
Calculate a measure of effect size (R2) Draw a conclusion from the analysis Relate the result of the analysis to the research hypothesis Summary
11.4 A Second Example: The Parking Lot Study Revisited State the null and alternative hypotheses (H0 and H1) Make a decision about the null hypothesis
Calculate the degrees of freedom (df) Set alpha (α), identify the critical value, and state a decision rule Calculate a statistic: F-ratio for one-way ANOVA Make a decision whether to reject the null hypothesis Determine the level of significance
Calculate a measure of effect size (R2) Draw a conclusion from the analysis The relationship between the t-test and the F-ratio
11.5 Analytical Comparisons Within the One-Way ANOVA Planned versus unplanned comparisons Between-group and within-group variance in analytical comparisons Conducting planned comparisons
State the null and alternative hypotheses (H0 and H1) Make a decision about the null hypothesis Draw a conclusion from the analysis Relate the result of the analysis to the research hypothesis
Conducting unplanned comparisons Concerns regarding unplanned comparisons Methods of controlling familywise error Controlling familywise error: the Tukey test
11.6 Looking Ahead 11.7 Summary 11.8 Important Terms 11.9 Formulas Introduced in This Chapter 11.10 Using SPSS 11.11 Exercises
649
This chapter returns to the discussion of a statistical procedure designed to test research hypotheses. The statistical procedure we will discuss, the one-way analysis of variance (ANOVA), is used to compare differences between the means of three or more groups that comprise one independent variable. Conceptually, the one-way ANOVA is a simple extension of the t-test (Chapter 9) that tests the difference between the means of two groups. However, having more than two groups makes the calculation and interpretation of the analysis somewhat more complicated. In discussing the one-way ANOVA, this chapter incorporates many of the issues presented thus far in this book, such as descriptive statistics, inferential statistics, errors in hypothesis testing, and measures of effect size.
650
11.1 An Example from the Research: It's your Move
What are effective ways of learning the fundamental aspects of a complicated new topic? A team of researchers at Erasmus University Rotterdam in the Netherlands, led by Anique de Bruin, focused their attention on the process of “principled understanding,” through which the learning of basic principles helps people develop strategies for solving problems within a specific domain (de Bruin, Rikers, & Schmidt, 2007). The researchers believed that learning principles allows the person to “apply the principle to diverse situations, and thereby revise and expand knowledge of the skill in which the principle is based” (p. 190).
The researchers examined the development of principled understanding using the example of learning how to play chess. They emphasized that learning the rules of chess is distinct from learning the principles of chess. “The basic rules cover the factual knowledge that is necessary to learn the principles of the domain … [whereas] the principles describe the procedures necessary to reach certain subgoals in the game … [in other words], the rules explain how to play a chess game, whereas the principles instruct how to win a chess game” (de Bruin et al., 2007, p. 190).
The main goal of the study, which we will refer to as the chess study, was to compare the effectiveness of three strategies used by beginning players to develop a principled understanding of chess. These strategies differed in the extent to which players had to explain their thinking to others. The first strategy, observation, consisted of having people watch chess games being played without explaining what they were looking at or why. The second strategy, prediction, also involved having people watch chess games. However, these people were asked to make verbal predictions regarding what they believed should be a player's next move. The third strategy, explanation, not only involved observing and predicting but also required the players to explain the reasoning behind their predictions. The researchers believed that the more people have to verbalize and explain their predictions, the more they have to assess and monitor their comprehension of the principles, which in turn leads to higher levels of principled understanding.
In the chess study, a sample of college students who had never played chess were randomly assigned to one of three conditions (Observation, Prediction, and Explanation). Given the complexity and length of chess games, rather than have students watch entire games being played, the researchers had students watch scenarios in a computer software program in which only three chess pieces were on the board: a white rook and king and the black king. In these scenarios, the pieces were placed in different positions on the board and students watched a researcher play against the computer. The researcher moved the white pieces with the goal of winning the game by checkmating the computer's black king. The students first watched a series of scenarios while employing one of the three assigned strategies (Observation, Prediction, or Explanation). Next, each student played five scenarios with
651
the goal of winning as many of the five games as possible.
In the published report of the findings, de Bruin et al. (2007) described the study's research hypothesis in the following way: “We hypothesized that participants in the [explanation] condition would … checkmate the King more often … than the other two conditions … [we also] hypothesized that the prediction condition would outperform the observation condition” (p. 192).
In our example, we will mirror the study's findings using a sample of 42 students equally divided into the Observation, Prediction, and Explanation conditions. The number of checkmates in five games for each student is listed in Table 11.1(a). Frequency distribution tables, constructed for each group in Table 11.1(b), suggest that students in the Explanation condition had a greater number of checkmates than did students in the Observation and Prediction conditions. For example, 50% of the students in the Explanation condition had four or five checkmates; on the other hand, the majority of students in the Observation (64%) and Prediction (79%) conditions had only zero or one checkmates.
The next step is to calculate descriptive statistics (mean [ X ¯ i ] and standard deviation [si]) of the dependent variable for each group (Table 11.2). Figure 11.1 presents a bar graph comparing the group means, with the T-shaped lines in each bar representing the distance of one standard error of the mean above and below the mean (±1 s X ¯ ). The standard errors have been calculated using Formula 7–5 in Chapter 7.
Looking at the descriptive statistics and bar graph, we observe that the mean number of checkmates for the Explanation group (M = 3.00) is the highest of the three groups. Contrary to the research hypothesis, however, the mean of the Prediction group (M = .86) is less than that of the Observation group (M = 1.36). Although these differences indicate initial support for the research hypothesis that the Explanation group should show the highest level of principled understanding, a first glance at the data does not appear to support the other hypothesis that principled understanding is higher when people make verbal predictions while watching chess games rather than simply watching. However, it is important to keep in mind that any conclusions about the research hypotheses are premature; later in this chapter, we will discuss statistical analyses designed to test differences between groups.
The next step in analyzing the data is to calculate an inferential statistic that tests the differences between the means of the groups. The following section introduces both a family of statistical procedures used to test differences between group means and a statistic specifically used to compare the means of three or more groups that make up one independent variable.
Table 11.1 The Number of Checkmates for Students in the Observation,
652
Prediction, and Explanation Conditions of the Chess Study Table 11.1 The Number of Checkmates for Students in the Observation,
Prediction, and Explanation Conditions of the Chess Study
(a) Raw Data
Observation Prediction Explanation
Student # Checkmates Student # Checkmates Student # Checkmates
1 0 1 3 1 1
2 5 2 0 2 5
3 0 3 1 3 2
4 0 4 0 4 1
5 2 5 1 5 5
6 1 6 2 6 4
7 5 7 1 7 1
8 1 8 0 8 5
9 2 9 2 9 5
10 0 10 0 10 1
11 0 11 1 11 5
12 1 12 0 12 4
13 0 13 0 13 1
14 2 14 1 14 2 Table 11.1 The Number of Checkmates for Students in the Observation, Prediction,
and Explanation Conditions of the Chess Study
(b) Frequency Distribution Tables
Observation Prediction Explanation
# Checkmates
f % # Checkmates
f % # Checkmates
f %
5 2 14% 5 0 0% 5 5 36%
4 0 0% 4 0 0% 4 2 14%
3 0 0% 3 1 7% 3 0 0%
2 3 21% 2 2 14% 2 2 14%
1 3 21% 1 5 36% 1 5 36%
0 6 43% 0 6 43% 0 0 0%
Total 14 100% Total 14 100% Total 14 100%
653
11.2 Introduction to Analysis of Variance (ANOVA)
As we mentioned at the start of this chapter, a research situation containing an independent variable that has three or more groups is an extension of the situations discussed in Chapter 9, which involved comparing the means of two groups. Therefore, we will introduce the statistical procedure that will be used in this chapter by reexamining the t-test for independent means.
654
Review: Testing the Difference between the Means of Two Samples
In Chapter 9, the difference between the means of two samples was evaluated using the t- test for independent means: t = X ¯ 1 − X ¯ 2 s X ¯ 1 − X ¯ 2
Table 11.2 Descriptive Statistics of Number of Checkmates for the Three Conditions in the Chess Study
Table 11.2 Descriptive Statistics of Number of Checkmates for the Three Conditions in the Chess Study
Mean ( X ¯ i ) Standard Deviation (si)
X ¯ i = ∑ X N s i = ∑ X − X ¯ 2 N − 1
a. Observation
X ¯ 1 = 0 + 5 + 0 + ⋯ + 1 + 0 + 2 14 = 19 14 = 1.36
s 1 = ( 0 − 1.36 ) 2 + ⋯ + ( 2 − 1.36 ) 2 14 − 1 = 3.02 = 1.74
b. Prediction
X ¯ 1 = 3 + 0 + 1 + ⋯ + 0 + 0 + 1 14 = 12 14 = .86
s 2 = ( 3 − .86 ) 2 + ⋯ + ( 1 − .86 ) 2 14 − 1 = .90 = .95
c. Explanation
X ¯ 3 = 1 + 5 + 2 + ⋯ + 4 + 1 + 2 14 = 42 14 = 3.00
s 3 = ( 1 − 3.00 ) 2 + ⋯ + ( 2 − 3.00 ) 2 14 − 1 = 3.38 = 1.84
Figure 11.1 Bar Graph of Number of Checkmates for the Three Conditions in the Chess Study
655
As we learned in Chapter 10, the calculated value of a t-statistic is a function of two factors that correspond to the numerator and denominator of this formula. The numerator, the difference between the means of the two groups ( X ¯ 1 − X ¯ 2 ), was referred to as between-group variability. Between-group variability was defined as differences among the means of the different groups that comprise an independent variable. The denominator, the standard error of the difference ( s X ¯ 1 − X ¯ 2 ), represents within-group variability, defined as the variability of the dependent variable within the groups comprising an independent variable.
We also learned in Chapter 10 that increasing between-group variability and/or decreasing within-group variability results in a larger calculated value of the t-statistic, which in turn increases the likelihood of rejecting the null hypothesis. In this chapter, we will turn our attention to the relationship between between-group and within-group variability.
Understanding Between-Group and Within-Group Variability
The numerator of the formula for the t-test, the difference between the means of the two groups ( X ¯ 1 − X ¯ 2 ), represents between-group variability. Why might groups in a study differ from each other? For example, imagine we have developed a program designed to increase students' reading skills. After putting one group of students through the program, we give them a test and compare their test scores with the scores of a group that did not go through the program.
Assuming the test scores of the two groups are not the same, how can we account for this difference? The difference between groups could be due to the hypothesized effect being tested (in this case, a student's participation or nonparticipation in the program). However, in addition to participation or nonparticipation in the program, some of this difference may also be due to random, chance factors, as well as other factors that have not been taken into account by the researcher. In Chapter 10, we referred to variability that cannot be explained or accounted for as “error.” Taking all this into consideration, we may state that between-group variability comprises two parts: the hypothesized “effect” and error.
The denominator of the formula for the t-test, within-group variability, is represented by the standard error of the difference ( X ¯ 1 − X ¯ 2 ). Formula 9–1 (see below) illustrates the formula for the standard error of the difference: s X ¯ 1 − X ¯ 2 = S 1 2 N 1 + S 2 2 N 2
We observe from this formula that one component of X ¯ 1 − X ¯ 2 the standard deviation of each group (s); the standard deviation measures the variability of scores within each group. Based on what we learned in Chapter 10, because all members of a group are assumed to be the same, there should theoretically be no variability within each group.
656
Consequently, any observed differences within a group must be described as unaccounted for. In other words, within-group variability may be defined as error.
To summarize, testing the difference between group means involves relating two types of variability to each other: between-group variability and within-group variability. Between- group variability (differences between groups) consists of the hypothesized effect and error, whereas within-group variability (differences within groups) is error. A statistic such as the t-statistic is calculated by dividing between-group variability by within-group variability in the following way: t = X ¯ 1 − X ¯ 2 s X ¯ 1 − X ¯ 2 = between - group variabiliy within - group variabiliy = difference between groups difference within groups = effect + error error
Relating Between-Group and Within-Group Variability
The calculated value of a statistic such as the t-statistic is a function of the relationship of between-group and within-group variability in a set of data. In essence, the amount of between-group variability (the numerator of the formula) must be sufficiently greater than the amount of within-group variability (the denominator of the formula) to reject the null hypothesis and conclude that there is a statistically significant difference between group means. Stated another way, the hypothesized effect must be large enough to overcome the amount of error that exists in the data.
Figure 11.2 illustrates the relationship between between-group variability and within-group variability. Figure 11.2(a) demonstrates the effect of within-group variability on the amount of overlap between distributions. Looking at the two sets of distributions in Figure 11.2(a), the difference between groups (between-group variability) is the same in the two situations: ( X ¯ 1 − X ¯ 2 = 40 − 20 = 20 ). However, the distributions in the right half of Figure 11.2(a) are flatter and wider than the distributions on the left, indicating a greater amount of within-group variability. Comparing the two sets of distributions, we see that the greater the amount of within-group variability, the greater the overlap between the two distributions. Consequently, as within-group variability increases, the ability to conclude that groups differ from each other decreases.
Figure 11.2(b) illustrates the consequence of within-group variability on the ability to conclude that groups differ. The two sets of distributions both reflect situations where there is a large amount of within-group variability in a set of data. Comparing the pair of distributions on the left with the pair on the right, we see that to reduce the amount of overlap between distributions, the difference between groups must increase, in this case from X ¯ 1 − X ¯ 2 = 40 − 20 = 20 to X ¯ 1 − X ¯ 2 = 50 − 10 = 40 . This shows that the greater the amount of within-group variability in a set of data, the greater the between- group variability must be to determine that groups differ from each other. Put another way, the more error there is in a set of data, the greater the effect must be in order to be detected.
657
Figure 11.2 The Relationship between Between-Group and Within-Group Variability
658
Introduction to Analysis of Variance (ANOVA) and the F- Ratio
Statistical procedures that test differences between group means involve relating between- group variability and within-group variability. In Chapter 4, we measured the variability in a set of data using a statistic known as variance. The t-test is a member of a larger family of statistical procedures known as the analysis of variance. The analysis of variance (ANOVA) is a family of statistical procedures used to test differences between two or more group means. It is called the “analysis” of variance because it involves analyzing different types of variance, such as between-group and within-group variance. This chapter will discuss one particular type of analysis of variance, known as a one-way ANOVA. The one- way ANOVA is a statistical procedure used to test differences between the means of three or more groups that comprise a single independent variable.
Conceptually, the one-way ANOVA is a simple extension of the t-test, in that both statistical procedures involve dividing between-group variance by wi thin-group variance. However, rather than calculating a t-statistic, in the one-way ANOVA, we will calculate an F-ratio, defined as a statistic used to test differences between group means. The word ratio indicates that the value of the statistic is determined by dividing one quantity by another, in this case dividing between-group variance by within-group variance: F = between - group variance within - group variance
Characteristics of the F-Ratio Distribution
Figure 11.3 provides an illustration of the F-ratio distribution. Keep in mind that the F- ratio, like the t-statistic, cannot be expressed in a single distribution. This is because the shape of the distribution changes as a function of sample size and the number of groups that comprise an independent variable.
Unlike the other theoretical distributions we have examined thus far in this book, the F- ratio distribution is not a symmetrical bell-curve shape with both positive and negative values. Instead, it is positively skewed, starting with the value 0 at the left and extending to infinity (∞) on the right.
The F-ratio distribution contains only positive values because calculating F-ratios involves calculating variances, whose values can only be positive numbers. Here, for example, is the formula for the sample variance (s2), defined in Chapter 4 as the average squared deviation from the mean: s 2 = ∑ X − X ¯ 2 N − 1
The numerator of this formula is based on squared deviations of scores from the mean ( X
659
− X ¯ 2 ). Because squared deviations can only be positive numbers, the F-ratio can only be a positive number.
In terms of the modality of the F-ratio distribution, we can see by looking at Figure 11.3 that the mode of the F-ratio distribution is approximately equal to 1. To explain this, recall that a statistic such as the F-ratio may be characterized by the formula (effect + error)/error.
Figure 11.3 Example of the F-Ratio Distribution
If, as stated in the null hypothesis, the groups do not differ, this implies there is no (meaning zero) “effect.” Consequently, the expected value of the F-ratio is (0 + error)/error, which is error/error—any number divided by itself is equal to 1.
660
Learning Check 1: Reviewing what you've Learned So Far
1. Review questions a. Conceptually, what are between-group and within-group variability? What do they consist
of or represent? b. In testing the difference between group means, how are between-group and within-group
variability related to each other? What has to exist for the null hypothesis to be rejected? c. In testing the difference between group means, what is the implication of a large amount of
within-group variability (error) on the amount of between-group variability (effect) needed to reject the null hypothesis?
d. What does ANOVA stand for, and why? e. In testing the difference between means, why is the inferential statistic called the F-ratio? f. What is the shape of the theoretical distribution of F-ratios? Why is it this shape? g. What is the modality of the theoretical distribution of F-ratios?
661
11.3 Inferential Statistics: One-Way Analysis of Variance (ANOVA)
Having introduced the analysis of variance, we will now discuss how to test a research hypothesis regarding differences between group means by calculating and evaluating the F- ratio for the one-way ANOVA. The process of hypothesis testing in this situation consists of the same steps outlined in the earlier chapters of this book:
state the null and alternative hypotheses (H0 and H1); make a decision about the null hypothesis; draw a conclusion from the analysis; and relate the result of the analysis to the research hypothesis.
As we move through these four steps, differences between the one-way ANOVA and the t- test discussed in Chapter 9 will be identified and discussed.
662
State the Null and Alternative Hypotheses (H0 and H1)
In the parking lot study in Chapter 9, the null hypothesis used in testing the difference between the means of two samples was stated as H0: μIntruder = μNo intruder. But what is the null hypothesis in an analysis involving more than two groups? In the chess study, the null hypothesis may be stated as H0: μObservation = μPrediction = μExplanation, which implies that the mean number of checkmates for the three groups in the population is equal to each other. Another, more generic, way of stating the null hypothesis for the one-way ANOVA is the following: H 0 : all μ s are equal
Stating that all of the population means (μs) are equal is simply another way of stating that there are no differences among the means of the groups.
If the null hypothesis is stated as “all μs are equal,” does it therefore follow that the alternative hypothesis (H1) is “all μs are not equal”? The answer is no. Stating the alternative hypothesis this way, which in the chess study is the same as stating H1: μObservation ≠ μPrediction μExplanation, is actually incorrect. It is incorrect because it requires every mean to be different from every other mean in order for the null hypothesis to be rejected. In fact, all that's needed to reject the null hypothesis is for just one of the means to be different from the others, as this would imply that the null hypothesis (“all μs are equal”) is not supported by the data. Therefore, the correct form of the alternative hypothesis is as follows: H 1 : not all μ s are equal
To more clearly comprehend the difference between the incorrect “all μs are not equal” and the correct “not all μs are equal,” consider the following null hypothesis: “All people have the same birthday.” An incorrect alternative hypothesis for this null hypothesis would be, “All people do not have the same birthday.” This is incorrect because the only way to reject the null hypothesis would be if everyone has a different birthday Instead, the correct alternative hypothesis is “Not all people have the same birthday” because, in this case, all that's needed to reject the null hypothesis “All people have the same birthday” is to find two people whose birthdays are not the same.
663
Make a Decision about the Null Hypothesis
Once the null and alternative hypotheses have been stated, we move on to making the decision whether to reject the null hypothesis that all of the means are equal. As we have already observed, this decision is made using the following steps:
calculate the degrees of freedom (df); set alpha (α), identify the critical value, and state a decision rule; calculate a statistic: F-ratio for the one-way ANOVA; make a decision whether to reject the null hypothesis; determine the level of significance; and calculate a measure of effect size.
Except for the addition of the last step, these are the same steps used in our calculations in Chapters 7 and 9. As we learned in Chapter 10, measures of effect size are statistics used to estimate the magnitude of the hypothesized effect in the population. These measures provide information that is useful in interpreting the results of an inferential statistical procedure and will be included throughout the remainder of this book.
Calculate the Degrees of Freedom (df)
The first step in making the decision to reject the null hypothesis is to calculate the degrees of freedom in the set of data. Unlike the t-test, the F-ratio for a one-way ANOVA contains not one but two degrees of freedom: one for between-group variance (dfBG) and one for within-group variance (dfBG). Because a research study may have any number of groups and any number of scores within a group, these two degrees of freedom are distinct from each other and must be calculated separately.
First, the degrees of freedom for between-group variance (dfBG) in the one-way ANOVA is equal to the number of groups that comprise the independent variable, minus one:
(11-1) df BG = ≠ groups − 1
To understand the reasoning behind dfBG, recall from Chapter 7 that the concept of degrees of freedom was defined as the number of values or quantities that are free to vary when a statistic is used to estimate a parameter. Because between-group variance involves differences between group means, we are using these means ( X ¯ i ) to estimate the population mean μ. In the case of the chess study, if we were to combine the means of the three groups to estimate μ, how many of the group means are free to vary? The answer to this question is two; this is because once the first two group means have been calculated,
664
the third one is fixed. For the chess study: d f B G = ≠ groups − 1 = 3 − 1 = 2
Next, the degrees of freedom for within-group variance (dfwG) in the one-way ANOVA is equal to
(11-2) d f W G = ∑ ( N i − 1 )
where Ni is the number of scores in each group. The formula for dfWG is a simple extension of the formula for the degrees of freedom for the t-test for two sample means (df = (N1 – 1) + (N2 – 1)). To understand the reasoning behind Formula 11–2, it is useful to note that within-group variance in a one-way ANOVA is based on the deviation of an individual score from its particular group mean. For each group, the number of degrees of freedom is the number of scores in the group minus one (Ni – 1). For example, because the Observation group contains 14 scores, this group has 14 – 1 (or 13) degrees of freedom. Calculating dfWG involves combining the degrees of freedom of all of the groups; this is accomplished by summing N – 1 across the number of groups. For the chess study: d f W G = ∑ ( N i − 1 ) = ( 14 − 1 ) + ( 14 − 1 ) + ( 14 − 1 ) = 13 + 13 + 13 = 39
To summarize, for the chess study, the two degrees of freedom are dfBG = 2 and dfWG = 39.
Set Alpha (α), Identify the Critical Value, and State a Decision Rule
Alpha (α) is the probability of a statistic necessary to reject the null hypothesis. As in earlier chapters of this book, alpha for the one-way ANOVA is traditionally set at .05 (α = .05). In Chapters 7 and 9, stating alpha for the t-test also required indicating whether alpha was one-tailed or two-tailed. However, it is not necessary to make this distinction in an ANOVA because, as we have already observed, F-ratios can only be positive numbers.
The next step in our calculations is to identify the critical value of the F-ratio that separates the region of rejection from the region of non-rejection. As with the t-test, the critical value changes depending on the size of the sample. Furthermore, given that an independent variable may consist of any number of groups, there is a different distribution of F-ratios and different critical values for every combination of sample sizes and number of groups.
Table 4 in the back of this book provides a table of critical values for the F-ratio; a portion of this table is listed in Table 11.3. The table consists of a series of columns under the label “Degrees of Freedom for Numerator” and a series of rows next to the label “Degrees of Freedom for Denominator.” The degrees of freedom correspond to the numerator and the denominator of the F-ratio. For a one-way ANOVA, the degrees of freedom for the
665
numerator is dfBG and the degrees of freedom for the denominator is dfWG. In the chess study example, the degrees of freedom are dfBG = 2 and dfWG = 39.
Table 11.3 Example of Critical Values of the F-Ratio Table 11.3 Example of Critical Values of the F-Ratio
Degrees of Freedom for Numerator
1 2 3 4 5 6 7 8 9
Degrees of Freedom for Denominator
1 161.4 199.5 215.7 224.6 230.2 234.0 236.8 238.9 240.5
4052 4999 5403 5625 5764 5159 5928 5981 6021
2 18.51 19.00 19.16 19.25 19.30 19.33 19.36 19.37 19.38
98.49 99.00 99.17 99.25 99.30 99.33 99.36 99.37 99.39
3 10.13 9.55 9.28 9.12 9.01 8.94 8.88 8.84 8.81
34.12 30.82 29.46 28.71 28.24 27.91 27.67 27.49 27.34
4 7.71 6.94 6.59 6.39 6.26 6.16 6.09 6.04 6.00
21.20 18.00 16.69 15.98 15.52 15.21 14.98 14.80 14.66
5 6.61 5.79 5.41 5.19 5.05 4.95 4.88 4.82 4.78
16.26 13.27 12.06 11.39 10.97 10.67 10.45 10.29 10.15
6 5.99 5.14 4.76 4.53 4.39 4.28 4.21 4.15 4.10
13.74 10.92 9.78 9.15 8.75 8.47 8.26 8.10 7.98
To identify the critical value for the chess study, we turn to the table of critical values in Table 4. Based on our knowledge that dfBG = 2, we move to the 2 column under “Degrees of Freedom for Numerator.” Next, because dfWG = 39, we move down the rows of the table looking for a row corresponding to 39 for the “Degrees of Freedom for Denominator.” Here we find that the table does not provide critical values for dfWG = 39 but only for df = 30 and df = 40. For the chess study, we use the smaller of these two values: df = 30. It would be inappropriate to use the critical value associated with df = 40 as doing so would imply the sample was larger than it actually was.
For the combination of dfBG = 2 and dfWG = 30, we find two critical values: a boldfaced 3.32, with the number 5.39 immediately below it. The boldfaced number is the critical value for α = .05, and the number below it is the critical value for α = .01. Therefore, the critical value for the chess study may be stated as follows: For α = .05 and d f = 2 , 39 , critical value = 3.32
The critical value and regions of rejection and non-rejection for the chess study are illustrated in Figure 11.4.
666
Once the critical value has been identified, the next step is to state a decision rule that specifies the values of the F-ratio resulting in the rejection of the null hypothesis. For the chess study, the following decision rule may be stated: If F > 3.32 , reject H 0 ; otherwise, do not reject H 0
The region of rejection for this example is represented by the shaded area of the distribution in Figure 11.4.
Calculate a Statistic: F-Ratio for the One-Way ANOVA
The next step in hypothesis testing is to calculate a value of an inferential statistic. Because calculating the F-ratio for the one-way ANOVA involves calculating variances, it's useful to review the formula for the sample variance (s2), which was defined in Chapter 4 as the average squared deviation of a score from its mean: s 2 = ∑ ( X − X ¯ ) 2 N − 1 = S S d f
The numerator of this formula, the sum of squared deviations ( ∑ ( X − X ¯ ) 2 ), is also known as the Sum of Squares (SS). The denominator of this formula (N – 1) is the degrees of freedom (df) for a set of scores.
Calculating the F-ratio for a one-way ANOVA involves modifying the above formula to calculate two types of variance: between-group variance (MSBG) and within-group variance (MSWG). Variances calculated within an ANOVA have a specific label, Mean Square (MS), because the literal definition of variance is the average squared, or mean squared, deviation. The remainder of this section demonstrates how these two variances and the F-ratio are calculated.
Figure 11.4 Critical Value and Regions of Rejection and Non-Rejection for the Chess Study
Calculate Between-Group Variance (MSBG)
Between-group variance refers to differences among the means of the groups comprising an
667
independent variable; it measures the hypothesized effect being tested. In the t-test, between-group variance was simply the difference between the means of two groups ( X ¯ 1 − X ¯ 2 ). However, in the one-way ANOVA, between-group variance is based on the difference between the mean of each group and the mean for the total sample ( X ¯ i − X ¯ T ).
For a simple illustration of between-group variance in the one-way ANOVA, imagine a situation where, in calculating the means of two groups, you find they do not differ from each other: X ¯ 1 = 10 and X ¯ 2 = 10 . If you were to combine the data from the two groups and calculate the mean of the total sample ( X ¯ T ), you would find X ¯ T is also equal to 10. This situation is illustrated in Figure 11.5(a). Looking at the figure, you will observe that the lack of difference between group means is reflected in the lack of difference between the group means and the total mean ( X ¯ 1 − X ¯ T = 10 − 10 = 0 and X ¯ 2 − X ¯ T = 10 − 10 = 0 ).
Figure 11.5(b) illustrates the situation where the means of two groups do in fact differ from each other ( X ¯ 1 = 5 and X ¯ 2 = 15 ). With X ¯ T again equal to 10, this figure shows that the difference between group means can be represented by the differences between group means and the total mean ( X ¯ 1 − X ¯ T = − 5 and X ¯ 2 − X ¯ T = 5 ).
Based on differences between group means and the total mean, MSBG for the one-way ANOVA is calculated using Formula 11–3:
(11-3) M S B G = S S B G d f B G = N i ∑ ( X ¯ i − X ¯ T ) 2 # groups − 1
Figure 11.5 Between-Group Variance as Differences between Group Means ( X ¯ i ) and the Total Mean ( X ¯ T )
where Ni is the number of scores in each group, X ¯ i is the mean for each group, and X ¯ T is the mean for the total sample.
The first step in calculating MSBG is to calculate the mean of the total sample ( X ¯ T ). When all of the groups have the same sample size (Ni), the mean of the total sample in a
668
one-way ANOVA is the mean of the group means:
(11-4) X ¯ T = ∑ X ¯ i # groups
where X ¯ i is the mean for each group. For the chess study, using the group means in Table 11.2, X ¯ T is calculated below: X ¯ T = ∑ X ¯ i # groups = 1.36 + .86 + 3.00 3 = 1.74 = 5.22 3
Once X ¯ T has been determined, MSBG may be calculated. As we discussed earlier, the degrees of freedom (dfBG) is the number of groups minus one. Because the chess study contains three groups, dfBG is equal to 3 – 1, or 2. Again using the group means from Table 11.2, we find that MSBG is equal to M S B G = N i ∑ ( X ¯ i − X ¯ Τ ) 2 # group − 1 = 14 [ ( 1.36 − 1.74 ) 2 + ( .86 − 1.74 ) 2 + ( 3.00 − 174 ) 2 ] 3 − 1 = 14 [ ( − .38 ) 2 + ( − .88 ) 2 + ( 1.26 ) 2 ] 2 = 14 [ .14 + .77 + 1.59 ) 2 = 14 ( 2.50 ) 2 = 35.00 2 = 17.50
Calculate Within-Group Variance (MSWG)
Within-group variance (MSWG), the variance of scores within groups for the one-way ANOVA, is calculated using Formula 11–5:
(11-5) M S W G = S S W G d f W G = ( N i − 1 ) ∑ s i 2 ∑ ( N i − 1 )
where Ni is the number of scores in each group and si is the standard deviation for each group.
Starting with the denominator of Formula 11–5, we have already calculated the degrees of freedom (dfWG). Because the chess study example contains three groups, each of which has 14 scores (Ni = 14), dfWG is equal to (14 – 1) + (14 – 1) + (14 – 1) or 39. Next, using the standard deviations calculated in Table 11.2, we may calculate MSWG for the chess study in the following way: M S W G = ( N i − 1 ) ∑ s i 2 ∑ ( N i − 1 ) = ( 14 − 1 ) [ ( 1.74 ) 2 + ( .95 ) 2 + ( 1.84 ) 2 ] ( 14 − 1 ) + ( 14 − 1 ) + ( 14 − 1 ) = 13 [ 3.03 + .90 + 3.39 ] 13 + 13 + 13 = 13 ( 7.32 ) 39 = 95.16 39 = 2.44
Students sometimes make computational errors in calculating MSBG and MS,if/G. For example, in calculating MSBG, students may forget to square the difference between a group mean and the total mean, such as (1.36 – 1.74) or (.86 – 1.74), which may result in a negative value for MSBG. Keep in mind that variance, because it is based on squared deviations, must always be a positive number.
669
Students might also calculate the degrees of freedom (dfBG and (dfWG) incorrectly. To avoid incorrect calculations, remember that the degrees of freedom for between-group variance (dfBG) is based on the number of groups (# groups – 1), whereas the degrees of freedom for within-group variance (dfWG) is based on the number of scores in each group (N – 1).
Calculate the F-Ratio (F)
In a one-way ANOVA, the F-ratio (F) is calculated using Formula 11–6:
(11-6) F = M S B G M S W G
where MSBG is between-group variance and MSWG is within-group variance. In the chess study example: F = M S B G M S W G = 17.50 2.44 = 7.17
Create an ANOVA Summary Table
To organize and present the calculation of the F-ratio, researchers sometimes construct what is known as an ANOVA summary table. An ANOVA summary table is a table that summarizes the calculations of an analysis of variance (ANOVA).
We have constructed three ANOVA summary tables in Table 11.4. The summary table in Table 11.4(a) includes the symbols and notation used to designate the SS, df MS, and F- ratio. The middle summary table (Table 11.4(b)) contains formulas and calculations for the different parts of the one-way ANOVA. The bottom summary table (Table 11.4(c)) is the type of table we would use to report the results of the analysis for the chess study. The first column in each table uses the label “Source” to describe the hypothesized explanation or cause of each type of variance. In Table 11.4(c), the word Condition refers to the independent variable in this particular study, and the word Error is used to describe variance that is not explained or accounted for.
You may have noticed that the bottom row of each ANOVA summary table is labeled Total. This row corresponds to the total sample and consists of the total Sum of Squares (SST) and the total degrees of freedom (dfT). In a one-way ANOVA, the total Sum of Squares is the sum of the between-group and within-group Sums of Squares (SSBG + SSWG); in this example, SST = 35.00 + 95.16 = 130.16. The total degrees of freedom is calculated in a similar manner, where dfT = dfBG + dfWG; in the chess study, dfT = 2 + 39 = 41. Although not used in the calculation of the F-ratio, SST will be used later in this chapter to calculate a measure of effect size associated with the F-ratio.
670
Table 11.4 Summary Tables for the One-Way ANOVA Table 11.4 Summary Tables for the One-Way
ANOVA
a. Notation and Symbols
1 Source SS df MS F
Between-group (BG) SS BG df BG MS BG F
Within-group (WG) SS WG df WG MS WG
Total (T) SS T df T Table 11.4 Summary Tables for the One-Way ANOVA
b. Formulas
Source SS df MS F
Between-group (BG)
N i ∑ ( X ¯ i − X ¯ T ) 2
# groups – 1
S S B G d f B G
M S B G M S W G
Within-group (WG)
( N i − 1 ) ∑ s i 2 Σ(Ni – 1) S S W G d f W G
Total (T) SSBG + SSWG dfBG + dfWG
Table 11.4 Summary Tables for the One- Way ANOVA
c. Chess Study Example
Source SS df MS F
Condition 35.00 2 17.50 7.17
Error 95.16 39 2.44
Total 130.16 41
671
Learning Check 2: Reviewing what you've Learned So Far
1. Review questions a. If you were conducting a one-way ANOVA involving four groups (Group 1 to Group 4),
what are the two ways you could state the null hypothesis (H0)? b. Which of these is the correct way to state the alternative hypothesis: H1: all us are not equal
or H1: not all us are equal? Why is one way correct and the other way incorrect? c. Why does hypothesis testing in the one-way ANOVA involve two degrees of freedom? d. Why is dfBG = # groups = 1 and dfBG = Σ(Ni – 1)? e. How is stating the critical value and decision rule for the F-ratio different from the t-test? f. In measuring between-group and within-group variance, what does the symbol MS
represent?
2. For each of the following, calculate the degrees of freedom (df) and identify the critical value of F (assume α = .05).
1. # groups = 4, Ni = 11 2. # groups =5, Ni = 6 3. # groups = 6, Ni = 8
3. For each of the following, calculate between-group variance (MSBG). a. Ni = 9, X ¯ 1 = 4.00 , X ¯ 2 = 9.00 , X ¯ 3 = 2.00 b. Ni = 10, X ¯ 1 = 27.00 , X ¯ 2 = 39.00 , X ¯ 3 = 33.00 , X ¯ 4 = 37.00 c. Ni = 7, X ¯ 1 = 3.75 , X ¯ 2 = 4.25 , X ¯ 3 = 3.85 , X ¯ 4 = 2.90 , X ¯ 5 = 5.25
4. For each of the following, calculate within-group variance (MSWG). a. Ni = 18, s1 = 2.00, s2 = 4.00, s3 = 5.00 b. Ni = 5, s1 = 12.00, s2 = 9.00, s3 = 11.00, s4 = 8.00 c. Ni= 11, s1 = .27, s2 = .51, s3 = .46, s4 = .34, s5 = .40
5. For each of the following, calculate the F-ratio (F) and create an ANOVA summary table. a. Ni = 12, X ¯ 1 = 3.50 , s1 = 1.25, X ¯ 2 = 4.50 , s2 = 1.15, X ¯ 3 = 4.00 , s3 = .80 b. Ni = 8, X ¯ 1 = 2.74 , s1 = .49, X ¯ 2 = 2.26 , s2 = .38, X ¯ 3 = 2.90 , s3 = .58, X ¯ 4 = 2.03 ,
s4 = .45
Make a Decision Whether to Reject the Null Hypothesis
The decision regarding the null hypothesis is made by comparing the value of the F-ratio calculated from the data with the critical value. The null hypothesis is rejected if the calculated value is greater than the critical value. For the chess study, this decision is stated as follows: F = 7.17 > 3.32 ∴ reject H 0 ( p < .05 )
Because the F-ratio value of 7.17 calculated from the data is greater than the critical value of 3.32, it falls in the region of rejection, and the decision is made to reject the null
672
hypothesis.
Determine the Level of Significance
If and when the decision is made to reject the null hypothesis, it is useful to provide a precise assessment of the probability of the calculated F-ratio by comparing it with the .01 critical value. For dfBG = 2 and dfWG = 39, the .01 critical value in Table 4 is equal to 5.39. For the chess study, the level of significance may be stated as follows: F = 7.17 > 5.39 ∴ p < .01
The level of significance for the chess study is illustrated in Figure 11.6. As we can see from the figure, the calculated F-ratio value of 7.17 exceeds both the α = .05 and .01 critical values. This implies that its probability is not only less than .05 but also less than .01.
Calculate a Measure of Effect Size (R2)
As was learned in Chapter 10, a measure of effect size is an index or estimate of the size or magnitude of a hypothesized effect. More specifically, the effect size may be measured as the percentage of variance in the dependent variable accounted for by the independent variable, ranging from .00 to 1.00 (0% to 100%). In contrast to inferential statistics such as the t-test or F-ratio, measures of effect size are useful in interpreting statistical analyses because they are not affected by sample size and also because they are comparable across different samples and different studies.
There are several measures of effect size appropriate for a one-way ANOVA, each of which has its advantages and disadvantages. For the purposes of our discussion, we will use the relatively simple measure of effect size known as R2. The R2 statistic is the percentage of variance in the dependent variable associated with differences between the groups that comprise the independent variable. The R2 statistic is similar to the r2 measure of effect size (used to accompany the t-test) that was discussed in Chapter 9.
The measure of effect size R2 for the one-way ANOVA is calculated using Formula 11–7:
(11-7) R 2 = S S B G S S T
Figure 11.6 Determining the Level of Significance for the Chess Study
673
where SSBG is the sum of squares for the between-group variance and SST is the sum of squares for the total amount of variance. Here, SSBG corresponds to variance in the dependent variable attributed to differences between groups, while SST represents the total amount of variance in the dependent variable. By dividing SSBG by SST, we are calculating a percentage of the total amount of variance in the dependent variable accounted for by differences between groups.
Using the values for SSBG and SST provided in the ANOVA summary table in Table 11.4(c), we can calculate R2 for the independent variable Condition in the chess study as follows: R 2 = S S B G S S T = 35.00 130.16 = .27
Because a measure of effect size such as R2 is meant to supplement the results of inferential statistical analyses, the results of the analysis in the chess study may be reported as
The R2 in this example may be described in the following way: “27% of the variance in the number of checkmates is explained by the different strategies used by students.” To interpret the value of R2, we can use Cohen's (1988) rules of thumb that we previously encountered in Chapter 10:
A “small” effect produces an R2 of .01.
A “medium” effect produces an R2 of .06.
A “large” effect produces an R2 of .15 or greater.
Interpreted in this manner, an R2 value of .27 would be considered a relatively large effect.
674
Draw a Conclusion from the Analysis
What conclusion may be drawn from rejecting the null hypothesis in a one-way ANOVA? One way to report the results of the analysis in the chess study is the following:
The mean number of checkmates for the 14 students in each of the Observation (M = 1.36), Prediction (M = .86), and Explanation (M = 3.00) conditions was not all equal to each other, F(2, 39) = 7.17, p < .01, R2 = .27.
Notice that this sentence contains a great deal of relevant information about the analysis:
the dependent variable (“number of checkmates”), the sample size for each group (“14 students”), the independent variable (“Observation … Prediction … and Explanation … conditions”), descriptive statistics (“M = 1.36 … M = .86 … M = 3.00”), the nature and direction of the findings (“were not all equal to each other”), and information about the inferential statistic (“F(2, 39) = 7.17, p < .01, R2 = .27”), which includes the inferential statistic calculated (F), the degrees of freedom (2, 39), the value of the statistic (7.17), the level of significance (p < .05), and the measure of effect size (R2 = .27).
675
Relate the Result of the Analysis to the Research Hypothesis
The last step in hypothesis testing is to relate the result of the analysis to the study's research hypothesis. The research hypothesis in the chess study was that the number of checkmates would be greater for students in the Explanation condition than for students in the Prediction condition, which in turn would be greater than for students in the Observation condition. At this point in the analysis, can we draw a conclusion regarding the extent to which this research hypothesis has been supported? The answer, quite simply, is no.
As it turns out, rejecting the null hypothesis in the one-way ANOVA does not indicate the specific nature and direction of any differences between the means of the groups. Looking back at the null (H0) and alternative (H1) hypothesis for the one-way ANOVA, we see that the alternative hypothesis is stated as “not all μs are equal.” Consequently, when the null hypothesis is rejected, the most specific conclusion that can be made is that the means are not all equal to each other—we cannot state whether any of the group means are significantly different from the others. As a result, we are not yet able to draw conclusions regarding whether a study's research hypothesis has been supported.
To test a study's research hypotheses, it is necessary to conduct additional, more specific analyses. For example, one of the hypotheses in the chess study was that the Explanation strategy is more effective than Prediction. To test this hypothesis, we need to conduct an analysis that involves only these two groups, excluding the Observation condition. Section 11.5 of this chapter describes these analyses, known as “analytical comparisons.” Because we have not yet conducted analytical comparisons, the most appropriate statement we can make at this point regarding a study's research hypothesis is the following:
At this point in the analysis, we are unable to determine whether the research hypothesis in the study is supported until analytical comparisons are conducted.
Later in this chapter, we will introduce and illustrate the process of conducting and interpreting analytical comparisons.
676
Summary
The process of conducting a one-way ANOVA is summarized in Table 11.5, using the chess study as an example. As we mentioned earlier, because the results of the one-way ANOVA do not allow conclusions to be made regarding the extent of support for a study's research hypothesis, it is necessary to conduct additional analyses known as analytical comparisons. However, before moving on to a discussion of analytical comparisons, a second example of a one-way ANOVA will first be presented and analyzed.
Table 11.5 Summary, Conducting the One-Way ANOVA (Chess Study Example)
677
Learning Check 3: Reviewing what you've Learned So Far
1. Review questions
a. What does R2 in a one-way ANOVA represent? b. What is the most specific conclusion that can be drawn when the null hypothesis for a one-
way ANOVA has been rejected? c. Does rejecting the null hypothesis for a one-way ANOVA support or not support a research
hypothesis?
2. Although research has examined factors women take into consideration when making birth control decisions, relatively few studies have looked at men. As male oral contraceptives were starting to become available, Jaccard, Hand, Ku, Richardson, and Abella (1981) conducted a study to investigate males' concerns regarding these contraceptives. In the study, male college students were asked to rate the importance of one of three aspects in deciding whether they would use an oral contraceptive: effectiveness as a birth control device, convenience (time and effort needed to use the device), and potential risks to their own health—the higher the rating, the greater the importance. Jaccard et al. hypothesized that males would rate risks to their health as more important than either the contraceptives effectiveness or convenience. Imagine the following descriptive statistics are reported for the ratings of the three aspects: effectiveness (N = 9, M = 3.18, s = 1.21), convenience (N = 9, M = 3.43, s = 1.14), and health risks (N = 9, M = 4.51, s = 1.03).
a. State the null and alternative hypotheses (H0 and H1).
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate the F-ratio for the one-way ANOVA and create a summary table. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
6. Calculate a measure of effect size (R2). c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
678
11.4 A Second Example: The Parking Lot Study Revisited
This section presents a second example of conducting a one-way ANOVA. There are two main purposes of this example. The first purpose is to more quickly illustrate the process of conducting a one-way ANOVA than in the first example. Because the first example explained both why the specific calculations were being made as well as how to perform the calculations, we think it would be useful to provide a second example that focuses primarily on the calculations. The second purpose of this example is to compare the one-way ANOVA with the t-test. Because both the one-way ANOVA and the t-test involve comparing the means of groups, its important for you to understand the relationship between these two statistical procedures.
To accomplish these two purposes, we will once again use the parking lot study from Chapter 9. As you may recall, this example involved testing the difference in the amount of time taken to leave a parking space by two groups of drivers: those intruded upon and those not intruded upon by another driver. Table 11.6 summarizes the descriptive statistics for the data from the parking lot study.
When this example was first presented in Chapter 9, we used the t-test to reach the following conclusion: “The mean departure time for the 15 drivers in the Intruder group (M = 40.73 s) is significantly greater than the mean departure time for the 15 drivers in the No intruder group (M = 31.67 s), t(28) = 2.42, p < .05.”
In this section, we will learn how the parking lot study may be analyzed using a one-way ANOVA, following the same four steps that we have used in our earlier calculations:
state the null and alternative hypotheses (H0 and H1), make a decision about the null hypothesis, draw a conclusion from the analysis, and relate the result of the analysis to the research hypothesis.
Table 11.6 Descriptive Statistics, Parking Lot Study (Chapter 9) Table 11.6 Descriptive Statistics, Parking Lot
Study (Chapter 9)
Driver Group
Intruder No intruder
Ni 15 15
Mean ( X i ¯ ) 40.73 31.67
Standard deviation (si) 10.42 10.08
679
State the Null and Alternative Hypotheses (H0 and H1)
As we have already learned, the null hypothesis states that the means of the groups in the population are all equal to each other, whereas the alternative hypothesis states that the means are not all equal. If we analyze the parking lot study using a one-way ANOVA, the two statistical hypotheses may be stated as follows: H 0 : all μs are equal H 1 : not all μs are equal
680
Make a Decision about the Null Hypothesis
The decision whether to reject the null hypothesis for the parking lot study is made by repeating the same steps that we used in the earlier example of the chess study Each of these steps is summarized below.
Calculate the Degrees of Freedom (df)
The two degrees of freedom for this one-way ANOVA, dfBG and dfWG, correspond to between-group and within-group variance. Using the information in Table 11.6, the two degrees of freedom are calculated as follows: d f B G = # groups − 1 d f W G = ∑ ( N i − 1 ) = 2 − 1 = ( 15 − 1 ) + ( 15 − 1 ) = 1 = 28
The degrees of freedom for the parking lot study may be reported as 1, 28.
Set Alpha (α), Identify the Critical Value, and State a Decision Rule
Assuming α = .05 and using the degrees of freedom calculated above, the critical value of the F-ratio is identified using Table 4. Because in this example dfBG = 1 and dfWG = 28, we locate the 1 column under the label “Degrees of Freedom for Numerator” and then move down the column until we reach the row corresponding to 28 degrees of freedom for the denominator. For this combination of alpha and degrees of freedom, we find a critical value of 4.20. Therefore, For α = . 05 and d f = 1 , 28 , critical value = 4 . 20
For the parking lot study, the decision rule for the F-ratio may be stated as follows: If F > 4.20 , reject H 0 ; otherwise, do not reject H 0
If the F-ratio calculated from the data exceeds 4.20, it falls in the region of rejection and will lead to the rejection of the null hypothesis.
Calculate a Statistic: F-Ratio for the One-Way ANOVA
Three steps are used to calculate the F-ratio for the one-way ANOVA: calculate between- group variance (MSBG), calculate within-group variance (MSWG), and calculate the F-ratio (F). Each of these steps is completed below for the parking lot study.
Calculate Between-Group Variance (MSBG)
681
The first step in calculating the F-ratio for the one-way ANOVA is to calculate between- group variance (MSBG). To calculate MSBG, we must first calculate X ¯ T , which is the mean for the total sample of data. Using the group means provided in Table 11.6, X ¯ T is calculated below: X T ¯ = ∑ X ¯ i # groups = 40.73 + 31.67 2 = 72.40 2 = 36.20
Once X ¯ T has been calculated, we can calculate MSBG: M S B G = N i ∑ ( X ¯ i − X ¯ T ) 2 # groups − 1 = 15 [ ( 40.73 − 36.20 ) 2 + ( 31.67 − 36.20 ) 2 ] 2 − 1 = 15 [ ( 4.53 ) 2 + ( + 4.53 ) 2 ] 1 = 15 ( 20.52 ) + ( 20.52 ) 1 = 15 ( 41.04 ) 1 = 615.60 1 = 615.60
Calculate within-group variance (MSWG). The second step in calculating the F-ratio is to calculate within-group variance (MSWG). Again using the information in Table 11.6, we find that M S W G = ( N i − 1 ) ∑ s i 2 ∑ ( N i − 1 ) = ( 15 − 1 ) [ ( 10.42 ) 2 + ( 10.08 ) 2 ] ( 15 − 1 ) + ( 15 − 1 ) = 14 [ 108.58 + 101.61 ] 14 + 14 = 14 ( 210.19 ) 28 = 2942.66 28 = 105.10
Calculate the F-Ratio (F)
Once the two variances have been calculated, the next step is to calculate a value for the F- ratio by dividing MSBG by MSWG For the parking lot study, we find the following: F = M S B G M S W G = 615.60 105.10 = 5.86
An ANOVA summary table for this analysis is presented in Table 11.7.
Make a Decision Whether to Reject the Null Hypothesis
We make the decision whether to reject the null hypothesis by comparing the value of the F-ratio calculated from the data with the critical value. For the parking lot study, the following decision is made: F = 5.86 > 4.20 ∴ reject H 0 ( p<.05 )
Because the calculated F-ratio of 5.86 is greater than the critical value of 4.20, we reject the null hypothesis and conclude that the means are not all equal to each other.
Table 11.7 Summary Table, One-Way ANOVA, Parking Lot Study Table 11.7 Summary Table, One-Way
ANOVA, Parking Lot Study
Source SS df MS F
Driver group 615.60 1 615.60 5.86
Error 2942.66 28 105.10
682
Total 3558.26 29
Figure 11.7 Determining the Level of Significance for the Parking Lot Study
Determine the Level of Significance
Because we have rejected the null hypothesis, it is appropriate to determine whether the probability of the F-ratio is less than .01. Comparing the F-value of 5.86 (df = 1, 28) with the .01 critical value of 7.64 in Table 4, the following conclusion may be made: F = 5.86 < 7.64 ∴ p < .05 ( but not < .01 )
The level of significance for this example is illustrated in Figure 11.7. Looking at the figure, we observe that the F-ratio value of 5.86 falls between the α = .05 (4.20) and .01 (7.64) critical values. Therefore, the appropriate level of significance for this analysis is p < .05.
Calculate a Measure of Effect Size (R2)
To supplement the information provided by the results of the one-way ANOVA, a measure of effect size (R2) is calculated. Obtaining values of SSBG and SST from the ANOVA summary table in Table 11.7, the value of R2 for the parking lot study may be stated as follows: R 2 = S S B G S S T = 615.60 3358.26 = .17
From the value of R2, it may be concluded that 17% of the variance in the departure times of drivers may be accounted for by the presence or absence of another driver. Using Cohen's (1988) rules of thumb presented earlier in this chapter, we can see that this is a relatively large effect.
683
Draw a Conclusion from the Analysis
What conclusion can be drawn from the analysis in this example? In the chess study example discussed earlier in this chapter, the independent variable Condition consisted of more than two groups. Consequently the most specific conclusion that could be drawn from rejecting the null hypothesis was that the means of the three groups were not all equal to one another. However, because the parking lot study's independent variable (Driver group) consists of only two groups, rejecting the null hypothesis enables us to make a more specific conclusion regarding the nature and direction of the difference between groups. This conclusion may be stated as follows:
The mean departure time for the 15 drivers in the Intruder group (M = 40.73 s) is significantly greater than the mean departure time for the 15 drivers in the No intruder group (M = 31.67 s), F(1, 28) = 5.86, p < .05, R2 = .17.
684
The Relationship between the t-Test and the F-Ratio
The parking lot study has now been analyzed with two different statistical procedures: the t-test and the one-way ANOVA. Although both analyses resulted in the decision to reject the null hypothesis, the statistics do not look the same: t(28) = 2.42 and F(1, 28) = 5.86. Based on our knowledge that the t-test and one-way ANOVA both involve testing differences between group means, what is the relationship between the t-statistic and the F- ratio?
From Formula 11–8, we see that the squared value for the t-statistic is equal to the value of the F-ratio:
(11-8) t 2 = F
In other words, for the same set of data, the t-statistic and the F-ratio are algebraically equivalent. To equate the t-statistic and F-ratio, the calculated value of the t-statistic must be squared to convert any negative t-statistics to positive numbers.
The values of t2 and F for the parking lot study are compared below: t 2 = F ( 2.42 ) 2 = 5.86 5.86 = 5.86
This example illustrates that either the t-test or the one-way ANOVA can be used to test the difference between the means of two groups. This choice is typically based on such things as personal preference, ease of calculation, and familiarity. The important thing to remember is that the conclusion drawn from either method will be the same. We will not get different results by using one statistical procedure instead of the other.
685
11.5 Analytical Comparisons within the One-Way ANOVA
This section returns to the chess study analyzed earlier in this chapter. In that example, the most specific conclusion we could draw from the decision to reject the null hypothesis was that the mean number of checkmates of the three groups of learners (Observation, Prediction, Explanation) was not all equal to one another. Because the result of a one-way ANOVA doesn't indicate which groups differ or do not differ from each other, to test the study's research hypotheses, we must conduct additional statistical analyses. Because these statistical analyses involve comparing groups that are part of a larger study, they are known as analytical comparisons.
Analytical comparisons are comparisons between groups that are part of a larger research design. This section will discuss how to calculate two types of analytical comparisons: planned comparisons and unplanned comparisons. As we will see, these two types of analytical comparisons differ more in terms of their rationale and evaluation than in how they are conducted.
686
Planned versus Unplanned Comparisons
The heart of many studies is the research hypothesis, which specifies the nature and direction of expected relationships between variables. When researchers design a study, they typically have a plan regarding the comparisons between groups needed to test their research hypotheses. These comparisons are known as planned comparisons, defined as comparisons that researchers build into the research design prior to the data collection process. Planned comparisons are sometimes referred to as “a priori” comparisons, which literally means “prior to.”
The purpose of planned comparisons is to test research hypotheses. For example, one of the research hypotheses in the chess study was that the number of checkmates would be greater for students in the Explanation condition than for students in the Prediction condition. Even before the actual research begins, the researcher has already anticipated the importance of comparing the means of these two groups, excluding the Observation condition, for testing the research hypothesis.
In addition to planned comparisons, there is a second type of analytical comparison known as unplanned comparisons, which are comparisons that have not been built into the research design prior to the data collection process. The decision to conduct unplanned comparisons is typically made after the F-ratio for the one-way ANOVA has been calculated and the null hypothesis has been rejected. These types of analytical comparisons are also known as “post hoc” comparisons, meaning “after the fact.”
There are two common reasons for conducting unplanned comparisons. First, the research study may be examining an issue or question that has not been extensively studied, making it difficult to state research hypotheses before data collection has been conducted. Second, the results of initial statistical analyses may suggest something the researcher did not anticipate when designing the study, leading the researcher to conduct additional analyses to investigate these unexpected results.
The following sections will discuss between-group and within-group variance in analytical comparisons, as well as how to conduct and evaluate planned and unplanned comparisons for the one-way ANOVA. Although both planned and unplanned comparisons involve the calculation of an F-ratio, they differ in terms of how the F-ratio is evaluated to make the decision about the null hypothesis.
687
Between-Group and Within-Group Variance in Analytical Comparisons
Because analytical comparisons are part of a larger research design involving differences between group means, conducting these comparisons requires the calculation of an F-ratio. We will represent the F-ratio in an analytical comparison by the symbol Fcomp. Similar to the method for calculating the F-ratio for the one-way ANOVA, calculating Fcomp consists of dividing between-group variance by within-group variance.
Between-group variance is different in an analytical comparison than it is in the one-way ANOVA. This is because we are comparing a subset of the groups, rather than all of the groups, that comprise the independent variable. As a result, the “effect” we are evaluating is not the same. For example, although the chess study consisted of three groups (Observation, Prediction, and Explanation), the “effect” we may be interested in testing is the difference between Explanation and Prediction. Consequently, between-group variance in an analytical comparison is not MSBG, which is based on all of the groups, but instead is represented by M S B G comp , which is based only on the groups being compared.
Within-group variance represents an estimate of the amount of “error” in a set of data. In calculating Fcomp, we could calculate within-group variance using just the groups involved in the analytical comparison. For example, if we were interested in comparing the Explanation and Prediction conditions, we could calculate within-group variance using the data for these two groups. However, in a set of data, we have an estimate of error based not just on some of the data but on all of the data: MSWG from the one-way ANOVA. Because it provides the best estimate of error, MSWG is the measure of within-group variance used in conducting analytical comparisons.
To summarize, the F-ratio for an analytical comparison, Fcom, is calculated by dividing M S B G comp (between-group variance [“effect”] based on the difference between the two groups being compared) by MSWG (within-group variance [“error”] based on all of the groups):
(11-9) F comp = M S B G c o m p M S W G
Below we illustrate how to calculate and evaluate Fcomp for both planned and unplanned comparisons.
688
Conducting Planned Comparisons
This section will illustrate how to conduct planned comparisons for the difference between the means of two groups. (For those who would like to explore the topic in greater depth and detail than this book can accommodate, more complicated planned comparisons are discussed in Keppel, Saufley & Tokunaga [1992] and Keppel and Wickens [2004].) You will recall that researchers in the chess study hypothesized that “participants in the [explanation] condition would … checkmate the King more often … than the other two conditions” (de Bruin et al., 2007, p. 192). To test one of the study's research hypotheses, we will conduct a planned comparison between the Explanation and Prediction conditions.
Although planned comparisons may be a new concept, they are conducted following the same steps as the one-way ANOVA, each of which is discussed below:
state the null and alternative hypotheses (H0 and H1), make a decision about the null hypothesis, draw a conclusion from the analysis, and relate the result of the analysis to the research hypothesis.
State the Null and Alternative Hypotheses (H0 and H1)
The statistical hypotheses for a planned comparison are very similar to those in earlier examples. In testing the difference between the number of checkmates for the Explanation and Prediction groups, the null and alternative hypotheses may either be stated as H 0 : μ Explanation = μ Prediction H 1 : μ Explanation ≠ μ Prediction
or in the more generic form: H 0 : all μs are equal H 1 : not all μs are equal
Make a Decision about the Null Hypothesis
The decision about the null hypothesis is made following the same steps as before. Each of these steps is discussed below.
Calculate the Degrees of Freedom (df)
There are two degrees of freedom for a planned comparison that correspond to the numerator and denominator of Fcomp: d f B G c omp = # groups − 1 d f W G = ∑ ( N i − 1 )
In terms of the numerator, the degrees of freedom is the number of groups being compared
689
minus one. Because planned comparisons involve two groups, the between-group degrees of freedom ( d f B G c omp ) is equal to 2 – 1, or 1. However, the degrees of freedom for the denominator (dfWG) are based on all of the groups in the study and are associated with MSWG from the one-way ANOVA. Looking at the ANOVA summary table for the chess study in Table 11.4(c), we see that dfWG = 39. To summarize, the degrees of freedom for the Explanation versus Prediction comparison is d f B G c omp = # groups - 1 d f W G = ∑ ( N i − 1 ) = 2 − 1 = ( 14 − 1 ) + ( 14 − 1 ) + ( 14 − 1 ) = 1 = 39
The two degrees of freedom for this analytical comparison may be represented as df = 1, 39.
Set Alpha (α), Identify the Critical Value, and State a Decision Rule
Alpha, the probability of Fcomp needed to reject the null hypothesis, can be set to the traditional value of .05 (α = .05). The critical value for an F-ratio, whether for a one-way ANOVA or a planned comparison, may be found in Table 4. Because the degrees of freedom for the Explanation versus Prediction comparison are d f B G c o m p = 1 and dfWG = 39, we identify the critical value by locating the dfnum = 1 column of the table and moving down until we reach the d f d e c o m p row of 30. Here, we find a critical value of 4.17 for α = .05. Therefore, For α = . 05 and d f =1, 39, critical value =4 . 17
Now that the critical value has been identified, a decision rule may be stated regarding the conditions leading to rejection of the null hypothesis that the two means are equal. For the Explanation versus Prediction planned comparison: If F comp > 4.17 , reject H 0 ; otherwise, do not reject H 0
Calculate a Statistic: F-Ratio for Analytical Comparisons (Fcomp)
The steps used to calculate Fcomp involve two of the three steps used to calculate the F-ratio for the one-way ANOVA: calculate between-group variance and calculate the F-ratio.
The between-group variance for an analytical comparison ( M S B G comp ) is calculated using the following formula:
(11-10) M S B G comp = N i ( X ¯ i − X ¯ j ) 2 2
where Ni is the number of scores for each group, X ¯ i is the mean for the first group involved in the comparison, and X ¯ j is the mean for the other group in the comparison. Notice that the difference between the two means ( X ¯ i − X ¯ j ) must be squared to make
690
M S B G comp a positive number.
Obtaining the means of the Explanation and Prediction groups from Table 11.2, each of which is based on a sample size of Ni = 14, M S B G comp is calculated as follows: M S B G comp = N i ( X ¯ i − X ¯ j ) 2 = 14 ( 2.14 ) 2 2 = 14 ( 4.58 ) 2 = 32.06 = 64.12 2
Once M S B G comp has been calculated, we are ready to calculate Fcomp by dividing M S B G comp by MSWG (Formula 11–9). The value for MSWG may be obtained from the summary table for the one-way ANOVA. Looking at Table 11.4(c), a MSWG value of 2.44 was calculated for the chess study. Consequently, Fcomp for the analytical comparison between the Explanation and Prediction groups may now be calculated: F comp = M S B G comp M S W G = 32.06 2.44 = 13.14
Make a Decision Whether to Reject the Null Hypothesis
The decision to reject the null hypothesis is made by comparing the calculated value of Fcomp with the critical value for the analytical comparison. For the Explanation versus Prediction comparison, we may state that F comp = 13.14 > 4.17 ∴ reject H 0 ( p < .05 )
By making the decision to reject the null hypothesis, we have decided that the mean number of checkmates for the Explanation and Prediction conditions significantly differs from each other.
Determine the Level of Significance
Because we have made the decision to reject the null hypothesis, it is appropriate to determine whether the probability of Fcomp is less than .01. For this example, using the .01 critical value of 7.56 for df = 1, 39, we determine that F comp = 13.14 > 7.56 ∴ p < .05
Because the value of 13.14 for Fcomp exceeds the .01 critical value of 7.56, its probability is not only less than .05 but also less than .01.
Calculate a Measure of Effect Size ( R c o m p 2 )
Just as we did with the one-way ANOVA, it is possible to calculate a measure of effect size for analytical comparisons. Calculating R2 for an analytical comparison ( R c o m p 2 ) is a simple modification of the formula for R2 presented earlier:
(11-11) R c o m p 2 = S S B G comp S S T
691
where S S B G comp is the between-group sum of squares for the comparison and SST is the sum of squares for the total amount of variance.
In calculating R c o m p 2 , the value for S S B G comp is the same as the value for M S B G comp : M S B G c o m p = S S B G c o m p d f B G c o m p = S S B G c o m p 1 = S S B G c o m p
As indicated by the calculations above, S S B G comp and M S B G comp are the same. For the Explanation versus Prediction analytical comparison, S S B G comp is equal to the M S B G comp value of 32.06 calculated earlier.
The next step in calculating R c o m p 2 is to locate the value for SST, which can be found in the summary table for the one-way ANOVA. For the chess study (Table 11.4(c)), SST = 130.16. Now, inserting our values of S S B G c o m p and SST into Formula 11–11, R c o m p 2 for the Explanation versus Prediction analytical comparison is calculated as follows:
Again using Cohen's (1988) guidelines, we observe that the difference in the number of checkmates between the Explanation and Prediction conditions represents a relatively large effect.
Draw a Conclusion from the Analysis
The following conclusion may be drawn from the Explanation versus Prediction comparison:
The mean number of checkmates for the 14 students in the Explanation condition (M = 3.00) is significantly greater than the number of checkmates for the 14 students in the Prediction condition (M = .86), F(1, 39) = 13.14, p < .01, R2 = .25.
Note that, like the t-test, we can indicate the specific nature and direction of differences between the groups because analytical comparisons involve only two groups.
Relate the Result of the Analysis to the Research Hypothesis
In light of the chess study's research hypotheses, what is the implication of finding a
692
significantly greater number of checkmates for students in the Explanation condition than we found in the Prediction condition? The researchers answered this question in the following manner:
The results of the study suggest that predicting the next move combined with self-explaining the predictions positively contributed to the development of principled understanding of a chess endgame. Participants in the prediction and self-explanation condition showed better understanding of the principles that underlie the KRK [king rook king] endgame than the prediction only condition, (de Bruin et al., 2007, p. 202)
It is worth noting that the researchers did not imply that their research hypothesis had been “proven.” Instead, they wrote that their findings “suggest” that having students explain their reasoning leads to higher levels of principled understanding.
Table 11.8 summarizes the steps involved in conducting planned comparisons, using the Explanation versus Prediction comparison as an example. The next section discusses how to conduct the second type of analytical comparison, unplanned comparisons, as part of a one-way ANOVA.
693
Conducting Unplanned Comparisons
Whereas planned comparisons are designed at the beginning of the study and are guided by theory and research hypotheses, the decision to conduct unplanned (post hoc) comparisons is not made until after the F-ratio for the one-way ANOVA has been calculated and found to be statistically significant. This section will first discuss concerns related to unplanned comparisons, followed by a presentation of one type of unplanned comparison in which each group is compared with each of the other groups.
Table 11.8 Summary, Conducting Planned Comparisons (Explanation vs. Prediction in the Chess Study Example)
Concerns regarding Unplanned Comparisons
As we learned earlier, two types of situations may result in conducting unplanned comparisons. First, a research study may be of an exploratory nature, and researchers consequently may not be able to formulate research hypotheses regarding specific
694
differences between groups. Second, initial statistical analyses may have unexpected findings, resulting in the need for further analyses.
Conducting analytical comparisons in these types of situations raises two main concerns. First, because unplanned comparisons are conducted in response to the results of statistical analyses rather than specific theory and hypotheses, the results and interpretation of these comparisons may be a function of the characteristics and idiosyncrasies of the samples. In other words, the fact that significant differences between groups are found in one study doesn't imply that they will be found in another study. This observation is particularly relevant when the samples are small or perhaps not representative of the larger populations.
A second concern about unplanned comparisons involves the probability of committing a Type I error over a set of analyses. Familywise error, defined as the probability of making at least one Type I error across a set of comparisons, is a concern when multiple unplanned comparisons are conducted to determine the source of a significant F-ratio. Type I error, as we learned in Chapter 10, consists of rejecting the null hypothesis when it is in fact true, implying that one has concluded that an effect exists when in fact it does not. The probability of making a Type I error in any individual comparison is equal to the stated alpha level, typically .05. However, the more comparisons that are made within a set of data, the greater the probability of making at least one Type I error.
Let's use a hypothetical example to illustrate the concept of familywise error as well as to demonstrate how unplanned comparisons are conducted and evaluated. Imagine that we conduct a study consisting of five groups, each having nine participants (Ni = 9). In analyzing the data for this hypothetical study, the following ANOVA summary table is created:
Source SS df MS F
Between-group 365.60 4 91.40 3.52
Within-group 1040.00 40 26.00
Total 1405.60 44
In making a decision about the null hypothesis, for α = .05 and df = 4, 40, we find a critical value of 2.61 in Table 4. Because the calculated F-ratio of 3.52 is greater than the critical value of 2.61, the null hypothesis is rejected.
Because the only conclusion we can draw from rejecting the null hypothesis is that the means of the five groups are not all equal to each other, analytical comparisons could be conducted to determine significant differences among the groups. In the absence of specific hypotheses regarding which groups should be included in these comparisons, we decide to make all possible comparisons between the groups. In other words, we conduct unplanned comparisons in which each group is compared with each of the other groups. For the five
695
groups in our study, there are 10 possible unplanned comparisons: X ¯ 1 vs . X ¯ 2 X ¯ 1 vs . X ¯ 3 X ¯ 1 vs . X ¯ 4 X ¯ 1 vs . X ¯ 5 X ¯ 2 vs . X ¯ 3 X ¯ 2 vs . X ¯ 4 X ¯ 2 vs . X ¯ 5 X ¯ 3 vs . X ¯ 4 X ¯ 3 vs . X ¯ 5 X ¯ 4 vs . X ¯ 5
In this situation, making 10 comparisons increases the likelihood of familywise error (i.e., making at least one Type I error among these comparisons). In other words, the greater the number of comparisons made in a set of data, the greater the likelihood that groups will differ from each other as a function of random, chance factors rather than the hypothesized effect.
Methods of Controlling Familywise Error
A number of different methods may be used to control familywise error, a full discussion of which is beyond the scope of this book. (See Keppel and Wickens [2004] for a detailed description and comparison of a number of these methods.) For the purposes of our current discussion, it is important for you to understand two aspects of these methods. First, methods for controlling familywise error lower the probability of making a Type I error across a set of analytical comparisons by making it more difficult to reject the null hypothesis in each individual comparison. Second, the choice of a specific method depends on a variety of factors, such as the goals of the research study, the number and types of analytical comparisons to be conducted, and the extent to which the researcher wishes to reduce the probability of making a Type I error.
To illustrate several of the more commonly used methods for controlling familywise error, imagine that a researcher conducts a study consisting of four groups: three different types of treatments (Treatments I, II, and III) and a control group that does not receive any type of treatment (Control). Furthermore, the researcher is only interested in comparing each of the three treatments with the control group (Treatment I vs. Control, Treatment II vs. Control, Treatment III vs. Control). In a situation like this one, which involves comparing each group with a single reference group, a method known as the Dunnett test (Dunnett, 1955) would be the appropriate method for controlling familywise error.
Another type of research situation requiring the control of familywise error may involve combining groups before conducting analytical comparisons. For example, suppose a study consists of three groups (A, B, and C). The researcher may want to compare each group with each other (A vs. B, A vs. C, B vs. C). These are examples of what are called simple comparisons, defined as analytical comparisons between two groups. However, the researcher may also want to compare each group with the combined average of the other two groups (A vs. the average of B and C, B vs. the average of A and C, C vs. the average of A and B). These are examples of complex comparisons, which are analytical comparisons involving more than two groups. In this research situation, the Scheffé test (Scheffé, 1953)
696
may be used to control the probability of making a Type I error across this set of comparisons. The Scheffé test is defined as a statistical procedure used to control familywise error in situations in which a researcher conducts all possible simple and complex comparisons in a set of data.
Compared to these first two examples, the most typical research situation involving unplanned comparisons consists of comparing each group with each of the other groups— in other words, conducting all possible simple comparisons (but no complex comparisons) in a set of data. In this situation, the Tukey test (Tukey, 1953) is often used. The next section describes the Tukey test in more detail and illustrates how it can be used in a set of data to control familywise error.
Controlling Familywise Error: The Tukey Test
The Tukey test (Tukey, 1953), sometimes referred to as the Tukey Honestly Significant Difference (HSD) Test, is a statistical procedure used to limit the probability of familywise error in situations in which a researcher conducts all possible simple comparisons between groups. The Tukey test limits familywise error by increasing the critical value of the statistic needed to reject the null hypothesis for each comparison, thereby making it more difficult to reject the null hypothesis.
In terms of the steps to be followed in conducting unplanned comparisons using the Tukey test, it is important to note that they are (with one exception to be discussed momentarily) exactly the same as for planned comparisons:
state the null and alternative hypotheses (H0 and H1), make a decision about the null hypothesis, draw a conclusion from the analysis, and relate the result of the analysis to the research hypothesis.
The only difference between planned and unplanned comparisons occurs within the step, “Make a decision about the null hypothesis.” Rather than beginning this step with “Set alpha (α), identify the critical value, and state a decision rule,” we instead “Set the desired probability of familywise error (αFW), identify the critical value (FT), and state a decision rule.” This is illustrated below using the hypothetical study of five groups with Ni = 9 introduced earlier.
Set the Desired Probability of Familywise Error (αFW), Identify the Critical Value (FT), and State a Decision Rule
The first step in making a decision about the null hypothesis using the Tukey test is to set the probability of familywise error. Typically, this probability is set to .05, meaning there is
697
a combined .05 probability of making a Type I error across the set of unplanned comparisons. This probability, symbolized by αFW = .05, is very different from the probabilities calculated in planned comparisons, in which each individual comparison has a .05 probability of Type I error (α = .05).
Once αFW has been set, the critical value used to evaluate Fcomp (FT) is calculated using the following formula:
(11-12) F T = ( q T ) 2 2
where qT is a statistic known as the Studentized Range statistic (Table 5 in the back of this book). Looking at the Appendix, three pieces of information are needed to determine qT: k (the number of groups that make up the independent variable), dferror (dfWG in a one-way ANOVA), and αFW (the probability of familywise error).
In the hypothetical example described above (F(4, 40) = 3.52, p < .05), k is equal to 5, dferror is equal to 40, and αFW has been set at the traditional α = .05. Using Table 5, we identify qT for this example by locating the k = 5 column and then moving down this column until we reach the dferror = 40 row. For αFW = 05, we find qT is equal to 4.04. Entering qT = 4.04 into Formula 11–12: F T = ( q T ) 2 2 = ( 4.04 ) 2 2 = 8.16 = 16.32 2
Therefore, the critical value for each and all of the unplanned comparisons in this example (five groups and Nii = 9) is 8.16.
Once the critical value FT has been calculated, we can state a decision rule used to decide whether the result of an unplanned comparison is statistically significant. For the hypothetical example, the following decision rule may now be stated: I f F comp > 8.16 reject H 0 ; otherwise, do not reject H 0
To reject the null hypothesis for any of the 10 unplanned comparisons in this example, the value of Fcomp in the comparison must be greater than 8.16.
Once the decision rule is stated, we are ready to move onto the next two parts of the step, “Make a decision about the null hypothesis”, “Calculate a statistic: F-ratio for analytical comparison (Fcomp)” and “Make a decision whether to reject the null hypothesis.” First, using a procedure such as the Tukey test does not affect the value of the F-ratio for the analytical comparison—in fact, Fcomp is calculated exactly the same way regardless of whether an analytical comparison is planned or unplanned. Furthermore, making the decision whether to reject the null hypothesis is the same for both planned and unplanned comparisons: If Fcomp is greater than the critical value, the null hypothesis is rejected.
698
However, comparing the Tukey critical value (FT) with the critical value for planned comparisons highlights the consequence of conducting unplanned comparisons.
For the hypothetical example described earlier, the critical value for unplanned comparisons (FT) is 8.16. However, if these comparisons had been planned, the critical value for = .05 and df = 1, 40 would have been only 4.08. Rejecting the null hypothesis for unplanned comparisons requires a larger value of Fcomp than for planned comparisons. Using a method to control familywise error such as the Tukey test makes it more difficult to reject the null hypothesis, thereby lowering the probability of Type I error.
In summary, methods such as the Dunnett, Scheffé, and Tukey test are designed to reduce the possibility of concluding that an effect exists in the population when it does not. By making it more difficult to reject the null hypothesis, we take a conservative approach, one that minimizes the effects of chance factors that may exist in a particular sample of data.
699
Learning Check 4: Reviewing what you've Learned So Far
1. Review questions a. What is the purpose of conducting analytical comparisons? b. What are two types of analytical comparisons? In what research situations would you use
one versus the other? c. What is between-group and within-group variance in an analytical comparison? d. What are concerns researchers have about unplanned comparisons? e. What is familywise error? How can it be controlled?
2. Earlier in this chapter, a study was described examining males' concerns regarding oral contraceptives (Jaccard et al., 1981). The study hypothesized that males would rate risks to their health as more important to them than either the contraceptives effectiveness or convenience. Below is the ANOVA summary table for the study:
Source SS df MS F
Aspect of contraception 9.00 2 4.50 3.54
Error 30.56 24 1.27
Total 39.56 26
To test one of the study's research hypotheses, you conduct a planned comparison of Health risks (N = 9, M = 4.51) versus Effectiveness (N = 9, M = 3.18). Complete the following steps for this planned comparison.
a. State the null and alternative hypotheses (H0 and H1).
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate the F-ratio for an analytical comparison (Fcomp). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
6. Calculate a measure of effect size (R2). c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
3. Imagine that a researcher conducts a study consisting of five groups, each of which consists of seven participants (Ni = 7). The researcher conducts a one-way ANOVA and rejects the null hypothesis. Because her research addresses a new topic, the researcher decides to compare each group with each of the other groups, controlling familywise error using the Tukey test.
a. How many unplanned comparisons will be made in this situation? b. What is the value of qT? c. Assuming that the probability of familywise error is set at .05 (αFW = .05), what is the
critical value for each comparison (FT)?
700
11.6 Looking Ahead
This chapter illustrates both the advantages and disadvantages of increasing the complexity of a research study. Compared with the t-test discussed in Chapter 9, research studies analyzed with the one-way ANOVA and analytical comparisons are somewhat more difficult to conduct and analyze, but they provide the opportunity to test theories and research hypotheses in a more complex, comprehensive, and efficient way. The next chapter extends the discussion of comparing group means by discussing studies that have not one but two independent variables. As we will see, conducting this type of study not only allows researchers to examine more than one effect in the same study but also provides the opportunity to see how these effects may combine or interact with each other.
701
11.7 Summary
The analysis of variance (ANOVA) is a family of statistical procedures designed to test differences between two or more group means. It is called the “analysis” of “variance” because it involves analyzing different types of variance such as between-group and within- group variance.
The one-way ANOVA tests the differences between the means of two or more groups that comprise a single independent variable.
Testing differences between group means using a one-way ANOVA involves calculating a statistic known as the F-ratio, which is calculated by dividing between-group variance (MSBG) by within-group variance (MSWG). The theoretical distribution of F-ratios is positively skewed with a mean approximately equal to 1. To organize and summarize the calculations of a one-way ANOVA, researchers construct a table known as an ANOVA summary table.
Researchers often calculate a measure of effect size as an index or estimate of the size or magnitude of the hypothesized effect. Within the one-way ANOVA, one measure of effect size is the R2 statistic, which is the percentage of variance in the dependent variable associated with differences between the groups that comprise the independent variable.
Rejecting the null hypothesis in the one-way ANOVA does not indicate the specific nature and direction of the differences between the means of the groups. Therefore, to determine which of the groups are different or not different from one another, researchers conduct analytical comparisons, which are comparisons between groups that are part of a larger research design.
There are two types of analytical comparisons. Planned (a priori) comparisons are comparisons the researcher builds into the research design before data collection has begun; the purpose of these comparisons is to test research hypotheses. The second type, unplanned (post hoc) comparisons, are comparisons that have not been built into the research design prior to the data collection process; the decision to conduct these comparisons is made after the F-ratio for the one-way ANOVA has been calculated and the null hypothesis has been rejected. Although both types of analytical comparisons involve the calculation of an F-ratio (Fcomp), they differ in terms of how this F-ratio is evaluated to make the decision about the null hypothesis.
Two main concerns have been expressed about unplanned comparisons. First, because they are based on the results of statistical analyses rather than theory and hypotheses, the results and interpretation of unplanned comparisons may be a function of the characteristics and idiosyncrasies of the sample. Second, the more comparisons that are made within a set of
702
data, the greater the probability of making at least one Type I error (i.e., rejecting the null hypothesis when it is in fact true). Familywise error is the probability of committing a Type I error over a set of analyses.
Researchers use a number of different methods used to control familywise error, including the Dunnett test, the Scheffé test, and the Tukey test. These methods lower the probability of making a Type I error across a set of analytical comparisons by making it more difficult to reject the null hypothesis in each individual comparison. The choice of method depends on such factors as the goals of the research study, the number and types of analytical comparisons to be conducted, and the extent to which the researcher wishes to reduce the probability of making Type I errors across a set of comparisons.
703
11.8 Important Terms
analysis of variance (ANOVA) (p. 429) one-way ANOVA (p. 429) F-ratio (p. 429) ANOVA summary table (p. 439) R2 (p. 442) analytical comparisons (p. 453) planned (a priori) comparisons (p. 453) unplanned (post hoc) comparisons (p. 453) familywise error (p. 460) Dunnett test (p. 461) simple comparisons (p. 461) complex comparisons (p. 461) Scheffé test (p. 461) Tukey test (p. 462)
704
11.9 Formulas Introduced in this Chapter
Degrees of Freedom for Between-Group Variance, One-Way ANOVA (dfBG)
(11-1) d f B G = # groups-1
Degrees of Freedom for Within-Group Variance, One-Way ANOVA (dfWG)
(11-2) d f WG = Σ ( N i − 1 )
Between-Group Variance, One-Way ANOVA (MSBG)
(11-3) M S BG = S S BG d f BG = N i Σ ( X ¯ i − X ¯ T ) 2 # group − 1
Total Sample Mean, One-Way ANOVA ( X ¯ T )
(11-4) X ¯ T = Σ X ¯ i # group
Within-Group Variance, One-Way ANOVA (MSWG)
(11-5) M S WG = S S WG d f WG = ( N i − 1 ) Σ s i 2 Σ ( N i − 1 )
F-Ratio, One-Way ANOVA (F)
(11-6) F = M S BG M S WG
Measure of Effect Size, One-Way ANOVA (R2)
(11-7) R 2 = S S BG S S T
Relationship Between t-Test and F-Ratio for the One-Way ANOVA
(11-8) t 2 = F
F-Ratio, Analytical Comparison (Fcomp)
705
(11-9) F c o m p = M S B G c o m p M S W G
Between-Group Variance, Analytical Comparison ( M S BG comp )
(11-10) M S BG comp = N i ( X ¯ i − X ¯ j ) 2 2
Measure of Effect Size, Analytical Comparison ( R c o m p 2 )
(11-11) R c o m p 2 = S S BG comp S S T
Critical Value, Tukey Test for Unplanned Comparisons (FT)
(11-12) F T = ( q T ) 2 2
706
11.10 Using SPSS
707
One-Way Analysis of Variance (ANOVA): The Chess Study (11.1)
1. Define independent and dependent variables (name, # decimals, labels for the variables, labels for values of the independent variable) and enter data for the variables.
NOTE: Numerically code values of the independent variable (i.e., 1 = Observation, 2 = Prediction, 3 = Explanation) and provide labels for these values in Values box within Variable View.
2. Select the One-way ANOVA procedure within SPSS.
How? (1) Click Analyze menu, (2) click Compare Means, and (3) click One-Way ANOVA.
3. Identify the dependent variable and the independent variable, and ask for descriptive statistics.
How? (1) Click dependent variable and → Dependent List, (2) click independent variable and → Factor, (3) click Options and click Descriptive, (4) click Continue, and (5) click OK.
708
4. Examine output.
Analytical Comparisons: Planned Comparisons—Explanation vs. Prediction (11.5)
1. Within the one-way ANOVA procedure,
How? (1) Click Analyze menu, (2) click Compare Means, (3) click One-Way ANOVA, (4) click Contrasts, (5) assign 1 and −1 coefficients to the two groups being compared and assign 0 to group(s) not being compared, (6) click Continue, and (7) click OK.
NOTE: Because 1 = Observation, 2 = Prediction, and 3 = Explanation, the Observation group is assigned a coefficient = 0, the Prediction group is assigned a coefficient = 1, and the Explanation group is assigned a coefficient = −1.
709
2. Examine output.
710
11.11 Exercises
1. Researchers Latané and Rodin (1969) were interested in studying helping behavior. In one of their experiments, they set up a situation in which the participant was brought into the waiting room either by themselves or with two or four confederates. The researcher says she will be right back, but as she walks into the other room, she pretends to sprain her ankle. The number of seconds from her cry until the participant offers help is recorded as the dependent variable. The following data are representative of their findings:
Alone: 26, 25, 30, 20, 32
Two confederates: 30, 33, 29, 40, 36
Four confederates: 32, 39, 35, 41, 44 a. Calculate the sample size (Ni), mean ( X ¯ i ), and standard deviation (si) for
each group. b. Create a bar graph for the data with the number of confederates on the X-axis
and the mean number of seconds on the Y-axis.
2. For each of the following, calculate the degrees of freedom (df) and identify the critical value of F (assume α = .05).
a. # groups = 3, Ni = 10
b. # groups = 5, Ni = 26
c. # groups = 4, Ni = 15
d. There are five groups, with 20 participants in each group.
3. For each of the following, calculate the degrees of freedom (df) and identify the critical value of F (assume α = .05).
a. # groups = 3, Ni = 8 b. # groups = 6, Ni = 5 c. # groups = 4, Ni = 9 d. There are three groups, with five participants in each group.
4. For each of the following, calculate between-group variance (MSBG). a. Ni = 13, X ¯ 1 = 3.00 , X ¯ 2 = 5.00 , X ¯ 3 = 10.00 b. Ni = 11, X ¯ 1 = 19.00 , X ¯ 2 = 13.00 , X ¯ 3 = 21.00 , X ¯ 4 = 15.00 c. Ni = 15, X ¯ 1 = 1.13 , X ¯ 2 = 1.97 , X ¯ 3 = 2.60 , X ¯ 4 = 2.16 , X ¯ 5 = 3.82 d. Ni = 9, X ¯ 1 = 57.96 , X ¯ 2 = 64.68 , X ¯ 3 = 61.72 , X ¯ 4 = 68.53 , X ¯ 5 =
711
62.04 , X ¯ 6 = 56.09
5. For each of the following, calculate between-group variance (MSBG). a. Ni = 20, X ¯ 1 = 5.50 , X ¯ 2 = 3.75 , X ¯ 3 = 4.25 b. Ni = 8, X ¯ 1 = 14.85 , X ¯ 2 = 18.33 , X ¯ 3 = 19.92 , X ¯ 4 = 12.61 c. Ni = 9, X ¯ 1 = 40.81 , X ¯ 2 = 38.74 , X ¯ 3 = 33.29 , X ¯ 4 = 34.76 , X ¯ 5 =
41.15 d. Ni = 6, X ¯ 1 = 3.97 , X ¯ 2 = 4.56 , X ¯ 3 = 4.05 , X ¯ 4 = 3.31 , X ¯ 5 = 3.74 ,
X ¯ 6 = 2.91
6. For each of the following, calculate within-group variance (MSWG). a. Ni = 7, s1 = 1.00, s2 = 2.00, s3 = 6.00 b. Ni = 6, s1 = 13.00, s2 = 16.00, s3 = 12.00, s4 = 15.00 c. Ni = 9, s1 = 23.72, s2 = 18.50, s3 = 19.97, s4 = 16.39, s5 = 20.06 d. Ni = 23, s1 = 4.53, s2 = 4.18, s3 = 3.70, s4 = 3.46, s5 = 3.03, s6 = 2.71
7. For each of the following, calculate within-group variance (MSWG). a. Ni = 15, s1 = .75, s2 = .48, s3 = .63 b. Ni = 19, s1 = 4.52, s2 = 4.86, s3 = 4.28, s4 = 4.47 c. Ni = 22, s1 = 17.76, s2 = 14.92, s3 = 16.20, s4 = 20.01, s5 = 19.57 d. Ni = 12, s1 = 2.38 s2 = 1.91, s3 = 1.37, s4 = 3.06, s5 = 2.87, s6 = 2.54
8. For each of the following, calculate the F-ratio (F) and create an ANOVA summary table.
a. Ni = 20, X ¯ 1 = 20.00 , s1 = 8.00, X ¯ 2 = 25.00 , s2 = 10.00, X ¯ 3 = 15.00 , s3 = 9.00
b. Ni = 5, X ¯ 1 = 40.00 , s1 = 10.70, X ¯ 2 = 35.00 , s2 = 8.00, X ¯ 3 = 50.00 , s3 = 12.50, X ¯ 4 = 55.00 , s4 = 12.00
c. Ni = 11, X ¯ 1 = 7.00 , s1 = 2.50, X ¯ 2 = 12.00 , s2 = 4.00, X ¯ 3 = 16.00 , s3 = 5.50, X ¯ 4 = 9.00 , s4 = 3.75, X ¯ 5 = 11.00 , s5 = 4.00
9. For each of the following, calculate the F-ratio (F) and create an ANOVA summary table.
a. Ni = 10, X ¯ 1 = 14.00 , s1 = 2.50, X ¯ 2 = 12.50 , s2 = 2.10, X ¯ 3 = 17.00 , s3 = 3.00
b. Ni = 16, X ¯ 1 = 7.50 , s1 = 1.70, X ¯ 2 = 8.00 , s2 = 3.00, X ¯ 3 = 6.75 , s3 = 2.50, X ¯ 4 = 8.10 , s4 = 1.89
c. Ni = 14, X ¯ 1 = 65.24 , s1 = 10.43, X ¯ 2 = 71.52 , s2 = 9.86, X ¯ 3 = 75.06 , s3 = 12.79, X ¯ 4 = 69.62 , s4 = 11.94, X ¯ 5 = 61.91 , s5 = 11.43
712
10. Using the means and standard deviations calculated in Exercise 1 for the helping behavior study, determine if the number of seconds until the participant helps the person in distress is the same for the three experimental conditions.
a. State the null and alternative hypotheses (H0 and H1).
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate the F-ratio (F) and create an ANOVA summary table. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (R2).
c. Draw a conclusion from the analysis.
11. To evaluate the effectiveness of four different teaching methods, algebra student's are randomly assigned to one of four methods (A, B, C, or D) and then given a test. The descriptive statistics for each of the four methods are provided below. Determine whether there are any differences in the test scores of the four methods.
Method A: N1 = 20, X ¯ 1 = 76.00 , s1 = 10.80
Method B: N2 = 20, X ¯ 2 = 75.00 , s2 = 9.50
Method C: N3 = 20, X ¯ 3 = 85.00 , s3 = 11.00
Method D: N4 = 20, X ¯ 4 = 87.00 , s4 = 10.80 a. State the null and alternative hypotheses (H0 and H1).
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate the F-ratio (F) and create an ANOVA summary table. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (R2).
c. Draw a conclusion from the analysis.
12. Although men and women have been found to perform differently on tests of mental ability, less is known about possible reasons for these differences. A researcher hypothesizes that one's beliefs play a role, such that women who believe they should not perform well on these tests as men will in fact not perform well. To test this hypothesis, a group of women are given a written test of spatial abilities. Three different instructions were included with this test: (1) Women perform better on the
713
test than men, (2) men perform better on the test than women, or (3) women and men perform equally well on the test. It was hypothesized that women who were told that women perform better than men would score higher than women who were told men were better or that women and men were equal. Here are descriptive statistics of the test scores: Women better: N = 11, M = 6.91, s = 2.63; Men better: N = 11, M = 7.45, s = 2.07; Women and men equal: N = 11, M = 9.82, s = 2.23.
a. State the null and alternative hypotheses (H0 and H1).
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate the F-ratio (F) and create an ANOVA summary table. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (R2).
c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
13. Past research has examined relationships between people's backgrounds and their personality. For example, Eysenck (1982) found a relationship between blood type and introversion, and Pellegrini (1973) concluded that astrological signs were related to level of femininity. Gupta (1992) examined the relationship between season of birth and impulsivity (the lack of ability or desire to control one's behavior). Gupta hypothesized that people born in the seasons with more extreme temperatures (winter and summer) are more impulsive than those born in the more mild seasons (spring and autumn). She asked a sample of adults to report their season of birth and then administered a personality measure of impulsivity. Imagine the following descriptive statistics are reported for the test scores: Winter: N = 6, M = 8.31, s = 1.74; Spring: N = 6, M = 6.52, s = 1.87; Summer: N = 6, M = 9.09, s = 2.35; Fall: N = 6, M = 5.86, s = 1.95.
a. State the null and alternative hypotheses (H0 and H1).
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate the F-ratio (F) and create an ANOVA summary table. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (R2).
c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
714
14. Employers seek ways to improve the performance of their employees. Gardner, Van Dyne, and Pierce (2004) hypothesized that performance is influenced by organizational self-esteem, defined as an employee's evaluation of his or her personal adequacy as an organizational member. More specifically, they hypothesized that the higher one's organizational self-esteem, the higher will be one's job performance. The data below (representative of the study's findings) represent the employees' job performance level on a scale from 1 (low performance) to 5 (high performance). Determine whether job performance varies as a function of organizational self- esteem.
Organizational Self-Esteem Organizational Self-
Esteem
Low Medium High
3.0 4.0 5.0
2.5 3.5 4.0
3.0 4.0 4.5
3.0 3.0 4.0
2.0 4.0 5.0
3.5 4.0 4.5
a. Calculate the sample size (Ni), mean ( X ¯ i ), and standard deviation (si) for each group.
b. State the null and alternative hypotheses (H0 and H1).
c. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate the F-ratio (F) and create an ANOVA summary table. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (R2).
d. Draw a conclusion from the analysis. e. Relate the result of the analysis to the research hypothesis.
15. A researcher studying the effects of different learning techniques upon memory hypothesizes people learn better when training is distributed over time rather than massed all at once. She conducts an experiment in which 32 participants are randomly assigned to one of four conditions. Giving each subject several lessons to learn, she systematically varies the period of time between each lesson: 0 minutes
715
between lessons (a massed learning condition), 10 minutes, 20 minutes, or 30 minutes. After the participants receive all of the lessons, they complete a 25-item test measuring comprehension of the material. Their scores on the test are given below— conduct a one-way ANOVA on these data.
Time Between Lessons Time Between Lessons
0 min 10 min 20 min 30 min
4 10 14 21
5 11 12 18
11 9 8 16
4 6 5 13
9 4 15 20
3 5 11 14
7 4 7 12
5 5 6 10
a. Calculate the sample size (Ni), mean ( X ¯ i ), and standard deviation (si) for each group.
b. State the null and alternative hypotheses (H0 and H1).
c. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate the F-ratio (F) and create an ANOVA summary table. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (R2).
d. Draw a conclusion from the analysis. e. Relate the result of the analysis to the research hypothesis.
16. What is the relationship between the t-test and an ANOVA? In Chapter 9, we looked at two different types of instruction in reading comprehension (imagery and repetition). A t-test for independent means found a significant difference between the effectiveness of the two types, t(18) = 2.85, p < .05. To see the relationship between the t-test and the ANOVA, calculate the F-ratio for the one-way ANOVA for the same data and compare the results.
Imagery: N = 10, X ¯ 1 = 12.10 , s1 = 1.60
716
Repetition: N = 10, X ¯ 2 = 9.90 , s2 = 1.85 a. State the null and alternative hypotheses (H0 and H1).
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate the F-ratio (F) and create an ANOVA summary table. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (R2).
c. Draw a conclusion from the analysis. d. Compare your value of the F-ratio with the t-test value of 2.85 to confirm their
mathematical equivalence.
17. An important issue in court cases is the accuracy of eyewitness testimony. Behavioral scientists have suggested eyewitnesses can be influenced by how a question is phrased. A researcher conducts a study where 15 people watch a film of an accident in which Car A runs a stop sign and hits Car B at a speed of 20 miles per hour. After watching the film, she asks each of the 15 people to estimate Car A's speed at the moment of impact (they do not know the actual speed). However, the question is phrased three different ways. Five people are asked, “How fast was Car A going at the time of the accident with Car B?” Another five are asked, “How fast was Car A going when it hit Car B?” The last five are asked, “How fast was Car A going when it smashed into Car B?” The researcher hypothesizes estimates of speed will vary as a function of the wording of the question, with more extreme wordings leading to higher estimates.
Wording of Question Wording of Question
“Accident” “Hit” “Smashed”
18 27 28
23 22 24
17 24 30
19 19 22
21 25 27
a. Calculate the sample size (Ni), mean ( X ¯ i ), and standard deviation (si) for each group.
b. State the null and alternative hypotheses (H0 and H1).
c. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df).
717
2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate the F-ratio (F) and create an ANOVA summary table. 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (R2).
d. Draw a conclusion from the analysis.
18. Looking back at Exercise 17, the most specific conclusion that could be drawn from rejecting the null hypothesis was that the estimates of speed for the three types of wording were not all equal to each other. However, the researcher hypothesized that more extreme wordings lead to higher estimates of speed such that participants would report a higher speed when the word “smashed” was used rather than “accident”. Conduct the appropriate analytical comparison to test this specific hypothesis.
a. State the null and alternative hypotheses (H0 and H1).
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate the F-ratio for an analytical comparison (Fcomp). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size ( R c o m p 2 )
c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
19. The hypothetical study in Exercise 15 examined the effects of different learning techniques upon memory. Given the hypothesis is that people learn better when training is distributed over time rather than massed all at once, conduct an analytical comparison testing the difference in comprehension test scores for the 0-minute versus 10-minute conditions.
a. State the null and alternative hypotheses (H0 and H1).
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate the F-ratio for an analytical comparison (Fcomp). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size ( R c o m p 2 )
c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
718
20. The chess study discussed in this chapter (de Bruin et al., 2007) hypothesized that the number of checkmates should be greater for students asked to make predictions while learning to play chess than students who simply observe games of chess being played. Using the ANOVA summary table in Table 11.4(c), conduct an analytical comparison of the number of checkmates for the Prediction (N = 14, M = .86) versus Observation (N = 14, M = 1.36) conditions.
a. State the null and alternative hypotheses (H0 and H1).
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate the F-ratio for an analytical comparison (Fcomp). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size ( R c o m p 2 ).
c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
21. An experimenter ran a study that consisted of four groups, each with 16 participants. This research study was exploratory research, and the experimenter did not have any specific hypotheses or expectations about the results of the analyses. He found a significant F-ratio for the one-way ANOVA and now would like to compare each group to each of the other groups to locate the differences.
a. How many unplanned comparisons will he need to conduct? b. What are the concerns about conducting so many comparisons? c. What is the value of qT? d. Assuming familywise error is set at the traditional value of .05 (αFW = .05),
what is the critical value for each of the unplanned comparisons (FT)?
22. Imagine the F-ratio for a one-way ANOVA is rejected for a study involving six groups, each with nine participants. A researcher plans on comparing each group with every other group, controlling familywise error using the Tukey test.
a. How many unplanned comparisons will be conducted? b. What is the value of dferror? c. What is the value of qT? d. Assuming familywise error is set at .05 (αFW = .05), what is the critical value
for each comparison (FT)?
719
Answers to Learning Checks
Learning Check 2
2.
a. dfBG = 3, dfWG = 40, critical value = 2.84
b. dfBG = 4, dfWG = 25, critical value = 2.76
c. dfBG = 5, dfWG = 42, critical value = 2.45
3.
a. X ¯ T = 5.00 , SSBG = 234.00, dfBG = 2, MSBG = 117.00
b . X ¯ T = 34.00 , SSBG = 840.00, dfBG = 3, MSBG = 280.00
c. X ¯ T = 4.00 , SSBG = 20.37, dfBG = 4, MSBG = 5.09
4.
a. SSWG = 765.00, dfWG = 51, MSWG = 15.00
b. SSWG = 1640.00, dfWG = 16, MSWG = 102.50
c. SSWG = 8.20, dfWG = 50, MSWG =.16
5. a.
Source SS df MS F
Between-group 6.00 2 3.00 2.56
Within-group 38.72 33 1.17
Total 44.72 35
b.
Source SS df MS F
Between-group 4.00 3 1.33 5.78
Within-group 6.44 28 .23
Total 10.44 31
720
Learning Check 3
2. a. H0: all μs are equal; H1: not all μs are equal b.
1. dfBG = 2, dfWG = 24
2. If F > 3.40, reject H0; otherwise, do not reject H0
Source SS df MS F
Aspect of contraception 9.00 2 4.50 3.54
Error 30.56 24 1.27
Total 39.56 26
3. F = 3.54 > 3.40 ∴ reject H0 (p < .05) 4. F = 3.54 < 5.61 ∴ p < .05 (but not < .01) 5. R2 = .23
c. The mean importance ratings of three aspects of oral contraceptives in samples of nine males (Effectiveness, M = 3.18; Convenience, M = 3.43; Health risks, M = 4.51) were not all equal to each other, F(2, 24) = 3.54, p < .05, R2 = .23.
d. We are unable to determine whether the research hypothesis in the study is supported until analytical comparisons are conducted.
Learning Check 4
2. a. H: μHealth risks = μConvenience; H: μHealth risks? μConvenience b.
1. dfBGcomp = 1, dfWG = 24 2. If Fcomp > 4.26, reject H0; otherwise, do not reject H0 3. MSBGcomp = 7.96; MSWG = 1.27; Fcomp = 6.27 4. Fcomp = 6.27 > 4.26 ∴ reject H0 (p < .05)
721
5. Fcomp = 6.27 < 7.82 ∴ p < .05 (but not < .01) 6. R c o m p 2 = 2.0
c. In rating the importance of different aspects of oral contraceptives, the samples of nine males rated health risks (M = 4.51) as being significantly more important than its effectiveness (M = 3.18), F(1, 24) = 6.27, p < .05, R2 = .20.
d. The result of this analytical comparison supports the research hypothesis that males would rate risks to their health as more important to them than the contraceptive's effectiveness.
3. a. Ten comparisons b. dferror = 30; qT = 4.10 c. FT = 8.41
722
Answers to Odd-Numbered Exercises
1. a.
Alone: N = 5, X ¯ = 26.60 , s = 4.67
Two confederates: N = 5, X ¯ = 33.60 , s = 4.51
Four confederates: N = 5, X ¯ = 38.20 , s = 4.76
b.
3. a. dfBG = 2, dfWG = 21, critical value = 3.47 b. dfBG = 5, dfWG = 24, critical value = 2.84 c. dfBG = 3, dfWG = 32, critical value = 2.92 d. dfBG = 2, dfWG = 12, critical value = 3.89
5. a.
X ¯ T = 4.50 , SSBG = 32.40, dfBG = 2, MSBG = 16.20
b.
X ¯ T = 16.43 , SSBG = 263.04, dfBG = 3, MSBG = 87.68
X ¯ T = 37.75 , SSBG = 456.57, dfBG = 4, MSBG = 114.14
X ¯ T = 3.76 , SSBG = 10.08, dfBG = 5, MSBG = 2.02
7.
a. SSWG = 16.66, dfWG = 42, MSWG = .40
b. SSWG = 1482.30, dfWG = 72, MSWG = 20.59
c. SSWG = 33260.85, dfWG = 105, MSWG = 316.77
723
d. SSWG = 387.64, dfWG = 66, MSWG = 5.87
9. a.
Source SS df MS F
Between-group 105.00 2 52.50 8.02
Within-group 176.94 27 6.55
Total 281.94 29
b.
Source SS df MS F
Between-group 18.40 3 6.13 1.13
Within-group 325.65 60 5.43
Total 344.05 63
c.
Source SS df MS F
Between-group 1502.34 4 375.59 2.92
Within-group 8356.14 65 128.56
Total 9858.48 69
11. a.
H0: all μs are equal; H1: not all μs are equal
b. 1. dfBG = 3, dfWG = 76
2. If F > 2.76, reject H0; otherwise, do not reject H0
724
Source SS df MS F
Teaching method 2254.80 3 751.60 6.76
Error 8446.07 76 111.13
Total 10700.87 79
3. F = 6.76 > 2.76 ∴ reject H0 (p < .05) 4. F = 6.76 > 4.13 ∴ p < .01 5. R2 = .21
c. This analysis found that the mean examination scores for the 20 students in the four different teaching methods (Method A, M = 76.00; Method B, M = 75.00; Method C, M = 85.00; and Method D, M = 87.00) were not all equal, F(3, 76) = 6.76, p < .01, R2 = .21.
13. a.
H0: all μs are equal; H1: not all μs are equal
b. 1. dfBG = 3, dfWG = 20 2. If F > 3.10, reject H0; otherwise, do not reject H0
3.
Source SS df MS F
Season of birth 40.92 3 13.64 3.44
725
Error 79.25 20 3.96
Total 120.17 23
4. F = 3.44 > 3.10 ∴ reject H0 (p < .05) 5. F = 3.44 < 4.94 ∴ p < .05 (but not < .01) 6. R2 = .34
c. The mean levels of impulsivity in samples of six people in the four seasons of birth (winter, M = 8.31; spring, M = 6.52; summer, M = 9.09; fall, M = 5.86) were not all equal to each other, F(3, 20) = 3.44, p < .05, R2 = .34.
d. We are unable to determine whether the research hypothesis in the study is supported until analytical comparisons are conducted.
15. a.
0 minutes: N1 = 8, X ¯ 1 = 6.00 , s1 = 2.78
10 minutes: N2 = 8, X ¯ 2 = 6.75 , s2 = 2.82
20 minutes: N3 = 8, X ¯ 3 = 9.75 , s3 = 3.77
30 minutes: N4 = 8, X ¯ 4 = 15.50 , s4 = 3.93
b.
H0: all μs are equal; H1: not all μs are equal
c. 1. dfBG = 3, dfWG = 28 2. If F > 2.95, reject H0; otherwise, do not reject H0 3.
Source SS df MS F
Time between lessons 446.96 3 148.99 13.15
Error 317.31 28 11.33
Total 764.37 31
4. F = 13.15 > 2.95 ∴ reject H0 (p < .05) 5. F = 13.15 > 4.57 ∴ p < .01 6. R2 = .58
d. The mean comprehension test scores of eight employees with different amounts of time between lessons (0 minutes, M = 6.00; 10 minutes, M = 6.75; 20 minutes, M = 9.75; 30 minutes, M = 15.50) were not all equal to each
726
other, F(3, 28) = 13.15, p < .01, R2 = .58. e. We are unable to determine whether the research hypothesis in the study is
supported until analytical comparisons are conducted. 17.
a.
“Accident”: N = 5, X ¯ = 19.60 , s = 2.41
“Hit”: N = 5, X ¯ = 23.40 , s = 3.05
“Smashed”: N = 5, X ¯ = 26.20 , s = 3.19
b. H0: all μs are equal; H1: not all μs are equal c.
1. dfBG = 2, dfWG = 12 2. If F > 3.88, reject H0; otherwise, do not reject H0 3.
Source SS df MS F
Wording 109.75 2 54.88 6.51
Error 101.16 12 8.43
Total 210.91 14
4. F = 6.51 > 3.88 ∴ reject H0 (p < .05) 5. F = 6.51 < 6.93 ∴ p < .05 (but not < .01) 6. R2 = .52
d. This analysis reveals that the estimate of the car's speed at the time of the accident for the five people exposed to different wording of the question (accident, M = 19.60; hit, M = 23.40; smashed, M = 26.20) were not all equal to each other, F(2, 12) = 6.51, p < .05, R2 = .52.
19. a. H0: μ0 min = μ10 min; H1: μ0 min ≠ μ10 min b.
1. d f B G c o m p , dfWG = 28 2. If Fcomp > 4.20, reject H0; otherwise, do not reject H0 3. M S B G c o m p = 2.25 ; MSWG = 11.33; Fcomp = .20 4. Fcomp = .20 < 4.20 ∴ do not reject H0 (p > .05) 5. Not applicable (H0 not rejected) 6. R c o m p 2 = .01
c. The mean comprehension test scores of eight employees with 0 minutes between lessons (M = 6.00) and 10 minutes between lessons (M = 6.75) did
727
not significantly differ, F(1, 28) = .20, p > .05, R2 = .003. d. The results of this analytical comparison do not support the research
hypothesis that people learn better when training is distributed over time rather than massed all at once.
21. a. Six comparisons b. Familywise error (the probability of making at least one Type I error across a
set of comparisons) increases as the number of comparisons made increases. c. qT = 3.74 d. FT = 6.99
728
Sharpen your skills with SAGE edge at edge.sagepub.com/tokunaga
SAGE edge for students provides a personalized approach to help you accomplish your coursework goals in an easy-to-use learning environment. Log on to access:
eFlashcards Web Quizzes Chapter Outlines Learning Objectives Media Links SPSS Data Files
729
Chapter 12 Two-Way Analysis of Variance (ANOVA)
730
Chapter Outline 12.1 An Example From the Research: Vote—or Else! 12.2 Introduction to Factorial Research Designs
Testing interaction effects An example of interaction effect An example of no interaction effect
Testing main effects The relationship between main effects and interaction effects Factorial designs vs. single-factor designs: advantages and disadvantages
12.3 The Two-Factor (A × B) Research Design Notational system for components of the two-factor research design Analyzing the two-factor research design: Descriptive statistics
Notational system for means in the two-factor research design Examining the descriptive statistics
12.4 Introduction to Analysis of Variance (ANOVA) for the Two-Factor Research Design F-ratios in the two-factor research design
12.5 Inferential Statistics: Two-Way Analysis of Variance (ANOVA) State the null and alternative hypotheses (H0 and H1)
Main effect of Factor A Main effect of Factor B A × B interaction effect
Make decisions about the null hypotheses Calculate the degrees of freedom (df) Set alpha (α), identify the critical values, and state decision rules Calculate statistics: F-ratios for the two-way ANOVA Make decisions whether to reject the null hypotheses Determine the levels of significance
Calculate measures of effect size (R2) Draw conclusions from the analyses Relate the results of the analyses to the research hypothesis
12.6 Investigating a Significant A × B Interaction Effect: Analysis of Simple Effects 12.7 Looking Ahead 12.8 Summary 12.9 Important Terms 12.10 Formulas Introduced in This Chapter 12.11 Using SPSS 12.12 Exercises
Several chapters in this book have discussed statistical procedures used to test differences between groups. In Chapter 9, we compared two groups of drivers (Intruder and No intruder) in terms of how long they took to leave a shopping mall parking space. Chapter 11 included an example comparing the effectiveness of three strategies (Observation, Prediction, and Explanation) used to develop a principled understanding of chess. These research studies, which consisted of a single independent variable, provided useful information regarding a research topic. However, studying the effects of independent variables one at a time has limitations of a conceptual, statistical, and pragmatic nature. In
731
this chapter, we will introduce the analysis of studies consisting of two independent variables.
732
12.1 An Example from the Research: Vote—or Else!
Throughout your life, people have attempted to influence your behavior. When you were younger, your parents probably tried to get you in the habit of regularly brushing your teeth. If so, what type of strategy did they employ? Did they tell you what a nice smile you would have if you brushed regularly? Or did they frighten you with horror stories about cavities and torturous visits to the dentist? A group of researchers headed by Howard Lavine of the State University of New York at Stony Brook made the following observation about the role of fear in modifying behavior: “The use of fear appeals has long played a central role in attempts to change and shape attitudes by means of persuasive messages” (Lavine et al., 1999, p. 337).
Lavine and his fellow researchers were interested in assessing the effectiveness of two types of messages for persuading people to vote in elections. The first type of message they examined emphasized rewards one may receive or feel by voting, whereas the second type of message focused on threatening personal or political consequences for failing to vote. But rather than simply proposing that one type of message is more or less persuasive than the other, the researchers instead hypothesized that the effectiveness of a particular type of message depends on characteristics of the person receiving the message. For some voters, they suggested, a rewarding message might be more effective than a threatening message, whereas for other types of voters, a threatening message may be more persuasive.
The characteristic of voters that the researchers chose to examine was authoritarianism, defined as the extent to which people “perceive a great deal of threat from their immediate and broader environment … [and are] predisposed to view the world as dangerous” (Lavine et al., 1999, p. 338). The researchers hypothesized that the effectiveness of a particular type of message depends on the person's level of authoritarianism. As they wrote,
Our key prediction was that … high authoritarian recipients would perceive the threat message as more valid and persuasive than the reward message. Because high authoritarians perceive the world as a dangerous and threatening place, they should resonate to a message highlighting the potential negative consequences of not voting. … We also explored the possibility
that a threat-related persuasive message would lead to less persuasion among low authoritarian recipients, relative to the reward message. (Lavine et al., 1999, p. 340)
The research hypothesis in the study was that the effectiveness of the type of message depends on the person's level of authoritarianism. More specifically, the researchers
733
maintained that a threatening message will be more persuasive than a message emphasizing rewards for people with a high level of authoritarianism, but a message focusing on rewards will be more effective than a threatening message for people with low authoritarianism.
The sample in the study consisted of “86 voting-eligible (i.e., > 17 years old and American citizens) students (n = 34 men; n = 52 women) at the University of Minnesota, who participated in the study for extra credit in their introductory psychology course” (Lavine et al., 1999, p. 341). The critical criterion for inclusion in the study was that the person be old enough to vote.
In conducting the study, the first independent variable was the type of message (rewarding or threatening) used to persuade people to vote. Research participants in the study were randomly assigned to receive one of two booklets. One version of the booklet contained statements emphasizing rewards gained from voting, such as “Voting allows one to play an active role in the direction taken by your government.” The other version of the booklet described threats or punishments that may occur from failing to vote, such as “Not voting allows others to take away your right to express your values.”
The second independent variable was the person's level of authoritarianism. To measure this variable, the researchers asked participants to complete a 10-item scale that included statements designed to gauge their level of authoritarianism, such as “Our country will be destroyed someday if we do not smash the perversions eating away at our moral fiber and traditional beliefs.” On the basis of their responses, participants were classified as being either “low” or “high” in authoritarianism; the more the person agreed with the statements, the more they were considered to be “high” in authoritarianism.
The dependent variable in the study was the perceived quality of the message conveyed by each type of booklet. To gauge perceptions of quality, researchers asked the participants to complete a 12-item questionnaire in which they indicated their level of agreement with a range of statements related to the booklet's quality and effectiveness, such as “I found the material to be convincing” and “The material did not contain persuasive arguments.” The average rating across the 12 items was calculated; the higher the average rating, the higher was the perceived message quality.
In summary, the study contained two independent variables: Message type, which consisted of two types (Reward and Threat), and Authoritarianism, which comprised two levels (Low and High). The dependent variable in the study, Message quality, was a continuous variable ranging from 1.00 to 7.00. This study is the first one in this book that includes more than one independent variable. The characteristics of this type of research design, known as a factorial research design, are discussed in greater detail in the next section.
734
12.2 Introduction to Factorial Research Designs
In Chapters 9 and 11, each research study was an example of a single-factor research design, which is a research design consisting of one independent variable. In a single-factor research design, each research participant is associated with one level of an independent variable. For example, in the parking lot study (Chapter 9), drivers were randomly assigned to either the Intruder or No intruder group; in the chess study (Chapter 11), students were assigned to one of three strategies (Observation, Prediction, or Explanation).
Unlike previous chapters, the study in this chapter, which will be referred to as the voting message study, contains two independent variables: Message type and Authoritarianism. Consequently, each participant belongs to one combination of the two independent variables: Reward/Low, Reward/High, Threat/Low, or Threat/High. Combining the levels of the two independent variables, the voting message study may be represented the following way:
Message type
Authoritarianism
Reward Threat
Low
High
The voting message study is an example of a factorial research design, which is a research design consisting of all possible combinations of two or more independent variables. This section discusses two critical features of factorial research designs: the ability to test interaction effects and the ability to test main effects.
735
Testing Interaction Effects
A critical reason for including two or more independent variables in the same study is the opportunity to determine whether these variables may combine, or interact, to influence the dependent variable. That is, factorial research designs allow researchers to examine whether an interaction effect exists between variables. An interaction effect occurs when the effect of one independent variable on the dependent variable changes at the different levels of another independent variable. Put another way, an interaction effect exists when the effect of one independent variable depends on the level of another independent variable.
In the voting message study, the researchers hypothesized an interaction effect between message type and authoritarianism such that the effect of message type (reward vs. threat) on perceived message quality depends on the participant's level of authoritarianism (low vs. high). For people with a low level of authoritarianism, the reward message was hypothesized to be more persuasive than was the threat message; however, the threat message was predicted to be more persuasive than the reward message for people with a high level of authoritarianism.
The ability to examine interaction effects between independent variables is not attainable from single-factor designs. For this reason, identifying interaction effects represents a unique characteristic of factorial research designs and an important advantage for researchers wishing to examine more than one factor in their analyses. Below, a hypothetical research study is used to illustrate the presence and absence of an interaction effect.
An Example of Interaction Effect
To further understand the concept of interaction effects, suppose you were interested in studying the effect of physical exercise on one's mood. More specifically, for people who exercise on a regular basis, what happens to their psychological well-being when they are unable to work out? In a study headed by Gregory Mondin at the University of Wisconsin, the researchers asked a group of regular exercisers to refrain from exercising for a 5-day period (Mondin et al., 1996). Dr. Mondin and his associates found that the level of anxiety in the research subjects increased from the first day (Monday) to the third day (Wednesday), but their anxiety dropped from Wednesday to Friday as the period of exercise deprivation ended.
Let's say you wonder whether the change in anxiety level is the same for all regular exercisers. For example, your examination of research might suggest that the gender of the exerciser should be taken into account. It could be hypothesized, for example, that women
736
are better able to maintain a steady mood state during periods of exercise deprivation than are men, whose level of anxiety rises and falls more sharply during the same period. On the basis of such an assumption, you might hypothesize that the effect of exercise deprivation on anxiety level depends on the gender of the exerciser. In this case, you would be hypothesizing an interaction effect between Day and Gender.
A table of hypothetical means for this study is provided in Figure 12.1(a). Looking at this figure, we see that combining three levels of Day (Monday, Wednesday, Friday) with two levels of gender (Male, Female) results in six possible combinations of the two variables. We will use the term cell to refer to a combination of independent variables; the exercise deprivation study therefore consists of a total of six cells. The numbers in the cells of this table are cell means, defined as the mean of the dependent variable for a particular combination of independent variables. For example, in Figure 12.1(a), the upper-left cell mean of 2.20 is the mean anxiety level (measured on a 1 [low] to 5 [high] scale) on Monday for the Male exercisers.
Figure 12.1(b) is a line graph of the cell means for the six Day × Gender combinations. The cell means have been portrayed using a line graph because the independent variable “Day,” located along the horizontal (X) axis, is a continuous variable with the same distance (2 days) between the three values (Monday, Wednesday, Friday). In the line graph, the means for the three levels of Day are plotted separately for each level of Gender (Male, Female), which is the other independent variable. If Day had not been a continuous variable, it would not be appropriate to “connect the dots” using a line graph, and we would have instead used a bar chart to represent the cell means.
Looking at the two lines in Figure 12.1(b), we see that changes in anxiety level during the period of exercise deprivation are not the same for the two genders. For women, the level of anxiety changes only slightly during this period. The level of anxiety for men, however, rises sharply from Monday to Wednesday before dropping back down on Friday. This hypothetical study is an illustration of an interaction effect. This is because the effect of one independent variable (Day) on the dependent variable (Anxiety level) is not the same at the different levels of the other independent variable (Gender). That is, the effect of Day depends on the level of Gender.
Figure 12.2 presents other examples of possible interaction effects of Day and Gender. As we can see from these figures, there are a wide variety of possible interaction effects between variables. Graphs such as those in Figures 12.1 and 12.2 are extremely useful in inspecting data for the presence of interaction effects. However, the simple fact that the lines in a graph may not be the same does not necessarily mean there is a statistically significant interaction effect. Because of chance factors that operate in every study, we cannot make a definitive statement about an interaction effect until we have conducted the appropriate statistical analysis.
737
Figure 12.1 Example of an Interaction Effect (Day by Gender).
Figure 12.2 Other Examples of an Interaction Effect (Day by Gender)
An Example of No Interaction Effect
The hypothetical data for the exercise data provided in Figure 12.3 illustrate a research situation where no interaction effect is present between two variables. Looking at the cell means and the line graph, we observe that, even though the two genders differ in their anxiety levels, the changes in anxiety level across days are the same for men as for women. For example, the level of anxiety for the male participants increased from 2.30 to 3.10 from Monday to Wednesday (a difference of .80 [3.10 – 2.30]) and decreased from 3.10 to 2.20 from Wednesday to Friday (a difference of .90 [3.10 – 2.20]). Similar changes were found for the female participants (2.90 – 2.10 = .80 and 2.90 – 2.00 = .90).
Unlike the research situations in Figure 12.2, the lines of the two groups in Figure 12.3 are parallel. The parallel lines reflect the fact that the effect of the independent variable (Day) is the same at the different levels of the other independent variable (Gender). When the effect of one variable does not depend on the levels of the other variable, there is no interaction effect between variables.
Figure 12.4 presents other examples of two variables that do not interact. Similar to the
738
presence of interaction, the absence of an interaction effect can take a wide variety of shapes. As we mentioned earlier, whether or not an interaction effect is present or absent is ultimately determined by conducting statistical analyses rather than by simply examining a table or figure.
Figure 12.3 Example of No Interaction Effect (Day by Gender)
Figure 12.4 Other Examples of No Interaction Effect (Day by Gender)
739
Testing Main Effects
In addition to the ability to determine whether interaction effects exist between independent variables, a second critical aspect of factorial research designs is the ability to examine main effects, defined as the effect of an independent variable on the dependent variable within a factorial research design; the main effect of an independent variable is separate from the effects of the other independent variables. In the exercise deprivation study, for example, we can determine whether anxiety levels change as a function of Day, regardless of the gender of the research participant. We can also test whether men and women differ in their level of anxiety, regardless of the day of the week in which they do not exercise. These effects of Day and Gender, distinct from any interaction effect between them, are known as main effects.
Based on the table of cell means in Figure 12.1(a), Table 12.1 illustrates two main effects for the exercise deprivation study. The boldfaced means in the margins of the table to the right of and below the cell means represent the two main effects of Day and Gender, respectively. These means are called marginal means, defined as the means in a table that represent a main effect.
When each cell is based on the same number of scores, the marginal means for an independent variable are calculated by averaging the results over the different levels of the other independent variable. For example, assuming the exercise deprivation study contained an equal number of males and females, the marginal means for the main effect of Day are calculated by computing the mean of the males and females for each of the 3 days: X ¯ Monday = 2.20 + 2.00 2 = 4.20 2 = 2.10 X ¯ Wednesday = 3.60 + 2.40 2 = 6.00 2 = 3.00 X ¯ Friday = 2.40 + 2.10 2 = 4.50 2 = 2.25
By combining the means of the two genders for each day, we are in fact acting as though the independent variable Gender does not exist.
The marginal means for the main effect of Gender are determined by calculating the average anxiety level across the 3 days separately for the male and female exercisers:
Table 12.1 Examples of Main Effects (Day and Gender) Table 12.1 Examples of Main Effects (Day and Gender)
Day
Monday Wednesday Friday
Gender Male 2.20 3.60 2.40 2.73
Female 2.00 2.40 2.10 2.17
2.10 3.00 2.25
740
X ¯ Male = 2.20 + 3.60 + 2.40 3 = 8.20 3 = 2.73 X ¯ Female = 2.00 + 2.40 + 2.10 3 = 6.50 3 = 2.17
It is important to understand that main effects are equivalent to the information that would be obtained by conducting separate single-factor research designs for each of the independent variables. For example, the three marginal means of 2.10, 3.00, and 2.25 constitute a single-factor research design in which we can determine whether anxiety levels differ as a function of Day. Similarly, the marginal means for males (M = 2.73) and females (M = 2.17) represent a single-factor research design for Gender. Therefore, rather than having to conduct two separate studies with two different samples to determine the effects of Day and Gender on anxiety level, in a factorial research design both effects can be tested within the same study.
741
The Relationship between Main Effects and Interaction Effects
Factorial research designs contain both main effects and interaction effects, each of which may either be present or absent in any given research study. This section discusses two aspects of the relationship between main effects and interaction effects: (1) The presence or absence of main effects provides no indication of whether an interaction effect is present or absent, and vice versa, and (2) whether and how one interprets main effects depends on the presence or absence of interaction effects.
To illustrate these two aspects of the relationship between main effects and interaction effects, Figure 12.5 provides hypothetical data for eight different research designs. In a research design with two independent variables, there are eight possible combinations of present or absent main effects and interaction effect. In this figure, we will assume that any difference between means represents the presence of an effect; in reality, however, we would need to conduct statistical analyses on the data to determine whether this is true.
The top four line graphs (Figures 12.5(a–d)) all have an absent interaction effect, which is represented by two parallel lines. However, looking at the tables of means, we see these four scenarios differ in whether either or both of the main effects are present. For example, in Figure 12.5(a), neither of the two main effects is present; in Figure 12.5(d), both main effects are present.
In the bottom four line graphs (Figures 12.5(e–h)), in which an interaction effect is present, it is possible to have the same four pairings of present or absent main effects as in Figures 12.5(a–d). As a whole, these eight scenarios illustrate a critical aspect of the relationship between main effects and interaction effects:
The presence or absence of main effects provides no indication of whether an interaction effect is either present or absent, and vice versa.
That is to say, whether or not a particular effect is present has no direct influence on the other effects.
The second aspect of the relationship between main effects and interaction effects pertains to the order in which they are examined. Even though main effects and interaction effects are independent of each other, there is typically a certain order to be followed in interpreting the results of a factorial research design:
Figure 12.5 Combinations of Absent and Present Main Effects and Interaction Effect
742
Whether and how one interprets main effects depends on the presence or absence of the interaction effect.
If an interaction effect is present, this implies that the effects of one independent variable depend on the level of the other independent variable. As a result, we cannot assess the effect of one independent variable in isolation of the other but must rather consider both of them simultaneously. Even if one or both of the main effects are present, the interaction effect takes precedence; attention is focused on the cell means to understand the source and nature of the interaction.
743
Let's assume that there was an interaction effect between Day and Gender in the exercise deprivation study. An appropriate answer to the question, “Does anxiety level change across days?” would be, “It depends on the exerciser's gender.” Similarly, the question, “What is the difference between male and female exercisers in their anxiety levels?” would be answered by saying, “The difference depends on the day of exercise deprivation.”
But what if an interaction effect is not present? Because this implies that the effect of one variable remains the same across the different levels of the other variable, we would not need to consider both factors at the same time. Instead, we could think of the two factors as being distinct from each other, allowing us to shift our focus from the interaction effect to the main effects. In essence, we may now treat this design as if it were two separate single- factor designs, one for each main effect. When an interaction effect is not present, we turn our attention to the marginal means that represent the main effects.
Figure 12.6 illustrates the basic process for considering and analyzing factorial research designs. Note that the first question addresses the presence of an interaction effect. If an interaction effect is present, attention turns to the cell means to determine the nature of the interaction effect. If, on the other hand, the interaction effect is not present, we focus instead on the main effects. If a main effect is present, the marginal means are examined to determine which groups that comprise an independent variable differ from each other. If neither of the main effects is present, the analysis of the design stops; no other comparisons are necessary. Keep in mind, however, that this is a basic strategy; the particular demands and research hypotheses contained in a study must always be taken into account.
744
Factorial Designs vs. Single-Factor Designs: Advantages and Disadvantages
Compared with single-factor designs, factorial research designs have several distinct advantages of both a conceptual and pragmatic nature. Conceptually, factorial research designs allow researchers to determine whether an interaction effect exists between independent variables. Pragmatically, these designs provide researchers with the opportunity to examine multiple effects in a single study. Each of these advantages is discussed below.
A critical conceptual benefit of including two or more independent variables in the same study is the opportunity to determine whether these variables may combine or interact to influence the dependent variable. This is particularly important because although a variable may have an effect on another variable, this effect rarely occurs in isolation. In the real world, it is typically the case that a variable operates at the same time as other variables. Consequently, the most appropriate way to represent the complexity of phenomena studied by researchers is to include multiple variables in the same study and determine the presence of interaction effects.
Figure 12.6 Basic Analysis Strategy for Factorial Research Designs
Factorial research designs not only allow researchers to determine whether interaction effects are present among independent variables but also test the significance of each independent variable separately from the others. This leads to a second, pragmatic, advantage of factorial research designs over single-factor designs: the ability to study multiple main effects in one study.
The ability to study multiple effects in one study results in savings of time and energy. To understand this, imagine that in conducting the exercise deprivation study described earlier, we had 10 participants in each of the six combinations of Day and Gender. Observe in Figure 12.7 that we are able to test the effects of Day and Gender with a total of 60
745
participants. On the other hand, if we had chosen to use two separate single-factor research designs to study the effects of Day and Gender, we would have needed a total of 120 participants (60 participants in each of the two studies).
As you can see, factorial research designs have distinct advantages over their single-factor counterparts. However, two disadvantages of factorial research designs are an increase in the complexity of the calculations needed to analyze the data and an increase in the complexity of interpreting the results of the statistical analyses.
Compared with a single-factor design, in which the effect of only one variable is tested, conducting the analysis of a factorial research design requires analyzing the main effect of each independent variable as well as any interaction effects that exist between independent variables. The interpretation of factorial research designs is complicated by the fact that the presence or absence of interaction effects must be determined before examining any main effects that may be present. Also, when an interaction effect is present, the researcher must appropriately analyze and interpret the interaction effect in light of the study's research hypothesis.
Figure 12.7 Illustration of Sample Size in a Factorial Design (10 Participants in Each Combination) Figure 12.7 Illustration of Sample Size in a Factorial
Design (10 Participants in Each Combination)
Day
Monday Wednesday Friday
Gender Male 10 10 10 30
Female 10 10 10 30
20 20 20 60
Testing and interpreting the effects in a factorial research design is somewhat more complicated than analyzing single-factor designs. Hopefully, this will not discourage you from using factorial research designs when such designs will benefit your research. Rather, the goal of this book is to help you make informed decisions regarding the type and nature of a particular research design, to assist you in matching the benefits of different designs to the needs of specific research situations, and to provide you with the information and skills necessary to analyze and employ these designs effectively.
746
Learning Check 1: Reviewing what you've Learned So Far
1. Review questions a. What are differences between the single-factor research design and factorial research designs? b. What does it mean for there to be (or not be) an interaction effect between independent
variables? c. Looking at a figure of cell means, what would indicate an interaction effect is present? d. What is the difference between cell means and marginal means? What is represented by each
of these two types of means? e. What is the relationship between the presence or absence of main effects and the presence or
absence of interaction effects? f. Why does a determination of an interaction effect precede an examination of main effects? g. Compared with the single-factor research design, what are the advantages and disadvantages
of factorial research designs?
2. Create a line graph for each of the following tables of cell means (assume the scores have a possible range of 0 to 10 and put Factor A along the X (horizontal) axis).
a.
b.
c.
d.
3. In each of the following figures, is an interaction effect present or not present?
a.
747
b.
c.
d.
4. For the figures in Exercise 3, develop a table of cell means and marginal means and determine whether the main effects and the interaction effect are present or absent.
748
12.3 The Two-Factor (A × B) Research Design
Now that the concepts of factorial research designs, main effects, and interaction effects have been introduced, let's return to the voting message study introduced earlier in this chapter. This study is an example of the simplest factorial research design, one consisting of two independent variables.
A factorial research design consisting of two independent variables is referred to as a two- factor research design or an A × B research design (read “A-by-B”), in which “A” and “B” are represented by the numbers of levels in each of the two independent variables. For example, because Message type (Reward vs. Threat) and Authoritarianism (Low vs. High) consist of two levels, the voting message study can be referred to as a “2 × 2 research design.” Note that the total number of combinations in a factorial research design may be calculated by multiplying the number of levels of the independent variables. In the voting message study, there are (2 × 2) or four combinations of the two independent variables.
749
Notational System for Components of the Two-Factor Research Design
By including a second independent variable, the two-factor research design necessitates additions to this book's notational system. First, the two independent variables will be labeled Factor A and Factor B. For the voting message study, Message type will be called Factor A and Authoritarianism will be labeled Factor B. It does not matter which independent variable is given which label; however, it is important to remember the label for each variable to prevent errors in later calculations. Next, the lowercase letters a and b are used to indicate the number of levels of each factor. In this example, Message type (Factor A) consists of a = 2 levels (a1 = Reward and a2 = Threat), and Authoritarianism (Factor B) also consists of two levels (b1 = Low and b2 = High). The following figure illustrates the designation of factors and levels of factors:
Message type (Factor A)
Reward (a1) Threat (a2)
Authoritarianism (Factor B) Low (b1) a 1 b 1 a2 b1 High (b2) a 1 b 2 a 2 b 2
The number of scores in each combination of the two independent variables is represented by the symbol NAB. For the voting message study, if you were to assign eight participants to each of the four combinations, NAB would be equal to 8 (NAB = 8). Table 12.2 summarizes the notational symbols for the A × B two-factor research design and illustrates these symbols using the voting message study.
750
Analyzing the Two-Factor Research Design: Descriptive Statistics
Although the original voting message study contained a sample of 86 participants, the study's findings will be reproduced in this chapter using a smaller sample of 32 participants, with 8 participants assigned to each of the four combinations (NAB = 8). The ratings of message quality for the eight participants in each of the four combinations are presented in Table 12.3(a).
Table 12.2 Notational Symbols for the Two-Factor Research Design Table 12.2 Notational Symbols for the Two-Factor Research Design
Aspect of Design Symbol(s) Voting Message Study Example
Factors (Independent variables) Factor A Factor A = Message type
Factor B Factor B = Authoritarianism
# of levels within each factor a, b a = 2, b = 2
Level within a factor ai a1 = Reward, a2 = Threat
bi b1 = Low, b2 = High
# of combinations of factors a × b 2 × 2 = 4
Combination of factors ai bi a1 b1 = Reward/Low
a1 b2 = Reward/High
a2 b1 = Threat/Low
a2 b2 = Threat/High
# of scores for each combination N AB NAB = 8
As was the case in analyzing single-factor designs, the first step in the process of analyzing the two-factor research design is to calculate descriptive statistics. Table 12.3(b) provides descriptive statistics of the ratings of message quality for each combination in the voting message study. This table introduces the notation for cell means ( X ¯ A B ) and standard deviations (SAB); the letters AB in the subscript indicate that these statistics represent a particular combination of Factors A and B.
Notational System for Means in the Two-Factor Research Design
To properly summarize the data and to prepare for the calculation of inferential statistics, the cell means in Table 12.3 have been used to create the tables in Table 12.4. Table
751
12.4(a) introduces further additions to this book's notational system. In this table, the symbols X ¯ A and X ¯ B are used to represent the marginal means of the dependent variable for the different levels of the two main effects (Factor A and Factor B, respectively). In Table 12.4(b), the marginal means represent the two main effects of Message type and Authoritarianism. Also, the bottom right-hand corner of Table 12.4(a) contains the mean of the dependent variable for the total sample ( X ¯ T ) . X ¯ T is the mean of all of the data, combined across all of the cells; it can be determined by calculating the mean of the marginal means for either of the independent variables: X ¯ T = Σ X ¯ A α = 4.31 + 4.56 2 = 8.87 2 = 4.44 o r X ¯ T = Σ X ¯ B b = 4.37 + 4.50 2 = 8.87 2 = 4.44
Table 12.3 Rating of Message Quality for Participants in the Message Voting Study
Table 12.3 Rating of Message Quality for Participants in the Message Voting Study
(a) Raw Data
Message Type: Reward Reward Threat Threat
Authoritarianism: Low High Low High
4 5 3 5
4 3 5 4
5 4 5 6
6 4 4 5
4 5 3 5
5 3 5 6
5 3 4 4
4 5 4 5 Table 12.3 Rating of Message Quality for Participants in the
Message Voting Study
(b) Descriptive Statistics
Message Type: Reward Reward Threat Threat
Authoritarianism: Low High Low High
N AB 8 8 8 8
Mean ( X ¯ A B ) 4.62 4.00 4.13 5.00
Standard deviation (SAB) .74 .93 .83 .76
The notation for means in the two-factor research design is provided in Table 12.5.
752
Examining the Descriptive Statistics
Figure 12.8 provides a bar graph of the cell means that will be useful in examining the descriptive statistics. In this graph, Authoritarianism has been placed upon the X-axis, with the light-and dark-colored bars representing the different levels of Message type (Reward and Threat). Authoritarianism has been placed along the X-axis to highlight how participants low and high in authoritarianism differ in their reaction to the two types of messages, which is the study's primary research hypothesis. In displaying the data, a bar graph has been used rather than a line graph because the independent variable Authoritarianism consists of distinct groups (Low, High) that are not on a numeric continuum.
From an examination of Figure 12.8, we can begin to determine whether there is an interaction effect. It appears from the graph that ratings of message quality for the two types of messages are different for participants with low versus high authoritarian personality characteristics. Low authoritarian participants reacted more positively to the Reward message (M = 4.62) than they did to the Threat message (M = 4.13); however, the opposite was true for the high authoritarian participants, where the Threat mean (M = 5.00) is higher than the Reward mean (M = 4.00).
Table 12.4 Table of Means for the Two-Factor Research Design, Voting Message Study
Table 12.5 Notation for Means in the Two-Factor Research Design Table 12.5 Notation for Means in the Two-Factor Research Design
Aspect of Design Notation Voting Message Study Example
Level within a factor X ¯ A X ¯ A = 4.31 , X ¯ A = 4.56 X ¯ B X ¯ B = 4.37 , X ¯ B = 4.50 Combination of factors X ¯ AB X ¯ AB = 4.62 , X ¯ AB = 4.00 X ¯ AB = 4.13 , X ¯ AB = 5.00 Total sample X ¯ T X ¯ T = 4.44
From our examination of Figure 12.8, it appears that the effect of message type on ratings of message quality may depend on the participant's level of authoritarianism. The descriptive statistics provide initial support for the study's research hypothesis, which proposed this interaction effect between message type and authoritarianism. However, we
753
must still determine whether the interaction effect is statistically significant.
In addition to the interaction effect, we are also interested in examining the descriptive statistics to begin to understand the main effects. Looking back at Table 12.4(b), we observe that the mean ratings for the two message types (4.31 and 4.56) are somewhat similar to one another. This suggests that, when combined over the two levels of Authoritarianism, the two message types do not appear to greatly differ in their ratings of message quality. A small difference also exists for the main effect of Authoritarianism (4.37 vs. 4.50). This indicates that, when message type is ignored, participants with low and high levels of authoritarianism rate the quality of the messages in a similar manner. As we mentioned earlier, however, it is important to remember that the interpretation of main effects depends on whether or not an interaction effect is present.
Figure 12.8 Bar Graph of Message Quality by Message Type and Authoritarianism in the Voting Message Study
754
12.4 Introduction to Analysis of Variance (ANOVA) for the Two-Factor Research Design
To introduce the analysis of the two-factor research design, let's return to the one-way ANOVA covered in Chapter 11. The one-way ANOVA, you will recall, is used to test differences between the means of three or more groups that comprise a single independent variable. Conducting the one-way ANOVA involves calculating a statistic known as the F- ratio, which is the ratio of two types of variance: between-group variance and within-group variance.
The numerator of the F-ratio, between-group variance (MSBG), refers to differences among the means of the groups comprising an independent variable. This variance is attributable to two factors: the hypothesized effect of the independent variable as well as factors not taken into account by the researcher, referred to as “error.” The denominator of the F-ratio, within-group variance (MSWG), is the variability of the dependent variable within the groups comprising an independent variable. It represents variance in the dependent variable not explained or accounted for by the independent variable and therefore represents error. To summarize: F = between group variane within − group variance = effect + error error = M S BG M S WG
755
F-Ratios in the Two-Factor Research Design
As with the single-factor design, the goal in analyzing the two-factor research design is to determine whether a particular effect is statistically significant by calculating the F-ratio statistic, a statistic that contains between-group and within-group variance. Before implementing the two-factor research design, however, we must first redefine the term group. Rather than being a level of one independent variable, in the two-factor research design, a “group” represents a combination of two independent variables. For example, in the voting message study, one group, the Low/Reward combination, consists of the low- authoritarian participants who receive the Reward message.
In the voting message study, between-group variance refers to differences in the ratings of message quality among the four combinations of Message type and Authoritarianism. Why might these four combinations differ from each other? As it turns out, differences between the combinations may be due to three factors: the effect of Message type (Reward vs. Threat), the effect of Authoritarianism (Low vs. High), and the Message type by Authoritarianism interaction effect.
From this, we can assert that between-group variance in the two-factor research design comprises three elements: the main effect of Factor A, the main effect of Factor B, and the A × B interaction effect. This may also be expressed using the following formula: Between − group variance = main effect of Factor A + main effect of Factor B + A × B interaction effect
To fully understand between-group variance in the two-factor research design, we must analyze each of these three effects.
What is within-group variance in the two-factor research design? In Chapter 11, within- group variance was based on the variability of scores within the groups that comprise a single independent variable. However, because a group in the two-factor research design is a combination of two independent variables, within-group variance in this design is based on the variability of scores within the different combinations of the two factors.
To summarize, the goal in analyzing the two-factor research design is to determine the statistical significance of each of the three effects that comprise between-group variance: the main effects of Factor A and Factor B, and the A × B interaction effect. This is accomplished by calculating and evaluating three F-ratios, each of which divides its appropriate source of between-group variance by within-group variance. The next section discusses the steps needed to calculate and interpret these F-ratios.
756
12.5 Inferential Statistics: Two-Way Analysis of Variance (ANOVA)
Although the descriptive statistics in Table 12.3(b) appear to indicate support for the voting message study's primary research hypothesis, we must still conduct inferential statistical analyses to determine whether any observed effects are statistically significant. These analyses are conducted using essentially the same steps as in earlier chapters:
state the null and alternative hypotheses (H0 and H1), make decisions about the null hypotheses, draw conclusions from the analyses, and relate the results of the analyses to the research hypothesis.
Note from the preceding list that the main difference between these steps and those presented in earlier chapters is the use of plural nouns rather than singular. For example, rather than “make a decision about the null hypothesis,” we must now “make decisions about the null hypotheses” for the two main effects and the interaction effect.
757
State the Null and Alternative Hypotheses (H0 and H1)
In the two-way ANOVA, three effects are tested: the main effect of Factor A, the main effect of Factor B, and the A × B interaction effect. These effects, each of which has its own set of null and alternative hypotheses, will be described below.
Main Effect of Factor A
For the main effect of Message type in the voting message study, we are testing whether ratings of message quality of those exposed to the Reward message are different from those given the Threat message, ignoring the persons level of authoritarianism. In other words, Factor A is treated as if it were a single-factor design. For that reason, the statistical hypotheses for Factor A look exactly the same as in the one-way ANOVA discussed in Chapter 11: H 0 : all μ s are equal H 1 : not all μ s are equal
Main Effect of Factor B
Similar to the main effect of Factor A, testing the main effect for Factor B (Authoritarianism) involves ignoring the other factor. In this example, we are testing whether ratings of message quality are different for people who exhibit low versus high authoritarianism, ignoring the type of message the person receives. Because Factor B is also treated as if it were a single-factor design, the null and alternative hypotheses are the same as for the main effect of Factor A: H 0 : all μ s are equal H 1 : not all μ s are equal
A × B Interaction Effect
The A × B interaction effect, unlike the main effects, takes into consideration the joint influence of the two factors. In the voting message study, we are interested in seeing whether or not an interaction effect exists between Message and Authoritarianism. Therefore, the statistical hypotheses for the A × B interaction effect are different from those for the main effects: H 0 : an interaction effect does not exist H 1 : an interaction effect exists
The null hypothesis for the interaction effect states that no interaction effect exists between the two independent variables. The alternative hypothesis, which is mutually exclusive from the null hypothesis, states that some type of interaction effect does in fact exist. You may have noticed that the alternative hypothesis does not specify the precise nature of the interaction effect. As we illustrated in Figure 12.2, a significant interaction effect can take
758
on a wide variety of different shapes. For this reason, the alternative hypothesis does not attempt to state any particular form of the interaction effect.
759
Make Decisions about the Null Hypotheses
Once the three sets of null and alternative hypotheses have been stated, the next step is to make decisions regarding whether to reject any of the null hypotheses. These decisions are made using the following steps:
calculate the degrees of freedom (df); set alpha (a), identify the critical values, and state decision rules; calculate statistics: F-ratios for the two-way ANOVA; make decisions whether to reject the null hypotheses; determine the levels of significance; and calculate measures of effect size.
Each of these steps is discussed below; within each step, the main effects and the interaction effect are discussed separately.
Calculate the Degrees of Freedom (df)
The first step in calculating and evaluating the F-ratios in a two-way ANOVA is to calculate the appropriate degrees of freedom (df). There are four degrees of freedom in the two-factor research design, corresponding to the three effects that make up between-group variance (main effect of Factor a, main effect of Factor B, and the A × B interaction effect) and within-group variance.
Main Effect of Factor A (dfA)
The number of degrees of freedom for the main effect of Factor A is essentially the same as the degrees of freedom for the single-factor design. The degrees of freedom for Factor A (dfA) is the following:
(12-1) d f A = a − 1
where a is the number of levels of Factor A. In this example, Factor A (Message type) consists of two levels (Reward and Threat). Therefore, we may state that dfA is equal to d f A = a − 1 = 2 − 1 = 1
Main Effect of Factor B (dfB)
The formula for the degrees of freedom associated with Factor B (dfB) is a simple modification of that for dfA:
760
(12-2) d f B = b − 1
where b is the number of levels of Factor B. Authoritarianism, which has been labeled Factor B, consists of two levels: Low and High. We may then conclude that d f B = b − 1 = 2 − 1 = 1
A × B Interaction Effect (dfA B)
The number of degrees of freedom for the A × B interaction effect is the product of the degrees of freedom associated with the two factors represented in this interaction. More specifically, we may state the following:
(12-3) d f A × B = ( a − 1 ) ( b − 1 )
where a is the number of levels of Factor A and b is the number of levels of Factor B. Explaining the precise logic behind the formula for dfA × B is beyond the framework of this book and will not be described here. (See Keppel and Wickens [2004] for a detailed explanation of the rationale for the formula.) Because both Message type and Authoritarianism in the voting message study consist of two groups (a = 2 and b = 2), dfA × B may be formulated as d f A × B = ( a − 1 ) ( b − 1 ) = ( 2 − 1 ) ( 2 − 1 ) = ( 1 ) ( 1 ) = 1
Within-Group Variance (dfWG)
Finally, we must determine the degrees of freedom associated with within-group variance. In the single-factor design, the degrees of freedom was based on the number of scores in each level of the independent variable minus 1 (Ni – 1) multiplied by the number of groups that comprised the independent variable. In the two-factor research design, the degrees of freedom for within-group variance is based on the number of scores in each combination of the two independent variables minus one (NAB – 1), which is multiplied by the number of combinations of the two independent variables. Therefore, the formula for dfWG is the following:
(12-4) d f WG = ( a ) ( b ) ( N AB − 1 )
where a is the number of levels of Factor A, b is the number of levels of Factor B, and NAB is the number of scores in each A × B combination.
In the voting message study, dfWG is equal to the number of combinations of Message type
761
and Authoritarianism (a = 2, b = 2) multiplied by the number of scores in each combination (NAB = 8) minus 1. Therefore, dfWG is equal to d f WG = ( a ) ( b ) ( N AB − 1 ) = ( 2 ) ( 2 ) ( 8 − 1 ) = ( 4 ) ( 7 ) = 28
Set Alpha (α), Identify the Critical Values, and State Decision Rules
As in earlier chapters, in the two-way ANOVA, alpha is traditionally set at .05 (α = .05). Once alpha has been determined, the critical values of the F-ratio for the two main effects and the interaction effect may be identified.
Main Effect of Factor A
For the main effect of Factor A, the two degrees of freedom associated with the between- group and within-group variance of its F-ratio are dfA and dfWG. For the voting message study, we have determined that the degrees of freedom are dfA = 1 and dfWG = 28. Using the table of critical values for the F-ratio in the Appendix, we move to the 1 column under “Degrees of freedom for Numerator” and then down the rows of the table until we reach the one corresponding to 28 for the “Degrees of freedom for Denominator.” For α = .05, we find a critical value of 4.20. Consequently, the critical value for the main effect of Message type may be stated as follows: For α = .05 and d f = 1 , 28 , critical value = 4.20
Using the identified critical value of 4.20, the following decision rule may be stated for the main effect of Message type: If F A > 4.20 , reject H 0 ; otherwise, do not reject H 0
The critical value and regions of rejection and non-rejection for this main effect are illustrated in Figure 12.9.
Main Effect of Factor B
The two degrees of freedom associated with the main effect of Factor B are dfB and dfWG. For the voting message study, these degrees of freedom are dfB = 1 and dfWG = 28. Therefore, the critical value for the main effect of Authoritarianism may be stated as For α = .05 and d f = 1 , 28 , critical value = 4.20
For the main effect of Authoritarianism, the decision rule regarding the null hypothesis may be stated as If F B > 4.20 , reject H 0 ; otherwise, do no reject H 0
A × B Interaction Effect
762
For the interaction effect between Factor A and Factor B, the two degrees of freedom are dfA × B and dfWG. For the voting message study, these degrees of freedom are dfA × B = 1 and dfWG = 28. Consequently, the critical value for the Message type × Authoritarianism interaction effect may be stated as For α = .05 and d f = 1 , 28 , critical value = 4.20
Using the critical value of 4.20 for the Message type × Authoritarianism interaction effect, the decision rule regarding the null hypothesis is If F A × B > 4.20 , reject H 0 ; otherwise, do not reject H 0
Figure 12.9 Critical Value and Regions of Rejection and Non-Rejection for the Main Effect of Message Type in the Voting Message Study
In this example, which is a 2 × 2 design, the degrees of freedom for the numerator of the two main effects (dfA and dfB) and the A × B interaction effect (dfA × B) all happen to be equal to 1. This makes the identification of the critical values a relatively straightforward matter. However, what if we were using a 3 × 4 research design such as the one illustrated below in Table 12.6? Assuming there are three scores in each A × B combination (NAB = 3), Table 12.6 demonstrates the calculation of the degrees of freedom and the identification of the critical value for the two main effects and the interaction effect. Note that the three critical values in this table are all different from each other. For this reason, it is essential to correctly calculate the appropriate degrees of freedom for each effect.
Calculate Statistics: F-Ratios for the Two-Way ANOVA
The next step in making decisions about the null hypotheses is to calculate values of an inferential statistic. Conducting a two-way ANOVA requires calculating the F-ratio statistic, in which a source of between-group variance is divided by within-group variance: F = b e t w e e n − group variane within − group variance
Three F-ratios are calculated as part of the two-way ANOVA, one for each source of between-group variance: main effect of Factor A (FA), main effect of Factor B (FB), and the A × B interaction effect (FA × B). In this section, we first calculate four types of variance
763
(three related to between-group variance and one representing within-group variance) before calculating values for the three F-ratios.
Calculate the Mean Squares (MS)
The first step in conducting a two-way ANOVA is to calculate the mean squares (MS) that comprise the three F-ratios. This section discusses how to calculate four mean squares: the main effect of Factor A (MSA), the main effect of Factor B (MSB), the A × B interaction effect (MSA × B), and within-group variance (MSWG).
Main effect of Factor A (MSA). One source of between-group variance in the two-factor research design is associated solely with Factor A, without considering Factor B and the A × B interaction effect. The between-group variance associated with Factor A (MSA) is calculated as follows:
(12-5) M S A = S S A d f A = ( b ) ( N AB ) [ Σ ( X ¯ A − X ¯ T ) 2 ] a − 1
where b is the number of levels of Factor B, NAB is the number of scores within each A × B combination, X ¯ A is the mean for each level of Factor A, X ¯ T is the mean for the total sample, and a is the number of levels of Factor A.
Although it may not be apparent at first glance, the formula for MSA is very similar to the formula for MSBG in the one-way ANOVA (Formula 11–3): M S BG = N i Σ ( X ¯ i − X ¯ T ) 2 # groups − 1
Table 12.6 Degrees of Freedom and Critical Values for 3 × 4 Design (NAB = 3 and α = .05)
764
765
Learning Check 2: Reviewing what you've Learned So Far
1. Review questions a. What do the letters A and B represent in an “A × B factorial research design”? b. In the two-factor research design, what is between-group variance composed of? c. In the two-factor research design, what is within-group variance? d. Why are there three sets of statistical hypotheses in a two-factor research design? e. Why don't the null and alternative hypotheses for the interaction effect specify the nature of
the interaction effect? f. Why are the degrees of freedom (df), critical values, and decision rules for main effects
similar to those in the single-factor research design?
2. For each of the research situations below, identify the two independent variables and indicate the type of A × B design (e.g., 2 × 2, 3 × 3).
a. An instructor teaches two courses: one taught in the typical classroom setting and one taught online. She believes students taking the classroom course are more likely to keep up with assigned course readings than students in the online course; furthermore, she believes this difference is greater for students living off-campus than on-campus.
b. When a child misbehaves, who is responsible for the child's bad behavior? A family therapist believes that parents are more likely to attribute responsibility to the child than to themselves, particularly when the child is a teenager rather than an adolescent.
c. Before taking an exam, do you engage in rituals like kissing a luck charm or wearing a particular piece of clothing? Two researchers (Rudski & Edwards, 2007) hypothesized that students engage in rituals more when an exam is important (vs. unimportant) to students' grade in the course, as well as when an exam was seen as “very difficult” (as opposed to “easy” or “somewhat difficult”).
3. For each of the following tables of cell means, calculate marginal means ( X ¯ A and X ¯ B and the total mean ( X ¯ T ) .
a.
b.
c.
d.
766
4. For each of the following, calculate four degrees of freedom (dfA, dfB, dfA × B, dfWG) and identify the three critical values of the three F-ratios (FA, FB, FA × B) (assume α = .05).
a. a = 2, b = 2, NAB = 5
b. a = 3, b = 2, NAB = 6
c. a = 3, b = 3, NAB = 10
d. a = 4, b = 3, NAB = 7
The numerator in both formulas represents the difference between the mean of a level of an independent variable and the total mean: X ¯ i − X ¯ T for the one-way ANOVA and X ¯ A − X ¯ T for Factor A in the two-way ANOVA. The critical change that must be made in analyzing the main effect of Factor A is the addition of the symbol b in the numerator, which takes into account the fact that this design includes a second independent variable (Factor B).
To calculate MSA for the main effect of Message type, b (the number of levels of Authoritarianism) is equal to 2 and NAB is equal to 8; the marginal means for the two levels of Message type ( X ¯ A ) , as well as the total mean ( X ¯ T ) , are calculated in Table 12.4(b), and a (the number of levels of Message type) is equal to 2. Using this information, the calculation of MSA is provided below: M S A = ( b ) ( N AB ) [ Σ ( X ¯ A − X ¯ T ) 2 ] a − 1 = ( 2 ) ( 8 ) [ ( 4.31 − 4.44 ) 2 + ( 4.56 − 4.44 ) 2 2 − 1 = 16 [ ( − .13 ) 2 + ( .12 ) 2 ] 1 = 16 ( .02 + .01 ) 1 = 16 ( .03 ) 1 = .48 1 = .48
Main effect of Factor B (MSB). The second effect tested in the two-way ANOVA is the main effect of Factor B. The formula for the mean squares for Factor B (MSB) is a simple modification of the formula for MSA, substituting the letter A for the letter B, and vice versa:
(12-6) M S B = S S B d f B = ( a ) ( N A B ) [ Σ ( X ¯ B − X ¯ T ) 2 ] b − 1
where a is the number of levels of Factor A, NAB is the number of scores within each A × B combination, X ¯ B is the mean for each level of Factor B, X ¯ T is the mean for the total sample, and b is the number of levels of Factor B.
For the voting message study, Factor B represents Authoritarianism. Based on the information used to calculate MSA, as well as the marginal means for low and high Authoritarianism ( X ¯ B ) from Table 12.4(b), MSB is calculated as follows: M S B = ( a ) ( N AB ) [ Σ ( X ¯ B − X ¯ T ) 2 ] b − 1 = ( 2 ) ( 8 ) [ ( 4.37 − 4.44 ) 2 + ( 4.50 − 4.44 ) 2 2 − 1 = 16 [ ( − .07 ) 2 + ( .06 ) 2 ] 1 = 16 ( .005 + .004 ) 1 = 16 ( .009 ) 1 = .14
767
1 = .14
A × B interaction effect (MSA × B). How do you calculate the between-group variance associated with the A × B interaction effect? As we have already learned, between-group variance in the two-factor research design consists of three factors: the main effect of Factor A, the main effect of Factor B, and the A × B interaction effect. In terms of the sum of squares (SS), the composition of between-group variance may be expressed using the following formula: S S BG = S S A + S S B + S S A × B
With a little algebraic manipulation, the sums of squares for the A × B interaction (SSA × B) can be represented by the following: S S A × B = S S BG − S S A − S S B
The sums of squares for the A × B interaction (SSA × B) is equal to the total amount of between-group variance (SSBG) minus that associated with each of the two main effects (SSA and SSB).
Based on the above formula, calculating SSA × B requires first calculating SSBG. Formula 12–7 provides the computational formula for SSBG:
(12-7) S S BG = N AB Σ ( X ¯ AB − X ¯ T ) 2
where NAB, is the number of scores within each A × B combination, X ¯ B is the mean for each combination, and X ¯ T is the mean for the total sample. As was the case in a single- factor design, we can see from Formula 12.7 that between-group variance in the two-factor research design is based on the difference between the mean of each group (a combination of the two factors) and the total mean ( X ¯ AB − X ¯ T ) .
The above formula for SSBG can be used to help create the following computational formula for MSA × B:
(12-8) M S A × B = S S A × B d f A × B = [ N AB Σ ( X ¯ AB − X ¯ T ) 2 ] − S S A − S S B ( a − 1 ) ( b − 1 )
where NAB is the number of scores within each A × B combination, X ¯ AB is the mean for each combination of the two factors, X ¯ T is the mean for the total sample, SSA is the sums of squares for Factor A, SSB is the sums of squares for Factor B, a is the number of levels of Factor A, and b is the number of levels of Factor B.
768
Much of the information needed to calculate MSA × B can be obtained from the earlier calculation of MSA and MSB, and the values for X ¯ AB are found in the descriptive statistics calculated in Table 12.3(b). In the voting message study, MSA × B for the Message type × Authoritarianism interaction effect is calculated as follows: M S A × B = [ N AB Σ ( X ¯ AB − X ¯ T ) 2 ] − S S A − S S B ( a − 1 ) ( b − 1 ) = 8 [ ( 4.62 − 4.44 ) 2 + ( 4.00 − 4.44 ) 2 + ( 4.13 − 4.44 ) 2 + ( 5.00 − 4.44 ) 2 ] − .48 − .14 ( 2 − 1 ) ( 2 − 1 ) = 8 [ ( .18 ) 2 + ( − .44 ) 2 + ( − .31 ) 2 + ( .56 ) 2 ] − .48 − .14 1 = 8 ( .03 + .19 + .10 + .31 ) − .48 − .14 1 = 8 ( .63 ) − .48 − .14 1 = 5.04 − .48 − .14 1 = 4.42 1 = 4.42
Within-group variance (MSWQ). Within-group variance is an estimate of error variance and is based on the variability of scores within the combinations of Factors A and B. The formula for within-group variance for the two-factor research design (MSWG) is provided in Formula 12–9:
(12-9) M S WG = S S WG d f WG = ( N AB − 1 ) Σ s AB 2 ( a ) ( b ) ( N AB − 1 )
where NAB is the number of scores in each combination, SAB is the standard deviation for each combination, a is the number of levels of Factor A, and b is the number of levels of Factor B.
Consistent with what we've already observed with between-group variance, the formula for MSWG in the two-factor research design is very similar to the formula for MSWG in the one-way ANOVA (Formula 11–5): M S WG = ( N i − 1 ) Σ s i 2 Σ ( N i − 1 )
Within-group variance in both types of research designs is based on the variability of scores within each group, represented by the sample standard deviation (s).
For the voting message study, the values for NAB, a and b can be found in earlier calculations; the standard deviations for the combinations are found in the descriptive statistics (Table 12.3(b)). Using this information, MSWG for the voting message study can be stated in the following way: M S WG = ( N AB − 1 ) Σ s AB 2 ( a ) ( b ) ( N AB − 1 ) = ( 8 − 1 ) [ ( .74 ) 2 + ( .93 ) 2 + ( .83 ) 2 + ( .76 ) 2 ] ( 2 ) ( 2 ) ( 8 − 1 ) = 7 [ .55 + .86 + .69 + .58 ] 28 = 7 ( 2.68 ) 28 = 18.76 28 = .67
Calculate the F-Ratios (F)
Once the four variances have been calculated, the next step in a two-way ANOVA is to calculate three F-ratios, one for each of the two main effects and one for the interaction effect. These F-ratios are calculated by dividing a particular source of between-group
769
variance by the within-group variance.
Main Effect of Factor A (FA)
To calculate the F-ratio for the main effect of Factor A (FA), we divide the between-group variance associated with this effect (MSA) by within-group variance (MSWG):
(12-10) F A = M S A M S WG
For the voting message study, the F-ratio for the main effect of Message type is equal to F A = M S A M S WG = .48 .67 = .72
Main Effect of Factor B (FB)
The F-ratio for Factor B (FB) is calculated in a similar manner as the main effect of Factor A:
(12-11) F B = M S B M S WG
The F-ratio for the main effect of Authoritarianism in the voting message study is equal to F B = M S B M S WG = .14 .67 = .21
A × B Interaction Effect (FA × B)
Finally, the F-ratio for the A × B interaction effect (FA × B) is calculated using the following formula:
(12-12) F A × B = M S A × B M S WG
For the Message type × Authoritarianism interaction effect, the F-ratio is equal to F A × B = M S A × B M S WG = 4.42 .67 = 6.60
The symbols and computational formulas for a two-way ANOVA are provided in the ANOVA summary tables in Table 12.7. The ANOVA summary table in Table 12.7(c) is the one you would use to report the results of the analysis for the voting message study. As we observed in our discussion of the one-way ANOVA, the word Error is used to represent within-group variance as this is variance in the dependent variable that cannot be explained or accounted for by the independent variables and is therefore attributed to random, chance factors.
770
Make Decisions Whether to Reject the Null Hypotheses
Once values for the three F-ratios (FA FB, FA × B) have been calculated, the next step is to make the decision whether to reject the null hypothesis for each of the three effects. Each of these decisions for the voting message study is described below.
Main Effect of Factor A
The decision of whether to reject the null hypothesis regarding the main effect of Factor A is made by comparing the calculated value of FA with its critical value. For the main effect of Message type, we may state that F A = .72 < 4.20 ∴ do not reject H 0 ( p > .05 )
Table 12.7 Summary Tables for the Two-Way ANOVA Table 12.7 Summary Tables for the Two-Way ANOVA
a. Notation and symbols
Source SS df MS F
Factor A (A) SS A df A MS A F A
Factor B (B) SS B df B MS B F B
A × B interaction (A × B) SS A × B df A × B MS A × B F A × B
Within-group (WG) SS WG df WG MS WG
Total (T) SS T df T Table 12.7 Summary Tables for the Two-Way ANOVA
b. Formulas
Source SS df MS F
Factor A ( b ) ( N AB ) [ Σ ( X ¯ A − X ¯ T ) 2 ]
a – 1 S S A d f A M S A M S WG
Factor B ( a ) ( N AB ) [ Σ ( X ¯ B − X ¯ T ) 2 ]
b – 1 S S B d f B M S B M S WG
A × B interaction
N AB Σ ( X ¯ i − X ¯ T ) 2 − S S A − S S B
(a – 1) (b) – 1) S S A × B d f A × B
M S A × B M S WG
Within- group
( N AB − 1 ) Σ S AB 2 (a)(b)(NAB – 1)
S S WG d f WG
Total SSA + SSB + SSA × B + SSWG
dfA + dfB + dfA × B + dfWG
Table 12.7 Summary Tables for the Two-Way ANOVA
771
c. Voting Message Study Example
Source SS df MS F
Message type .48 1 .48 .72
Authoritarianism .12 1 .12 .21
Message type × Authoritarianism 4.42 1 4.42 6.60
Error 18.76 28 .67
Total 23.78 31
Because the calculated value of .72 for the main effect of Message type is less than the critical value of 4.20, it falls in the region of non-rejection and the null hypothesis is not rejected. This is illustrated in Figure 12.10. The main effect of Message type is not significant, meaning that the message quality ratings of those receiving the Reward (M = 4.31) and Threat (M = 4.56) messages (regardless of their level of authoritarianism) do not significantly differ.
Main Effect of Factor B
The decision regarding the statistical significance of the main effect of Factor B is made by comparing the value of FB with its critical value. In the voting message study, the decision regarding the main effect of Authoritarianism is represented as follows: F B = .21 < 4.20 ∴ do not reject H 0 ( p > .05 )
Because the null hypothesis for Factor B was not rejected, we may conclude that the message quality ratings for the High (M = 4.37) and Low (M = 4.50) authoritarianism participants are not significantly different.
A × B Interaction Effect
In this step, we make the decision regarding whether there is a significant interaction effect between Factor A and Factor B. In the voting message study, we may state the following: F A × B = 6.60 > 4.20 ∴ reject H 0 ( p < .05 )
In this example, because the value of 6.60 for FA × B is greater than the critical value of 4.20, we make the decision to reject the null hypothesis and conclude that a significant interaction effect exists between Message type and Authoritarianism.
Figure 12.10 Making the Decision Regarding the Null Hypothesis for the Main Effect of Message Type in the Voting Message Study
772
In the voting message example, although neither of the two main effects was statistically significant, a significant interaction effect was nevertheless found. This reinforces our earlier observation that the absence of main effects provides no indication of whether an interaction effect is either present or absent.
Determine the Levels of Significance
If the decision is made to reject the null hypothesis, it is useful for the sake of precision to determine whether the probability of a calculated value of a statistic is not only less than .05 (p < .05) but also less than .01 (p < .01). In this step, we determine the appropriate level of significance for the two main effects and the interaction effect.
Main Effect of Factor A
For the main effect of Factor A in the voting message study, the null hypothesis was not rejected because the probability of the calculated value of FA = .72 was greater than the alpha value of .05 (p > .05). Consequently, because the probability of FA is greater than .05, it is not necessary to determine whether its probability is less than .01.
Main Effect of Factor B
Because the probability of FB for the main effect of Authoritarianism (FB = .21) was not less than .05, the null hypothesis was not rejected. As a result, there is no reason to determine whether the probability of FB is less than .01.
A × B Interaction Effect
Unlike the two main effects, the Message type × Authoritarianism interaction effect was statistically significant. As a result, it is appropriate to make a precise assessment of the probability of the calculated F-ratio by comparing it with the .01 critical value. Looking at the Appendix, we observe that for dfA × B = 1 and dfWG = 28, the .01 critical value for the F-ratio is equal to 7.64. The level of significance for the Message type × Authoritarianism interaction effect may be stated as follows:
773
F A × B = 6.60 < 7.64 ∴ p < .05 ( but not < . 01 )
Because the F-ratio value of 6.60 is between the α = .05 (4.20) and .01 (7.64) critical values, the probability of the F-ratio is less than .05 but not less than .01. Therefore, the appropriate level of significance for the F-ratio for the Message type × Authoritarianism interaction effect is p < .05.
Calculate Measures of Effect Size (R2)
The purpose of calculating inferential statistics such as the F-ratio is to make one of two decisions: reject or not reject the null hypothesis. However, in addition to determining whether or not an effect is statistically significant, it is useful to calculate a measure of effect size that estimates the size or magnitude of a particular effect. As we discussed in Chapter 11, there are different measures of effect size. Although debate continues regarding the appropriate measure of effect size for factorial research designs, because of its relative simplicity, this book will rely on R2, which is the percentage of the total variance in the dependent variable accounted for by a particular effect.
Main Effect of Factor A
The main effect of Factor A is treated as though it were a single-factor design. Consequently, the formula for R2 for Factor A is essentially the same as that presented in Chapter 11:
(12-13) R 2 A = S S A S S T
where SSA is the sum of squares for Factor A and SST is the sum of squares for the total amount of variance. Using the values for SSA and SST provided in the ANOVA summary table in Table 12.7(c), we may calculate the R2 for Message type in the voting message study in the following way: R A 2 = S S A S S T = .48 23.78 = .02
The R2 for Message type implies that 2 % of the variance in ratings of message quality is explained by the type of message, reward versus threat, presented to the participant. Using Cohen's (1988) rules of thumb, we can see that this is considered a relatively small effect.
Main Effect of Factor B
As you might expect, the formula for R2 for the main effect of Factor B is a simple modification of the formula used for Factor A:
(12-14)
774
R B 2 = S S B S S T
where SSB is the sum of squares for Factor B and SST is the sum of squares for the total amount of variance. Again relying on the summary table in Table 12.7(c), the R2 for the main effect of Authoritarianism may be stated as R B 2 = S S B S S T = .12 23.78 = .01
Therefore, 1% of the variance in ratings of message quality is explained by the participants level of authoritarianism, which is also a relatively small effect.
A × B Interaction Effect
The R2 for the A × B interaction effect is calculated in a manner similar to that for the two main effects:
(12-15) R 2 A × B = S S A × B S S T
where SSA × B is the sum of squares for the A × B interaction effect and SST is the sum of squares for the total amount of variance. In the voting message study, the R2 for the Message type × Authoritarianism interaction effect may be stated as
775
Learning Check 3: Reviewing what you've Learned So Far
1. Review questions a. Which descriptive statistic is used to calculate between-group variance in the two-way
ANOVA? Which descriptive statistic is used to calculate within-group variance? b. Why is the word Error included in an ANOVA summary table?
2. For each of the following, calculate the F-ratios (F) for the two-way ANOVA, create an ANOVA
summary table, and calculate measures of effect size (R2). a. SSA = 20.00, dfA = 1, SSB = 10.00, dfB = 1, SSA × B = 15.00, dfA × B = 1, SSWG = 80.00,
dfWG = 20 b. SSA = 54.00, dfA = 1, SSB = 25.00, dfB = 1, SSA × B = 16.00, dfA × B = 1, SSWG =
400.00, dfWG = 44 c. SSA = 4.60, dfA = 2, SSB = 9.79, dfB = 1, SSA × B = 10.31, dfA × B = 2, SSWG = 43.87,
dfWG = 42
3. Does the type of car you own influence how attractive you are to someone of the opposite sex? Two researchers (Dunn & Searle, 2010) showed men and women a picture of a person of the opposite sex driving one of two cars: a mid-priced sedan or a luxury car. Each participant then rated the attractiveness of the driver on a 1 to 10 scale (the higher the rating, the higher the attractiveness). As research suggested that men focus on a woman's physical attractiveness but women are concerned with a man's wealth or status, it was hypothesized that men would give similar ratings of attractiveness to the drivers of the two types of cars, whereas women would give lower ratings to the driver of the mid-priced car than the luxury car. The ratings of attractiveness given by men and women participants to the drivers of the two types of cars are summarized below:
Sex (Factor A): Type of Car (Factor B):
Man Mid- priced
Man Luxury
Woman Mid- priced
Woman Luxury
NAB 7 7 7 7
Mean ( x ¯ AB ) 6.43 6.86 4.29 7.29
Standard deviation (SAB) 1.51 1.57 1.11 1.38
a. Create a figure to represent the cell means (place Sex on the X-axis). b. Conduct the two-way ANOVA and create an ANOVA summary table. c. Report the decisions regarding the null hypotheses and the level of significance for the main
effects and the interaction effect.
d. Calculate measures of effect size (R2) for the main effects and the interaction effect.
An R2 value of .19 is an effect of relatively large magnitude.
776
Draw Conclusions from the Analyses
Given the level of complexity of the analysis of the two-factor research design, it is important to describe the results of the analysis as accurately and completely as possible. Below is one way of describing the two-way ANOVA for the voting message study:
Participants' ratings of message quality were analyzed using a 2 (Message type: threat vs. reward) × 2 (Authoritarianism: low vs. high) ANOVA. The main effect of Message type was not significant, F(1, 28) = .72, p > .05, R2 = .02, implying that the message quality ratings of the Reward (M = 4.31) and Threat (M = 4.56) conditions were not significantly different. In terms of the main effect of Authoritarianism, a nonsignificant difference was found in the ratings of participants low (M = 4.37) or high (M = 4.50) in authoritarianism, F(1, 28) = .21, p > .05, R2 = .01. However, a significant Message type × Authoritarianism interaction effect was found, F(1, 28) = 6.60, p < .05, R2 = .19, implying that the effect of receiving a message emphasizing threat versus reward on the perceived quality of the message depends on the participants level of authoritarianism.
The above paragraph starts by informing the reader of the dependent variable that was analyzed (“ratings of message quality”), the statistical procedure used to analyze the dependent variable, the independent variables, and the groups that comprise each independent variable (“a 2 (Message type: threat vs. reward) × 2 (Authoritarianism: low vs. high) ANOVA”). The next sentences summarize the tests of significance for both the main effects and the interaction effect. For each effect tested, information is provided regarding descriptive statistics, the nature and direction of the differences between groups, and the inferential statistic. Reporting the inferential statistic involves providing the statistic that has been calculated, the degrees of freedom, the calculated value of the statistic, the level of significance of the statistic, and the measure of effect size.
777
Relate the Results of the Analyses to the Research Hypothesis
Table 12.8 summarizes the steps followed to conduct the two-way ANOVA, using the voting message study to illustrate these steps. The last step in this process is to relate the results of the analyses to the study's research hypothesis. In the voting message study, the researchers hypothesized an interaction effect between message type and authoritarianism such that the reward message was hypothesized to be more persuasive than the threat message for people low on authoritarianism, whereas the threat message was hypothesized to be more persuasive than the reward message for people high on authoritarianism. Given that a significant Message type × Authoritarianism interaction was found in conducting the two-way ANOVA, has the research hypothesis been supported? As you may have guessed, the answer to this question is no.
Table 12.8 Summary, Conducting the Two-Way ANOVA, Voting Message Study Example
778
779
Rejecting the null hypothesis for an A × B interaction effect does not provide either support or a lack of support for a research hypothesis. Based on the statement of the alternative hypothesis (H1: an interaction effect exists), the most specific conclusion we can draw from rejecting the null hypothesis for the A × B interaction effect is simply that an interaction effect is present. Rejecting the null hypothesis does not provide insight into the precise nature or source of the interaction effect. Therefore, at this point in the analysis, we are unable to determine whether the study's research hypothesis has been supported until additional statistical analyses are conducted.
In Chapter 11, the most specific conclusion that could be drawn from rejecting the null hypothesis in a one-way ANOVA is that the means for the groups are not all equal to each other. To determine the precise nature and source of the significant effect, additional analyses need to be conducted. These analyses, referred to as analytical comparisons, involve making comparisons between groups that are part of a larger research design. A similar situation exists in the two-factor research design when a significant A × B interaction effect is found. The next section introduces, in a conceptual manner, analyses designed to more determine the source and nature of a significant A × B interaction effect.
780
12.6 Investigating a Significant A × B Interaction Effect: Analysis of Simple Effects
A significant A × B interaction implies that the effect of one independent variable changes at the different levels of the other independent variable. However, to more precisely determine the source of a significant interaction effect, additional analyses must be conducted that break down a larger factorial research design into a number of smaller designs.
As a way of introducing the logic that underlies these analyses, let's return to the research hypothesis in the voting message study:
High authoritarian recipients would perceive the threat message as more valid and persuasive than the reward message. … A threat-related persuasive message would lead to less persuasion among low authoritarian recipients, relative to the reward message. (Lavine et al, 1999, p. 340)
What specific comparisons are needed to test this research hypothesis? It appears that two specific comparisons are necessary: the difference between the Reward and Threat conditions for those high in authoritarianism and the difference between the Reward and Threat conditions for those low in authoritarianism. This analysis plan, illustrated in Figure 12.11(a), shows we are testing the effect of one independent variable (Message type) separately for each level of the other independent variable (Authoritarianism).
The analyses in Figure 12.11(a) are examples of what is known as the analysis of simple effects. A simple effect is the effect of one independent variable at one of the levels of another independent variable. For the voting message study, we would be testing the simple effect of Message type at each level of Authoritarianism.
The researchers in the voting message study did in fact compare the perceived quality of the message for the two message types (reward vs. threat) separately for the low- and high- authoritarian participants. Below is their conclusion regarding their analysis of the simple effects:
Figure 12.11 Illustration of Simple Effects Analyses, Voting Message Study
781
Follow-up contrasts revealed that high authoritarians did indeed perceive the threat message as more persuasive than the reward message … low authoritarians perceived the reward message as containing more persuasive arguments than the threat message. (Lavine et al., 1999, p. 343)
On the basis of the results of their simple effects analyses, the researchers concluded that their research hypothesis had been supported.
The above example tested the simple effect of Factor A (Message type) at each level of Factor B (Authoritarianism). As Figure 12.11(b) illustrates, we might instead have chosen to test the simple effect of Authoritarianism at each level of Message type. Because it is possible to test the simple effects of either of two factors, which of the two should you choose? The choice of which simple effects to test does not alter the statistical significance of the interaction. Rather, the different simple effects provide different perspectives on the same set of data.
The simple effect of Message type at each level of Authoritarianism (Figure 12.11(a)) focuses on the effectiveness of different types of message (reward vs. threat) for people who have a particular level of authoritarianism. On the other hand, the simple effect of Authoritarianism at each level of Message type (Figure 12.11(b)) focuses on how different types of people (people who are low vs. high in authoritarianism) respond to a particular type of message. The simple effects of Message type emphasize the effectiveness of different
782
types of messages, whereas the simple effects of Authoritarianism emphasize the responses of different types of people. Ultimately, the choice of which simple effects to test depends on the purpose of the study and the way in which research hypotheses are stated.
In summary, rejecting the null hypothesis for the A × B interaction effect does not reveal the precise source and nature of the interaction. To directly test one's research hypotheses, additional analyses are necessary. A common way to conduct these analyses is to examine the simple effects of one variable at the different levels of the other variable. Due to space considerations, this book will not discuss the calculations necessary to conduct these analyses. Those who are interested in exploring this topic in greater detail are referred to Keppel & Wickens (2004) and Keppel et al. (1992).
The two-factor research design and two-way ANOVA are perhaps the most challenging research methodology and statistical procedures covered in this book. As you have seen, the inclusion of more than one independent variable not only increases the number of calculations necessary to carry out the statistical analyses but also requires a greater understanding of different effects that can exist in the same study, as well as the relationship between these effects. Although this may be challenging, you have hopefully begun to gain an appreciation of the ability of factorial designs to capture the complexity of phenomena that researchers choose to study.
783
12.7 Looking Ahead
As you have seen, much of the presentation and discussion of the two-way ANOVA builds off of the information provided in earlier chapters. Chapter 9 included a discussion of the t- test, which is used to compare the means of two groups. In Chapter 11, the one-way ANOVA was presented to compare the means of three or more groups that comprise one independent variable. The t-test, one-way ANOVA, and two-way ANOVA are used when the independent variable or variables are categorical in nature, consisting of groups. The goal of these three statistical procedures is to compare the means of groups on a continuous dependent variable, a variable that is measured at the interval or ratio level of measurement. The next chapter introduces a statistical procedure used when both the independent and dependent variables are continuous in nature.
784
Learning Check 4: Reviewing what you've Learned So Far
1. Review questions a. What conclusion can you draw from a significant A × B interaction? What conclusion can
you not draw? b. What is the difference between a main effect and a simple effect? c. What is the purpose of analyzing simple effects in the A × B research design?
2. For each of the research situations below, draw a figure such as those in Figure 12.11 that illustrate simple effects analyses appropriate to test the study's research hypothesis.
a. An instructor teaches two courses: one taught in the typical classroom setting and one taught online. She believes students taking the classroom course are more likely to keep up with assigned course readings than students in the online course; furthermore, she believes this difference is greater for students living off-campus than on-campus.
b. When a child misbehaves, who is responsible for the child's bad behavior? A family therapist believes that parents are more likely to attribute responsibility to the child than to themselves, particularly when the child is a teenager rather than an adolescent.
c. Before taking an exam, do you engage in rituals like kissing a luck charm or wearing a particular piece of clothing? Two researchers (Rudski & Edwards, 2007) hypothesized that students engage in rituals more when an exam is important (vs. unimportant) to students' grade in the course, as well as when an exam was seen as “very difficult” (as opposed to “easy” or “somewhat difficult”).
785
12.8 Summary
Unlike single-factor research designs, which consist of a single independent variable, factorial research designs consist of all possible combinations of the levels of two or more independent variables. A factorial research design consisting of two independent variables is referred to as a two-factor (A × B) research design.
Conceptually, factorial research designs allow researchers to determine whether an interaction effect exists between independent variables. An interaction effect is present when the effect of one of the independent variables on the dependent variable changes (is not the same) at the different levels of another independent variable.
In addition to testing for interaction effects, a factorial research design also involves testing for main effects; a main effect is the effect of a particular independent variable on the dependent variable of a particular factor. The two-factor research design contains two main effects, each of which is treated as a separate single-factor design.
There are two important aspects of the relationship between main effects and interaction effects. First, the presence or absence of main effects provides no indication of whether an interaction effect is present or absent, and vice versa. Second, whether and how one interprets main effects depends on the presence or absence of interaction effects. If an interaction effect is present, the effects of both independent variables must be considered simultaneously rather than in isolation of each other. If an interaction effect is not present, the two-factor design may be treated as if it were two separate single-factor designs, one for each main effect.
A significant A × B interaction effect implies that the effect of one independent variable changes at the different levels of the other independent variable. However, to more precisely determine the source of a significant interaction effect, additional analyses must be conducted that break down a larger factorial research design into a number of smaller designs. For example, simple effects analyses may be conducted to test the effect of one independent variable at each of the different levels of another independent variable.
786
12.9 Important Terms
single-factor research design (p. 488) factorial research design (p. 489) interaction effect (p. 489) cell (p. 490) cell mean (p. 490) main effect (p. 494) marginal mean (p. 494) two-factor (A × B) research design (p. 502) simple effect (p. 531)
787
12.10 Formulas Introduced in this Chapter
Degrees of Freedom for Main Effect of Factor A, Two-Way ANOVA (dfA)
(12-1) d f A = a − 1
Degrees of Freedom for Main Effect of Factor B, Two-Way ANOVA (dfB)
(12-2) d f B = b − 1
Degrees of Freedom for A × B Interaction Effect, Two-Way ANOVA (dfA × B)
(12-3) d f A × B = ( a − 1 ) ( b − 1 )
Degrees of Freedom for Within-Group Variance, Two-Way ANOVA (dfWG)
(12-4) d f WG = ( a ) ( b ) ( N AB − 1 )
Between-Group Variance for Main Effect of Factor A, Two-Way ANOVA (MSA)
(12-5) M S A = S S A d f A = ( b ) ( N AB ) [ Σ ( X ¯ A − X ¯ T ) 2 ] a − 1
Between-Group Variance for Main Effect of Factor B, Two-Way ANOVA (MSB)
(12-6) M S B = S S B d f B = ( b ) ( N AB ) [ Σ ( X ¯ B − X ¯ T ) 2 ] b − 1
Sums of Squares for Between-Group Variance, Two-Way ANOVA (SSBG)
(12-7) S S BG = N AB Σ ( X ¯ AB − X ¯ T ) 2
Between-Group Variance for A × B Interaction, Two-Way ANOVA (MSA×B)
(12-8) M S A × B = S S A × B d f A × B = [ N AB Σ ( X ¯ AB − X ¯ T ) 2 ] − S S A − S S B ( a − 1 ) ( b − 1 )
788
Within-Group Variance, Two-Way ANOVA (MSWG)
(12-9) M S WG = S S WG d f WG = ( N AB − 1 ) Σ s 2 AB ( a ) ( b ) ( N AB − 1 )
F-Ratio for Main Effect of Factor A, Two-Way ANOVA (FA)
(12-10) F A = M S A M S WG
F-Ratio for Main Effect of Factor B, Two-Way ANOVA (FB)
(12-11) F B = M S B M S WG
F-Ratio for A × B Interaction, Two-Way ANOVA (FA × B)
(12-12) F A × B = M S A × B M S WG
Measure of Effect Size for Main Effect of Factor A, Two-Way ANOVA ( R A 2 )
(12-13) R A 2 = S S A S S T
Measure of Effect Size for Main Effect of Factor B, Two-Way ANOVA ( R B 2 )
(12-14) R B 2 = S S B S S T
Measure of Effect Size for A × B Interaction, Two-Way ANOVA ( R A × B 2 )
(12-15) R A × B 2 = S S A × B S S T
789
12.11 Using SPSS
790
Two-Way Analysis of Variance (ANOVA): The Voting Message Study (12.1)
1. Define independent and dependent variables (name, # decimals, labels for the variables, labels for values of the independent variable) and enter data for the variables.
NOTE: Numerically code values of the independent variables (i.e., Message type [1 = Reward, 2 = Threat], Authoritarianism [1 = Low, 2 = High]) and provide labels for these values in Values box within Variable View.
2. Select the two-way ANOVA procedure within SPSS.
How? (1) Click Analyze menu, (2) click General Linear Model, and (3) click Univariate.
3. Identify the dependent variable, the independent variables, and ask for descriptive statistics.
How? (1) Click dependent variable and Dependent Variable, (2) click independent variable and Fixed Factor(s), (3) click and click Descriptive
791
Statistics, (4) click , and (5) click
4. Examine output.
792
12.12 Exercises
1. Create a line graph for each of the following tables of cell means (assume the scores have a possible range of 0 to 10 and put Factor A along the X (horizontal) axis).
a.
b.
c.
d.
2. Create a line graph for each of the following tables of cell means (assume the scores have a possible range of 0 to 10 and put Factor A along the X (horizontal) axis).
a.
b.
c.
d.
793
3. Create a table of cell means for each of the following graphs.
a.
b.
c.
d.
4. Create a table of cell means for each of the following graphs.
a.
794
b.
c.
d.
5. For each of the tables from Exercise 1, calculate the marginal means ( X ¯ A and X ¯ B ) and determine whether the main effects and A × B interaction are significant (in this exercise, assume any difference between means is significant).
a.
795
b.
c.
d.
6. For each of the tables from Exercise 2, calculate the marginal means ( X ¯ A and X ¯ B ) and determine whether the main effects and A × B interaction are significant (in this exercise, assume any difference between means is significant).
a.
b.
796
c.
d.
7. For each of the tables created in Exercise 3, calculate the marginal means ( X ¯ A and X ¯ B ) and determine whether the main effects and the A × B interaction are significant (in this exercise, assume any difference between means is significant).
a.
b.
c.
797
d.
8. For each of the tables created in Exercise 4, calculate the marginal means ( X ¯ A and X ¯ B ) and determine whether the main effects and the A × B interaction are significant (in this exercise, assume any difference between means is significant).
a.
b.
c.
798
d.
9. For each of the research situations below, identify the two independent variables and indicate the type of A × B design (e.g., 2 × 2, 3 × 3).
a. An instructor gives a final exam in which she randomly assigns students to one of two conditions: handwritten or typed. She also assesses and classifies students' typing ability: low or high typing ability. She hypothesizes that students with low typing ability will do better on the handwritten exam than the typed, but the opposite will be true for those with high typing ability.
b. A researcher examines how reading to young children affects their interest in books. The researcher believes that the age of the child being read to and the type of book being read might influence their interest. She conducts a study in which groups of 3-year-old, 5-year-old, and 7-year-old children either read a fiction book or a nonfiction book. She hypothesizes that for the youngest children, reading fiction books rather than nonfiction books will lead to a greater interest in reading. However, as children get older, this difference between the two types of books will grow smaller such that for the oldest children, the type of book will not make a difference.
c. A researcher hypothesizes that the effect of alcohol consumption on one's motor skills depends on the person's weight (underweight, normal, or overweight), but that this differs for men versus women. More specifically, for underweight people, men and women are equally affected. However, the more a male weighs, the greater the effect of alcohol, whereas for females, the effect stays the same no matter how much females weigh.
d. A researcher believes that the influence of drinking coffee on the ability to fall asleep depends on the amount of coffee one drinks and the time of day one drinks the coffee. She has groups of students drink either 1, 2, or 3 cups of coffee in the morning, afternoon, or evening. She hypothesizes that drinking coffee in the morning only has a slight effect on the later ability to fall asleep, and this is the same regardless of the amount of coffee one drinks. However, drinking coffee later in the day has a greater effect on the ability to fall asleep; furthermore, the more coffee one drinks later in the day, the greater the impact on this ability.
e. A researcher is interested in studying the relationship between students' motivation and their performance on different types of tests. She hypothesizes that students who have low motivation to achieve their goals will perform at a
799
level similar to students with high achievement motivation when a test is easy. However, on a hard test, students with low achievement motivation will do worse than they did on the easy test, whereas students high on achievement motivation will do better than they did on the easy test.
10. For each of the research situations below, identify the two independent variables and indicate the type of A × B design (e.g., 2 × 2, 3 × 3).
a. A gardener reads an article that states that lawns grow better when they are watered once a week for 30 minutes as opposed to three times a week for 10 minutes. From his experience, the gardener hypothesizes that which watering schedule is better depends on the size of the lawn (small vs. large). For small lawns, the type of watering schedule does not make a difference. However, for large lawns, it is better to water once a week for 30 minutes than three times a week for 10 minutes.
b. A teacher is interested in seeing how attending class and reading the assigned chapters is related to her students' performance on tests. She assesses each student in terms of how often they attend class (rarely, sometimes, regularly) and how often they read the assigned chapters (rarely, sometimes, regularly). She hypothesizes that students who attend class rarely will do poorly on tests, and they will do poorly regardless of how often they read the assigned chapters. She also believes students who attend class regularly will do well on tests, and they will do well regardless of how often they do the readings. However, for students who attend occasionally, she hypothesizes that the more they do the required readings, the better they will do on tests.
c. A fire chief reads research about how automobile accidents may be related to the color of the vehicles involved. Consequently, he wonders whether he should change the color of fire engines from red to yellow. Although he thinks both colors will be equally visible during the day, at night the yellow trucks will be easier to see than red trucks.
d. Let's say you wish to test the effects of consuming alcohol; more specifically, there is a psychological effect such that the expectation that one is drinking alcohol can influence aggression. To test this hypothesis, you have a group of people drink either a nonalcoholic beer or a regular beer. Regardless of what they are actually drinking, some of these people are told they are drinking nonalcoholic beer and some are told they are drinking regular beer. You later measure them on their level of aggression. You hypothesize that, regardless of which type of beer they are drinking, those who are told they are drinking regular beer will be more aggressive than those told they are drinking nonalcoholic beer.
800
a.
b.
c.
11. For each of the following tables of cell means, calculate marginal means ( X ¯ A and X ¯ B ) and the total mean ( X ¯ T ) .
a.
b.
c.
12. For each of the following tables of cell means, calculate marginal means ( X ¯ A and X ¯ B ) and the total mean ( X ¯ T ) ).
13. For each of the following, calculate four degrees of freedom (dfA, dfB, dfA × B, dfWG) and identify the three critical values of the three F-ratios (FA, FB, FA B) (assume α = .05).
a. a = 2, b = 2, NAB = 4
b. a = 2, b = 2, NAB = 18
c. a = 3, b = 2, NAB = 9
d. a = 2, b = 4, NAB = 6
801
14. For each of the following, calculate four degrees of freedom (dfA, dfB, dfA × B, dfWG) and identify the three critical values of the three F-ratios (FA, FB, FA × B) (assume α = .05).
a. a = 2, b = 2, NAB = 10
b. a = 2, b = 3, NAB = 5
c. a = 3, b = 3, NAB = 12
d. a = 3, b = 4, NAB = 5
15. For each of the following, calculate the F-ratios (F) for the two-way ANOVA, create an ANOVA summary table, and calculate measures of effect size (R2).
a. SSA = 12.00, dfA = 1, SSB = 24.00, dfB = 1, SSA × B = 8.00, dfA × B = 1, SSWG = 64.00, dfWG = 16
b. SSA = 3.91, dfA = 1, SSB = 5.64, dfB = 1, SSA × B = 3.12, dfA × B = 1, SSWG = 67.04, dfWG = 92
c. SSA = 4.23, dfA = 1, SSB = 3.72, dfB = 2, SSA × B = 1.81, dfA × B = 2, SSWG = 8.62, dfWG = 18
d. SSA = 52.78, dfA = 3, SSB = 18.58, dfB = 1, SSA × B = 78.06, dfA × B = 3, SSWG = 297.21, dfWG = 32
e. SSA = 78.43, dfA = 1, SSB = 61.90, dfB = 3, SSA × B = 27.39, dfA × B = 3, SSWG = 695.32, dfWG = 56
16. For each of the following, calculate the F-ratios (F) for the two-way ANOVA, create an ANOVA summary table, and calculate measures of effect size (R2).
a. SSA = 16.50, dfA = 1, SSB = 9.75, dfB = 1, SSA × B = 31.26, dfA × B = 1, SSWG = 337.82, dfWG = 56
b. SSA = 174.37, dfA = 2, SSB = 387.02, dfB = 1, SSA × B = 250.87, dfA × B = 2, SSWG = 2011.65, dfWG = 54
c. SSA = 21.90, dfA = 2, SSB = 13.29, dfB = 2, SSA × B = 36.09, dfA × B = 4, SSWG = 143.12, dfWG = 63
d. SSA = 125.87, dfA = 3, SSB = 176.31, dfB = 3, SSA × B = 359.00, dfA × B = 9, SSWG = 812.59, dfWG = 48
17. For the examples in Exercise 15, indicate the type of A × B design (e.g., 2 × 2, 3 × 3) and the number of scores in each combination of the two factors (NAB).
18. For the examples in Exercise 16, indicate the type of A × B design (e.g., 2 × 2, 3 × 3) and the number of scores in each combination of the two factors (NAB).
19. One study examined parents' use of discipline with their children (McKee et al.,
802
2007). They asked, “Do rates of harsh verbal and physical discipline differ by gender of parent and child?” (p. 188). They wrote that “with regard to frequency of use of harsh discipline, we propose that boys will receive more harsh discipline than girls, particularly from fathers” (p. 188). To examine this hypothesis, they asked a sample of boys and girls to indicate how often their mother or father used harsh discipline with them. The frequency of harsh discipline by mothers and fathers reported by the boys and girls is summarized below:
Child (Factor A): Boy Boy Girl Girl
Parent (Factor B): Mother Father Mother Father
N AB 6 6 6 6
Mean ( X ¯ AB ) 2.83 4.67 2.17 2.00
Standard deviation (SAB) 1.17 .52 1.17 1.10
a. Create a figure to represent the cell means. b. Conduct the two-way ANOVA and create an ANOVA summary table. c. Report the decisions regarding the null hypotheses and the level of significance
for the main effects and the interaction effect. d. Calculate measures of effect size (R2) for the main effects and the interaction
effect.
20. Does playing violent video games lead to greater aggressive behavior? A team of researchers believed greater amount of time playing these games leads to higher feelings of hostility (Barlett, Harris, & Baldassaro, 2007). Furthermore, they hypothesized this increase in hostility is greater for people playing games with realistic-looking weapons rather than standard game controllers. They had groups of people play a violent video game for either 15 minutes of 30 minutes; some of these people used a standard game controller and others used a gun-shaped controller; this is a 2 (Controller: Standard vs. Gun) × 2 (Time: 15 min vs. 30 min) design. After playing the game, their level of hostility was measured.
Time (Factor A): 15 mins 15 mins 30 mins 30 mins
Controller (Factor B): Standard Gun Standard Gun
N AB 5 5 5 5
Mean ( X ¯ AB ) ) 10.60 13.60 16.40 25.60
Standard deviation (SAB) 3.51 3.05 2.70 3.36
a. Create a figure to represent the cell means. b. Conduct the two-way ANOVA and create an ANOVA summary table.
803
c. Report the decisions regarding the null hypotheses and the level of significance for the main effects and the interaction effect.
d. Calculate measures of effect size (R2) for the main effects and the interaction effect.
21. A researcher is interested in examining how teachers' expectations of their students' scholastic abilities might affect students' self-perceptions of their own ability and whether this effect varies as a function of a student's age. She has a sample of teachers rate their students as having either low or high expectations of success; these teachers are from the first, third, and fifth grades. She then has students rate their own abilities on a scale of 1 (low) to 15 (high). She hypothesizes that students with low expectations by the teacher would display lower ratings of their own ability than students with high teacher expectations. Furthermore, she believes this difference will increase as students progress through the educational system, growing larger from first to third to fifth grade. Below are descriptive statistics of students' ratings of their ability.
Grade (Factor A): 1st 1st 3rd 3rd 5th 5th
Expectation (Factor B): Low High Low High Low High
N AB 3 3 3 3 3 3
Mean ( X ¯ AB ) 5.00 6.33 3.67 8.67 2.33 9.67
Standard deviation (SAB) 1.00 1.53 1.53 1.53 1.15 1.53
a. Create a figure to represent the cell means. b. Conduct the two-way ANOVA and create an ANOVA summary table. c. Report the decisions regarding the null hypotheses and the level of significance
for the main effects and the interaction effect. d. Calculate measures of effect size (R2) for the main effects and the interaction
effect.
22. Concerns regarding the possibility of getting skin cancer have influenced beliefs regarding the positive benefits of getting a suntan. One study examined whether men and women have similar beliefs regarding whether a woman's suntan influences perceptions of her physical attractiveness (Banerjee, Campo, & Greene, 2008). They showed men and women one of three photographs: a woman with no tan, a medium tan, or a dark tan; next, they rated the woman in terms of her physical attractiveness. The researchers hypothesized that, for men (but not for women), the darker a woman's tan, the more she is seen as physically attractive.
Gender (Factor Male Male Male Female Female Female
804
A):
Tan (Factor B): No Tan
Medium Dark No Tan
Medium Dark
N AB 4 4 4 4 4 4
Mean ( X ¯ A B ) )
4.75 5.00 7.50 7.25 6.75 6.50
Standard deviation (SAB)
.96 .82 1.29 1.26 .96 1.29
a. Create a figure to represent the cell means. b. Conduct the two-way ANOVA and create an ANOVA summary table. c. Report the decisions regarding the null hypotheses and the level of significance
for the main effects and the interaction effect. d. Calculate measures of effect size (R2) for the main effects and the interaction
effect.
23. Characteristics of students may influence how favorably they're perceived by their teachers. One study examined whether teachers' perceptions of students may be influenced by the student's gender (male, female) and socioeconomic status (SES) (lower class, upper class) (Auwarter & Aruguete, 2008). They had a sample of teachers read a passage about a troubled student; in these passages, the researchers varied the student's gender and SES. After reading the passage, the teacher rated the student on his or her personal characteristics on a 1 to 5 scale; the higher the rating, the more positively the teacher viewed the student. The researchers hypothesized that male students would be rated more highly if they were of upper rather than lower SES, but the reverse was true for female students.
SES (Factor A): Lower Lower Upper Upper
Gender (Factor B): Male Female Male Female
3 4 4 4
4 2 3 2
3 4 4 2
2 5 3 4
5 3 3 3
3 4 4 2
2 5 5 3
a. Calculate descriptive statistics ( N AB , X ¯ AB , s AB ) for each combination of
805
the two factors. b. Create a figure to represent the cell means. c. Conduct the two-way ANOVA and create an ANOVA summary table. d. Report the decisions regarding the null hypotheses and the level of significance
for the main effects and the interaction effect. e. Calculate measures of effect size (R2) for the main effects and the interaction
effect.
24. When children see adults arguing, what do they think or expect will happen? Two researchers had elementary school children watch a videotape of a man and a woman arguing (El-Sheikh & Elmore-Staton, 2007). In some of the videotapes, the adults were intoxicated, but in some, the adults were sober. After watching the videotape, each child was asked how likely either the man or the woman would be verbally or physically aggressive (the higher the score, the greater the likelihood of aggression). These researchers hypothesized that children expect higher levels of aggression when they thought the adults were intoxicated versus if they were sober, particularly if the intoxicated adult was a man as opposed to a woman.
Condition (Factor A): Intoxicated Intoxicated Sober Sober
Sex of Adult (Factor B): Man Woman Man Woman
2 5 2 3
4 4 1 5
3 5 3 3
6 6 2 2
3 3 3 3
4 5 3 5
a. Calculate descriptive statistics ( N AB , X ¯ AB , s AB ) for each combination of the two factors.
b. Create a figure to represent the cell means. c. Conduct the two-way ANOVA and create an ANOVA summary table. d. Report the decisions regarding the null hypotheses and the level of significance
for the main effects and the interaction effect. e. Calculate measures of effect size (R2) for the main effects and the interaction
effect.
25. One way of improving job performance in organizations is through the use of “self- managed work teams”: teams of employees given a certain amount of control (autonomy) over how they conduct their work activities. Previous research has suggested too little autonomy leads to bored, unmotivated employees; however, too
806
much autonomy may result in a lack of structure and supervision. A researcher hypothesizes that the effect of autonomy on team performance may in fact depend on the size of the work team. She designs an experiment whereby teams are given a simple task to perform (the construction of a toy boat). Participants are randomly assigned to one of three sized work teams: Small (4 members), Medium (7 members), or Large (10 members). Each group is given either low or high autonomy in determining how the boats are to be constructed. She counts the number of toy boats built by each team in the designated time.
Team Size (Factor A):
Small Small Medium Medium Large Large
Autonomy (Factor B):
Low High Low High Low High
2 4 4 6 7 1
3 5 2 2 5 2
1 5 4 5 4 2
0 3 3 3 3 3
a. Calculate descriptive statistics ( N AB , X ¯ AB , s AB ) for each combination of the two factors.
b. Create a figure to represent the cell means. c. Conduct the two-way ANOVA and create an ANOVA summary table. d. Report the decisions regarding the null hypotheses and the level of significance
for the main effects and the interaction effect. e. Calculate measures of effect size (R2) for the main effects and the interaction
effect.
26. Earlier in this chapter, we mentioned a study looking at students' use of rituals related to taking exams (Rudski & Edwards, 2007). The researchers believed that two factors related to students' use of rituals are the importance of the exam to the student's grade in the course and the perceived difficulty of the exam. More specifically, they hypothesized that when an exam is low in importance, students will only engage in rituals if the exam is very difficult; however, when an exam is high in importance, students will engage in a high level of rituals regardless of the difficulty of the exam. Listed below are the number of rituals performed by students preparing to take exams either low or high in importance to their course grades and believed to be either easy, somewhat difficult, or very difficult:
Importance (Factor A):
Low Low Low High High High
807
Difficulty (Factor B):
Easy Somewhat Very Easy Somewhat Very
0 2 4 3 4 3
2 3 3 5 3 4
1 0 5 3 5 4
0 2 4 2 4 5
3 2 3 5 4 4
2 1 4 4 5 6
a. Calculate descriptive statistics ( N AB , X ¯ AB , s AB ) for each combination of the two factors.
b. Create a figure to represent the cell means. c. Conduct the two-way ANOVA and create an ANOVA summary table. d. Report the decisions regarding the null hypotheses and the level of significance
for the main effects and the interaction effect. e. Calculate measures of effect size (R2) for the main effects and the interaction
effect. 27. For each of the research situations described in Exercises 19 to 22, draw a figure that
illustrates simple effects analyses appropriate to test the study's research hypothesis. 28. For each of the research situations described in Exercises 23 to 25, draw a figure that
illustrates simple effects analyses appropriate to test the study's research hypothesis.
808
Answers to Learning Checks
Learning Check 1
2.
a.
b.
c.
d.
3. a. Present b. Absent c. Absent d. Present
4.
809
a.
b.
c.
d.
Learning Check 2
2. a. Type of course (Classroom, Online) and Living arrangement (Off-campus,
On-campus) (2 × 2) b. Attribution of responsibility (Child, Parent) and Age of Child (Adolescent,
Teenager) (2 × 2) c. Perceived difficulty of exam (Easy, Somewhat difficult, Very difficult) and
Importance of exam (Important, Unimportant) (3 × 2) 3.
810
a.
b.
c.
d.
4.
a. Degrees of freedom: dfA = 1 dfB = 1 dfA × B = 1 dfWG = 16
Critical values: FA = 4.49 FB = 4.49 FA × B = 4.49
b. Degrees of freedom: dfA = 2 dfB = 1 dfA × B = 2 dfWG = 30
Critical values: FA = 3.32 FB = 4.17 FA × B = 3.32
c. Degrees of freedom: dfA = 2 dfB = 2 dfA × B = 4 dfWG = 81
Critical values: FA = 3.11 FB = 3.11 FA × B = 2.48
d. Degrees of freedom: dfA = 3 dfB = 2 dfA × B = 6 dfWG = 72
Critical values: FA = 2.76 FB = 3.15 FA × B = 2.25
Learning Check 3
811
2. a.
Source SS df MS F R2
Factor A 20.00 1 20.00 5.00 .16
Factor B 10.00 1 10.00 2.50 .08
A × B interaction 15.00 1 15.00 3.75 .12
Error 80.00 20 4.00
Total 125.00 23
b.
Source SS df MS F R2
Factor A 54.00 1 54.00 5.94 .11
Factor B 25.00 1 25.00 2.75 .05
A × B interaction 16.00 1 16.00 1.76 .03
Error 400.00 44 9.09
Total 495.00 47
c.
Source SS df MS F R2
Factor A 4.60 2 2.30 2.20 .07
Factor B 9.79 1 9.79 9.37 .14
A × B interaction 10.31 2 5.16 4.94 .15
Error 43.87 42 1.04
Total 68.57 47
3.
812
a.
b.
Source SS df MS F
Sex 5.12 1 5.12 2.60
Type of car 20.59 1 20.59 10.45
Sex × Type of car 11.55 1 11.55 5.86
Error 47.29 24 1.97
Total 84.55 27
c. Main effect of Sex: FA = 2.60 < 4.26 ∴ do not reject H0 (p > .05)
Main effect of Type of car: FB = 10.45 > 7.82 ∴ reject H0 (p < .01)
Sex × Type of car interaction: FA × B = 5.86 > 4.26 ∴ reject H0 (p < .05)
d. Main effect of Sex: R A 2 = .06
Main effect of Type of car: R B 2 = .24
Sex × Type of car interaction: R A × B 2 = .14
Learning Check 4
2.
a.
813
b.
c.
814
Answers to Odd-Numbered Exercises
1.
a.
b.
c.
d.
3.
a.
b.
815
c.
d.
5.
a.
b.
c.
d.
7.
816
a.
b.
c.
d.
9. a. Exam (Hand-written, Typed) and Typing ability (Low, High) (2 × 2) b. Age (3, 5, 7) and Type of book (Fiction, Nonfiction) (3 × 2) c. Weight (Underweight, Normal, Overweight) and Gender (Male, Female) (3 ×
2) d. Amount of coffee (1, 2, or 3 cups) and Time of day (Morning, Afternoon,
Evening) (3 × 3) e. Achievement motivation (Low, High) and Type of test (Easy, Hard) (2 × 2)
11.
817
a.
b.
c.
13. a.
Degrees of freedom:
dfA = 1 dfB = 1 dfA × B = 1 dfWG = 12
Critical values: FA = 4.75 FB = 4.75 FA × B = 4.75
b.
Degrees of freedom:
dfA = 1 dfB = 1 dfA × B = 1 dfWG = 68
Critical values: FA = 4.00 FB = 4.00 FA × B = 4.00
c.
Degrees of freedom:
dfA = 2 dfB = 1 dfA × B = 2 dfWG = 48
Critical values: FA = 3.23 FB = 4.08 FA × B = 3.23
d.
Degrees of freedom:
dfA = 1 dfB = 3 dfA × B = 3 dfWG = 40
818
Critical values: FA = 4.08 FB = 2.84 FA × B = 2.84
15. a.
Source SS df MS F R2
Factor A 12.00 1 12.00 3.00 .11
Factor B 24.00 1 24.00 6.00 .22
A × B interaction 8.00 1 8.00 2.00 .07
Error 64.00 16 4.00
Total 108.00 19
b.
Source SS df MS F R2
Factor A 3.91 1 3.91 5.37 .05
Factor B 5.64 1 5.64 7.74 .07
A × B interaction 3.12 1 3.12 4.28 .04
Error 67.04 92 .73
Total 79.71 95
c.
Source SS df MS F R2
Factor A 4.23 1 4.23 8.83 .23
Factor B 3.72 2 1.86 3.88 .20
A × B interaction 1.81 2 .91 1.89 .10
Error 8.62 18 .48
Total 18.38 23
d.
Source SS df MS F R2
Factor A 52.78 3 17.59 1.89 .12
Factor B 18.58 1 18.58 2.00 .04
A × B interaction 78.06 3 26.02 2.80 .17
819
Error 297.21 32 9.29
Total 446.63 39
e.
Source SS df MS F R2
Factor A 78.43 1 78.43 6.32 .09
Factor B 61.90 3 20.63 1.66 .07
A × B interaction 27.39 3 9.13 .74 .03
Error 695.32 56 12.42
Total 863.04 63
17. a. 2 × 2 design, NAB = 5 b. 2 × 2 design, NAB = 24 c. 2 × 3 design, NAB = 4 d. 4 × 2 design, NAB = 5 e. 2 × 4 design, NAB = 8
19.
a.
b.
Source SS df MS F
Factor A 16.67 1 16.67 15.88
Factor B 4.17 1 4.17 3.97
A × B interaction 6.00 1 6.00 5.17
Error 12.00 20 1.05
820
Total 47.83 23
c. Main effect of Child: FA = 15.88 > 8.10 ∴ reject H0 (p < .01)
Main effect of Parent: FB = 3.97 < 4.35 ∴ do not reject H0 (p > .05)
Child × Parent interaction: FA × B = 5.71 > 4.35 ∴ reject H0 (p < .05)
d. Main effect of Child: R A 2 = .35
Main effect of Parent: R B 2 = .09
Child × Parent interaction: R A × B 2 = .13 21.
a.
Source SS df MS F
Grade .78 2 .39 .20
Expectation 93.39 1 93.39 48.04
Grade × Expectation 27.44 2 13.72 7.06
Error 23.33 12 1.94
Total 144.94 17
c.
Main effect of Grade: FA = .20 < 3.89 ∴ do not reject H0 (p > .05)
Main effect of Expectation: FB = 48.04 > 9.33 ∴ reject H0 (p < .01)
Grade × Expectation interaction: FA × B = 7.06 > 6.93 ∴ reject H0 (p < .01) d.
821
Main effect of Grade: R A 2 = .01
Main effect of Expectation: R B 2 = .64
Grade × Expectation interaction: R A × B 2 = .19 23.
a.
SES (Factor A): Lower Lower Upper Upper
Gender (Factor B): Male Female Male Female
NAB 7 7 7 7
Mean ( X ¯ A B ) ) 3.14 3.86 3.71 2.86
Standard deviation (SAB) 1.07 1.07 .76 .90
b.
c.
Source SS df MS F
SES .32 1 .32 .35
Gender .04 1 .04 .04
SES × Gender 4.32 1 4.32 4.71
Error 22.00 24 .92
Total 26.68 27
d. Main effect of SES: FA = .35 < 4.23 ∴ do not reject H0 (p > .05)
Main effect of Gender: FB = .04 < 4.23 ∴ do not reject H0 (p > .05)
SES × Gender interaction: FA × B = 4.71 > 4.23 ∴ reject H0 (p < .05)
e. Main effect of SES: R A 2 = .01
822
Main effect of Gender: R B 2 = .001
SES × Gender interaction: R A × B 2 = .16 25.
a.
Team size (Factor A):
Small Small Medium Medium Large Large
Autonomy (Factor B):
Low High Low 4 High 4 Low 4
High
NAB 4 4 4 4 4 4
Mean ( X ¯ A B ) )
1.50 4.25 3.25 4.00 2.00 3.33
Standard deviation (SAB)
1.29 .96 .96 1.83 1.71 1.82
b.
c.
Source SS df MS F
Team size 2.33 2 1.17 .67
Autonomy .38 1 .38 .22
Team size × Autonomy 31.00 2 15.50 8.93
Error 13.25 18 1.74
Total 64.96 23
d. Main effect of Team size: FA =.67 < 3.55 ∴ do not reject H0 (p > .05)
Main effect of Autonomy: FB = .22 < 4.41 ∴ do not reject H0 (p > .05)
Team size × Autonomy interaction: FA × B = 8.93 > 6.01 ∴ reject H0 (p < .01)
823
e. Main effect of Team size: R A 2 = .04
Main effect of Autonomy: R B 2 = .01
Team size × Autonomy interaction: R A × B 2 = .48 27.
a. Simple effect of Child (Factor A) at each level of Parent (Factor B).
b. Simple effect of Time (Factor A) at each level of Controller (Factor B).
c. Simple effect of Expectation (Factor B) at each level of Grade (Factor A).
d. Simple effect of Tan (Factor B) at each level of Grade (Factor A).
824
825
Sharpen your skills with SAGE edge at edge.sagepub.com/tokunaga
SAGE edge for students provides a personalized approach to help you accomplish your coursework goals in an easy-to-use learning environment. Log on to access:
eFlashcards Web Quizzes Chapter Outlines Learning Objectives Media Links SPSS Data Files
826
Chapter 13 Correlation and Linear Regression
827
Chapter Outline 13.1 An Example From the Research: Snap Judgment 13.2 Introduction to the Concept of Correlation
Describing the relationship between variables Nature of the relationship Direction of the relationship Strength of the relationship
Measuring the relationship between variables: correlational statistics Introduction to the Pearson correlation coefficient (r) Calculating correlational statistics: the role of variance and covariance
13.3 Inferential Statistics: Pearson Correlation Coefficient State the null and alternative hypotheses (H0 and H1) Make a decision about the null hypothesis
Calculate the degrees of freedom (df) Set alpha (α), identify the critical values, and state a decision rule Calculate a statistic: Pearson correlation coefficient (r) Make a decision whether to reject the null hypothesis Determine the level of significance
Calculate a measure of effect size (r2) Draw a conclusion from the analysis Relate the result of the analysis to the research hypothesis A computational formula for the Pearson r
Represent the covariance between X and Y (SPXY) Represent the variance of Variable X (SSX) Represent the variance of Variable Y (SSY) Calculate the Pearson correlation coefficient (r)
13.4 Predicting One Variable From Another: Linear Regression The linear regression equation Calculating the linear regression equation Calculate the slope of the equation (b) Calculate the Y-intercept of the equation (a) Report the linear regression equation Drawing the linear regression equation A second example: put down your pencils Calculating the linear regression equation Drawing the linear regression equation
13.5 Correlating Two Sets of Ranks: The Spearman Rank-Order Correlation Introduction to ranked variables An example from the research: do students and teachers think alike? Inferential statistics: the Spearman rank-order correlation coefficient State the null and alternative hypotheses (H0 and H1) Make a decision about the null hypothesis Draw a conclusion from the analysis Relate the result of the analysis to the research hypothesis
13.6 Correlational Statistics vs. Correlational Research Does ANOVA = causality? “Correlation is not causality”—what does that mean?
13.7 Looking Ahead 13.8 Summary
828
13.9 Important Terms 13.10 Formulas Introduced in This Chapter 13.11 Using SPSS 13.12 Exercises
The research examples analyzed in earlier chapters of this book share one feature: They have all involved comparing the means of groups. With independent variables such as Group (Intruder, No intruder), Condition (Observation, Prediction, Explanation), and Message type (Reward, Threat), the research and statistical hypotheses in these studies addressed whether any differences in group means on the dependent variable were statistically significant. The data in these studies were analyzed with different versions of the analysis of variance (ANOVA), which is used when the independent variable or variables are categorical in nature. This chapter introduces statistical procedures that are used when the independent variable is continuous rather than categorical.
829
13.1 An Example from the Research: Snap Judgment
There's an old saying that describes what might happen when you meet someone for the first time: “You don't get a second chance to make a first impression.” People often make distinct and long-lasting judgments about others based on their initial exposure to them. Researchers, in turn, have studied two questions related to these judgments. First, how much time do people need to form first impressions of others? Second, how accurate are these impressions?
For example, researchers at Harvard University videotaped college instructors teaching their classes and created three 2-second clips of these videotapes with the sound removed (Ambady & Rosenthal, 1993). The researchers showed these clips to students who were not taking classes from these instructors and asked them to rate the instructors on a number of personality characteristics. The researchers found that the ratings of these students, based on three 2-second silent exposures to a person they had never met, were similar to those of students who had taken classes from these instructors for an entire semester.
Following up on this and other related research, a team of researchers at Harvard Medical School led by Moshe Bar conducted an experiment to determine how quickly people can form initial impressions (Bar, Neta, & Linz, 2006). More specifically, they wanted to discover how quickly people can form impressions of another person based strictly on the sight of the other person's face. In addition, they wished to determine if these impressions are similar to those made by people with a greater amount of exposure to the same face.
In this study, which we'll call the first impression study, research participants were shown the faces of 24 people on a computer screen. To explore how quickly first impressions may be made, the amount of time each face was shown to a participant was either 39 milliseconds (39 ms or 39/1,000ths of a second) or 1700 milliseconds (1700 ms or 1.7 seconds). To give a sense of the duration involved in 39 ms, imagine you're a baseball player batting against a pitcher who throws a ball 90 miles per hour. Once the ball is thrown, you have 4/10ths of a second to decide whether to swing your bat. Now imagine that the pitcher can throw the ball at 900 miles per hour—that's how fast the ball would have to be thrown for you to have only 39 ms to decide whether to swing.
The participants in this study rated each face in terms of the extent to which they believed it belonged to a threatening person (1 = least threatening to 5 = most threatening). The 24 faces were shown to groups of people in the 39 ms and 1700 ms conditions; for each face, the average threatening rating was calculated for each of the two groups. Therefore, the data in this study consist of two pieces of information for each of the 24 faces: the average threatening rating for those shown the face for 39 ms and the average threatening rating for those shown the face for 1700 ms. These two threatening ratings (which we will refer to simply as “39 ms” and “1700 ms”) are both continuous in nature, with a possible range of
830
values falling along a numeric continuum (1.00 [least threatening] to 5.00 [most threatening]). Note that our current example differs from the examples in previous chapters, in which at least one of the variables has been categorical or qualitative in nature, with the two variables comprising different groups.
The threatening ratings for the 24 faces for the 39 ms and 1700 ms variables are listed in Table 13.1(a). Each of the participants (in this case, the 24 faces) has two scores. For example, the first face had a threatening rating of 2.00 for participants who saw the face for 39 ms and a threatening rating of 2.40 for those seeing the face for 1700 ms.
Descriptive statistics for the two variables are provided in Table 13.1(b). In earlier chapters, descriptive statistics were calculated for a single variable for each of the groups that make up an independent variable. In Chapter 9, for example, the average time taken to leave a parking space was calculated for the Intruder and No intruder conditions. In the first impression study in this chapter, however, data have been collected not on one variable but rather on two: the threatening rating for the 39 ms condition and the threatening rating for the 1700 ms condition. As such, the two variables must be distinguished from each other. In Table 13.1(b), you see that the 39 ms variable has been labeled Variable X and the 1700 ms variable has been labeled Variable Y. The means for the two variables are represented by the symbols X ¯ and Y ¯ , and the two standard deviations (sX and sY) have the subscripts X and Y.
It is useful to create a figure to gain an initial understanding of data that have been collected. In previous chapters, a bar graph was used to compare the means of groups that comprise an independent variable or a combination of two independent variables. However, the goal of the first impression study is not to compare groups on a dependent variable but rather to see whether scores on one variable are associated with scores on another variable. We might ask, for example, whether the threatening ratings for a face shown for the relatively shorter duration of 39 ms are similar to the ratings for the same face when it is shown for the relatively longer duration of 1700 ms.
When both variables being analyzed are continuous in nature, which means they are measured along a numeric continuum, the relationship between scores on the variables may be illustrated by creating a figure known as a scatterplot: a graphical display of paired scores on two variables. Figure 13.1 presents a scatterplot of the data for the 39 ms and 1700 ms variables: 39 ms (the X variable) is located along the horizontal axis, and 1700 ms (the Y variable) has been placed along the vertical axis. The location of each face within the scatterplot is based on its two respective pieces of information, with each dot in the scatterplot representing a particular face. For example, the dot in the furthest lower left of the scatterplot represents the pair of scores for the first face in Table 13.1(a), which had a threatening rating of 2.00 for the 39 ms condition and an threatening rating of 2.40 for the 1700 ms condition.
831
Table 13.1 Threatening Ratings for Faces in the First Impression Study Table 13.1 Threatening Ratings for Faces in the First Impression Study
(a) Raw Data
Face 39 ms 1700 ms
1 2.00 2.40
2 2.53 2.80
3 3.17 3.38
4 2.43 3.43
5 4.02 4.00
6 3.01 3.18
7 2.80 2.77
8 3.44 3.88
9 2.77 2.30
10 3.27 3.56
11 2.48 3.25
12 3.00 3.34
13 2.55 2.35
14 2.27 3.08
15 2.40 2.06
16 3.45 3.75
17 2.36 3.30
18 4.76 4.34
19 2.83 3.00
20 3.17 3.59
21 2.38 2.83
22 3.15 3.19
23 3.26 3.06
24 2.45 2.60 Table 13.1 Threatening Ratings for Faces in the
First Impression Study
(b) Descriptive Statistics
39 ms (X) 1700 ms (Y)
832
N 24 24
Mean X ¯ = 2.92 Y ¯ = 3.14
Standard deviation sX = .62 sY = .57
Examining a figure such as a scatterplot provides an initial indication of how scores on two variables are associated with each other. Looking at the scatterplot in Figure 13.1, we observe that there is a general tendency for the ratings of a face to be similar for the 39 ms and 1700 ms conditions. For example, a low (or high) threatening rating for a face in the 39 ms condition is associated with a low (or high) threatening rating for the same face in the 1700 ms condition. The next section discusses how the relationship between two variables may be described and characterized.
Figure 13.1 Scatterplot of Threatening Ratings for the 39 ms and 1700 ms Conditions in the First Impression Study
833
13.2 Introduction to the Concept of Correlation
A scatterplot provides an initial indication of whether scores on one variable are related to scores on another variable. The purpose of this section is to introduce the concept of correlation, which may be defined as a mutual or reciprocal relationship between two variables such that systematic changes in the values of one variable are accompanied by systematic changes in the values of another variable. This section will discuss how the relationship between variables may be described.
834
Describing the Relationship between Variables
In this book, the relationship between variables will be described along three aspects: the nature of the relationship, the direction of the relationship, and the strength of the relationship. Each of these three aspects is described and illustrated below.
Nature of the Relationship
The nature of the relationship between variables pertains to the manner in which changes in scores on one variable correspond to changes in scores on another variable. The nature of the relationship between two variables may take on different forms. For example, a linear relationship is a relationship between variables that is appropriately represented by a straight line, such that increases or decreases in scores for one variable are associated with corresponding increases or decreases in scores for another variable. The scatterplot in Figure 13.1 for the first impression study illustrates a linear relationship, in that increases in the threatening ratings for the faces shown for 39 ms are associated with increases in the threatening ratings for the faces shown for 1700 ms.
Figure 13.2(a) also illustrates a linear relationship between two variables. It is considered a linear relationship because the data in the scatterplot are located close to a line that has been generated to represent the relationship. Note that the line is tilted at an angle, meaning it is not parallel to either the X- or the Y-axis. The angle of the line indicates that changes in one variable are associated with changes in the other variable, which, as you will recall, is the definition of correlation.
What would be the nature of the relationship between two variables if scores on one of the variables do not change? For example, what if we wanted to relate students' scores on a midterm exam to their scores on a final exam, but every student had exactly the same score on the midterm? Due to the lack of differences in midterm scores (the X variable), we can't use this variable to understand any differences on the final exam (the Y variable). In this situation, a line through the data would be perfectly vertical, meaning it is parallel to the Y- axis.
If, on the other hand, every student had exactly the same score on the final exam (the Y variable), we would not need to know any student's score on the midterm exam (the X variable) to predict how well he or she did on the final. Instead, we would simply make the same prediction for every student, regardless of his or her midterm score. As a result, a line through the data would be perfectly horizontal, meaning it is parallel to the X-axis. From these two examples, we are able to observe an important principle to be used in comparing the relationships between two variables: To relate one variable to another, there must be differences in the scores of both variables.
835
Figure 13.2 Nature of the Relationship between Two Variable
It is also possible for the nature of the relationship between two variables to not be linear. A nonlinear relationship is a relationship between variables that is not appropriately represented by a straight line. For example, how many hours of sleep do you need each night to be at your best? One hour? Fourteen hours? For most of us, the answer lies somewhere between these two extremes, such that too little sleep makes one inattentive but too much sleep makes one lethargic. A proposed nonlinear relationship between “number of hours of sleep” and “alertness” is portrayed in Figure 13.2(b).
Figure 13.3 Direction of the Relationship between Two Variables
Although nonlinear relationships certainly do exist, linear relationships are more common. For this reason, the correlational statistical procedures discussed later in this chapter assume that the relationship between two variables is linear in nature. Statistics have been developed to measure nonlinear relationships between two variables; discussion of these statistics may be found in advanced books on correlational procedures (see, e.g., Cohen & Cohen, 1983, chap. 6; Hays, 1988, pp. 698–701; Keppel & Zedeck, 1989, pp. 50–54; McNemar, 1969, pp. 315–316).
Direction of the Relationship
The second aspect of the relationship between variables refers to the direction in which changes in one variable are associated with changes in another. A positive relationship is a relationship in which increases in the scores for one variable are associated with increases in the scores of another variable. The scatterplot in Figure 13.3(a) is an example of a positive relationship. Note that the word “positive” is not used to evaluate the relationship—it indicates the direction of the relationship rather than its “value” or “goodness.”
836
One study found a positive relationship between television viewing and diet such that the greater the number of hours a day children spent watching television, the more they expressed a preference for unhealthy foods such as cakes, candy, cookies, and soda (Signorielli & Staples, 1997). The combined use of the words “greater” and “more” describes a positive relationship, one in which the scores on two variables move in the same direction.
A negative relationship, on the other hand, is a relationship in which increases in the scores for one variable are associated with decreases in the scores of another variable. Consistent with the use of the word “positive” in our previous example, the word “negative” does not imply that the relationship is somehow “worse” or “less desirable” than a “positive relationship”. Referring back to the Signorielli and Staples (1997) study, what if the Y variable was measured as the preference for healthy foods rather than unhealthy foods? If so, the researchers would have observed a negative relationship: As the amount of TV watching increases, the preference for healthy foods decreases. In a negative relationship, such as the one depicted in Figure 13.3(b), the scores on two variables move in opposite directions.
Figure 13.4 Strength of the Relationship between Two Variables
Strength of the Relationship
The third way in which the relationship between variables may be described is in terms of its strength. The strength of a relationship is the extent to which scores on one variable are associated with scores on another variable. The scatterplot in Figure 13.4(a) illustrates what is known as a “perfect” relationship, which is a relationship in which each score for one variable is associated with one and only one score for the other variable. You will note from this figure that all of the data in a perfect relationship fall exactly on the line. What if, for example, a perfect relationship existed between people's height and their weight? If so,
837
everyone who was 4 feet tall would weigh exactly the same amount (perhaps 90 pounds), all people who were 6 feet tall would weigh exactly 205 pounds, and so on. In such a case, knowing a person's height would enable one to perfectly predict that person's weight, and vice versa.
As you probably suspect, perfect relationships are the incredibly rare exception rather than the rule. The scatterplot in Figure 13.4(b) represents what may be referred to as a “strong” relationship, which is a relationship in which a score on one variable is associated with a relatively small range of scores on another variable. If we were to encapsulate all of the points in this scatterplot, it would assume a narrow elliptical shape. When there is a strong relationship between two variables, knowing the score on one variable allows us to predict the score on the other variable within a small range. The frequently observable correlation between height and weight (e.g., the expectation that a man who is 5'10? will typically weigh between 170 and 190 pounds) is an example of a strong relationship.
In contrast, Figures 13.4(c) demonstrates a “moderate” relationship—note that as a relationship becomes weaker, the encapsulating representation of the relationship becomes less elliptical and more rounded. Finally, the last scatterplot (Figure 13.4(d)) illustrates a zero relationship between two variables. A “zero” relationship exists when all of the scores on one variable are associated with a wide range of scores on another variable. If we were to encapsulate all of the points in this scatterplot, it would take on the shape of a circle. For example, what if we were asked to predict a person's IQ from his or her height? Because there is no relationship between people's heights and their intelligence, any particular height will inevitably be associated with a wide range of IQs. When there is no relationship between two variables, knowing the score on one variable does not allow us to predict the score on the other variable with any degree of precision.
838
Measuring the Relationship between Variables: Correlational Statistics
Scatterplots are very useful in gaining an initial sense of the relationship between variables. To measure and test these relationships, however, researchers calculate statistics known as correlational statistics, which are statistics designed to measure relationships between variables. There are a number of different correlational statistics, with the choice of statistic depending on how variables in a study have been measured. The most commonly used correlational statistic was developed in 1895 by Karl Pearson, a very influential mathematician who, among his many achievements, founded the world's first university statistics department at University College in London in 1911. This statistic is introduced and defined in the next section.
Introduction to the Pearson Correlation Coefficient (r)
The Pearson correlation coefficient, represented by the symbol r, is a statistic that measures the linear relationship between two continuous variables measured at the interval and/or ratio level of measurement. The Pearson r is designed to measure the three previously mentioned aspects of a relationship mentioned earlier: nature, direction, and strength.
First, in terms of the nature of the relationship, the Pearson r assesses the degree of linear relationship between two variables. It assumes that the relationship may be represented by a straight line, such that increases (or decreases) in one variable correspond to increases (or decreases) in the other variable (see Figure 13.2(a)). If the relationship between the variables is not linear, the Pearson correlation coefficient would not be the appropriate statistic to measure the relationship. This highlights the importance of creating and examining a scatterplot of data that have been collected before calculating any statistics on the data.
Second, the sign (+ or –) of a Pearson correlation coefficient indicates the direction of the relationship, with the possible values for the Pearson r being either positive or negative. For example, a positive value of r (e.g., .47, .10, or .64) indicates a positive relationship like the one depicted in Figure 13.3(a). Negative values of r (e.g., –.53, –.25, or –.09) represent negative relationships like the one in Figure 13.3(b). Similar to the z-score and t-test, the sign of the correlation is typically only reported if it is a negative (–) value.
Finally, the strength of the relationship is represented by the numeric value of the correlation, with the possible values of the Pearson r ranging from −1.00 to +1.00. A correlation of either +1.00 or −1.00 represents the strongest relationship possible: a perfect relationship. Note that both +1.00 and −1.00 represent perfect relationships, in that the sign (+ or –) indicates the direction of the relationship rather than its strength. Figure
839
13.4(a) presents a perfect positive relationship in which r = 1.00. Between the two maximum values of −1.00 and 1.00, the specific value of the Pearson r represents the strength of the relationship. For example, a correlation with the value of .35 represents a stronger relationship than a correlation of .29, a weaker relationship than a correlation of –.46, and the same strength as a correlation of –.35. The midpoint of the range of possible values of the Pearson r is zero (.00), which implies the complete absence of a relationship between the two variables (see Figure 13.4(d)).
Jacob Cohen (1988), who also pioneered the measures of effect size discussed in Chapter 10, provided the following guidelines for interpreting values of the Pearson r in terms of its strength:
Pearson r
Strength of Relationship Negative Positive
Weak –.10 to –.30 .10 to .30
Moderate –.30 to –.50 .30 to .50
Strong –.50 to −1.00 .50 to 1.00
According to these guidelines, a relationship between two variables in a set of data is considered “weak” if the relationship has a calculated value of the Pearson r between .10 and .30 (ignoring the ± sign). A relationship is considered “moderate” if r is between .30 and .50 and “strong” if r exceeds ± .50.
It is critical to keep in mind that the above guidelines are designed to aid in the interpretation of a calculated value of r; they do not take into account critical factors such as the size of the sample or whether the relationship is statistically significant. Later in this chapter, a comparison will be made between Cohen's guidelines and the statistical significance of a correlation.
Calculating Correlational Statistics: The Role of Variance and Covariance
The Pearson correlation coefficient is a statistic measuring the linear relationship between two variables. Later in this chapter, formulas used to calculate the Pearson r will be presented and illustrated using the data from the first impression study. This section will introduce the conceptual basis of these formulas to explain what's involved in measuring the relationship between variables.
Correlational statistics are based on variance—a concept to which you have already been introduced in the earlier chapters of this book. In relating two variables with each other, we begin with the understanding that both variables contain a certain amount of variance,
840
meaning there are differences among the scores for the two variables. In the first impression study, for example, there were differences among the ratings of the faces shown at 39 ms and differences among the ratings of the faces shown at 1700 ms. However, in addition to variance, the relationship between two variables is based on the extent to which differences in the scores for the two variables are systematically connected to each other. In other words, measuring the relationship between variables involves both the extent to which two variables vary on their own and the extent to which they vary together.
Covariance refers to the extent to which two variables vary together such that they have shared variance (i.e., variance in common with each other). Height and weight provide a useful example of variables that have both variance and covariance. There is variance in each of these two variables, such that people differ in how tall they are and in how much they weigh. In addition, height and weight have covariance. Generally speaking, the taller (or shorter) one is, the more (or less) one weighs. An individual who is taller (or shorter) than the average person generally weighs more (or less) than the average person. For this reason, we can say that height and weight vary together, or covary.
The following conceptual formula for the Pearson correlation coefficient takes into account both the amount of covariance between Variables X and Y (the extent to which X and Y vary together) and the variance of Variables X and Y (the extent to which X and Y vary on their own): r = covariance ( X,Y ) variance ( X ) variance ( Y ) = extent to which X and Y vary together extent to which X and Y vary on their own
From the above formulas, we see that the Pearson correlation coefficient r is a ratio of the covariance between two variables to the variance of the two variables. When the amount of covariance is high relative to the variance of the two variables, a strong relationship will exist between the two variables, and the value for r will be high (closer to ±1.00). On the other hand, if the amount of covariance is relatively low, there will be a weak relationship between the two variables, reflected in a smaller value of r. If there is absolutely no covariance between two variables, the Pearson r will be equal to zero (.00).
841
Learning Check 1: Reviewing what you've Learned So Far
1. Review questions a. For what reasons would you create a scatterplot for a set of data? b. What three aspects can be used to describe the relationship between variables? c. What is the difference between a linear relationship and a nonlinear relationship? d. What is the difference between a positive relationship and a negative relationship? Is a
positive relationship better than a negative relationship? e. What is the difference between a perfect relationship, a moderate relationship, and a zero
relationship? f. What does the Pearson correlation coefficient (r) measure? g. What are the main features of the Pearson r? h. What is covariance? What roles do covariance and variance play in calculating the Pearson r?
2. For each of the following situations, create a scatterplot of the data and describe the nature, direction, and strength of the relationship.
a. A teacher hypothesizes that the more days of school a student misses, the worse the student will do on the final exam (possible scores on the exam range from 0 to 20).
Student # Days Missed Final Exam Score
1 3 16
2 2 18
3 5 13
4 8 7
5 4 12
6 7 11
7 6 14
b. A researcher predicts that the more often a student raises his or her hand in class, the more favorably the teacher will view the student. She counts the number of times students raise their hands during a week and then has the teacher rate each student on a 1 (low) to 10 (high) scale.
Student # Times Raise Hands Teacher Rating
1 9 4
2 4 2
3 8 9
4 5 8
5 6 1
6 7 8
7 4 7
8 3 5
9 6 6
10 7 3
842
843
13.3 Inferential Statistics: Pearson Correlation Coefficient
Returning to the first impression study introduced at the beginning of the chapter, the scatterplot in Figure 13.1 indicates there is a relationship between threatening ratings for faces shown for 39 ms and for 1700 ms, which suggests that people can form impressions of another person quickly and that these impressions are similar to those impressions made after a greater amount of exposure to the same person. To determine whether the relationship is statistically significant, an inferential statistic must be calculated. This requires following the same steps used in earlier chapters:
state the null and alternative hypotheses (H0 and H1), make a decision about the null hypothesis, draw a conclusion from the analysis, and relate the result of the analysis to the research hypothesis.
The difference between this chapter and previous ones is that the inferential statistic calculated is not a version of ANOVA but rather a correlation coefficient, which measures the relationship between variables rather than differences between groups.
844
State the Null and Alternative Hypotheses (H0 and H1)
In previous chapters, we've examined research studies that involved testing differences between the means of groups. In each study, the null hypothesis was that the means of the groups in the larger population were equal to each other, implying that there was no (zero) difference between the groups. Alternatively, the alternative hypothesis stated that the population means were not equal to each other, implying that differences between groups were not equal to zero. A similar logic will be used as the basis for the null and alternative hypotheses for the correlation between two variables.
In testing a correlation between two variables, the null hypothesis is that the correlation in the population (represented by the Greek letter ρ [rho]) is equal to zero: H 0 : ρ = 0
This hypothesis implies that there is no, or zero, relationship between the two variables. What is a mutually exclusive alternative to the null hypothesis that the population correlation ρ is equal to zero? Similar to what we observed in previous chapters, one alternative hypothesis is that the correlation in the population (r) is not equal to zero: H 0 : ρ ≠ 0
Because the above alternative hypothesis ρ ≠ 0 is non-directional (two-tailed), the null hypothesis will be rejected if the samples value for the correlation is of sufficient size in either a positive or negative direction. As in the case of the t-test, it is possible to state a directional (one-tailed) alternative hypothesis in which the null hypothesis will only be rejected if the calculated correlation is in the hypothesized direction. For example, if we had sufficient reason to believe that the relationship between the 39 ms and 1700 ms scores could only be positive, we could state the alternative hypothesis as “H1: ρ > 0.” As in the case of the t-test, however, it is more customary to use a two-tailed test to allow for the possibility of a result in the unexpected direction.
845
Make a Decision about the Null Hypothesis
Once the null and alternative hypotheses have been stated, the next step is to make a decision whether to reject the null hypothesis that the population correlation is equal to zero. This decision is made using the following steps:
calculate the degrees of freedom (df); set alpha (α), identify the critical values, and state a decision rule; calculate a statistic: Pearson correlation coefficient (r); make a decision whether to reject the null hypothesis; determine the level of significance; and calculate a measure of effect size (r2).
Each of these steps is discussed below, using the first impression study to illustrate how each step may be completed.
Calculate the Degrees of Freedom (df)
The first step in making the decision about the null hypothesis is to calculate the degrees of freedom (df) associated with the correlation coefficient. In previous examples, the degrees of freedom was equal to the number of scores for a variable minus 1. For example, in Chapter 7 (the test of one mean), the degrees of freedom was equal to the number of scores for the variable minus 1, or N – 1. However, in the first impression study, data have been collected on two variables rather than one. Therefore, the formula for degrees of freedom for the correlation coefficient is as follows:
(13-1) df = N − 2
where N equals the number of pairs of scores involved in the computation of the correlation.
In the first impression study, there were 24 pairs of ratings of faces shown for 39 ms and 1700 ms. Therefore, the degrees of freedom is equal to d f = N − 2 = 24 − 2 = 22
Set Alpha (α), Identify the Critical Values, and State a Decision Rule
The next step is to set alpha (α), the probability of the correlation needed to reject the null hypothesis, and to identify the critical values of the correlation. As we have seen in earlier
846
examples, alpha is traditionally set at .05, meaning that the null hypothesis is rejected when the probability of obtaining the value of the correlation when the null hypothesis is true is less than .05. Because the alternative hypothesis in the first impression study is non- directional(H1: ρ ≠ 0), alpha may be stated as “α = .05 (two-tailed).”
Table 6 in the back of this book provides a table of critical values for r. This table contains a series of rows corresponding to different degrees of freedom (df). For each df the critical values for r have been provided for both a directional (one-tailed) and non-directional(two- tailed) alternative hypothesis. Like the t-test, the critical values of r do not include the sign + or – because the sign merely represents the direction of the relationship.
For the first impression study we identify the critical values by moving down the df column of Table 6 until we reach the row for df = 22. Moving to the right until we reach the α = .05 column for a two-tailed test, we find a critical value of .404. Therefore, the critical values may be stated as follows: For α = .05 two − tailed and df = 22 , critical values = ± . 404
Figure 13.5 illustrates the critical values and the regions of rejection and non-rejection for the first impression study. Note that the theoretical distribution for the Pearson r resembles that of the t-test in that both are symmetrical with a hypothetical mean of zero (0).
Once the critical values of r have been identified, we may state a decision rule specifying the values of the correlation resulting in the rejection of the null hypothesis. For the first impression study, the following decision rule is stated: If r < − . 404 or > .404, reject H 0 ; otherwise, do not reject H 0
This decision rule implies that the null hypothesis will be rejected if the value of r calculated from the sample is either less than –.404 or greater than +.404.
Figure 13.5 Critical Values and Regions of Rejection and Non-Rejection for the First Impression Study
Calculate a Statistic: Pearson Correlation Coefficient (r)
The next step in making the decision about the null hypothesis is to calculate a value of an inferential statistic. To introduce formulas that may be used to calculate the Pearson r, let's
847
examine the conceptual formula presented earlier: r = covariance X, Y variance X variance Y
Looking at the above formula, we find that three quantities must be represented to calculate the Pearson r: the covariance between Variables X and Y, the variance of Variable X, and the variance of Variable Y. Each of these is discussed below.
Represent the Covariance between X and Y: SPXY
Covariance refers to the extent to which the variance in the two variables is related to each other. Recall from Chapter 4 that variance is the average squared deviation of a score from the mean. Therefore, the variance of the X variable is based on the deviation of each X score from the mean of the X scores X − X ¯ , and the variance of the Y variable is based on the deviation of each Y score from its mean Y − Y ¯ .
Because covariance measures the degree to which two variables vary together, one way to think about covariance is to relate these two deviations to each other. Returning to our earlier discussion of height and weight, the covariance between height and weight implies that the difference between X – X ¯ for a person is related to the difference between Y − Y ¯ for that person. In other words, someone who is taller (or shorter) than the average person also weighs more (or less) than the average person.
Measuring the amount of covariance in a set of data consists of connecting the two deviations X − X ¯ and Y − Y ¯ for each participant. One way to do this is to multiply the two deviations with each other X − X ¯ Y − Y ¯ . As the number created by multiplying two numbers together is called a “product,” the covariance between two variables may be represented by the sum of these products, represented by the symbol SPXY (sum of products between variables X and Y). The definitional formula for SPXY is provided in Formula 13– 2:
(13-2) SP XY = ∑ X − X ¯ Y − Y ¯
where X is a score on the X ¯ variable, X is the mean of the X variable, Y is a score on the Y variable, and Y ¯ is the mean of the Y variable.
The value for SPXY for the first impression study is calculated in Table 13.2. The first step in calculating SPXY is to calculate the deviations of each participants scores on the X variable X − X ¯ and the Y variable Y − Y ¯ . For example, for the first face, X − X ¯ = 2.00 − 2.91 = − . 91 , and Y − Y ¯ = 2.40 − 3.14 = − . 74 . Once the two deviations have been determined, the next step is to calculate the product of the two deviations X − X ¯ Y − Y ¯ . For the first face, the product is equal to (–.91) (–.74) = .68. The products for the 24 faces in the first impression study are provided in the last column of Table 13.2.
848
Once the products have been calculated, the value for SPXY is the sum of the products. For the first impression study, SPXY is calculated as follows: SP XY = ∑ X − X ¯ Y − Y ¯
Table 13.2 Calculating the Sum of Products (SPXY) in the First Impression Study
Table 13.2 Calculating the Sum of Products (SPXY) in the First Impression Study
Face 39 ms (X) 1700 ms (Y) X − X ¯ Y − Y ¯ X − X ¯ Y − Y ¯
1 2.00 2.40 –.91 –.74 .68
2 2.53 2.80 –.38 –.34 .13
3 3.17 3.38 .26 .24 .06
4 2.43 3.43 –.48 .29 –.14
5 4.02 4.00 1.11 .86 .95
6 3.01 3.18 .10 .04 .00
7 2.80 2.77 –.11 –.37 .04
8 3.44 3.88 .53 .74 .39
9 2.77 2.30 –.14 –.84 .12
10 3.27 3.56 .36 .42 .15
11 2.48 3.25 –.43 .11 –.05
12 3.00 3.34 .09 .20 .02
13 2.55 2.35 –.36 –.79 .29
14 2.27 3.08 –.64 –.06 .04
15 2.40 2.06 –.51 –1.08 .56
16 3.45 3.75 .54 .61 .32
17 2.36 3.30 –.55 .16 –.09
18 4.76 4.34 1.85 1.20 2.21
19 2.83 3.00 –.08 –.14 .01
20 3.17 3.59 .26 .45 .11
21 2.38 2.83 –.53 –.31 .17
22 3.15 3.19 .24 .05 .01
23 3.26 3.06 .35 –.08 –.03
24 2.45 2.60 –.46 –.54 .25
X ¯ = 2.91 Y ¯ = 3.14 SPXY = 6.22
849
= . 68 + . 13 + … + − . 03 + . 25 = 6.22
Represent the Variance of Variable X: SSX
The second component of the Pearson r involves the variance of Variable X. Because variance is based on the squared deviation of scores from the mean, the variance of the X variable may be represented by the sum of the squared deviations between scores on the X variable and its mean ∑ X − X ¯ 2 . Formula 13–3 shows the definitional formula for this sum of squared deviations (represented by the symbol SSX):
(13-3) SS X = ∑ X − X ¯ 2
where X is a score on the X variable and X ¯ is the mean of the X variable.
For the first impression study data, SSX is calculated in the middle columns of Table 13.3. Calculating SSX requires squaring each of the deviations for the X variable calculated in Table 13.2. For the first face, for example, X − X ¯ 2 = − . 91 2 = . 84 . Once the deviations have been squared, they are summed to calculate SSX: S S X = ∑ ( X − X ¯ ) 2 = .83 + .14 + … + .12 + .21 = 8.71
Represent the Variance of Variable Y: SSy
In a manner similar to Variable X, the variance for the Y variable is represented by the sum of the squared deviations for scores on the Y variable ∑ Y − Y ¯ 2 . The definitional formula for the sum of squares for the Y variable (SSY) is presented in Formula 13–4:
(13-4) SS Y = ∑ Y − Y ¯ 2
where Y is a score on the Y variable, and Y − Y ¯ is the mean of the Y variable.
SSY is calculated for the first impression study in the final columns of Table 13.3: S S Y = ∑ ( Y − Y ¯ ) 2 = .55 + .12 + … + .01 + .29 = 7.41
Before moving on, its useful to point out something that may prevent errors in calculating the Pearson r. The first component of the Pearson r, SPXY, is calculated by multiplying together two deviations: X − X ¯ and Y − Y ¯ . Because both deviations can either be a positive or negative number, the product of these deviations, and ultimately the value for SPXY, can either be positive or negative.
Table 13.3 Calculating Deviations to Represent Variances in the First Impression Study
Table 13.3 Calculating Deviations to Represent Variances in the First 850
Table 13.3 Calculating Deviations to Represent Variances in the First Impression Study
Face 39 ms (X) 1700 ms (Y) X − X ¯ X − X ¯ 2 Y − Y ¯ Y − Y ¯ 2
1 2.00 2.40 –.91 .83 –.74 .55
2 2.53 2.80 –.38 .14 –.34 .12
3 3.17 3.38 .26 .07 .24 .06
4 2.43 3.43 –.48 .23 .29 .08
5 4.02 4.00 1.11 1.23 .86 .74
6 3.01 3.18 .10 .01 .04 .00
7 2.80 2.77 –.11 .01 –.37 .14
8 3.44 3.88 .53 .28 .74 .55
9 2.77 2.30 –.14 .02 –.84 .71
10 3.27 3.56 .36 .13 .42 .18
11 2.48 3.25 –.43 .18 .11 .01
12 3.00 3.34 .09 .01 .20 .04
13 2.55 2.35 –.36 .13 –.79 .62
14 2.27 3.08 –.64 .41 –.06 .00
15 2.40 2.06 –.51 .26 –1.08 1.17
16 3.45 3.75 .54 .29 .61 .37
17 2.36 3.30 –.55 .30 .16 .01
18 4.76 4.34 1.85 3.43 1.20 1.44
19 2.83 3.00 –.08 .01 –.14 .02
20 3.17 3.59 .26 .07 .45 .20
21 2.38 2.83 –.53 .28 –.31 .10
22 3.15 3.19 .24 .06 .05 .00
23 3.26 3.06 .35 .12 –.08 .01
24 2.45 2.60 –.46 .21 –.54 .29
X ¯ = 2.91 Y ¯ = 3.14 ∑ X − X ¯ 2 = 8.71 ∑ Y − Y ¯ 2 = 7.41
However, because SSX and SSY are based on squared deviations, which can only be positive numbers, the values for SSX and SSY must always be positive. Therefore, it is the sign (+ or –) of the value of SPXY that determines whether the value for the Pearson r is positive or negative.
851
Calculate the Pearson Correlation Coefficient (r)
Now that formulas for the covariance and variance of the two variables have been provided, they may be inserted into the conceptual formula for the Pearson r provided earlier: r = covariance X, Y variance X variance Y = SP XY SS X SS Y
The symbols SPXY, SS,X, and SS,Y may be replaced by the mathematical operations in Formulas 13–2 to 13–4 to create the definitional formula for the Pearson correlation coefficient:
(13-5) r = SP XY SS X SS Y ∑ X − X ¯ Y − Y ¯ ∑ X − X ¯ 2 ∑ Y − Y ¯ 2
where X is a score on the X variable, X ¯ is the mean of the X variable, Y is a score on the Y variable, and Y ¯ is the mean of the Y variable.
Using the values of SPXY, SSX, and SSY calculated earlier, we can determine the Pearson r for the first impression study in the following manner: r = S P XY ( S S X ) ( S S Y ) = 6.22 ( 8.71 ) ( 7.41 ) = 6.22 64.54 = 6.22 8.03 = .77
One part of the above calculations that may result in computational errors is the need to take the square root of the denominator SS X SS Y . The square root is calculated because, unlike SPXY in the numerator, the denominator is based on squared deviations rather than deviations. The purpose of the square root is to base both the numerator and denominator on deviations.
One way to check for errors in calculating r is to compare the calculated value of r with the scatterplot of the data. A value of r = .77 implies a strong positive linear relationship between two variables, which appears to match the scatterplot in Figure 13.1. If the square root of the denominator had not been calculated, the value of r would have been equal to 6.22 ÷ 64.54, or .10; a value of r = .10 implies a weak relationship, which is not suggested by the scatterplot.
852
Learning Check 2: Reviewing what you've Learned So Far
1. Review questions
a. In testing the Pearson r, what is implied by the null hypothesis (H0)? b. What three things are represented in calculating the Pearson r? c. Which of these determines whether the value for the Pearson r is a positive or negative
number: SPXY, SSX, or SSY? Why? d. Why is there a square root symbol in the denominator of the formula for the Pearson r?
2. For each of the following, calculate the degrees of freedom (df) and identify the critical values of r (assume a = .05 [two-tailed]).
a. Ni = 6 b. Ni = 9 c. Ni = 18
3. For each of the following, calculate the Pearson correlation (r). a. SP XY = 7.00, SS X = 11.00 , SS Y = 14.00 b. SP XY = − 1.25, SS X = 2 .00 , SS Y = 4.00 c. SP XY = 12.57, SS X = 135 . 31 , SS Y = 109.68
Make a Decision Whether to Reject the Null Hypothesis
Once the value for the Pearson r has been calculated, the decision is made regarding the null hypothesis by comparing the samples value of r with the identified critical value. For the first impression study, this decision is stated as follows: r = . 77>.404 ∴ reject H 0 p < . 05
Because the calculated value of .77 for the Pearson r is greater than the critical value .404, the decision is made to reject the null hypothesis that there is no relationship between the two variables. This decision implies that there is a statistically significant positive linear relationship between the average threatening rating for faces shown for 39 ms and for 1700 ms.
Determine the Level of Significance
Once the decision has been made to reject the null hypothesis, the next step is to determine whether the probability of the correlation is not only less than .05 but also less than .01. Turning to Table 6, for df = 22 and a two-tailed probability of .01, we find that the critical value of r is equal to .515. For the first impression example, we may state r = . 77 >. 515 ∴ p < . 01
Because the calculated value of r exceeds the .01 critical value of .515, we may conclude that the probability of this correlation is not only less than .05; it is less than .01 (p < .01).
853
The level of significance for the correlation in the first impression study is illustrated in Figure 13.6.
Calculate a Measure of Effect Size (r2)
Along with the inferential statistic, we learned in Chapter 10 that researchers often report a measure of effect size as an index or estimate of the size or magnitude of the effect. More specifically, the effect size is measured as the percentage of variance in one variable that may be accounted for by the other variable, ranging from .00 (0%) to 1.00 (100%).
For the Pearson correlation coefficient, a measure of effect size is r2, which is the square of the correlation coefficient r:
(13-6) Estimate of effect = r 2
If r2 looks familiar, that's because it is simply another version of the R2 statistic presented in earlier chapters.
In the first impression example, the measure of effect size is calculated as follows: r 2 = ( .77 ) 2 = .59
Therefore, we can conclude that 59% of the variability in the threatening ratings of faces shown for 1700 ms is explained by differences in threatening ratings of faces shown for 39 ms.
To help interpret or characterize this measure of effect size, Cohen (1988, pp. 79–81) has suggested the following classification for correlational research:
Figure 13.6 Determining the Level of Significance for the First Impression Study
Small effect: r 2 = . 01 r = ± . 10 Medium effect: r 2 = . 09 r = ± 30 Large effect: r 2 = . 25 r = ± . 50
In the first impression example, the r2 value of .59 implies the relationship between the ratings of faces shown for 39 ms and 1700 ms represents a large effect.
854
Draw a Conclusion from the Analysis
For the first impression study, we have now calculated the correlation coefficient and made the decision to reject the null hypothesis. The results of this analysis may be reported in the following way:
The threatening ratings for 24 faces shown for 39 ms and 1700 ms had a statistically significant strong positive linear relationship, r(22) = .77, p < .01, r2
= .59.
This sentence includes the following relevant information about the analysis:
the two variables being related to each other (“The threatening ratings for … faces shown for 39 ms and for 1700 ms”), the sample size (“24”), the nature and direction of the relationship (“a statistically significant strong positive linear relationship”), and information about the inferential statistic (“r(22) = .77, p < .01, r2 = .59”), which indicates the inferential statistic calculated (r), the degrees of freedom (22), the value of the statistic (.77), the level of significance (p < .01), and the measure of effect size (r2 = .59).
855
Relate the Result of the Analysis to the Research Hypothesis
The first impression study revealed a significant positive relationship between ratings of personality for faces shown for a very short time (39 ms) and a relatively longer amount of time (1700 ms). The authors of this study described the implications of their analysis for understanding how first impressions are formed in the following way:
In summary … the results indicate that people can form consistent threat impressions of faces with neutral expressions presented for as briefly as 39ms. (Bar et al., 2006, p. 271)
Table 13.4 lists the steps involved in calculating the Pearson correlation coefficient, using the first impression study to illustrative each step. The next section discusses an alternative way to calculate the Pearson r using formulas that require fewer calculations than Formulas 13–2 to 13–4.
856
A Computational Formula for the Pearson r
The formulas used to calculate the Pearson r presented in the previous section were referred to as definitional formulas. They were called definitional formulas because they literally reflect the definitions of variance and covariance as based on deviations from the mean. Definitional formulas have been emphasized throughout the book because they clearly represent the concepts being measured. However, because the calculation and squaring of deviations such as X − X ¯ and Y − Y ¯ can be tedious, this section will provide what is known as a computational formula for the Pearson correlation coefficient. A computational formula is a formula that is an algebraic manipulation of a definitional formula designed to minimize the complexity of mathematical calculations.
Table 13.4 Summary, Calculating the Pearson Correlation Coefficient Using Definitional Formulas (First Impression Study Example)
857
858
Learning Check 3: Reviewing what you've Learned So Far
1. Review questions
a. What does r2 for the Pearson correlation represent? 2. Students applying to graduate school may be asked to take an admissions test. To study how well
these tests predict performance in grad school, one study hypothesized a positive relationship between scores on the Graduate Management Admissions Test (GMAT) and performance in a graduate MBA (Masters of Business Administration) program such that higher GMAT scores are associated with higher levels of academic performance (Ahmadi, Raiszadeh, & Helms, 1997). To test their hypothesis, the researchers obtained the GMAT score (range: 200 to 800) and grade point average (GPA) (range: 0.00 to 4.00) from transcripts of students in an MBA program. The findings from the study are reproduced using the following sample of 26 students:
Student GMAT GPA
1 710 3.65
2 570 3.40
3 460 3.25
4 450 3.70
5 320 3.15
6 530 3.55
7 330 2.80
8 590 3.60
9 460 2.50
10 430 2.70
11 680 3.00
12 530 3.05
13 360 2.50
14 590 3.80
15 470 2.95
16 430 3.35
17 660 3.50
18 490 2.60
19 370 3.55
20 600 3.15
21 470 3.50
22 420 3.60
23 630 3.85
24 480 2.75
25 400 3.05
26 620 3.55
859
a. State the null and alternative hypotheses (H0 and H1). b. Make a decision about the null hypothesis
1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values, and state a decision rule. 3. Calculate the Pearson correlation (r). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
6. Calculate a measure of effect size (r2). c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
Calculating the Pearson r using computational formulas involves using the same three steps as we used in our earlier calculations: representing the covariance between Variables X and Y, representing the variance of Variable X, and representing the variance of Variable Y. Each of these steps is described below, using the first impression study as an example.
Represent the Covariance between X and Y (SPXY)
The first step in calculating the Pearson r is to represent the covariance between the X and Y variables by calculating the sum of products (SPXY). The computational formula for SPXY is provided below:
(13-7) SP XY = ∑ XY − ∑ X ∑ Y N
where X is a score on Variable X, Y is a score on Variable Y, and N is the number of pairs of scores.
Three quantities are needed to compute SPXY using the computational formula: ΣXY (the sum of the products of scores on Variables X and Y), ΣX (the sum of the scores for Variable X), and ΣY (the sum of the scores for Variable Y). Note that Formula 13–7 differs from the definitional formula for SPXY (Formula 13–2) in that it does not require calculating the deviations X − X ¯ and Y − Y ¯ .
The first part of SPXY, ΣXY, requires multiplying each of the scores on the X variable by its corresponding score on the Y variable. The products for the 24 faces in the first impression study are listed in the last column of Table 13.5. For example, multiplying the first faces threatening ratings of 2.00 (39 ms) and 2.40 (1700 ms) results in a product of 4.80. Connecting the two variables with each other by calculating products is similar to the definitional formula; in this case, however, the product is created by multiplying scores rather than deviations. At the bottom of this column, the products have been summed such that ΣXY is equal to 226.10.
860
In addition to calculating ΣXY, Table 13.5 also sums the scores for the X (ΣX = 69.95) and Y (ΣY = 75.44) variables as they are also part of Formula 13–7. Inserting the values of ΣXY, ΣX, and ΣY into the formula, we calculate SPXY for the first impression study data in the following manner: S P XY = ∑ XY − ( ∑ X ) ( ∑ Y ) N = 226.10 − ( 69.95 ) ( 75.44 ) 24 = 226.10 − 5277.03 24 = 226.10 − 219.88 = 6.22
Comparing the above value of 6.22 with the calculations in the previous section, we see that the same value is calculated for SPXY using either the definitional formula or the computational formula.
Represent the Variance of Variable X (SSX)
The sums of squares representing the variance of the X variable (SSX) may be calculated using the following computational formula:
(13-8) SS X = ∑ X 2 − ∑ X 2 N
where X is a score on Variable X, and N is the number of pairs of scores.
Table 13.5 Calculating the Sum of Products (SPXY) in the First Impression Study Using the Computational Formula
Table 13.5 Calculating the Sum of Products (SPXY) in the First Impression Study Using the
Computational Formula
Face 39 ms (X) 1700 ms (Y) (XY)
1 2.00 2.40 4.80
2 2.53 2.80 7.08
3 3.17 3.38 10.72
4 2.43 3.43 8.33
5 4.02 4.00 16.08
6 3.01 3.18 9.57
7 2.80 2.77 7.76
8 3.44 3.88 13.35
9 2.77 2.30 6.37
10 3.27 3.56 11.64
11 2.48 3.25 8.06
861
12 3.00 3.34 10.02
13 2.55 2.35 5.99
14 2.27 3.08 6.99
15 2.40 2.06 4.94
16 3.45 3.75 12.94
17 2.36 3.30 7.79
18 4.76 4.34 20.66
19 2.83 3.00 8.49
20 3.17 3.59 11.38
21 2.38 2.83 6.74
22 3.15 3.19 10.05
23 3.26 3.06 9.98
24 2.45 2.60 6.37
ΣX = 69.95 ΣY = 75.44 ΣXY = 226.10
Two quantities are needed to compute SSX: ΣX2 (the sum of squared scores for the X variable) and (ΣX)2 (the squared sum of scores for the X variable). The first quantity, ΣX2, is created by first squaring each of the scores for the X variable and then summing the squared scores (first squaring, then summing). The second quantity, (ΣX)2, is calculated by first summing the scores for the X variable and then squaring this sum (first summing, then squaring).
ΣX2 and (ΣX)2 for the 39 ms variable in the first impression study are calculated in Table 13.6. The fourth column of this table provides the squared value of each of the scores for the 39 ms variable (X2). The bottom of this column shows that the sum of the squared scores for the 39 ms variable, ΣX2, is equal to 212.59. Once the sum of the scores (ΣX = 69.95) is calculated, SSX for the 39 ms variable may be calculated as follows: S S X = ∑ X 2 − ( ∑ X ) 2 N = 212.59 − ( 69.95 ) 2 24 = 212.59 − 4893.00 24 = 212.59 − 203.88 = 8.71
Represent the Variance of Variable Y (SSY)
The formula for the sums of squares for the Y variable (SSY is a simple modification of the formula for SSX.
(13-9) SS Y = ∑ Y 2 − ∑ Y 2 N
862
where Y is a score on Variable Y and N is the number of pairs of scores. The last column of Table 13.6 lists the squared values for the 1700 ms variable; looking at the bottom of this column, we see that ΣY2 = 244.54. Using this number as well as the sum of the scores (ΣY = 75.44), SSY for the 1700 ms variable is calculated below: S S Y = ∑ Y 2 − ( ∑ Y ) 2 N = 244.54 − ( 75.44 ) 2 24 = 244.54 − 5691.19 24 = 244.54 − 237.13 = 7.41
Table 13.6 Calculating Squared Scores to Represent Variances in the First Impression Study
Table 13.6 Calculating Squared Scores to Represent Variances in the First Impression Study
Face 39 ms (X) 1700 ms (Y) X2 Y2
1 2.00 2.40 4.00 5.76
2 2.53 2.80 6.40 7.84
3 3.17 3.38 10.05 11.42
4 2.43 3.43 5.90 11.76
5 4.02 4.00 16.16 16.00
6 3.01 3.18 9.06 10.11
7 2.80 2.77 7.84 7.67
8 3.44 3.88 11.83 15.05
9 2.77 2.30 7.67 5.29
10 3.27 3.56 10.69 12.67
11 2.48 3.25 6.15 10.56
12 3.00 3.34 9.00 11.15
13 2.55 2.35 6.50 5.52
14 2.27 3.08 5.15 9.49
15 2.40 2.06 5.76 4.24
16 3.45 3.75 11.90 14.06
17 2.36 3.30 5.57 10.89
18 4.76 4.34 22.66 18.83
19 2.83 3.00 8.01 9.00
20 3.17 3.59 10.05 12.89
21 2.38 2.83 5.66 8.01
22 3.15 3.19 9.92 10.18
23 3.26 3.06 10.63 9.36
863
24 2.45 2.60 6.00 6.76
ΣX = 69.95 ΣY = 75.44 ΣX2 = 212.59 ΣY2 = 244.54
Calculate the Pearson Correlation Coefficient (r)
Once SPXY, SSX, and SSY have been calculated, the value for the Pearson r may be calculated. Formula 13–10 presents the computational formula for the Pearson r:
(13-10) r = SP XY SS X SS Y = ∑ XY − ∑ X ∑ Y N ∑ X 2 − ∑ X 2 N ∑ Y 2 − ∑ Y N
Using Formula 13–10, we calculate the correlation between the threatening ratings for the 39 ms and 1700 ms variables in the first impression study in the following way: r = S P XY ( S S X ) ( S S Y ) = 6.22 ( 8.71 ) ( 7.41 ) = 6.22 64.54 = 6.22 8.03 = .77
Note that the computational formula in Formula 13–10 results in the same value for the Pearson r as the definitional formula (Formula 13–5).
The steps needed to calculate the Pearson r using the computational formula are summarized in Table 13.7. We discussed the definitional formula to illustrate the conceptual basis of the Pearson correlation; the computational formula has been provided to assist you in your research efforts as you apply the statistical procedures covered in this book in your own work and study.
The research examples in this chapter have involved the calculation and interpretation of a correlation coefficient. The main purpose of the Pearson r is to measure and test the linear relationship between two variables. However, in addition to describing and understanding relationships between variables, another goal of science is prediction. Researchers frequently use and apply knowledge about identified relationships to predict one variable from another. Using the findings of the first impression study, for example, do you think it would be possible to predict ones later impressions about a person from ones first impressions? The next section will answer this question by demonstrating how researchers use the relationship between two variables to develop a means of predicting scores on one variable from another.
864
13.4 Predicting One Variable from another: Linear Regression
When a relationship is found to exist between variables, this provides an opportunity to predict one variable from another. For example, if college admissions counselors determine there's a relationship between students' high school grade point average (GPA) and their college GPA, this information can be used to predict an applicant's college GPA based on knowledge of his or her high school GPA. If political scientists find a relationship between ones income and ones attitudes about a presidential candidate, they may wish to use voters' incomes to predict the likelihood of voting for that candidate.
Table 13.7 Summary, Calculating the Pearson Correlation Coefficient Using Computational Formulas (First Impression Study Example)
For the purposes of our discussion, moving from “relationship” to “prediction” involves moving from “correlation” to “regression.” Regression may be defined as the use of a relationship between two or more correlated variables to predict values of one variable from values of other variables. Whereas the main purpose of correlation is to measure and test the relationship between variables, the goal of regression is to predict one variable from other variables.
865
Learning Check 4: Reviewing what you've Learned So Far
1. Review questions a. What is the purpose of a computational formula?
2. Calculate the Pearson correlation coefficient for each set of data below using the computational formulas:
a.
Participant X Y
1 8 7
2 3 1
3 7 6
4 6 2
5 8 3
6 4 5
b.
Participant X Y
1 17 11
2 23 7
3 16 5
4 22 9
5 29 4
6 14 10
7 20 3
8 27 7
c.
Participant X Y
1 4.09 51
2 3.13 28
3 2.58 43
4 4.26 86
5 1.19 75
6 4.82 88
7 1.63 93
8 3.67 31
9 4.70 74
10 1.52 82
866
11 4.47 77
12 2.07 61
Linear regression refers to a statistical procedure in which a straight line is fitted to a set of data to best represent the relationship between two variables. When we earlier introduced the concept of correlation (Section 13.2), we noted that the nature, strength, and direction of the relationship between two variables may be represented using a straight line. This notion of a straight line is the basis of how one variable can be used to predict another.
To introduce linear regression, let's examine the formula for a straight line. Many students learn the following formula in high school geometry classes: Y = mX + b
In this equation, Y is a value for a Y variable, X is a value for an X variable, m represents the slope of the line, and b is the Y-intercept. The slope of a line is the angle or tilt of the line relative to the X-axis (the horizontal axis), which is sometimes referred to as the “rise over run”—the rate of change in the Y variable divided by the rate of change in the X variable. A value of zero (0) for the slope represents a horizontal line that is parallel to the X-axis. Finally, the Y-intercept is the point at which the line crosses the Y-axis (the vertical axis) when X is equal to 0.
867
The Linear Regression Equation
Applying the concept of a straight line to linear regression involves creating a line that best represents, or fits, the relationship between two variables in a set of data. This “best-fit” line is known as the linear regression equation. The linear regression equation is a mathematical equation that predicts a score on one variable from a score on another variable based on the relationship between two variables.
The formula for the linear regression equation is a modification of the formula for a straight line:
(13-11) Y ′ = a + bX
where Y′ represents a predicted score for the Y variable, a is the Y-intercept, b is the slope, and X is a score for the X variable. One of the differences between Formula 13–11 and the formula for a line involves the use of Y′ rather than Y. Y′ represents a predicted value for the Y variable rather than an actual value. In other words, the purpose of a linear regression equation is to predict scores on one variable based on its relationship with another variable in a sample of data.
868
Calculating the Linear Regression Equation
There are two main steps in developing a linear regression equation for a set of data: calculating the slope (b) of the equation and calculating the Y-intercept (a) of the equation. This section provides formulas for b and a and calculates the linear regression equation for the first impression study.
Calculate the Slope of the Equation (b)
The first step in calculating the linear regression equation is to calculate the slope using the following formula:
(13-12) b = r S Y S X
where r is the correlation between Variables X and Y, SY is the standard deviation of Variable Y, and SX is the standard deviation of Variable X.
Looking at Formula 13–12, we see that the stronger the relationship between the two variables (represented by the Pearson r), the greater is the angle or slope of the equation (the value of b). Conversely, the weaker the relationship, the more the slope of the line approaches zero, which is represented by a horizontal line parallel to the X-axis. Because the Pearson r can either be a positive or negative number, b can also be positive or negative. A positive slope indicates that the equation angles upward from left to right, whereas a negative slope indicates that the equation angles downward from left to right. Also, the standard deviation of Variable Y (SY) is divided by the standard deviation of Variable X (SX) in order to represent the “rise over run” (rate of change in Variable Y divided by rate of change in Variable X) aspect of the slope described earlier.
To begin the calculation of the slope for the first impression study, recall that the correlation between threatening ratings for 39 ms and 1700 ms was r = .77. Obtaining the standard deviations for the two variables from the descriptive statistics in Table 13.1(b), the slope may be calculated as follows: b = r S Y S X = .77 .57 .62 = .77 ( .92 ) = .71
Calculate the Y-Intercept of the Equation (a)
In linear regression, the Y-intercept a is the predicted value for the Y variable when X is equal to zero. Formula 13–13 provides a computational formula for the Y-intercept:
(13-13)
869
a = Y ¯ − b X ¯
where Y ¯ is the mean of the Y variable, b is the slope of the equation, and X ¯ is the mean of the X variable.
Using the value of b = .71 calculated earlier and obtaining the means for the two variables from Table 13.1(b), we calculate the Y-intercept for the first impression in the following manner: a = Y ¯ − b X ¯ = 3.14 − .71 ( 2.92 ) = 3.14 − 2.07 = 1.07
Report the Linear Regression Equation
Once the values for the slope and the Y-intercept have been calculated, the linear regression equation for the variables involved in the analysis may be reported. For the first impression study, the following equation predicts the threatening rating for faces shown for 1700 ms from the threatening rating for faces shown for 39 ms: Y ' = a + b X 1700 ms′ = 1 . 07+ . 71 ( 39 ms )
Interpreting our regression equation, we obtain a predicted threatening rating for a face shown for 1700 ms (1700 ms′) by multiplying a threatening rating for the same face shown for 39 ms by .71 and then adding 1.07.
Note that the goal of a regression equation is not to make predictions within the sample that was used to create the regression equation. Instead, the purpose of the equation is to make predictions for future samples and participants. Returning to our earlier example of high school GPA and college GPA, college admissions counselors might use the regression equation created from the data in one sample to make predictions for future high school students applying to college.
870
Drawing the Linear Regression Equation
It is possible to draw a linear regression equation in a scatterplot of the data used to create the equation. One way to draw a line that represents the equation is to determine two points in the scatterplot and then draw a line through these two points. Determining these two points involves calculating the predicted score on the Y variable for any two possible scores on the X variable. For the first impression study, lets arbitrarily choose threatening ratings of 1.50 and 4.50 for the 39 ms variable. Inserting these two scores into the regression equation, we may determine the following two values of 1700 ms′:
39 ms = 1.50 39 ms =4.50
1700 ms′ = 1.07 + .71 ( 1.50 ) = 1.07 + .1.07 = 2.14
1700 ms′ = 1.07 + .71 ( 4.50 ) = 1.07 + 3.20 = 4.27
Once the two predicted scores on the Y variable have been calculated, the next step is to locate them in the scatterplot. In Figure 13.7, the coordinates of these two points ((1.50, 2.14) and (4.50, 4.27)) have been marked and a line that connects them has been drawn. There are several things to note about this line. First, the line has a positive or upward slope; this is because the relationship between the two variables is strong and positive (r = .77). Second, the line passes through the center of the scatterplot of scores, indicating that the equation “fits” the data as closely as possible. Because the relationship between the two variables is less than perfect, some data points fall above the line and some fall below. The equation minimizes as much as possible the differences between the data points and the line. In other words, the equation is the best fit for the data in this particular sample.
Table 13.8 summarizes the steps in calculating the linear regression equation, using the first impression study as the example. In the next section, we will discuss a second research example to accomplish two goals. The first goal is to further illustrate the process of calculating and interpreting linear regression equations. The second goal is to answer the following question: What happens when we try to predict one variable from another when there is no relationship between the two?
Figure 13.7 Linear Regression Equation for the First Impression Study
871
Table 13.8 Summary, Calculating the Linear Regression Equation (First Impression Study Example)
872
A Second Example: Put down your Pencils
Imagine you're sitting in a classroom taking an exam (this is probably not a very difficult thing for you to do). Working as quickly as you can, you glance up and see that several students have already turned in their exams. The following thoughts may come into your mind: “Am I going too slowly?” “What if I can't finish in time?” “Will I do worse than the students who've already finished?”
Is there a relationship between the amount of time students take to complete a test and their scores on the test? Although teachers and students have their own implicit beliefs about this relationship, William Herman at the State University of New York took the further step of examining the relationship empirically (Herman, 1997). The participants in this example were 32 students enrolled in one of Dr. Herman's psychology courses. The data from these students are included in the exercises at the back of the chapter.
The students in Dr. Herman's course were given a 100-item multiple-choice test. For each student in the course, Dr. Herman recorded information on two variables. The first variable, Time, was the number of minutes taken by each student to finish the examination. The second variable, Exam score, was the number of correct answers for each student, ranging from 0 to 100.
To save space, the calculation and testing of the Pearson correlation between Time and Exam score will not be shown here (a scatterplot of the data for the 32 students is provided in Figure 13.8). Instead, we will simply report what Dr. Herman found when he calculated the Pearson r between the two variables: r(30) = –.02, p > .05. From the results of his analysis, he concluded that the relationship between the amount of time students took to finish the exam and their scores on the exam was not significant. In fact, there was virtually zero relationship between the two variables! Dr. Herman described his findings in the following way: “This finding lends support for the conscious destruction of the myth that the rate of finishing such an examination is related to how well students will perform on these examinations” (Herman, 1997, p. 117).
Calculating the Linear Regression Equation
Even though no relationship was found between Time and Exam score, a linear regression equation can still be created between the two variables. To develop this equation, descriptive statistics of the two variables are needed:
Time (X) Exam Score (Y)
N 32 32
Mean X ¯ = 68.66 Y ¯ = 74.25
873
Standard deviation SX = 23.96 SY = 11.62
The first step in calculating the equation is to calculate the slope (b) using Formula 13–12. Using the correlation between the two variables (r = –.02), along with the standard deviations for the two variables, the slope is calculated as follows: b = r s Y s X = − .02 11.62 23.96 = − .02 ( .49 ) = − .01
In this example, the value for the slope (b = –.01) is a negative number because the relationship between the two variables is negative (r = –.02).
Once b has been calculated, the next step is to calculate the Y-intercept (a) using Formula 13–13. Using the value of b = –.01 and the means for the two variables, the Y-intercept is calculated as follows: a = Y ¯ − b X ¯ = 74.25 − ( − .01 ) ( 68.66 ) = 74.25 = 74.94 − ( − .69 )
Using the calculated values for b and a, we may report the following linear regression equation predicting a student's exam score from the number of minutes taken to complete the exam: Y ′ = a + b X Exam score′=74 . 94+ ( − .01 ) ( Time )
Drawing the Linear Regression Equation
To draw the linear regression equation into a scatterplot of the two variables, we need to identify two points along the line. Arbitrarily picking values of Time of 20 minutes and 120 minutes, two predicted exam scores are calculated as follows:
Time = 20 Time = 120
Exam score′ = 74.94 + ( − .01 ) ( 20 ) = 74.94 − .20 = 74.74
Exam score′ = 74.94 + ( − .01 ) ( 120 ) = 74.94 − 1.20 = 73.74
Figure 13.8 uses these two points to draw the regression line into the scatterplot of the data from the 32 students. Looking at the figure, note that the regression equation is nearly a horizontal line. The line has almost no slope because there is virtually no relationship between the two variables. The implication of a horizontal line is that, regardless of the amount of time it takes a student to complete the test, virtually the same exam score is predicted. For example, there is almost no difference between the predicted exam score of a student who takes 20 minutes and a student who takes 120 minutes (74.74 vs. 73.74).
Figure 13.8 Linear Regression Equation Predicting Exam Score from Time
874
From the regression equation, we see that the Y-intercept (a = 74.94) and the predicted exam scores for 20 and 120 minutes (74.74 and 73.74) are all almost identical to the mean of the Exam scores Y ¯ = 74.25 . When there is no relationship between two variables, the most accurate prediction we can make, regardless of the score on one variable, is the mean of the other variable. To further illustrate this, imagine you were asked to respond to the following question: “I'm thinking of a friend of mine who is 5′11″. What is his IQ?” Because there is no relationship between height and IQ, the best prediction you could make would be the mean IQ in the population (for example, 100). Furthermore, you would predict the mean IQ of 100 for every person regardless of his or her height. To improve accuracy of prediction beyond that provided by the mean, there must be a relationship between the two variables.
When the concept of correlational statistics was introduced earlier in this chapter, we noted that a number of different correlational statistics exist and that the choice of statistic depends on how variables in a study have been measured. The next section describes a correlational statistic that is used when both variables are measured at the ordinal level of measurement.
875
Learning Check 5: Reviewing what you've Learned So Far
1. Review questions a. What is the difference between correlation and regression? b. What is the purpose of a linear regression equation? c. What are the main components of a linear regression equation? d. What determines whether the slope (b) of a linear regression equation is positive or
negative? e. What do you need to do to draw a linear regression equation into a scatterplot? f. What does the linear regression equation look like when there is no relationship between
two variables? g. When there is no relationship between two variables, what is the most accurate prediction
you can make for the Y variable? 2. For each of the following, calculate the linear regression equation and draw it into a scatterplot.
a. Variable X: Range = 1–10, X ¯ i = 6.00 , SX= 2.00;
Variable Y: Range = 2–12, Y ¯ = 9.00 , sY = 3.00,
Pearson correlation: r = .30
b. Variable X: Range = 1–5, X ¯ i = 2.59 , SX = .88;
Variable Y: Range = 1–5, Y ¯ = 3.78 , SY = .93,
Pearson correlation: r = .47
c. Variable X: Range = 1–50, X ¯ i = 32.76 , SX = 4.60;
Variable Y: Range = 0–5, Y ¯ = 1.84 , SY = .42,
Pearson correlation: r = –.63
876
13.5 Correlating Two Sets of Ranks: The Spearman Rank- Order Correlation
In Chapter 1, we learned that variables in a research study may be measured at one of four levels of measurement: nominal, ordinal, interval, or ratio. In the examples described thus far in this chapter, both of the variables have been measured at the interval or ratio level of measurement. One critical feature of interval or ratio variables is that their values are equally spaced along a numeric continuum. For example, in the example in the previous section, the variable Time (in minutes) was measured at the ratio level of measurement, with its values (20, 21, 22, etc.) equally spaced from each other.
This section introduces a situation in which the goal is to correlate two variables measured at the ordinal level of measurement. The values of ordinal variables can be placed in an order relative to the other values, but the values are not assumed to be equally spaced. Some examples of ordinal variables are “size” (small, medium, large, extra large), “height” (short, average, tall), and “difficulty” (very easy, somewhat easy, somewhat difficult, very difficult). In ordinal variables, one value represents more or less of a variable than do other values, but it is not possible to specify the precise size or amount of the difference between values. For example, the difference between the number of ounces in a “medium” soda and a “large” soda may not be the same as the difference between “large” and “extra large.”
877
Introduction to Ranked Variables
A common form of ordinal data is ranks. A rank is defined as a relative position in a graded or evaluated group. The numeric value for a ranked variable reflects a relative location within an order. You may have been assigned a number as part of a ranked data set, for example, if your high school transcripts listed your class rank as “32nd out of 145” or if you finished “first,” “second,” or “third” in a race. Note that ranks are ordinal variables in that although you can say that a runner who finished “fourth” ran faster than the runner who finished “fifth,” you cannot specify the exact difference between the times of the two runners.
There are perhaps two primary situations in which ranked variables may be employed in a study. The first situation is one in which the collected data themselves consist of ranks. For example, a newspaper editor may ask a movie critic to rank eight movies from 1 (most favorite) to 8 (least favorite). A second situation involving ranked data is when data originally collected at the interval or ratio level of measurement are converted into ranks. For example, a teacher may give a 40-item quiz to 15 students such that each student's score ranges from 0 to 40 but then, on the basis of their quiz scores, rank orders the students from 1 (highest) to 15 (lowest).
This section below describes a research study whose goal was to correlate two sets of ranks to assess the degree of similarity between two groups of research participants. As you will note from the following discussion, calculating and interpreting a correlational statistic used with ordinal data has both similarities and differences from the Pearson correlation.
878
An Example from the Research: Do Students and Teachers Think Alike?
Imagine you're a college senior about to graduate and enter the workforce. What would be most important to you in selecting your first job after you graduate: a high salary, challenge and responsibility, or friendly coworkers? A study conducted by J. Stuart Devlin and Robin Peterson at New Mexico State University examined the extent to which college business students and their professors agreed on the desirability of various characteristics of students' first entry-level positions after graduation (Devlin & Peterson, 1994).
The 127 senior business majors and 28 business school professors in the study, which will be referred to as the business student study, were asked to rate the desirability of 13 possible characteristics of an entry-level position (see Table 13.9). Some of the characteristics reflected aspects of jobs that provide motivation for good job performance, such as the extent to which a job offers opportunities for growth and development, challenge, and freedom. The rest of the characteristics pertained to rewards and benefits that organizations give to employees, such as salary, job title, and job security.
The students and professors rated the desirability of the 13 characteristics on a 5-point scale (1 = low, 5 = high). For each characteristic, the average desirability rating was calculated separately for the students and professors. Based on each group's average ratings, the 13 characteristics were ranked from 1 = most important to 13 = least important. Table 13.9 provides the rankings for the business students and professors.
Looking at the two sets of ranks in Table 13.9, we see that the students and professors were in agreement on the relative desirability of some of the characteristics (freedom on the job, opportunity for advancement, type of work). However, the students ranked characteristics of the workplace and the job itself (e.g., opportunity for self-development, challenge and responsibility, and type of work) as more desirable than did faculty, who reported a higher preference for individual employee benefits and other personal concerns (e.g., salary, location of work, and job security).
Looking at the professors' rank ordering in the right column of Table 13.9, two of the characteristics (job title and training) have the same rank of 12.5. These two characteristics were assigned the same rank because they had the same average desirability. In a situation such as this, each of the tied scores is assigned a rank equal to the average of the tied positions. In this case, the average of the 12th and 13th ranked characteristics is 12.5. As another example of tied ranks, if two of the other characteristics had been tied for the 2nd and 3rd ranks, both characteristics would have been assigned a rank of 2.5, with the next highest ranked characteristic being assigned a rank of 4.
879
Inferential Statistics: The Spearman Rank-Order Correlation Coefficient
Table 13.9 Ranking of Importance of Entry-Level Job Characteristics, Students and Professors Table 13.9 Ranking of Importance of Entry-Level Job Characteristics,
Students and Professors
Job Characteristic Student Ranking Professor Ranking
Challenge and responsibility 3 6
Company reputation 12 9
Freedom on the job 10 10
Job security 8 5
Job title 13 12.5
Location of work 11 3
Opportunity for advancement 1 2
Opportunity for self-development 2 8
Salary 7 1
Training 6 12.5
Type of work 5 4
Working conditions 4 7
Working with people 9 11
The degree of similarity between the beliefs of the students and professors regarding the desirability of different job characteristics was measured and tested using the Spearman rank-order correlation (rs), named after the British psychologist Charles Spearman in 1904. The Spearman rank-order correlation is a statistic measuring the relationship between two variables measured at the ordinal level of measurement.
The Spearman rank-order correlation possesses many of the same qualities as the Pearson correlation coefficient. First, the possible values of the Spearman rs range from −1.00 to +1.00, with the sign (+ or –) indicating the direction of the relationship. A positive relationship indicates that the two sets of ranks are very similar to one another; a negative correlation implies the two sets of rankings are very dissimilar to one another. For the business student study, a negative value for rs would mean that characteristics ranked higher by the students were consistently ranked lower by the professors, and vice versa. Second, the numeric value of the Spearman rs reflects the strength of the relationship. The closer the
880
value rs is to +1.00 or −1.00, the greater is the similarity or difference between the two sets of ranks. A value of zero for rs indicates no relationship whatsoever between the two sets of ranks.
The primary distinction between the Spearman rs and the Pearson r is in how they are evaluated. Statistical procedures such as Pearson correlation, t-test, and ANOVA that use variables measured at the interval or ratio level of measurement are based on the assumption that the distribution of scores for the variables is normally distributed (bell shaped) in the larger population. This assumption influences the shape of the theoretical distributions of the Pearson r, t-statistic, and F-ratio, which in turn determines the critical values used to test the statistical significance of these statistics. However, statistics such as the Spearman rs that analyze ordinal variables do not make assumptions regarding the shape of the distribution. More specifically, the Spearman rs does not require variables to be normally distributed. For this reason, the critical values of the Spearman rs differ from those for the Pearson r.
The steps involved in calculating, evaluating, and interpreting the Spearman rank-order correlation are as follows:
state the null and alternative hypotheses (H0 and H1), make a decision about the null hypothesis, draw a conclusion from the analysis, and relate the result of the analysis to the research hypothesis.
Each of these steps is illustrated below using the business student study as the example. In moving through these steps, the similarities and differences between the Spearman rs and the Pearson r will be noted.
State the Null and Alternative Hypotheses (H0 and H1)
The null and alternative hypotheses for rs may be identical to those used to test the Pearson correlation coefficient: H 0 : ρ = 0 H 1 : ρ ≠ 0
where ρ (rho) is the correlation between the two sets of ranks in the population. The value of zero for ρ represents the absence of a relationship between the two sets of ranks. In the business student study, a non-directional alternative hypothesis (H1: ρ ≠ 0) has been employed to allow for the possibility that the ranks of the students and professors may either be similar (i.e., a positive relationship) or dissimilar (i.e., a negative relationship).
Make a Decision about the Null Hypothesis
881
Once the null and alternative hypotheses have been stated, the next step is to make a decision whether to reject the null hypothesis that the population correlation between the two sets of ranks is equal to zero. For the business student study, this involves asking whether a significant relationship exists between the rankings of students and professors. This question is answered by following these steps:
calculate the degrees of freedom (df); set alpha (α), identify the critical values, and state a decision rule; calculate a statistic: Spearman rank-order correlation (rs); make a decision whether to reject the null hypothesis; determine the level of significance; and calculate a measure of effect size.
Calculate the Degrees of Freedom (df)
The degrees of freedom for the Spearman rank-order correlation are the same as for the Pearson correlation: df = N − 2
where N is the number of pairs of scores being ranked. In the business student study, the rankings of 13 characteristics were paired to assess the similarity of students and professors. Therefore, the degrees of freedom is equal to d f = N − 2 = 13 − 2 = 11
Set Alpha, Identify the Critical Values, and State a Decision Rule
Alpha (α), which is the probability of a statistic needed to reject the null hypothesis, is once again set to the traditional .05 level. The critical values for the Spearman rank-order correlation are provided in Table 7. For the business student study, move down the dfcolumn until you reach the value 11. For α = .05 (two-tailed), we find the critical value is .560. This can be expressed in the following manner: For α = .05 two − tailed and df = 11 , critical values = ± . 560
The null hypothesis for rs is rejected when the calculated value of the statistic exceeds an identified critical value. For the business student study, we may state the following decision rule: If r s < − . 560 or >.560, reject H 0 ; otherwise, do not reject H 0
Calculate a Statistic: Spearman Rank-Order Correlation (rs)
The relationship between the two sets of ranks is measured by calculating a value for the Spearman rank-order correlation. The formula for rs is provided in Formula 13–14:
882
(13-14) r s = 1 − 6 ∑ D 2 N N 2 − 1
where D is the difference between the rankings for each variable and N is the number of pairs of variables. Note that this formula contains a constant (the number 6) used to calculate rs regardless of the number of pairs of variables involved in the analysis. Without going into detail, the number 6 in the numerator and N and N2 in the denominator are needed to represent all of the possible differences that can occur between the two rankings for each variable in a set of variables.
You may also wonder why calculating the Spearman rs involves subtracting a quantity from the number 1. The calculated value of the Spearman rs is based on the difference (D) between the two rankings on each variable. If there happened to be perfect agreement between two sets of rankings, what would the value of D be for each variable? The answer is zero (i.e., zero difference). The greater the differences between the rankings, the greater are the values for D and consequently the lower the value for rs. In essence, calculating the Spearman correlation involves starting with a perfect relationship (represented by the value of 1) and then subtracting any differences between the rankings.
The first step in calculating rs is to calculate ΣD2, which is the sum of the squared differences between the rankings for each variable. The difference between the rankings for each variable (D) is calculated by subtracting one ranking from the other. The fourth column in Table 13.10 lists the values for D for each of the 13 job characteristics in the business student study. For example, for the characteristic “Challenge and responsibility,” D is equal to the students' ranking (3) minus the professors' ranking (6): 3 – 6 = −3. Each value of D is then squared; for “Challenge and responsibility,” D2 = (–3)2 = 9. At the bottom of the last column, of Table 13.10, the squared differences are summed. For the business student study, the sum of the squared differences (ΣD2) is equal to 220.50.
Using the sum of the squared difference scores (ΣD2) from Table 13.10, we can calculate the value of rs for the 13 characteristics ranked by the students and professors as follows: r s = 1 − 6 ∑ D 2 N N ( N 2 − 1 ) = 1 − 6 ( 220.50 ) 13 ( 13 2 − 1 ) = 1 − 1323.00 13 ( 168 ) = 1 − 1323.00 2184 = 1 − .61 = .39
Table 13.10 Difference (D) of Rankings of Desirability of Entry-Level Job Characteristics, Students and Professors
Table 13.10 Difference (D) of Rankings of Desirability of Entry-Level Job Characteristics, Students and Professors
Job Characteristic Student Ranking
Professor Ranking
Difference (D)
Difference2
(D2)
Challenge and 3 6 –3 9
883
responsibility
Company reputation 12 9 3 9
Freedom on the job 10 10 0 0
Job security 8 5 3 9
Job title 13 12.5 .5 .25
Location of work 11 3 8 64
Opportunity for advancement
1 2 –1 1
Opportunity for self- development
2 8 –6 36
Salary 7 1 6 36
Training 6 12.5 –6.5 42.25
Type of work 5 4 1 1
Working conditions 4 7 –3 9
Working with people 9 11 –2 4
ΣD2 = 220.50
Therefore, the correlation between the student' and professors' rankings of the desirability of the 13 job characteristics was .39. Using Cohen's guidelines presented earlier in this chapter, we determine that there is a moderate relationship or degree of similarity between the rankings of the two groups.
Make a Decision Whether to Reject the Null Hypothesis
Do business students and their professors agree on the relative desirability of different entry-level job characteristics? Comparing the value of rs = .39 with the critical values, we reach the following conclusion: r s = . 39 is not < − . 560 or >.560 ∴ do not reject H 0 p > . 05
That is, the decision is made to not to reject the null hypothesis, and we subsequently conclude that the relationship between students' and professors' rankings of the 13 characteristics is not statistically significant. This implies that the students and professors did not significantly agree on the relative desirability of different entry-level job characteristics.
Determine the Level of Significance
In this example, the null hypothesis was not rejected. Therefore, it is not necessary to
884
compare the value of the statistic with the .01 critical value of .703.
Calculate Measure of Effect Size r s 2
As with the Pearson r, a measure of effect size may be calculated by squaring the Spearman rank order correlation:
(13-15) Estimate of effect = r s 2
For the business student study we may state that r s 2 = ( .39 ) 2 = .15
One way to interpret the value of r s 2 is to conclude that 15% of the variability in the students' rankings is explained by differences in the professors' rankings. Using Cohen's guidelines described earlier, we can see that an r s 2 value of .15 implies the relationship between the two sets of rankings represents a medium-sized effect.
Draw a Conclusion from the Analysis
What conclusion may be drawn about the nonsignificant relationship between the ranks of the students and professors? Following is one way that the results of the analysis may be reported:
The rankings of the 13 job characteristics for the 127 business students and the 28 business professors were not significantly related to each other, indicating that students and professors did not agree in their perceptions of the relative desirability of the characteristics, rs(11) = .39, p > .05, r s 2 = . 15 .
Relate the Result of the Analysis to the Research Hypothesis
The nonsignificance of the Spearman rank-order correlation suggests that business students and professors do not agree on the relative desirability of various entry-level job characteristics. Assuming that one part of being a professor is to understand students' work- related beliefs and values, what are the possible implications of this finding? The authors of this study reached the following conclusion about the implications of their research:
Apparently, professors were not very accurate in gauging student preferences…. This situation may be a result of infrequent contact on an informal basis between students and faculty…. Professors would be well advised to become more informed on student values so that they can be more effective in counseling
885
students about careers. (Devlin & Peterson, 1994, pp. 156–157)
Table 13.11 lists the steps followed in correlating two sets of ranks using the Spearman rank-order correlation, applying these steps to the business student study.
886
13.6 Correlational Statistics vs. Correlational Research
The main purpose of this chapter was to introduce what are known as correlational statistics. Correlational statistics, such as the Pearson r, are statistics designed to measure relationships between variables. However, in Chapter 1, we introduced another concept that included the word correlation: correlational research. The purpose of the final section of this chapter is to discuss the difference between correlational research, which concerns how data are collected, and correlational statistics, which concerns how data are analyzed.
Table 13.11 Summary, Calculating the Spearman Rank-Order Correlation (Business Student Study Example)
887
Learning Check 6: Reviewing what you've Learned So Far
1. Review questions a. How do variables measured at the ordinal level of measurement differ from variables
measured at the interval or ratio level of measurement? b. What research situations may involve the use of ranked variables? c. How are ranks assigned when there is a tie between two or more ranks? d. What are the main similarities and differences between the Spearman rank order correlation
(rs) and the Pearson correlation (r)? 2. Calculate the Spearman rank order correlation (rs) for each of the following pairs of ranks.
a.
Rank Order 1 Rank Order 2
4 4
1 2
3 1
2 3
b.
Rank Order 1 Rank Order 2
2 3
4 6
6 5
3 1
1 2
5 4
c.
Rank Order 1 Rank Order 2
9 9
7 6.5
2 5
6 8
3.5 2
8 6.5
3.5 3
1 1
5 4
3. One study examined the preferences of girls and boys for different musical instruments (Abeles,
888
2009). As part of the study the number of sixth-grade girls and boys playing eight different instruments was recorded. Based on these numbers, the instruments can be ranked from 1 (most preferred) to 8 (least preferred). Imagine the study hypothesized that boys and girls would display similar preferences for the different instruments.
Instrument Sixth-Grade Girls Sixth-Grade Boys
Cello 4 6
Clarinet 1.5 7
Drums 8 5
Flute 3 8
Saxophone 6 4
Trombone 7 2
Trumpet 5 1
Violin 1.5 3
a. State the null and alternative hypotheses (H0 and H1). b. Make a decision about the null hypothesis.
1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values, and state a decision rule. 3. Calculate the Spearman rank-order correlation (rs). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size r s 2 .
c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
Correlational research refers to methods of conducting research that examine the relationship between variables without the ability to infer cause-effect relationships. In Chapter 1, we provided examples of correlational research methods, including quasi- experiments, survey research, observational research, and archival research. Examples of the types of variables included in correlational research are biographical (e.g., age, gender, education, income), physiological (e.g., height, weight, reaction time), or psychological (e.g., intelligence, personality, attitudes and opinions). As diverse as these variables may be, they have in common the fact that they are typically measured by the researcher rather than controlled or manipulated. For example, a researcher cannot randomly assign someone to be “male” or “female.”
Correlational research methods differ from another type of research method, experimental research, which is used to infer cause-and-effect relationships between variables. The ability to infer causality is accomplished by isolating the effect of an independent variable through experimental control and manipulation, as well as random assignment to conditions. Chapter 11's chess study provided an example of experimental research. You will recall that in that study, participants were randomly assigned to one of three methods of learning chess devised by the researcher (Observation, Prediction, or Explanation). Unlike experimental research, correlational research methods do not involve experimental control
889
or random assignment. As a result, such methods are unable to test causal relationships between variables. For example, although researchers may find a relationship between high school GPA and college GPA, they may not infer from this that having a high GPA “causes” a student to have a high college GPA.
890
Does ANOVA = Causality?
An important distinction must be made between correlational research and correlational statistics. Correlational research refers to a type of research methodology, one that examines the relationship between variables without the ability to make causal inferences. On the other hand, correlational statistics refers to mathematical procedures that measure the relationship between variables. Although both correlational research and correlational statistics are interested in the relationship between variables, these two concepts are very different from each other. Correlational research refers to methods used to collect data, whereas correlational statistics refers to methods used to analyze data.
The distinction between correlational research and correlational statistics can be a source of confusion to students. Correlational research frequently, although not necessarily, involves the use of continuous variables such as Time and Exam score. As a result, data collected using correlational research are often analyzed using a correlational statistic. On the other hand, because experimental research frequently involves assigning participants to groups or conditions, the independent variables are typically categorical in nature. Consequently, researchers analyze differences between these groups using a version of ANOVA, such as a t-test or one-way ANOVA.
Students sometimes make the leap in inference that because ANOVA is typically used to analyze experiments and correlational statistics are employed in correlational research, using ANOVA enables one to make causal inferences while correlational statistics do not. This is simply not true—the ability to draw causal inferences is a function of how data are collected, not how they are analyzed.
The use of ANOVA to analyze data does not necessarily lead to the ability to make causal inferences. For example, in the voting message study discussed in Chapter 12 (two-way ANOVA), one of the independent variables was authoritarianism, defined as the extent to which people “perceive a great deal of threat from their immediate and broader environment” (Lavine et al., 1999, p. 338). Because authoritarianism is a personality characteristic, it could not be manipulated by the researcher. Consequently, the voting message study was an example of correlational research.
Just as ANOVA may be used to analyze correlational research, it is possible to analyze the results of an experiment using correlational statistics. For example, imagine a researcher wants to determine whether the temperature of a swimming pool affects swimmers' speed. Swimmers of equal ability are randomly assigned to swim 100 meters in 1 of 10 pools that have been manipulated to differ in temperature by two degrees each. In this situation, the researcher has used the critical features of experimental research: random assignment and experimental control. However, because the independent variable in this study, pool temperature, is measured at the interval level of measurement, a correlational statistical
891
procedure could be used to analyze the data.
It may surprise you to learn that ANOVA and correlational statistics are much more similar in nature than they are different. Entire textbooks have been written that demonstrate that any research situation that can be analyzed with ANOVA may also be analyzed with correlational statistics with exactly the same result (Keppel & Zedeck, 1989).
892
“Correlation is Not Causality”—What does that Mean?
Students are often taught that that “correlation is not causality.” This is true, in that a correlational research methodology does not allow one to make causal inferences. However, this assertion has absolutely nothing to do with correlational statistics. The ability to infer causal relationships between variables is a function of how data are collected, not how data are analyzed. The choice of statistical procedure is primarily a function of how variables are measured. ANOVA is used when the independent variables are categorical, whereas correlation is used when the independent variables are continuous. ANOVA is not “better” than correlation—it is simply different.
One issue that may confuse students in learning about correlational research is the common practice of labeling one of the variables the “independent variable” and the other the “dependent variable.” For example, researchers relating high school GPA and college GPA to each other may wish to call high school GPA the “independent variable” and college GPA the “dependent variable.” Although these labels are appropriate in experimental research, because changes in the dependent variable are perceived to be caused by or “dependent on” changes in the independent variable, in correlational research, these labels infer a direction of influence between the two variables that may be inappropriate. To avoid confusion, researchers sometimes refer to the X and Y variables in a correlational research study as the predictor variable and criterion variable, respectively, particularly when the goal of the study is to predict one variable from another.
893
13.7 Looking Ahead
The main purpose of this chapter was to introduce correlational statistical procedures. The Pearson correlation coefficient measures the linear relationship between two variables measured at the interval and/or ratio level of measurement, whereas the Spearman rank- order correlation measures the relationship between two ordinal variables. In discussing the difference between correlational statistics and correlational research, it is crucial to remember that the choice of statistical procedure is based on the nature of the variables included in the analysis. The next and final chapter of the book discusses statistical procedures used when both the independent and dependent variables are categorical in nature.
894
13.8 Summary
One way to visually examine the relationship between scores on two continuous variables measured at the interval or ratio level of measurement is to create what is known as a scatterplot, a graphical display of paired scores on two variables.
Correlation may be defined as a mutual or reciprocal relationship between two variables such that systematic changes in the values of one variable are accompanied by systematic changes in the values of another variable.
The relationship between variables may be described along three aspects: the nature of the relationship, the direction of the relationship, and the strength of the relationship.
The nature of the relationship between variables pertains to the manner in which differences in scores on one variable correspond to differences in scores on another variable. A linear relationship is a relationship between variables appropriately represented by a straight line, such that increases or decreases in one variable are associated with corresponding increases or decreases in another variable. A nonlinear relationship is a relationship between variables that is not appropriately represented by a straight line.
In terms of the direction of a relationship, when scores on variables move in the same direction, with increases in one variable associated with increases in another variable, this is referred to as a positive relationship. When scores on variables move in the opposite direction, with increases in one variable associated with decreases in another variable, this is called a negative relationship.
The strength of a relationship is the extent to which scores on one variable are associated with scores on another variable. A perfect relationship is a relationship in which each score for one variable is associated with only one score for the other variable, a strong relationship is a relationship in which a score on one variable is associated with a relatively small range of scores on another variable, and a zero relationship exists when all of the scores on one variable are associated with a wide range of scores on another variable.
Correlational statistics are statistics designed to measure the relationship between variables. There are a number of different correlational statistics; the choice of statistic depends on how variables in a study have been measured.
The most commonly used correlational statistic is the Pearson correlation coefficient (r), which is a statistic that measures the linear relationship between two continuous variables measured at the interval and/or ratio level of measurement. The possible values for the Pearson r range from −1.00 to +1.00. The sign (+ or –) of the correlation indicates the direction of the relationship, and the numeric value of the correlation represents the
895
strength of the relationship.
Correlational statistics are based on the concept of variance, which represents differences among the scores for variables, and covariance, which is the extent to which two variables have shared variance (i.e., variance in common with each other). The Pearson correlation coefficient r is a ratio of the covariance between two variables to the variance of the two variables.
Whereas the main purpose of correlation is to measure and test the relationship between variables, the main goal of regression is to use of a relationship between two or more correlated variables to predict values of one variable from values of other variables.
Linear regression refers to a statistical procedure in which a straight line is fitted to a set of data to best represent the relationship between the two variables. The linear regression equation (Y′ = a + bX) is a mathematical equation based on the relationship between two variables that predicts a score on one variable using a score on the other variable. In this equation, Y′ is a predicted score for the Y variable, a is an estimate of the Y-intercept, b is an estimate of the slope, and X is a score for the X variable.
Unlike the Pearson r, which measures the relationship between two variables measured at the interval or ratio level of measurement, the Spearman rank-order correlation (r) measures the relationship between two variables measured at the ordinal level of measurement.
Correlational statistics such as the Pearson r and the Spearman rs are distinct from correlational research methods, which are methods of conducting research designed to examine the relationship between variables without the ability to infer cause-effect relationships. Correlational research refers to methods used to collect data, whereas correlational statistics refers to methods used to analyze data. Researchers sometimes refer to the X and Y variables in a correlational research study as the predictor variable and criterion variable, respectively, particularly when the goal of the study is to predict one variable from another.
896
13.9 Important Terms
scatterplot (p. 572) correlation (p. 574) linear relationship (p. 574) nonlinear relationship (p. 575) positive relationship (p. 576) negative relationship (p. 576) correlational statistics (p. 578) Pearson correlation coefficient (r) (p. 578) covariance (p. 580) regression (p. 601) linear regression (p. 603) linear regression equation (p. 603) rank (p. 611) Spearman rank-order correlation (rs) (p. 612) predictor variable (p. 622) criterion variable (p. 622)
897
13.10 Formulas Introduced in this Chapter
898
Degrees of Freedom for Pearson Correlation (df)
(13-1) df = N − 2
Sum of Products (Definitional Formula) (SPxy)
(13-2) SP XY = ∑ X − X ¯ Y − Y ¯
Sum of Squares for the X Variable (Definitional Formula) (SSX)
(13-3) SS X = ∑ X − X ¯ 2
Sum of Squares for the Y Variable (Definitional Formula) (SSY)
(13-4) SS Y = ∑ Y − Y ¯ 2
Pearson Correlation (Definitional Formula) (r)
(13-5) r = ∑ X − X ¯ Y − Y ¯ ∑ X − X ¯ 2 ∑ Y − Y ¯ 2
Measure of Effect Size for the Pearson Correlation Coefficient (r2)
(13-6) Estimate of effect = r 2
Sum of Products (Computational Formula) (SPXY)
(13-7) SP XY = ∑ XY − ∑ X ∑ Y N
899
Sum of Squares for the X Variable (Computational Formula) (SSX)
(13-8) SS X = ∑ X 2 − ∑ X 2 N
Sum of Squares for the Y Variable (Computational Formula) (SSY)
(13-9) SS Y = ∑ Y 2 − ∑ Y 2 N
Pearson Correlation (Computational Formula) (r)
(13-10) r = ∑ XY − ( ∑ X ) ( ∑ Y ) N ( ∑ ( X 2 − ( ∑ X ) 2 N ) ( ∑ Y 2 − ( ∑ Y ) 2 N )
Linear Regression Equation
(13-11) Y ′ = a + bX
Estimate of Slope (b)
(13-12) b = r S Y S X
Estimate of Y-Intercept (a)
(13-13) a = Y ¯ − b X ¯
Spearman Rank-Order Correlation (rs)
(13-14) r s = 1 − 6 ∑ D 2 N N 2 − 1
Measure of Effect Size for the Spearman Rank-Order Correlation r s 2
900
(13-15) Estimate of effect = r s 2
901
13.11 Using SPSS
902
Pearson Correlation: The First Impression Study (13.1)
1. Define independent and dependent variables (name, # decimals, labels for the variables) and enter data for the variables.
2. Select the Pearson correlation procedure within SPSS.
How? (1) Click Analyze menu, (2) click Correlate, and (3) click Bivariate.
3. Select the variables to be correlated, and ask for descriptive statistics.
How? (1) Click variables and Variables, (2) click and click Means and standard deviations, (3) click , and (4) click .
903
4. Examine output.
Correlations
Linear Regression: The First Impression Study (13.4)
1. Begin the Linear regression procedure within SPSS.
How? (1) Click Analyze menu, (2) click Regression, and (3) click Linear.
2. Identify the dependent variable and the independent variable.
904
How? (1) Click dependent variable and Dependent, (2) Click independent variable and Independent(s), and (3) click OK.
3. Examine output.
Regression
Spearman Rank-Order Correlation: The Business Student Study (13.5)
1. Define the two sets of ranks (name, # decimals, labels for the variables) and enter the two sets of ranks.
2. Begin the Spearman rank order correlation procedure within SPSS.
905
How? (1) Click Analyze menu, (2) click Correlate, and (3) click Bivariate.
3. Select the variables to be correlated, and ask for the Spearman rank-order correlation.
How? (1) Click variables and Variables, (2) click Spearman, (3) click , and (4) click .
4. Examine output.
Nonparametric Correlations
906
13.12 Exercises
1. For each of the following situations, create a scatterplot of the data and describe the nature, direction, and strength of the relationship.
a. A researcher hypothesizes a positive relationship between the number of children in a family and the number of televisions a family owns.
Family # Children # Televisions
1 3 5
2 1 3
3 2 5
4 4 6
5 3 3
6 5 8
b. Is there a relationship between the amount of time spent watching sports on television and the amount of time actually playing sports? The following data are from a small sample of children who reported their average number of daily hours spent watching sports on television and playing sports.
Child # Hours Watching # Hours Playing
1 .5 .2
2 6 .3
3 1 1.0
4 3 3.0
5 1 .4
6 5 .5
7 2 2.0
8 2 1.8
9 4 1.6
c. A test developer gives people two versions of a survey (a longer version that contains 40 items and a shorter version that contains 20 items) and wants to see if scores on the two versions are related to each other.
Participant Longer Version Shorter Version
907
1 37 14
2 17 11
3 26 17
4 15 9
5 27 13
6 33 16
7 20 5
8 14 15
9 24 18
10 22 13
11 29 12
2. For each of the following situations, create a scatterplot of the data and describe the nature, direction, and strength of the relationship.
a. A set of parents hypothesize there is a negative relationship between the average number of hours their children spend on the Internet and their scores on a recent school quiz (range = 0–100).
Student # Hours Internet Quiz Score
1 7 65
2 2 96
3 4 92
4 8 49
5 5 73
6 3 90
7 6 57
8 3 78
9 1 81
10 4 59
11 6 76
12 1 95
13 2 66
b. A college counselor hypothesizes that the farther students live from campus (in miles), the less they feel they are part of the school community (1 = low, 10 = high).
908
Student Miles From Campus Part of Community
1 2 9
2 19 7
3 32 3
4 23 9
5 40 5
6 5 2
7 20 1
8 16 2
9 10 8
10 30 5
11 3 4
12 36 8
13 17 4
14 7 6
15 25 4
16 12 4
c. A publishing company develops a new college admissions test and hypothesizes that scores on the test are positively related to academic achievement in college. The company administers the test to a sample of college students (possible scores range from 0–100) and collects college grade point averages (GPAs) of the students.
Student Test Score College GPA
1 68 2.3
2 84 3.7
3 69 3.2
4 78 2.7
5 91 3.4
6 26 1.5
7 57 2.9
8 62 3.6
9 39 1.8
909
10 74 3.9
11 43 2.6
12 52 2.0
13 40 3.3
14 35 2.5
3. For each of the following, calculate the degrees of freedom (df) and identify the critical value of r.
a. N = 8, α = .05 (two-tailed) b. N = 19, α = .05 (two-tailed) c. N = 36, α = .05 (two-tailed) d. N = 50, α = .05 (one-tailed)
4. For each of the following, calculate the degrees of freedom (df) and identify the critical value of r.
a. N = 21, α = .05 (two-tailed) b. N = 40, α = .05 (two-tailed) c. N = 12, α = .05 (one-tailed) d. N = 120, α = .05 (one-tailed)
5. For each of the following, calculate the Pearson correlation (r). a. SPXY = −3.00, SSX = 8.00, SSY = 10.00 b. SPXY = 4.00, SSX = 9.00, SSY = 5.00 c. SPXY = −1.50, SSX = 7.50, SSY = 10.50 d. SPXY = 13.26, SSX = 5.73, SSY = 142.08
6. For each of the following, calculate the Pearson correlation (r). a. SPXY = .82, SSX = 6.91, SSY = 9.24 b. SPXY = −43.29, SSX = 411.89, SSY = 17.50 c. SPXY = 205.96, SSX = 485.73, SSY = 612.52
7. Calculate the Pearson correlation coefficient for each of the following sets of data a.
Variable X Variable Y
15 24
10 8
8 15
23 21
b.
Variable X Variable Y
910
3 1
6 8
7 1
1 8
5 7
8 2
5 1
c.
Variable X Variable Y
17 41
23 11
9 17
15 32
12 23
20 46
19 20
10 37
14 13
8 39
8. Calculate the Pearson correlation coefficient for each of the situations presented earlier in Exercise 1.
9. A teacher hypothesizes there is a positive relationship between her students' scores on a midterm examination and on the final examination.
Student Midterm Exam Final Exam
1 69 73
2 87 91
3 65 84
4 78 81
5 52 67
6 61 69
a. Draw a scatterplot of the data.
911
b. Calculate descriptive statistics (the mean and standard deviation) for the two variables.
c. State the null and alternative hypotheses (H0 and H1) (use a non-directional H1).
d. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values, and state a decision rule. 3. Calculate a statistic: Pearson correlation (r). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (r2).
e. Draw a conclusion from the analysis. f. Relate the result of the analysis to the research hypothesis.
10. What makes a person a “leader”? A social psychologist hypothesizes that the extent to which people are perceived to be leaders is positively related to the amount of talking they do, regardless of what they are actually saying. She conducts an experiment in which groups of participants solve a problem together. One of the participants in each group is a confederate who speaks a different number of sentences for different groups, half of which are designed to help the group solve the problem and half are not helpful. Following the group work, the confederate is rated by the other participants on his leadership ability (1 = low, 10 = high). The following data represent the findings of this fictional researcher.
Group # Sentences Spoken Rating of Leadership
1 4 5
2 10 6
3 29 7
4 22 9
5 6 2
6 7 3
7 18 8
8 15 4
a. Draw a scatterplot of the data. b. Calculate descriptive statistics (the mean and standard deviation) for the two
variables. c. State the null and alternative hypotheses (H0 and H1) (use a non-directional
H1). d. Make a decision about the null hypothesis.
912
1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values, and state a decision rule. 3. Calculate a statistic: Pearson correlation (r). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (r2).
e. Draw a conclusion from the analysis. f. Relate the result of the analysis to the research hypothesis.
11. What makes a joke funny? Sigmund Freud's research led him to theorize that people find jokes funnier when the jokes contain higher levels of sexuality or hostility. Kirsh and Olczak (2001) questioned whether gender affects this relationship between humor and hostility by measuring men's and women's reactions to comic books of differing levels of violence. The following data are representative of the women in their study. The researchers hypothesized that, because women find violence less humorous than do men, there is a negative relationship between the level of violence in a comic book (1 = low to 7 = high) and how humorous women rate the comic book (1 = not at all humorous to 6 = extremely humorous). In other words, the higher the level of violence in the comic book, the less humorous the comic is to women.
a. Draw a scatterplot of the data. b. Calculate descriptive statistics (the mean and standard deviation) for the two
variables.
Woman Level of Violence Humorous Rating
1 2 4
2 7 3
3 6 3
4 3 4
5 4 2
6 7 1
7 5 3
8 2 6
9 1 5
10 6 2
c. State the null and alternative hypotheses (H0 and H1) (use a non-directional H1).
d. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df).
913
2. Set alpha (a), identify the critical values, and state a decision rule. 3. Calculate a statistic: Pearson correlation (r). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (r2).
e. Draw a conclusion from the analysis. f. Relate the result of the analysis to the research hypothesis.
Student Time Exam Score
1 34 36
2 39 72
3 42 92
4 43 74
5 44 89
6 47 63
7 48 88
8 51 86
9 52 71
10 53 79
11 56 76
12 56 74
13 57 83
14 57 63
15 59 77
16 59 82
17 62 74
18 62 73
19 64 70
20 68 85
21 73 77
22 77 73
23 77 80
24 82 73
25 92 61
26 98 75
914
27 99 62
28 102 76
29 103 58
30 108 58
31 112 86
32 121 90
12. Earlier in this chapter, we discussed a study that examined the relationship between the amount of time students take to complete a test and their scores on the test (Herman, 1997). The two variables in this study were Time (the number of minutes taken by each student to finish the examination) and Exam score (the number of correct answers for each student, ranging from 0 to 100). The researcher in this study wished to test the hypothesis that the more time students take to finish an exam, the lower their score on the exam. In other words, there is a negative relationship between Time and Exam score. The data for the 32 students in this study are listed on page 636 (note that the scatterplot and descriptive statistics for these data were provided earlier in this chapter).
a. State the null and alternative hypotheses (H0 and H1) (use a non-directional H1).
b. Make a decision about the null hypothesis. 1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values, and state a decision rule. 3. Calculate a statistic: Pearson correlation (r). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (r2).
c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
13. For each of the sets of data in Exercise 1, calculate SPXY, SSX, SSY, and r using the computational formulas.
14. For each of the following, calculate the linear regression equation.
a. Variable X: X ¯ i = 5.00 , SX = 2.00;
Variable Y: Y ¯ = 4.00 , SX = 1.00, Pearson correlation: r = .40
b. Variable X: X ¯ i = 3.59 , SX = 1.32;
Variable Y: Y ¯ = 7.52 , SY = 2.46, Pearson correlation: r = –.28
915
c. Variable X: X ¯ i = 73.87 , SX = 12.65;
Variable Y: Y ¯ = 86.32 , SY = 16.80, Pearson correlation: r = .07 15. For each of the following, calculate the linear regression equation.
a. Variable X: X ¯ i = 45.00 , SX = 5.00;
Variable Y: Y ¯ = 11.00 , SY = 3.00, Pearson correlation: r = .30
b. Variable X: X ¯ i = 37.22 , SX = 11.76;
Variable Y: Y ¯ = 104.32 , SY = 15.92, Pearson correlation: r = .62
c. Variable X: X ¯ i = 4.32 , SX = .97;
Variable Y: Y ¯ = 2.51 , SY = .46, Pearson correlation: r = –.09 16. Calculate the linear regression equation for each of the situations presented earlier in
Exercise 1 and draw the regression equation into its scatterplot. 17. Calculate the linear regression equation for the data in Exercise 9 (midterm and final
exam scores) and draw the regression equation into its scatterplot. What final exam score would you predict for a student with a midterm exam score of 75?
18. Calculate the linear regression equation for the data in Exercise 11 (level of violence and humorous rating) and draw the regression equation into its scatterplot. What humorous rating would you predict for a comic book with a level of violence of 4?
19. Calculate the Spearman rank-order correlation (rs) for each of the following pairs of ranks.
a.
Rank Order 1 Rank Order 2
5 5
3 1
4 4
1 2
2 3
b.
Rank Order 1 Rank Order 2
1 1
2 3
916
3 4
4 6
5 2
6 5
7 7
8 9
9 8
10 10
c.
Rank Order 1 Rank Order 2
2.5 1
5 3
4 5.5
1 2
2.5 4
6 5.5
20. Calculate the Spearman rank-order correlation (rs) for each of the following pairs of ranks.
a.
Rank Order 1 Rank Order 2
1 5
2 3
3 1
4 8
5 2
6 7
7 4
8 6
9 7.5
10 10
11 7.5
917
b.
Rank Order 1 Rank Order 2
1 3
2 6
3.5 2
3.5 4
5 1
6 9
7.5 5
7.5 11
c.
Rank Order 1 Rank Order 2
6 8
2 4
10 10
3 1
8 12
4 5
12 13
5 3
13 14
9 6
1 2
11 7
7 9
14 11
21. Do students' needs change as they progress through their education? One study asked third-year and fourth-year college students preparing to be teachers what knowledges they would like to gain in order to effectively work with misbehaving students (Psunder, 2009). The seven knowledges below were rank ordered based on how often they were mentioned (1 = most often mentioned to 7 = least often mentioned). As the fourth-year students had gained more actual teaching experience than the third-
918
year students, it was believed their desired knowledges would be different from the third-year students.
Knowledge Third-Year Students
Fourth-Year Students
Working with problematic students 1 2
How to solve problems and conflicts 2 4
Cooperation with parents of problematic children
3 1
Discipline (practical suggestions) 4 5
How to make classes interesting 5 6
How to encourage students' activity 6 7
Taking measures in accordance with school legislation
7 3
a. State the null and alternative hypotheses (H0 and H1). b. Make a decision about the null hypothesis.
1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values, and state a decision rule. 3. Calculate a statistic: Spearman rank-order correlation (rs). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size r s 2 .
c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
22. As people do not agree on who are the most important artists of all time, one study examined whether textbooks written in English and French agree on who are the most important artists who were French born or spent a substantial amount of time in France (O'Hagan & Kelly, 2005). To measure the importance of each artist, the researchers counted the number of illustrations of each artist's work that appeared in a sample of English and French textbooks; the artists were then rank ordered from highest to lowest based on the number of illustrations. The rankings of the top 16 artists are provided below. Imagine that the researcher hypothesized that the rankings of artists are similar for English and French textbooks.
a. State the null and alternative hypotheses (H0 and H1). b. Make a decision about the null hypothesis.
1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical values, and state a decision rule.
919
Artist English Textbooks French Textbooks
Picasso 1 1
Matisse 2 2
Cezanne 3 3
Manet 4 6
Monet 5 4
Braque 6 7.5
van Gogh 7 5
Gauguin 8 7.5
Degas 9 10
Renoir 10 9
Duchamp 11 16
Courbet 12 11
Miro 13 15
Seurat 14 14
Leger 15 12
Toulouse-Lautrec 16 13
3. Calculate a statistic: Spearman rank-order correlation (rs). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size r s 2 .
c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
920
Answers to Learning Checks
Learning Check 1
2. a. There is a strong, negative linear relationship such that the more days a student
misses during a semester, the lower the student's score on the final exam.
b. There is a very weak relationship between the number of times a student raises his or her hand and how favorably the student is viewed by the teacher.
Learning Check 2
2. a. df = 4, critical value = ±.811 b. df = 7, critical value = ±.666 c. df = 16, critical value = ±.468
3. a. r = .56 b. r = –.44 c. r = .10
Learning Check 3
921
2. a. H0: ρ = 0, H1: ρ ≠ 0 b.
1. df = 24 2. If r < –.388 or > .388, reject H0; otherwise, do not reject H0 3. SPXY = 487.87, SSX = 298,203.85, SSY = 4.24, r = .43 4. r = .43 > .388 ∴ reject H0 (p < .05) 5. r = .43 < .496 ∴ p < .05 (but not < .01) 6. r2 = .18 (a medium effect)
c. There was a statistically significant positive relationship between students' GMAT scores and their grade point average (GPA), r(24) = .43, p < .05, r2 = .18.
d. This analysis supports the hypothesis that the higher the score on the GMAT, the greater the likelihood the student will achieve a high level of academic performance.
Learning Check 4
2. a. SP XY = 157 − 36 24 6 = 13.00 SS X = 238 − 36 2 6 = 22.00 SS Y = 124
− 24 2 6 = 28.00 r = 13.00 22.00 28.00 = 52 b. SP XY = 1131 − 168 56 8 = − 45.00 SS X = 3724 − 168 2 8 = 196.00 SS
Y = 450 − 24 2 8 = 58.00 r = − 45.00 196.00 58.00 = − . 42 c. SP XY = 2495.20 − 38.13 789 12 = − 11.85 SS X = 140.77 − 38.13 2 12
= 19.61 SS Y = 57459 − 789 2 12 = 5582.25 r = − 11.85 19.61 5582.25 = − . 04
Learning Check 5
2. a. Y′ = 6.30 + .45 (X)
b. Y′ = 2.49 + .50 (X)
922
c. Y′ = 3.72 – .06 (X)
Learning Check 6
2. a. r s = 1 − 6 6 4 4 2 − 1 = . 40 b. r = 1 − 6 12 4 6 2 − 1 = . 66 c. r s = 1 − 6 19 4 9 2 − 1 = . 84
3. a. H0: ρ = 0, H1: ρ ≠ 0 b.
1. df = 6 2. If rs < –.738 or > .738, reject H0; otherwise, do not reject H0 3. N = 8, ΣD2 =115.50, rs = –.38 4. rs = –.38 is not < –.738 or >. 738 ∴ do not reject H0 (p > .05) 5. Not applicable (H0 not rejected) 6. r s 2 = . 14
c. There was not a significant relationship between the rank ordering of boys and girls in their preferences for the eight musical instruments, rs (6) = –.38, p > .05, r2 = .14.
d. This analysis does not support the hypothesis that boys and girls would display similar preferences for the different types of musical instruments.
923
Answers to Odd-Numbered Exercises
1. a. There is a strong, positive linear relationship between the number of children
in a family and the number of televisions a family owns such that the larger the number of children, the more televisions a family owns.
b. There is a nonlinear relationship between the average number of daily hours spent watching sports on television and playing sports.
c. There is a moderate positive relationship between scores on the longer and shorter versions of the survey.
3. a. df = 6, critical value = ±.707 b. df = 17, critical value = ±.456 c. df = 34, critical value = ±.349 d. df = 48, critical value = .257 or –.257
5. a. r = –.34
924
b. r = .60 c. r = –.17 d. r = .46
7. a. SPXY = 91.00, SSX = 134.00, SSY = 150.00, r = .64 b. SPXY = −18.00, SSX = 34.00, SSY = 72.00, r = –.36 c. SPXY = −78.30, SSX = 228.10, SSY = 1434.90, r = –.14
9.
a.
b. Midterm exam (X): N = 6, X ¯ = 68.67 , sX = 12.44
Final exam (Y): N = 6, Y ¯ = 77.50 , SY = 9.38 c. H0: ρ = 0, H1: ρ ≠ 0 d.
1. df = 4 2. If r < –.811 or > .811, reject H0; otherwise, do not reject H0 3. SPXY = 495.00, SSX = 773.33, SSY = 439.50, r = .85 4. r = .85 > .811 ∴ reject H0 (p < .05) 5. r = .85 < .917 ∴ p < .05 but not < .01 6. r2 = .72 (a large effect)
e. There was a statistically significant positive relationship between students' midterm exam and final exam scores, r(4) = .85, p < .05, r2 = .72.
f. This analysis supports the teachers hypothesis that there is a positive relationship between her students' scores on a midterm examination and on the final examination.
11.
a.
925
b. Level of violence (X): N = 10, X ¯ = 4.30 , Sx = 2.21
Humorous rating (Y): N = 10, Y ¯ = 3.50 , SY = 1.58 c. H0: ρ = 0, H1: ρ ≠ 0 d.
1. df = 8 2. If r < –.632 or > .632, reject H0; otherwise, do not reject H0 3. SPXY = −24.50, SSX = 44.10, SSY = 22.50, r = –.78 4. r = –.78 < –.632 ∴ reject H0 (p < .05) 5. r = –.78< –.765. ∴ p < .01 6. r2 = .61 (a large effect)
e. There was a statistically significant negative relationship between the level of violence in comic books and women's ratings of humor, r(8) = –.78, p < .01, r2
= .61. f. This analysis supports the hypothesis that there is a negative relationship
between the level of violence in a comic book and how humorous women rate the comic book.
a. SP XY = 101 − 18 30 6 = 11.00 SS X = 64 − 18 2 6 = 10.00 SS Y = 168 − 30 2 6 = 18.00 r = 11.00 10.00 18.00 = . 82
b. SP XY = 28.80 − 24.50 10.80 9 = − . 60 SS X = 96.25 − 24.50 2 9 = 29.56 SS Y = 20.34 − 10.80 2 9 = 7.38 r = 6.22 8.71 7.41 = − . 04
c. SP XY = 3537 − 264 143 11 = 105.00 SS X = 6874 − 264 2 11 = 538.00 SS Y = 1999 − 143 2 11 = 140.00 r = 6.22 8.71 7.41 = . 38
15. a. Y' = 2.90 + .18(X) b. Y'=73.08 + .84(X) c. Y' = 2.69 + –.04(X)
Final′ = 33.55 + .64 (Midterm) a. For Midterm = 75, Final′ = 81.55
19. a. r s = 1 − 6 6 5 5 2 − 1 = . 70 b. r s = 1 − 6 18 10 10 2 − 1 = . 89 c. r s = 1 − 6 12 6 6 2 − 1 = . 66
926
21. a. H0: ρ = 0, H1: ρ ≠ 0 b.
1. df = 5 2. If rs < –.786 or > .786, reject H0; otherwise, do not reject H0 3. N = 7, ΣD2 = 28, rs = .50 4. rs = .50 is not < –.786 or > .786 ∴ do not reject H0 (p > .05) 5. Not applicable (H0 not rejected) 6. r s 2 = . 25
c. There was a nonsignificant relationship between the rank ordering of desired knowledges for the third-year and fourth-year students, rs (5) = .50, p > .05, r
2
= .25. d. This analysis supports the hypothesis that the desired knowledges of the third-
year and fourth-year students are different from each other.
927
Sharpen your skills with SAGE edge at edge.sagepub.com/tokunaga
SAGE edge for students provides a personalized approach to help you accomplish your coursework goals in an easy-to-use learning environment. Log on to access:
eFlashcards Web Quizzes Chapter Outlines Learning Objectives Media Links SPSS Data Files
928
Chapter 14 Chi-Square
929
Chapter Outline 14.1 An Example From the Research (One Categorical Variable): Are You My Type? 14.2 Introduction to the Chi-Square Statistic
Observed and expected frequencies Characteristics of the distribution of chi-square statistics Assumptions underlying the chi-square statistic
Independence of observations Minimum expected frequencies
14.3 Inferential Statistic: Chi-Square Goodness-of-Fit Test State the null and alternative hypotheses (H0 and H1) Make a decision about the null hypothesis
Calculate the degrees of freedom (df) Set alpha (α), identify the critical value, and state a decision rule Calculate a statistic: Chi-square (χ2) Make a decision whether to reject the null hypothesis Determine the level of significance Calculate a measure of effect size (Cramér's φ)
Draw a conclusion from the analysis Relate the result of the analysis to the research hypothesis Conducting the chi-square goodness-of-fit test with unequal hypothesized proportions
Calculate expected frequencies (fe)
Calculate the chi-square statistic (χ2) 14.4 An Example From the Research (Two Categorical Variables): Seeing Red 14.5 Inferential Statistic: Chi-Square Test of Independence
State the null and alternative hypotheses (H0 and H1) Make a decision about the null hypothesis
Calculate the degrees of freedom (df) Set alpha (α), identify the critical value, and state a decision rule Calculate a statistic: chi-square (χ2) Make a decision whether to reject the null hypothesis Determine the level of significance Calculate a measure of effect size (Cramér's φ)
Draw a conclusion from the analysis Relate the result of the analysis to the research hypothesis
14.6 Parametric and Nonparametric Statistical Tests Parametric statistical tests Introduction to nonparametric statistical tests Examples of nonparametric statistical tests
14.7 Looking Ahead 14.8 Summary 14.9 Important Terms 14.10 Formulas Introduced in This Chapter 14.11 Using SPSS 14.12 Exercises
Chapter 13 discussed correlational statistics, which are statistics designed to measure the relationship between variables. One correlational statistic, the Pearson correlation
930
coefficient, measures the relationship between two variables measured at the interval or ratio level of measurement. Another correlational statistic, the Spearman rank-order correlation, relates two ordinal variables with each other. The last chapter of this book introduces a statistic used when all of the variables in a study are categorical in nature, measured at the nominal level of measurement. Two research situations will be examined: one that consists of a single categorical variable and one that contains two categorical variables. As you have seen throughout this book, how variables in a study are measured determines which statistical procedure may be appropriate to analyze the data. However, the steps involved in analyzing the data remain the same.
931
14.1 An Example from the Research (One Categorical Variable): Are You my Type?
It is not surprising that the study of personality has received a great deal of attention for many years. People are very interested in understanding and identifying ways in which they are both similar to and different from others. One group of people for whom the topic of personality has been given particular emphasis is college students. You may have read that students' personalities are related to the colleges they attend, the majors they select, and the careers they pursue after graduation. Research has investigated the relationship between the personality characteristics of college students and behaviors that are both academic (i.e., exam scores, cheating on tests, the likelihood of graduating) and nonacademic (i.e., involvement in extracurricular activities, drug and alcohol use, sexual activity).
One line of research has examined whether the personalities of college students are different from other groups of people as well as from earlier generations of students. One such study, conducted by Kenneth Stewart and Paul Bernhardt at Frostburg State University, measured college students on two dimensions of personality (Stewart & Bernhardt, 2010). One dimension was the degree to which a person is extraverted or introverted. People who are extraverted tend to be assertive, talkative, and self-confident, whereas introverts are reserved, quiet, and reluctant to express their emotions or feelings. The second dimension was the degree to which one accepts or questions social norms, which are models or rules widely accepted within a group. People defined as “norm favoring” adhere to social norms and are described as efficient and self-disciplined; people who are “norm questioning” challenge traditional values and beliefs, tending more toward unconventionality and skepticism.
On the basis of their combination of extraversion/introversion and norm favoring/norm questioning, the college students in the study, which will be referred to as the personality study, were classified as one of four types: Alphas, Betas, Gammas, or Deltas. These four personality types were developed by psychologist Harrison Gough and his colleagues in their measure of personality known as the California Psychological Inventory (CPI) (Gough & Bradley, 2005). Alphas (extraverted and norm favoring) are task oriented, assertive, and talkative. They have the potential to be charismatic and selfless leaders, but they may also be authoritarian, manipulative, and self-centered. Betas (introverted and norm favoring) have a high degree of self-control, are more comfortable in the role of follower than leader, place others' needs before their own, but are also somewhat detached. Gammas (extraverted and norm questioning) question traditional beliefs and values, and freely express skepticisms and doubts. They may be creative and visionary, but they may also be reckless and rebellious. Finally, Deltas (introverted and norm questioning) are reflective, somewhat detached from others, and absorbed in their own thoughts. Deltas can be imaginative and artistic, but they may also be withdrawn and aloof.
932
The developers of the CPI believed that the four personality types are evenly distributed in the population, with approximately 25% of individuals classified as Alphas, Betas, Gammas, or Deltas. However, on the basis of their review of research, the researchers in the personality study hypothesized that the distribution of the four types among college students does not match or fit this distribution.
To test their hypothesis, the researchers collected CPI responses from 588 college students and classified each student as an Alpha, Beta, Gamma, or Delta. This variable, which will be referred to as Type, is a categorical variable measured at the nominal level of measurement. As it would require a great deal of space to list the data for all 588 students, Table 14.1 summarizes the data using a frequency distribution table. You may observe from the table, for example, that 251 of the 588 students were classified as Alphas, representing 42.7% (251/588) of the total sample.
Figure 14.1 presents a bar chart comparing the percentages of the four types. You will recall from Chapter 2 that a bar chart displays the frequencies (f) or percentages (%) of groups that comprise a categorical variable. The bars in a bar chart do not touch each other, indicating that the groups are qualitatively distinct from each other and cannot be placed along a numeric scale or continuum. The bar chart in Figure 14.1 displays percentages rather than frequencies because only reporting frequencies can be somewhat uninformative. For example, the 251 students classified as Alphas (f = 251) would have been evaluated differently if they represented 251 out of a total of 5,880 students rather than the 588 included in the study. Looking at Figure 14.1, you find that Alphas (extraverted and norm favoring) and Gammas (extraverted and norm questioning) were the most frequently occurring personality types among the sample of college students—combining these two types, three fourths of the students (75.7%) were extraverted rather than introverted. Among the introverts, students were almost equally likely to be a norm-favoring Beta (11.0%) or a norm-questioning Delta (13.3%).
Table 14.1 Frequency Distribution Table for Four Types in the Personality Study Table 14.1 Frequency Distribution Table for
Four Types in the Personality Study
Type f %
Alpha 251 42.7%
Beta 65 11.0%
Gamma 194 33.0%
Delta 78 13.3%
Total 588 100.0%
933
In earlier chapters, after creating tables and figures to examine data, descriptive statistics such as the mean and standard deviation were calculated. Measures of central tendency and variability are appropriate for continuous variables measured at the interval or ratio level of measurement; however, they do not describe categorical variables. For example, if you asked a group of people where they live, it would not make sense to calculate the “average” city because cities do not fall on a numeric continuum. Instead, you would summarize their responses by reporting the number and/or percentage of the sample living in each city. Consequently, analyzing categorical variables involves comparing frequencies and percentages rather than means and standard deviations.
For the sample of 588 students in the personality study, Table 14.1 and Figure 14.1 suggest that Alphas and Gammas are much more prevalent than Betas and Deltas. This provides initial support for the researchers' hypothesis that the distribution of personality types among college students is different from the distribution proposed by the developers of the CPI. The next section introduces a statistic that will be used to determine whether the difference between the two distributions is statistically significant.
934
14.2 Introduction to the Chi-Square Statistic
Figure 14.1 Bar Chart of Four Types for the Personality Study
The primary issue of concern illustrated by the personality study is whether the distribution of frequencies of groups that comprise a variable of interest is different from a proposed distribution. Are the frequencies of the four types in the sample of 588 college students different from frequencies expected to occur based on the beliefs of the developers of the CPI? Like many of the inferential statistics discussed in this book, answering this question requires calculating a statistic that evaluates the difference between a sample and a hypothesized population.
Although the research situation presented in this chapter may seem new to you, it's actually similar to the research situations examined in Chapter 7 (the test of one mean). In the research studies in that chapter, the goal was to evaluate the difference between a sample mean and a hypothesized population mean X ¯ − μ . To evaluate this difference, a sampling distribution was created of all of the possible sample means for samples of size N randomly selected from a population that has a mean equal to μ. Using this distribution of sample means, the difference between X ¯ − μ was transformed into a statistic known as the t- statistic. A distribution of t-statistics was then used to determine the probability of the t- statistic calculated from the sample of data. If this probability was low (p < .05), the null hypothesis was rejected, meaning that the difference between X ¯ − μ was statistically significant.
The situation illustrated by the personality study is very similar to those presented in Chapter 7, with one critical difference: The goal in this study was to evaluate differences between frequencies rather than means. The personality study was conducted to evaluate the difference between the distribution of frequencies of the four types in the sample of 588 students and the distribution expected to occur according to the developers of the CPI. Evaluating this difference requires a distribution of all of the possible combinations of four frequencies that can occur for samples of N = 588 drawn randomly from the population.
935
To illustrate how this distribution is created, imagine that from a population known to contain 25% of each of the four types, we randomly draw a sample of 588 students and count the number of Alphas, Betas, Gammas, and Deltas. If we were to do this an infinite number of times, we could develop a distribution of all of the possible combinations of the four frequencies for N = 588. This distribution can be used to transform the difference between a sample's frequencies and the proposed frequencies into a statistic. The probability of this statistic can then be determined to decide whether the difference between the distribution of frequencies in the sample and the distribution of proposed frequencies is statistically significant.
936
Observed and Expected Frequencies
To calculate the statistic used to analyze distributions of frequencies, two pieces of information must be obtained for each group in the study. The first piece of information is the observed frequency (f0), which is the frequency of a group in the sample of data. The observed frequencies for the personality study may be found in the frequency distribution table in Table 14.1. For example, the observed frequency of the Beta type is 65 (f0 = 65).
The second piece of information needed for each group is the expected frequency (fe), defined as the frequency of a group expected to occur in a sample of data under the assumption the null hypothesis is true. Expected frequencies are sometimes based on theory or research. For example, in the personality study, the expected frequencies were based on the CPI developers' belief that the four types are evenly distributed in the population. Based on this belief, 25% (.25) of the 588 students in the personality study would be expected to be classified as each type. The expected frequencies for the four types in the personality study are calculated below (along with the observed frequencies, which will be used for later calculations):
Type Observed Frequency (f0) Expected Frequency (fe)
Alpha 251 .25 ∗ 588 = 147
Beta 65 .25 ∗ 588 = 147
Gamma 194 .25 ∗ 588 = 147
Delta 78 .25 ∗ 588 = 147
Total 588 588
Based on the above calculations, according to the developers of the CPI, in a sample of 588 students, it is expected that 147 students belong to each of the four types. As a way of checking the accuracy of calculations, although the observed and expected frequencies for the groups differ from each other, they reach the same total (N = 588).
Once the observed and expected frequencies have been determined, the next step is to calculate the difference between these two frequencies for each group:
Type Observed Frequency (f0)
Expected Frequency (fe)
Observed – Expected (f0 – fe)
Alpha 251 147 251 – 147 = 104
Beta 65 147 65 – 147 = −82
Gamma 194 147 194 – 147 = 47
937
Delta 78 147 78 – 147 = −69
Total 588 588 0
To evaluate the data for the entire sample, we must combine the differences between the observed and expected frequencies across all of the groups. However, looking at the above table, we encounter a problem: The differences sum to 0 (104 + −82 + 47 + −69 = 0).
Similar to the differences between group means and the mean of the total sample in the one-way and two-way analysis of variance (ANOVA), the differences between observed and expected frequencies in any set of data will always sum to 0. Because it would be incorrect to say there are zero differences between the observed and expected frequencies, these differences must be squared to eliminate any negative differences (see table on page 654).
These squared differences between observed and expected frequencies are the basis of the chi-square statistic (χ2) (pronounced “kye,” which rhymes with “eye”). The chi-square statistic tests the difference between the distribution of observed frequencies in a sample and a proposed distribution of expected frequencies. The chi-square statistic calculated in this chapter is sometimes called the Pearson's chi-square statistic. It is named after Karl Pearson, who also developed the Pearson correlation coefficient discussed in Chapter 13. Later in this chapter, we will provide formulas needed to calculate different chi-square statistics.
Type Observed Frequency (f0)
Expected Frequency (fe)
(Observed – Expected)2
(f0 – fe) 2
Alpha 251 147 (104)2 = 10816
Beta 65 147 (–82)2 = 6724
Gamma 194 147 (47)2 = 2209
Delta 78 147 (–69)2 = 4761
Total 588 588
938
Characteristics of the Distribution of Chi-Square Statistics
In order for the probability of any particular value of the chi-square statistic to be determined, it must be located within a distribution of chi-square statistics; an example of a chi-square distribution is provided in Figure 14.2. Several aspects of this distribution are important to understand. First, like the t-statistic, F-ratio, and Pearson correlation, there is not just one chi-square statistic distribution; instead, there are different distributions depending on the number of groups that comprise the variable of interest. Second, because calculating the chi-square statistic involves squared differences between observed and expected frequencies ((f0 – fe)
2), the value for the chi-square statistic is always a positive number. Consequently, like the F-ratio, the distribution of chi-square statistics is positively skewed rather than symmetrical; as the number of groups comprising the variable of interest increases, the degree of skew decreases such that the chi-square distribution approaches a normal distribution. It's important for you to understand that the precise shape of this distribution changes somewhat dramatically as a function of the number of groups; however, for the sake of simplicity, we will use the distribution in Figure 14.2 throughout this chapter.
939
Assumptions Underlying the Chi-Square Statistic
In addition to the shape of the distribution of chi-square statistics, it is important to discuss two important assumptions on which this statistic is based: the independence of observations and minimum expected frequencies. Each of these two assumptions is discussed below.
Independence of Observations
In using the chi-square statistic, it is assumed that each observation in a set of data is independent of all other observations. In other words, an observation cannot be related to any other observation. This assumption, known as the assumption of independence of observations, is typically met by assigning each research participant to only one group. For example, in the personality study, the assumption of independence of observations is met by assigning each student to only one of the four types; a violation of this assumption would occur if a student could somehow be classified as belonging to more than one type at the same time.
Figure 14.2 Example Distribution of a Chi-Square Statistic
Minimum Expected Frequencies
As with all of the inferential statistics discussed in this book, the size of the sample on which the statistic is based plays a critical role: The smaller the sample size, the less confidence we have that the sample is representative of the population from which the sample was drawn. In using the chi-square statistic, two rules of thumb have been developed to address the effect of small sample sizes on expected frequencies: (1) None of the expected frequencies should be equal to zero, and (2) most if not all of the expected frequencies should be greater than 5. Failure to follow either or both of these two rules of thumb may limit the ability of the chi-square statistic to resemble or approximate the population from which the sample was drawn.
Different statistics have been developed to use with samples with small expected frequencies. For example, when any of the expected frequencies in a sample of data are less
940
than 5, it has been recommended that the Fisher's exact test be used rather than the chi- square statistic (Hays, 1988). On the other hand, if any of the expected frequencies are between 5 and 10, Yates's correction for continuity (Hays, 1988) may be used to test the differences between the observed and expected frequencies.
941
14.3 Inferential Statistic: Chi-Square Goodness-of-Fit Test
To test the differences between the observed and expected frequencies in a set of data, a chi-square statistic is calculated. When the data comprise a single variable, such as the Type variable in the personality study, the appropriate statistical procedure is known as the chi- square goodness-of-fit test. It has been given this name because it tests the “fit” between observed frequencies in a set of data and expected frequencies derived from theory or research.
942
Learning Check 1: Reviewing what you've Learned So Far
1. Review questions a. In creating a bar chart for a categorical variable, why would you choose to have each bar
represent a percentage (%) rather than a frequency (f)? b. Why is it inappropriate to calculate the mean and standard deviation for categorical
variables? c. What is the difference between an observed frequency and an expected frequency?
d. What is the main purpose of the chi-square statistic (χ2)? e. What are the main characteristics of the distribution of chi-square statistics? f. What are some assumptions underlying the chi-square statistic?
2. For each of the following situations, create a frequency distribution table and bar chart. a. A smartphone manufacturer wants to learn whether customers prefer a keyboard or a touch
screen:
Type of Phone Type of Phone Type of Phone Type of Phone
Touch screen Keyboard Touch screen Touch screen
Keyboard Touch screen Touch screen Keyboard
Touch screen Touch screen Keyboard Touch screen
Touch screen Touch screen Touch screen Touch screen
b. A café owner is interested in seeing the type of milk her customers add to their drinks: whole, low fat, or nonfat:
Type of Milk Type of Milk Type of Milk Type of Milk Type of Milk
Nonfat Low fat Whole Low fat Low fat
Low fat Low fat Low fat Whole Nonfat
Whole Low fat Whole Whole Low fat
Low fat Whole Nonfat Low fat Whole
The steps involved in calculating the chi-square statistic in the chi-square goodness-of-fit test are the same as in earlier chapters:
state the null and alternative hypotheses (H0 and H1), make a decision about the null hypothesis, draw a conclusion from the analysis, and relate the result of the analysis to the research hypothesis.
Each of these steps is presented below, using the personality study as an example.
943
State the Null and Alternative Hypotheses (H0 and H1)
What are the statistical hypotheses when the data being analyzed consist of frequencies rather than group means? Rather than stating differences between means, the null and alternative hypothesis for the chi-square goodness-of-fit test relate to the difference between the distributions of observed and expected frequencies. For example, the null hypothesis is stated as follows:
H0: The distribution of observed frequencies fits the distribution of expected frequencies.
For the personality study, this null hypothesis implies that the frequencies of the four types in the sample of 588 students are consistent with expected frequencies based on the developers of the CPI.
If the null hypothesis is that the distribution of observed frequencies fits the distribution of expected frequencies, what is the mutually exclusive alternative to this hypothesis? One way to state the alternative hypothesis is as follows:
H1: The distribution of observed frequencies does not fit the distribution of expected frequencies.
Rejecting the null hypothesis in the personality study implies that the distribution of frequencies of the four types in the sample of data does not fit the distribution of expected frequencies.
944
Make a Decision about the Null Hypothesis
The next step is to make the decision regarding whether to reject the null hypothesis. This consists of steps with which you are quite familiar:
calculate the degrees of freedom (df); set alpha (α), identify the critical value, and state a decision rule; calculate a statistic: chi-square (χ2); make a decision whether to reject the null hypothesis; determine the level of significance; and calculate a measure of effect size (Cramér's φ).
These steps are discussed below, using the personality study to demonstrate how each step may be completed.
Calculate the Degrees of Freedom (df)
The chi-square goodness-of-fit test involves differences between frequencies for the groups that comprise a variable. Similar to the one-way ANOVA discussed in Chapter 11, the degrees of freedom for this test is the number of groups that comprise the variable, minus 1:
(14-1) df = # groups − 1
In the personality study, there are four types (Alpha, Beta, Gamma, and Delta). Therefore, the degrees of freedom is equal to d f = # groups − 1 = 4 − 1 = 3
Set Alpha (α), Identify the Critical Value, and State a Decision Rule
Alpha (α), the probability of the statistic needed to reject the null hypothesis, is typically set to .05 (5%). Because the value for a chi-square statistic must always be a positive number, the theoretical distribution is positively skewed, with the region of rejection located only at the right end of the distribution. As such, it is not necessary to indicate whether alpha is one-tailed or two-tailed.
Table 8 in the back of this book provides a table of critical values for the chi-square statistic. This table consists of a series of rows corresponding to different degrees of freedom (df); the critical value for different values of alpha (α) are provided for each df. For the
945
personality study, the critical value is identified by moving down the df column until you reach the row associated with df = 3; moving to the right, for α = .05, a critical value of 7.81 is found. Therefore, the following critical value may be obtained for the personality study: For α = .05 and df = 3 , critical value = 7.81
Figure 14.3 illustrates the critical value and the regions of rejection and non-rejection for the personality study. The shape of the chi-square distribution resembles that of the F-ratio distribution used in ANOVA; this is because both statistics can have only positive values.
Figure 14.3 Critical Value and Regions of Rejection and Non-Rejection for the Personality Study
Once the critical value for the chi-square statistic has been identified, a decision rule may be stated regarding the values of the statistic that lead to the rejection of the null hypothesis. For the personality study, If χ 2 > 7.81, reject H 0 ; otherwise, do not reject H 0 .
Calculate a Statistic: Chi-Square (χ2) The next step in conducting the chi-square goodness-of-fit test is to calculate a value of the chi-square statistic. Two steps are involved in calculating this statistic: (1) calculate expected frequencies and (2) calculate the chi-square statistic. Although some of the calculations presented in this section will be the same as those presented earlier in introducing the chi- square statistic, they will be repeated here to formally introduce relevant formulas.
Calculate Expected Frequencies (fe)
The first step in calculating the chi-square statistic is to calculate the expected frequency (fe) for each of the groups comprising the variable of interest. The formula for an expected frequency in the chi-square goodness-of-fit test is provided in Formula 14–2:
(14-2) f e = hypothesized proportion N
946
where “hypothesized proportion” is the proportion of the population hypothesized to be in a particular group, and N is the total sample size.
In the personality study, the developers of the CPI believed that the four types are equally distributed in the population. Consequently, the hypothesized proportion for each group is .25 (25%). Applying these hypothesized proportions to the 588 students in the personality study (N = 588), the expected frequencies for the four types are calculated below:
Type f0 fe
Alpha 251 .25 (588) = 147.00
Beta 65 .25 (588) = 147.00
Gamma 194 .25 (588) = 147.00
Delta 78 .25 (588) = 147.00
Total 588 588.00
If the developers of the CPI are correct, in a sample of 588 students, each of the four types should contain 147.00 students. As a way of checking your calculations, note that the sum of the expected frequencies (Σfe) should be equal to the total sample size (N): Σ f e = N 147.00 + 147.00 + 147.00 + 147.00 = 588 588 = 588
Calculate the Chi-Square Statistic (χ2)
Once the expected frequencies have been determined, the chi-square statistic (χ2) that tests the differences between observed and expected frequencies can be calculated using the following formula:
(14-3) χ 2 = ∑ f 0 − f e 2 f e
where f0 is the observed frequency for each group and fe is the expected frequency for each group. As we discussed earlier in this chapter, the main component of the chi-square statistic is the squared difference between the observed and expected frequency for each group ((f0 – fe)
2). To calculate the chi-square statistic, the squared difference for each group is divided by each groups expected frequency (fe) before combining these calculations across the groups.
For the personality study, using the observed and expected frequencies calculated earlier, the chi-square statistic is calculated as follows: χ 2 = ∑ ( f 0 − f e ) 2 f e = ( 251 − 147.00 ) 2 147.00 + ( 65 − 147.00 ) 2 147.00 + ( 194 − 147.00 ) 2 147.00 + ( 78 − 147.00 ) 2 147.00 = ( 104.00 ) 2 147.00 + ( − 82.00 ) 2 147.00 + ( 47.00 ) 2 147.00 + ( − 69.00 ) 2 147.00 = 10816.00 147.00 + 6724.00 147.00 +
947
2209.00 147.00 + 4761.00 147.00 = 73.58 + 45.74 + 15.03 + 32.39 = 166.74
Looking at the above calculations, you may wonder why the squared difference between each observed and expected frequency ((f0 – fe)
2) is divided by its expected frequency (fe). The purpose of this division is to take into account the number of observations in each group when evaluating the difference between an observed and expected frequency. For example, a difference of 5 between f0 and fe should be given more weight if the expected frequency for a group is 10 than if it is 100. As a simple analogy, imagine you've decided to host a party. Your hosting would probably be more affected if you expected 5 people but 10 people showed up than if you expected 100 people but 105 showed up. In the personality study, dividing each group's difference by its expected frequency is not an issue because all of the expected frequencies are the same. However, later in this chapter, we will come across situations where this is not the case.
948
Learning Check 2: Reviewing what you've Learned So Far
1. Review questions a. In what type of research situation would the chi-square goodness-of-fit test be conducted? b. What is implied by the null and alternative hypotheses in the goodness-of-fit test? c. What are the two steps involved in calculating the chi-square statistic? d. In the goodness-of-fit test, what determines the values for the “hypothesized proportions”? e. In calculating the chi-square statistic, why is each squared difference divided by its expected
frequency? 2. For each of the following variables, calculate the degrees of freedom (df) and identify the critical
value (α = .05). a. Answer (Correct, Incorrect) b. Location (Top, Center, Bottom) c. Sport (Baseball, Football, Basketball, Soccer, Lacrosse)
3. One study compared two techniques designed to help children teach themselves how to do math problems (Grafman & Cates, 2010). The first method, Cover, Copy, and Compare (CCC), required students to read the problem, cover it with their hand, write their solution, and then compare their solution with one provided to them. The second method (MCCC) was identical to CCC except that students first copied down the problem before starting the CCC procedure. The researchers asked students which of the two methods they preferred: 34 chose CCC and 10 chose MCCC.
a. Calculate expected frequencies (fe) (assume the two hypothesized proportions are equal).
b. Calculate the chi-square statistic (χ2) for the goodness-of-fit test.
Make a Decision Whether to Reject the Null Hypothesis
Once the value of the chi-square statistic has been calculated, a decision is made regarding whether to reject the null hypothesis. This decision is made by comparing the chi-square statistic with the critical value. For the personality study, χ 2 = 166.74 > 7.81 ∴ reject H 0 p < .05
Because the obtained value of the chi-square statistic is greater than the α = .05 critical value of 7.81, it lies in the region of rejection. Consequently, the decision is made to reject the null hypothesis and conclude that the distribution of observed frequencies does not fit the distribution of expected frequencies.
Determine the Level of Significance
In the personality study, because the null hypothesis was rejected, the next step is to determine whether the probability of the chi-square statistic is less than .01 (1%). To do so, let's return to the table of critical values in Table 8. For df = 3, the .01 critical value is 11.34. Comparing the chi-square value of 166.74 with the .01 critical value,
949
χ 2 = 166.74 > 11.34 ∴ p < .01
Consequently, it can be concluded that the probability of the chi-square statistic in the personality study is not only less than .05 but also less than .01. The level of significance for the personality study is illustrated in Figure 14.4.
Figure 14.4 Determining the Level of Significance for the Personality Study
Calculate a Measure of Effect Size (Cramér's φ) As was discussed in earlier chapters, it is useful to supplement the results of an inferential statistic with a statistic that provides an estimate of the size or magnitude of the effect. There are several different measures of effect size for the chi-square statistic; in this book, we will use Cramér's φ (phi), sometimes referred to as Cramér's V. The formula for Cramér's φ for the chi-square goodness-of-fit test is provided as follows:
(14-4) Φ = χ 2 N # groups − 1
where χ2 is the calculated value of the chi-square statistic and N is the sample size.
For the personality study, N is equal to 588, and the number of groups is equal to 4. Using the calculated value of 166.74 for the chi-square statistic, Cramér's φ is calculated as follows: Φ = χ 2 N ( # groups − 1 ) = 166.74 588 ( 4 − 1 ) = 166.74 1764 = .09 = .30
As with other measures of effect size discussed in this book, the possible values for Cramér's φ range from .00 to 1.00. To assist in the interpretation of φ, Cohen (1988, pp. 224–226) defines small, medium, and large values of φ as the following: Small effect: ϕ = .10 Medium effect: ϕ = .30 Large effect: ϕ = .50
Accordingly, in the personality study, the value of φ of .30 indicates that the difference between the observed and expected frequencies for the four types represents a medium- sized effect.
950
Draw a Conclusion from the Analysis
For the personality study, the decision has been made to reject the null hypothesis. One way to communicate the result of this analysis is as follows:
The distribution of the four personality types in the sample of 588 students (Alpha [f = 251 (42.7%)], Beta [f = 65 (11.0%)], Gamma [f = 194 (33.0%)], Delta [f = 78 (13.3%)]) was significantly different from the CPI developers' expected distribution of f = 147 (25%) for each of the four types, χ2(3, N = 588) = 166.74, p < .01, φ = .30.
This sentence provides the following information about the analysis:
the variable being analyzed (“The distribution of the four personality types”), the sample (“588 students”), the groups that comprise the variable being analyzed (“Alpha … Beta … Gamma … Delta”), descriptive statistics (i.e., “f = 251 (42.7%)”), the nature and direction of the relationship (“was significantly different from the CPI developers' expected distribution”), and information about the inferential statistic (“χ2(3, N = 588) = 166.74, p < .01, φ = .30”), which indicates the inferential statistic calculated (χ2), the degrees of freedom (3), the value of the statistic (166.74), the level of significance (p < .01), and the measure of effect size (φ = .30).
In reporting the test of the chi-square statistic, the sample size (N = 588) is typically included. This is because, unlike other inferential statistics, the degrees of freedom for the chi-square statistic is based on the number of groups rather than the number of participants.
951
Relate the Result of the Analysis to the Research Hypothesis
What is the implication of rejecting the null hypothesis in the personality study? In terms of the purpose of this study, the following could be stated:
The result of this analysis supports the research hypothesis that the distribution of personality types among college students does not match or fit the distribution in the population proposed by the developers of the CPI.
Table 14.2 lists the steps involved in conducting the chi-square goodness-of-fit test. In the personality study, the hypothesized proportions of the four types were presumed to be the same (.25). The next section discusses how to conduct the goodness-of-fit test with unequal hypothesized proportions.
952
Learning Check 3: Reviewing what you've Learned So Far
1. Review questions a. What conclusion can be drawn when the null hypothesis for the goodness-of-fit test is
rejected? 2. The company that makes M&Ms candy (www.mms.com) once conducted a survey asking people
which new color they would prefer: purple, aqua, or pink. Below are the votes of a hypothetical sample that represents their actual findings:
Color Color Color Color Color Color
Pink Pink Aqua Purple Purple Aqua
Purple Purple Pink Pink Pink Purple
Pink Purple Purple Purple Aqua Aqua
Aqua Pink Pink Purple Pink Purple
Pink Purple Purple Aqua Pink Purple
Purple Pink Aqua Purple Purple
Conduct the chi-square goodness-of-fit test to determine whether there were any differences between the votes for three colors (assume equal hypothesized proportions).
a. Construct a frequency distribution table and bar chart for these data. b. State the null and alternative hypotheses (H0 and H1). c. Make a decision about the null hypothesis.
1. Calculate the degrees of freedom (df). 2. For α = .05, identify the critical value and state a decision rule. 3. Calculate a value of the chi-square statistic (χ2). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (Cramér's φ).
d. Draw a conclusion from the analysis.
Table 14.2 Summary, Conducting the Chi-Square Goodness-of-Fit Test (Personality Study Example)
953
954
Conducting the Chi-Square Goodness-of-Fit Test with Unequal Hypothesized Proportions
One purpose of the personality study was to compare the observed frequencies in the sample of 588 college students with expected frequencies derived from the developers of the CPI. In addition to comparing their sample of college students with the general population of adults, the researchers wished to see whether the distribution of the four personality types among college students may have changed over time. Consequently they compared their sample with the distribution of personality types found in a sample of 7,361 college students who completed the CPI during the 1980s.
In the 1980s sample, the four lifestyles were not equally distributed but instead were distributed the following way:
Type %
Alpha 35.2%
Beta 22.3%
Gamma 24.0%
Delta 18.5%
Total 100.0%
Looking at these percentages, it appears that Alpha (extraverted and norm favoring) was the most common personality type (35.2%), with similar percentages across the other three types.
The goal of this analysis was to test the difference between the observed frequencies of the sample of 588 college students with expected frequencies based on the proportions found in the 1980s sample. As it turns out, the chi-square goodness-of-fit test allows for the situation where the hypothesized proportions are not all the same. The purpose of this example is to demonstrate how to conduct the chi-square goodness-of-fit test with unequal hypothesized proportions and to highlight similarities and differences between this analysis and one with equal hypothesized proportions.
Many of the steps in conducting the chi-square goodness-of-fit test have the same result regardless of whether the hypothesized proportions are equal or unequal. Using the personality study as the example, the null and alternative hypotheses, degrees of freedom, alpha, critical value, and the decision rule are identical in both situations. Consequently, we will not report these steps below. What may change, however, is the calculated value of the chi-square statistic. Furthermore, if the calculated value of the statistic changes, what may also change is the decision that is made about the null hypothesis, the conclusions drawn
955
from the analysis, and whether the analysis supports the research hypothesis.
To compare the distribution of the four personality types for the sample of 588 students and the 1980s sample, a value of the chi-square statistic must be calculated. As before, there are two steps in calculating the chi-square statistic for the goodness-of-fit test with unequal expected frequencies: calculate expected frequencies (fe) and calculate the chi-square statistic (χ2).
Calculate Expected Frequencies (fe)
The expected frequency (fe) for each group is calculated by multiplying the group's hypothesized proportion by the total sample (N). Using the proportions for the 1980s sample reported above, the expected frequencies for the personality study data are calculated as follows:
Type f0 fe
Alpha 251 .352 (588) = 206.98
Beta 65 .223 (588) = 131.12
Gamma 194 .240 (588) = 141.12
Delta 78 .185 (588) = 108.78
Total 588 588.00
Note that these expected frequencies differ from those calculated when the hypothesized proportions were the same (.25) for all four types. However, the sum of the expected frequencies is again equal to the total sample size (206.98 + 131.12 + 141.12 + 108.78 = 588).
Calculate the Chi-Square Statistic (χ2) Once the expected frequencies have been determined, the next step is to calculate the chi- square statistic that tests the differences between the observed and expected frequencies. Regardless of whether or not the hypothesized proportions are equal, the chi-square statistic is calculated using the same formula (Formula 14-3): χ 2 = ∑ ( f 0 − f e ) 2 f e = ( 251 − 206.98 ) 2 206.98 + ( 65 − 131.12 ) 2 131.12 + ( 194 − 141.12 ) 2 141.12 + ( 78 − 108.78 ) 2 108.78 = ( 44.02 ) 2 206.98 + ( − 66.12 ) 2 131.12 + ( 52.88 ) 2 141.12 + ( − 30.78 ) 2 108.78 = 1937.76 206.98 + 4371.85 131.12 + 2796.29 141.12 + 947.41 108.78 = 9.36 + 33.34 + 19.82 + 8.71 = 71.23
In calculating the chi-square statistic, care must be taken to use the appropriate expected frequency (fe) for each group.
956
Learning Check 4: Reviewing what you've Learned So Far
1. Review questions a. What steps in conducting the chi-square goodness-of-fit test change when the hypothesized
proportions are unequal rather than equal? What steps do not change? 2. In Learning Check 3, preferences for three M&M colors (purple, aqua, and pink) were compared.
As it turns out, the final worldwide vote conducted by www.mms.com found that 42% chose purple, 38% chose aqua, and 20% chose pink. Conduct the chi square goodness-of-fit test comparing this sample's proportions with those reported by the M&Ms company.
a. State the null and alternative hypotheses (H0 and H1). b. Make a decision about the null hypothesis.
1. Calculate the degrees of freedom (df). 2. For α = .05, identify the critical value and state a decision rule. 3. Calculate a value of the chi-square statistic (χ2). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (Cramér's φ).
c. Draw a conclusion from the analysis.
Comparing the result of this analysis with the earlier one, the calculated value of χ2 changed from 166.74 to 71.23. Changing the expected frequencies may alter the decisions and conclusions resulting from the remaining steps in hypothesis testing (making the decision whether to reject the null hypothesis, drawing a conclusion from the analysis, etc.). For this reason, it is important that the hypothesized proportions are based on sound theoretical or empirical reasoning.
This section has demonstrated how to analyze a single categorical variable using the chi- square goodness-of-fit test. The next section introduces a somewhat more complicated research situation, one that consists of two categorical variables rather than one.
957
14.4 An Example from the Research (Two Categorical Variables): Seeing Red
You may have a favorite color you like to wear—perhaps it's blue or green or yellow. As such, you may wonder just how and when you developed this preference. Research regarding people's color preferences has found that young children apparently prefer the color red (Maier, Barchfeld, Elliot, & Peckrun, 2009). For example, 3–month-old infants have been found to spend more time looking at a red stimulus than at one that was yellow, blue, or green (R. J. Adams, 1987), and children in a daycare center preferred classrooms with red walls to walls that were purple, blue, green, yellow, orange, or gray (Read & Upington, 2009). To try to explain these findings, one study hypothesized that children prefer the color red because it seems to be associated with positive emotional states (Zentner, 2001). In that study, a sample of 3- and 4-year-olds not only exhibited a preference for the color red but also were more likely to assign it a happy face than a face that was sad or angry.
Believing that children may or may not like a color because it's associated with a particular experience or feeling, a team of researchers led by Markus Maier at Stony Brook University in New York wished to see whether these preferences change if these associations change (Maier et al., 2009). More specifically, they hypothesized that infants' preferences for the color red may be altered by associating the color with a face that is angry versus one that is happy:
Specifically, we predicted that in the presence of a happy (hospitable) face, red would be preferred because it signals a potentially desirable outcome … but that in the presence of an angry (hostile) face, red would not be preferred because it signals a potentially undesirable outcome. (Maier et al., 2009, p. 735)
To test their hypothesis, the researchers conducted an experiment with “a total of 40 infants (25 girls)…. The mean age of participants was 17.75 months (SD = 2.49), with a range of 15 to 26 months. All infants were members of a day care center” (Maier et al., 2009, pp. 736–737). In the study, infants were brought into a room and seated on the lap of a daycare worker in front of a table. On the table, covered by a white cloth, were three toys that were identical except for their color: red, green, or gray. Before the infant was shown the toys, he or she was shown one of two pictures of the face of a man. For half of the infants, the man had a happy expression; for the other half, the man was angry. Once the experimenters had determined that the infant had looked at the picture, they removed the cloth and showed the infant the toys. When the infant picked up one of the toys, the color of the toy was recorded.
958
The study, which will be referred to as the color preference study, features two variables of interest. The first variable, Face, is the expression of the face shown to the infant (Happy or Angry). The second variable, Toy color, is the color of the toy selected by the infant (Red, Green, or Gray). Both of these variables are categorical, consisting of distinct groups. The colors of the toys selected by the 20 infants in each of the Happy and Angry groups are listed in Table 14.3.
Table 14.3 The Color of the Selected Toy for Infants in the Happy and Angry Groups
Table 14.3 The Color of the Selected Toy for Infants in the Happy and Angry
Groups
Happy
Infant Toy Color
1 Red
2 Red
3 Green
4 Red
5 Gray
6 Red
7 Red
8 Red
9 Gray
10 Red
11 Red
12 Green
13 Red
14 Red
15 Red
16 Red
17 Red
18 Gray
19 Red
20 Red Table 14.3 The Color
959
Table 14.3 The Color of the Selected Toy for Infants in the Happy and Angry
Groups
Angry
Infant Toy Color
1 Green
2 Gray
3 Green
4 Red
5 Gray
6 Gray
7 Red
8 Green
9 Red
10 Green
11 Gray
12 Red
13 Green
14 Green
15 Red
16 Gray
17 Green
18 Gray
19 Gray
20 Red
Table 14.4 Tables of Frequencies and Percentages for the Color Preference Study
960
In order to organize the data, two tables have been constructed in Table 14.4. Table 14.4(a) is a contingency table, which is a table in which the rows and columns represent the values of categorical variables and the cells of the table contain observed frequencies for combinations of the variables. It is called a contingency table because the goal of the analysis is to determine whether the distribution of frequencies for one variable is contingent on (depends on) another variable. For the color preference study, we are examining whether the distribution of the colors of the toy is contingent on whether the child is shown a happy or angry face.
A contingency table may be described in terms of the number of rows and columns it contains. For example, the contingency table for the color preference study may be referred to as a “3 × 2” contingency table because it consists of three rows (Toy color: Red, Green, Gray) and two columns (Face: Happy, Angry). Each cell within the body of a contingency table contains the observed frequency for a particular combination of variables. For example, the number 15 in the upper-left cell of Table 14.4(a) implies that 15 infants exposed to the happy face selected the red toy.
Figure 14.5 Bar Chart of the Percentage of Toy Color (Red, Green, Gray) for Each Type of Face (Happy, Angry)
961
The main purpose of creating a contingency table is to prepare for the calculation of a chi- square statistic. However, in reporting the data for a study, it is useful to calculate and report percentages related to each observed frequency within the contingency table. Table 14.4(b) provides the percentage of infants selecting the three toy colors for each of the two faces. For example, the percentage 75.0% in the upper-left cell of this table represents the percentage of the 20 infants shown the happy face who selected the red toy (15/20 = 75.0%). It should be noted that, rather than having the two columns of percentages add up to 100.0%, percentages across the three rows instead could have been calculated to represent the two faces within each of the three toy colors. For example, for the 21 infants who selected the red toy, we could have calculated the percentage of infants who were shown the happy face (15/21 or 71.4%) versus the angry face (6/21 or 29.6%). However, column percentages were calculated because the research hypothesis of the color preference study addressed differences between color preferences rather than differences between faces.
To illustrate the percentages in Table 14.4(b), a bar chart is provided in Figure 14.5. The bars in this bar chart represent the three different toy colors for the two types of faces. The toy colors have been grouped together to illustrate differences in the distribution of color preferences for the two faces. Looking at the bar chart, we see that a wide majority (75%) of the infants shown the happy face chose the red toy. However, for the infants shown the angry face, there were very little differences between the three colors. This appears to provide initial support for the research hypothesis that red would be preferred when associated with a happy face but not preferred when associated with an angry face. However, because we need to calculate an inferential statistic to determine whether these differences are statistically significant, the next section introduces a chi-square statistic used to analyze two categorical variables.
962
Learning Check 5: Reviewing what you've Learned So Far
1. Review questions a. What is the purpose of a contingency table?
2. For each of the following situations, create a contingency table and a bar chart. a. A teacher wants to see whether preferences for chocolate versus vanilla ice cream are
different for boys versus girls.
Boy Girl
Ice Cream Ice Cream Ice Cream Ice Cream
Chocolate Chocolate Chocolate Vanilla
Chocolate Vanilla Vanilla Chocolate
Chocolate Chocolate Chocolate Chocolate
Vanilla Chocolate Chocolate Vanilla
Chocolate Chocolate Vanilla Chocolate
b. A college instructor examines whether freshmen, sophomores, juniors, and seniors differ in whether they tend to sit in the front or the back of the classroom:
Freshman Sophomore Junior Senior
Front Front Back Back
Front Back Back Back
Back Front Front Front
Front Front Front Back
Front Front Back Back
Back Back Front Back
Front Back Back Back
c. A political pollster examines whether opinions about a proposed law (in favor, against, unsure) are different for members of three political parties (Parties A, B, and C).
Party A Party B Party C
Opinion Opinion Opinion Opinion Opinion Opinion
Against Against In favor Against In favor In favor
In favor Unsure Against Unsure In favor In favor
Against Against Unsure In favor Unsure In favor
Against Unsure In favor Unsure In favor Against
Unsure Against Unsure Against In favor
Against Against Unsure In favor
963
14.5 Inferential Statistic: Chi-Square Test of Independence
To test differences between observed and expected frequencies in a set of data, a chi-square statistic is calculated. When the data comprise a single variable, such as the Type variable in the personality study, the chi-square goodness-of-fit test is conducted. However, when the data consist of two categorical variables, the appropriate statistical procedure is the chi- square test of independence. As its name suggests, the chi-square test of independence allows us to determine whether the distribution of frequencies for one variable is independent of another variable. Variables are not independent when the distribution of frequencies for one variable changes as a function of a second variable, which implies the two variables are related. For the color preference study, the question is whether the distribution of toy colors is the same for infants shown the happy face versus the angry face. If the two distributions are not the same, it can be concluded that Toy color and Face are related to each other, such that the distribution of frequencies for toy color changes as a function of whether the child is shown a happy or angry face.
The steps involved in conducting the chi-square test of independence are the same as the other inferential statistical procedures discussed in this book:
state the null and alternative hypotheses (H0 and H1), make a decision about the null hypothesis, draw a conclusion from the analysis, and relate the result of the analysis to the research hypothesis.
As you will see, adding a second categorical variable makes completing the steps for the chi- square test of independence slightly more complicated than the goodness-of-fit test.
964
State the Null and Alternative Hypotheses (H0 and H1)
The purpose of the chi-square test of independence is to determine whether two categorical variables are independent of each other, which is the same as determining whether the two variables are related. Therefore, the null and alternative hypotheses may be stated as the following:
H0: The two variables are independent (i.e., there is no relationship between the two variables).
H1: The two variables are not independent (i.e., there is a relationship between the two variables).
For the color preference example, the null hypothesis implies that an infant's color preference is independent of whether he or she is shown a happy or angry face. On the other hand, the alternative hypothesis implies that that an infant's color preference depends on the type of face he or she is shown.
965
Make a Decision about the Null Hypothesis
Once the null and alternative hypotheses have been stated, you move on to making a decision whether to reject the null hypothesis and conclude that the two variables are related. The steps used to make this decision for the chi-square test of independence are the same as for the goodness-of-fit test:
calculate the degrees of freedom (df); set alpha (α), identify the critical value, and state a decision rule; calculate a statistic: chi-square (χ2); make a decision whether to reject the null hypothesis; determine the level of significance; and calculate a measure of effect size (Cramér's φ).
Calculate the Degrees of Freedom (df)
Calculating the degrees of freedom (df) for two categorical variables is very similar to the situation presented in conducting the two-way ANOVA (Chapter 12). In that chapter, the degrees of freedom for the A × B interaction effect (dfA × B) was equal to (a – 1)(b – 1), where a and b were the number of groups that comprised two independent variables. Using similar logic, the degrees of freedom for the chi-square test of independence may be stated in the following way:
(14-5) df = # rows − 1 # columns − 1
where “# rows” and “# columns” relate to the contingency table created for the two variables. More specifically, the number of rows and columns are equal to the number of groups that comprise the two variables.
For the color preference study, the contingency table in Table 14.4(a) had three rows (Toy color: Red, Green, Gray) and two columns (Face: Happy, Angry). Using Formula 14-5, the degrees of freedom is equal to d f = ( # r o w s − 1 ) ( # c o l u m n s − 1 ) = ( 3 − 1 ) ( 2 − 1 ) = ( 2 ) ( 1 ) = 2
We may state, therefore, that there are two degrees of freedom in the color preference example.
Set Alpha (α), Identify the Critical Value, and State a Decision Rule
966
Alpha, the probability of the statistic needed to reject the null hypothesis, is again set at the traditional value of .05 or 5% (α = .05). Using the table of critical values for the chi-square statistic in Table 8, the critical value for the color preference example is identified by moving down the df column until we reach the df = 2 row. Moving to the right, for α = .05, we find a critical value of 5.99. Therefore, For α = .05 and df = 2 , critical value = 5.99
The decision rule explicitly describes the values of the statistic leading to the rejection of the null hypothesis. For the color preference example, If χ 2 > 5.99, reject H 0 ; otherwise, do not reject H 0
Calculate a Statistic: Chi-Square (χ2) Calculating the chi-square statistic for the chi-square test of independence involves the same two steps as for the goodness-of-fit test: calculating expected frequencies (fe) and calculating the chi-square statistic (χ2). However, the addition of a second variable changes how the expected frequencies are calculated.
Calculate Expected Frequencies (fe)
When there is one categorical variable, the expected frequency for a group is based on a hypothesized proportion based on theory or research. However, when there are two categorical variables, each expected frequency is based on the number of participants in the sample. More specifically, because each observed frequency represents a combination of the two variables in a sample of data, the formula for an expected frequency (fe) in the chi- square test of independence is based on these combinations:
(14-6) f e = row total column total N
where “row total” and “column total” refer to the different rows and columns in a contingency table, and N is equal to the total sample size. Looking at Table 14.4(a), for the color preference example, the “row total” refers to the number of infants selecting a particular toy color (regardless of type of face), the “column total” refers to the number of infants exposed to a particular face (regardless of toy color), and N refers to the total number of infants in the sample (N = 40).
To illustrate how expected frequencies are calculated for the chi-square test of independence, let's compute the expected frequency for the Happy/Red combination. Looking at the contingency table in Table 14.4(a), this combination is located in the top row, which has a total of 21 infants, and in the left column, which has a total of 20 infants. Using these two totals as well as the total sample size (N = 40), how many infants would be
967
expected to be in the Happy/Red combination? The expected frequency for the Happy/Red combination is f e = ( row total ) ( column total ) N = ( 21 ) ( 20 ) 40 = 420 40 = 10.50
Therefore, given that 21 of the 40 infants chose the red toy, and 20 of the 40 infants were exposed to the happy face, 10.50 of the 40 infants in this sample are expected to be in the Happy/Red combination.
Table 14.5 provides the expected frequencies for the six combinations in the color preference study. It should be noted that the calculation of the expected frequencies for the color preference study is made easier by the fact that the two column totals (20 and 20) are the same. The exercises at the end of this chapter contain examples where this is not the case in order to provide practice in using the correct row and column totals in your calculations.
To determine that the expected frequencies have been calculated correctly, the sum of the expected frequencies for each row or column should match the corresponding row or column total. For example, for the Red toy, the sum of the expected frequencies for the Happy and Angry faces is equal to the number of infants who chose the red toy (10.50 + 10.50 = 21). Another check of calculations is to ensure that the sum of the expected frequencies equals the total sample size (10.50 + 4.50 + 5.00 + 10.50 + 4.50 + 5.00 = 40).
Calculate the Chi-Square Statistic (χ2)
Once the expected frequencies have been determined, the chi-square statistic is calculated using the same formula (Formula 14–3) as for the goodness-of-fit test. Using the observed and expected frequencies in Table 14.5, the chi-square statistic for the color preference study is calculated as follows: χ 2 = ∑ ( f 0 − f e ) 2 f e = ( 15 − 10.50 ) 2 10.50 + ( 2 − 4.50 ) 2 4.50 + ( 3 − 5.00 ) 2 5.00 + ( 6 − 10.50 ) 2 10.50 + ( 7 − 4.50 ) 2 4.50 + ( 7 − 5.00 ) 2 5.00 = ( 4.50 ) 2 10.50 + ( − 2.50 ) 2 4.50 + ( − 2.00 ) 2 5.00 + ( − 4.50 ) 2 10.50 + ( 2.50 ) 2 4.50 + ( 2.00 ) 2 5.00 = 20.25 10.50 + 6.25 4.50 + 4.00 5.00 + 20.25 10.50 + 6.25 4.50 + 4.00 5.00 = 1.93 + 1.39 + .80 + 1.93 + .80 = 8.24
Table 14.5 Table of Observed Frequencies (fo) and Expected Frequencies (fe) for the Color Preference Study
968
In calculating the chi-square statistic, it is important to be sure to divide each squared difference between an observed and expected frequency ((f0 – fe)
2) by the correct expected frequency (fe).
Make a Decision Whether to Reject the Null Hypothesis
For the color preference study do you reject the null hypothesis that the two variables are independent of each other? Comparing the value of the chi-square statistic calculated from the sample with the identified critical value, χ 2 = 8.24 > 5.99 ∴ reject H 0 p < . 05
In this example, the decision to reject the null hypothesis implies that Toy color and Face are related to each other such that an infant's preference for the red, green, or gray toy depends on the type of face (happy or angry) the infant is shown.
Determine the Level of Significance
Because in this example the null hypothesis has been rejected, it is appropriate to determine whether the probability of the chi-square statistic is less than .01 (1%). Turning to Table 8, for df = 2 and α = .01, the critical value for the chi-square statistic is 9.21. Comparing the color preference study's chi-square value with this critical value, the following conclusion is drawn: χ 2 = 8.24 < 9.21 × p < . 05 but not < .01
The level of significance for the color preference study is illustrated in Figure 14.6. Because the value of the chi-square statistic falls between the .05 and .01 critical values, its probability is less than .05 but not less than .01. Consequently, the level of significance for this chi-square statistic would be reported as “p < .05.”
Calculate a Measure of Effect Size (Cramér's φ)
969
The Cramér's φ measure of effect size for the chi-square statistic was calculated earlier for the goodness-of-fit test. However, the formula for Cramér's φ for the chi-square test of independence is slightly different due to the existence of two variables rather than one variable and is presented as follows:
(14-7) Φ = χ 2 N k − 1
where χ2 is the calculated value of the chi-square statistic, N is the total sample size, and k is the smaller of the number of rows or number of columns in the contingency table. For the color preference example, the values for N and k are obtained by looking at the contingency table in Table 14.4(a). In this table, you find that N is equal to 40. Also, because the contingency table consists of three rows but only two columns, k is equal to 2. Using the value for the chi-square statistic calculated from this data, Cramér's φ is calculated as follows: Φ = χ 2 N ( k − 1 ) = 8.24 40 ( 2 − 1 ) = .46 = 8.24 40 = .21
Using Cohen's guidelines presented earlier in this chapter, the value of φ = .46 for the color preference study indicates that the relationship between Toy color and Face represents a large effect.
970
Draw a Conclusion from the Analysis
The results of a chi-square test of independence such as the one in the color preference study may be reported in the following way:
The frequencies representing the color preferences (red, green, or gray) of 40 infants shown either a happy or angry face were analyzed using the chi-square test of independence. This analysis found a significant relationship between Toy color and Face such that the distribution of the three color preferences is different for infants shown the happy face versus the angry face, χ2(2, N = 40) = 8.24, p < .05, φ = .46.
Figure 14.6 Determining the Level of Significance for the Color Preference Study
The above sentence provides the following information:
the variable being analyzed (“The frequencies representing the color preferences”), the sample (“40 infants”), the variables and groups involved in the analysis (“Toy color … red, green, or gray” and “Face … happy or angry face”), the statistical procedure used to analyze the variable (“chi-square test of independence”), the nature and direction of the relationship (“a significant relationship … such that the distribution of the three color preferences is different for infants shown the happy face versus the angry face”), and information about the inferential statistic (“χ2(2, N = 40) = 8.24, p < .05, φ =.46”), which indicates the inferential statistic calculated (χ2), the degrees of freedom (2), the sample size (N = 40), the value of the statistic (8.24), the level of significance (p < .05), and the measure of effect size (φ = .46).
971
Relate the Result of the Analysis to the Research Hypothesis
The results of the chi-square test of independence in the color preference study found that infants' color preferences depended on which type of face they were shown. Looking at the bar chart in Figure 14.5, although the red toy was preferred by the infants shown the happy face, there were very little differences between the three colors for the infants shown the angry face. As a result, the researchers concluded that their research hypothesis had been supported, summarizing their findings in the following way:
Our findings indicate that infants' preference for red changes with the context in which it is presented. Specifically, in a hospitable context, red is preferred, whereas in a hostile context, red is not preferred. (Maier et al., 2009, p. 737)
972
Learning Check 6: Reviewing what you've Learned So Far
1. Review questions a. What are the main differences between the chi-square goodness-of-fit test and the chi-
square test of independence? b. What is implied by the null and alternative hypotheses in the chi-square test of
independence? c. What determines the expected frequencies in the chi-square test of independence?
2. Each of the following situations consists of two variables; for each situation, calculate the degrees of
freedom (df) and determine the critical value of χ2 (assume α = .05). a. Price (Regular, Sale) and Buying decision (Buy Not buy) b. Size (Small, Medium, Large, Extra large) and Perceived value (Low, High) c. Location (Orchestra, Mezzanine, Balcony) and Rating (Excellent, Good, Fair, Poor)
3. Another study looking at children's color preferences examined whether color preferences may be related to one's sex (Walsh, Toma, Tuveson, & Sondhi, 1990). In this study, a group of 5- and 9- year-old children (60 boys and 60 girls) chose their favorite color of a Skittles candy (green, red, orange, or yellow). The following is a contingency table showing the distribution of preferred colors for the boys and girls in this sample:
Conduct the chi-square test of independence to address whether there is a difference in the distribution of color preferences for boys versus girls.
a. Construct a bar chart of these data. b. State the null and alternative hypotheses (H0 and H1). c. Make a decision about the null hypothesis.
1. Calculate the degrees of freedom (df). 2. Set alpha, identify the critical value, and state a decision rule.
3. Calculate a value of the chi-square statistic (χ2). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (Cramér's φ).
d. Draw a conclusion from the analysis.
Table 14.6 summarizes the steps involved in analyzing two categorical variables using the chi-square test of independence, using the color preference study as an example. The next section ends this chapter (and this book) by placing the chi-square statistic within a larger body of statistical procedures known as nonparametric statistics and illustrating the relationship between nonparametric statistics and statistics presented earlier in this book.
973
14.6 Parametric and Nonparametric Statistical Tests
Throughout this book, you've repeatedly seen that variables in a research study may be measured in different ways and at different levels of measurement. How a variable has been measured influences how it should be statistically analyzed. As discussed in the current chapter, when a study's variables are categorical in nature (measured at the nominal level of measurement), the chi-square statistic is used to analyze frequencies of groups that comprise the study's variables. On the other hand, Chapters 9, 11, and 12 discussed different versions of the ANOVA, which is used when the independent variable or variables are categorical but the dependent variable is continuous, measured at the interval or ratio level of measurement. The goal of ANOVA is to test differences between the means of groups. Furthermore, the Pearson correlation coefficient (Chapter 13), used when both variables are continuous, is designed to measure the linear relationship between variables. Based on their similarities and differences, these statistical procedures may be grouped into two broad categories known as parametric and nonparametric statistical tests. The following sections compare these two categories of tests and discuss situations in which nonparametric statistical tests might be used.
974
Parametric vs. Nonparametric Statistical Tests
Statistical procedures such as the t-test, ANOVA, and Pearson correlation are used to analyze continuous dependent variables measured at the interval or ratio level of measurement. These statistical procedures are examples of parametric statistical tests, which are statistical tests based on several assumptions about the populations from which samples are taken.
Table 14.6 Summary, Conducting the Chi-Square Test of Independence (Color Preference Study Example)
975
The first assumption of parametric statistical tests is that data from a sample are being used to estimate a population parameter. For example, the null hypothesis in the t-test, one-way AN OVA, and two-way ANOVA includes the population parameter μ, which is the mean of a variable in the population. The second assumption of parametric statistical tests pertains to the shape of the distribution of variables that are the basis of parameters. More
976
specifically, parametric statistics assume that the distribution of scores for a variable is normally distributed in the larger population. In the case of ANOVA, this assumption is taken one step further in that it is assumed that the dependent variable is normally distributed for each group involved in the analysis. The third assumption of parametric tests relates to the shape of the distribution of data in samples drawn from populations. Because parametric tests assume that variables are normally distributed in the population, it is also assumed that data collected on these variables from samples of the population are normally distributed.
In contrast to parametric tests, nonparametric statistical tests are statistical tests not based on the assumptions that underlie parametric statistical tests. One important feature of nonparametric statistical tests is that they do not involve the estimation of population parameters. For example, the null hypothesis for the chi-square goodness-of-fit test (H0: the distribution of observed frequencies fits the distribution of expected frequencies) pertains to a distribution of frequencies rather than a specific value of a parameter such as μ. A second key feature of nonparametric statistical tests is that they do not make assumptions regarding the shape of the distribution of variables in populations or samples. More specifically, nonparametric tests neither assume nor require variables to be normally distributed. In the chi-square goodness-of-fit test, for example, the distribution of expected frequencies can take on any shape because it is a function of hypothesized proportions derived from theory or research. Furthermore, in the chi-square test of independence, the expected frequencies are based on the row and column frequencies for a sample rather than on assumptions regarding how frequencies for the two variables are distributed in the population. For these reasons, the chi-square statistic is an example of a nonparametric statistical test. Because nonparametric tests do not make assumptions regarding the shape of the distribution, they are sometimes referred to as “distribution-free” tests.
977
Reasons for Using Nonparametric Statistical Tests
There are two primary reasons why nonparametric statistical tests may be used in a research study—these reasons relate to how variables have been measured as well as concerns regarding the shape of the distribution of scores in the sample. First, nonparametric statistical tests are used in research situations in which the dependent variable is measured at a level of measurement other than interval or ratio. We have introduced two nonparametric statistical tests in this book: the chi-square statistic discussed in this chapter and the Spearman rank-order correlation covered in Chapter 13. The chi-square statistic is used for categorical variables, which are variables measured at the nominal level of measurement. The Spearman correlation is used to measure the relationship between two variables measured at the ordinal level of measurement.
A second reason for using nonparametric statistical tests is when data have been collected on variables measured at the interval or ratio level of measurement, but there is concern that the sample data are not normally distributed. Parametric statistical tests such as the t- test, F-ratio, and Pearson r assume scores are normally distributed. If this assumption is not met, incorrect conclusions may be drawn from statistical analyses conducted on the data. For example, if one or both of the variables used to calculate a Pearson correlation are severely asymmetrical (skewed) rather than symmetrical, the relationship between the two variables may not be linear. As a result, the Pearson correlation does not provide an accurate indication of the relationship between the variables in this sample. In a situation such as this, a nonparametric statistical test rather than a parametric test may be used to analyze the data.
978
Examples of Nonparametric Statistical Tests
Several parametric statistical tests have a nonparametric alternative. Although these nonparametric tests are designed for use with variables measured at the ordinal or nominal levels of measurement, they can also analyze variables measured at the interval or ratio level of measurement. Researchers may choose to use a nonparametric test if they are concerned about the shape of the distribution of data in their samples, particularly when their sample is small or the sample sizes of groups differ greatly from one another. In this situation, nonparametric tests are used to lower the possibility of drawing inappropriate conclusions from statistical analyses.
Table 14.7 Parametric and Corresponding Nonparametric Statistical Tests Table 14.7 Parametric and Corresponding Nonparametric Statistical Tests
Parametric Test Corresponding Nonparametric Test
t-test for two independent means (Chapter 9)
Mann-Whitney test
t-test for dependent means (Chapter 9) Wilcoxon matched-pairs signed-rank test
One-way ANOVA (Chapter 11) Kruskal-Wallis test
Pearson correlation (Chapter 13) Spearman rank-order correlation
The left-hand column of Table 14.7 lists several parametric tests covered in this book; the right-hand column indicates the nonparametric alternative. Information regarding these nonparametric tests may be found in advanced statistical textbooks; these tests are mentioned at the conclusion of this book to give you a sense of the variety of statistical procedures available to researchers based on their particular needs.
Given the assumptions that must be met to use parametric statistical tests, you may wonder why researchers don't always use nonparametric tests, regardless of how the variables in a research study have been measured. As it turns out, when a continuous variable is in fact normally distributed, the nonparametric test has less statistical power than its parametric counterpart. Stated differently, this means that the nonparametric test is more likely to make a Type II error, which in Chapter 10 was defined as concluding an effect does not exist when it in fact does. Therefore, using nonparametric tests may be an overly conservative strategy that makes it difficult to detect effects.
979
14.7 Looking Ahead
This chapter introduced the analysis of research situations in which all of the variables in the research studies were categorical in nature. Although analyzing categorical variables differs in some ways from the analyses discussed in previous chapters, it is important to remember that the basic process of analyzing data does not change. Regardless of how the variables are measured or how the research question or hypothesis is stated, the same steps are used to conduct and interpret the results of statistical analyses. We have ended this chapter, and this book as a whole, by noting a few of the many other statistical procedures used by researchers to answer the various questions they have and the challenges they face. Looking ahead, we hope this book has given you a sense of how statistics can help answer questions you may have. More important, we wish you the best of luck in identifying and exploring these questions.
980
14.8 Summary
The chi-square statistic (χ2) tests the difference between observed frequencies, which is the frequency of a group in the sample of data, and expected frequencies, which is the frequency of a group expected to occur in a sample of data.
When the data comprise a single categorical variable, the appropriate statistical procedure is the chi-square goodness-of-fit test, which tests the “fit” between observed frequencies in a set of data and expected frequencies derived from theory or research.
There are two important assumptions of the chi-square statistic: independence of observations and minimum expected frequencies. Independence of observations implies that each observation is independent of all other observations in the set of data. Minimum expected frequencies means that none of the expected frequencies should be equal to zero, and most if not all of the expected frequencies should be greater than 5. When one or more of the expected frequencies are smaller than 5, Fisher's exact test may be used to analyze the data. If any of the expected frequencies falls between 5 and 10, Yates's correction for continuity may be used.
When there are two categorical variables, a contingency table is created, which is a table whose rows and columns represent the values of categorical variables and whose cells contain observed frequencies for combinations of the variables. In this situation, the appropriate statistical test is the chi-square test of independence, which tests whether the distribution of frequencies for one variable is either independent of or dependent on another variable.
Statistical procedures such as ANOVA and Pearson correlation that use continuous dependent variables measured at the interval or ratio level of measurement are examples of parametric statistical tests. Parametric statistical tests involve using data from a sample to estimate population parameters such as μ, and they assume that the distribution of scores for a variable is normally distributed in the larger population.
The chi-square statistic is an example of a nonparametric statistical test, which is a test that does not involve the estimation of population parameters or assume or require variables to be normally distributed. Two primary uses of nonparametric tests are when the dependent variable is measured at a level of measurement other than interval or ratio, as well as when data have been collected from interval or ratio variables but the distribution in the sample may not be normally distributed. Several parametric statistical tests such as the t-test or F- ratio for the one-way ANOVA have a nonparametric alternative.
981
14.9 Important Terms
observed frequency (fo) (p. 652) expected frequency (fe) (p. 652) chi-square statistic (Pearson's chi-square statistic) (X2) (p. 653) assumption of independence of observations (p. 654) Fisher's exact test (p. 655) Yates's correction for continuity (p. 655) chi-square goodness-of-fit test (p. 655) Cramér's φ (phi) (p. 662) contingency table (p. 671) chi-square test of independence (p. 674) parametric statistical tests (p. 684) nonparametric statistical tests (p. 684)
982
14.10 Formulas Introduced in this Chapter
983
Degrees of Freedom for the Chi-Square Goodness-of-Fit Test
(14-1) df = # groups − 1
Expected Frequency for the Chi-Square Goodness-of-Fit Test
(14-2) f e = hypothesized proportion N
Chi-Square Statistic (χ2) (14-3) χ 2 = ∑ f 0 − f e 2 f e
Cramér's φ for the Chi-Square Goodness-of-Fit Test (14-4) Φ = χ 2 N # groups − 1
Degrees of Freedom for the Chi-Square Test of Independence
(14-5) df = # rows − 1 # columns − 1
Expected Frequency for the Chi-Square Test of Independence
(14-6) f e = row total column total N
Cramér's φ for the Chi-Square Test of Independence (14-7) Φ = χ 2 N k − 1
984
14.11 Using SPSS
985
Chi-Square Goodness-of-Fit Test—Equal Hypothesized Proportions: The Personality Study (14.1)
1. Define variable (name, # decimals, label for the variable, labels for values of the variable) and enter data for the variable.
NOTE: Numerically code values of the variable (i.e., 1 = Alpha, 2 = Beta, 3 = Gamma, 4 = Delta) and provide labels for these values in Values box within Variable View.
2. Select the Chi-square procedure within SPSS.
How? (1) Click Analyze menu, (2) click Nonparametric tests, (3) click Legacy Dialogs, and (4) click Chi-square.
3. Select the variable to be analyzed.
How? (1) Click variable and , and (2) click .
986
4. Examine output.
Chi-Square Goodness-of-Fit Test—Unequal Hypothesized Proportions: The Personality Study (14.3)
1. Select the Chi-square procedure within SPSS.
How? (1) Click Analyze menu, (2) click Nonparametric tests, (3) click Legacy Dialogs, and (4) click Chi-square.
987
2. Select the variable to be analyzed and set hypothesized frequency for each group.
How? (1) Click variable and , (2) click Values, (3) enter hypothesized frequencies, and (4) click .
3. Examine output.
Chi-Square Test of Independence: The Color Preference Study (14.4)
1. Define variables (names, # decimals, labels for the variables, labels for values of the variables) and enter data for the variable.
988
NOTE: Numerically code values of the variables (i.e., Toy Color [1 = Red, 2 = Green, 3 = Gray], Face [1 = Happy, 2 = Angry]) and provide labels for these values in Values box within Variable View.
2. Select the Chi-square procedure within SPSS.
How? (1) Click Analyze menu, (2) click Descriptive Statistics, and (3) click Crosstabs.
3. Assign the variables to the rows and columns of the contingency table and ask for row or column percentages.
How? (1) Click row variable and Row(s), (2) click column variable and Column(s), (3) click , (4) click desired Percentages, and (5) click .
989
4. Ask for the chi-square statistic.
How? (1) Click , (2) click Chi-Square, (3) click , and (4) click .
5. Examine output.
990
14.11 Exercises
1. The owner of a car dealership interested in assessing young drivers' preferences asks 26 college students to indicate which type of automobile they are more likely to purchase: traditional gasoline powered or hybrid. Their responses are listed below:
Type of Automobile
Type of Automobile
Type of Automobile
Type of Automobile
Type of Automobile
Gasoline Gasoline Hybrid Hybrid Hybrid
Hybrid Hybrid Gasoline Hybrid Hybrid
Hybrid Gasoline Hybrid Gasoline
Gasoline Hybrid Hybrid Hybrid
Gasoline Hybrid Gasoline Hybrid
Hybrid Gasoline Hybrid Gasoline
a. Construct a frequency distribution table for these data and a bar chart of these data (put the percentage [%] of each type of automobile on the Y-axis).
2. A student interested in consumer behavior believes women are more concerned about their personal appearance than are men; consequently, she hypothesizes that women are less likely to go to a chain hairstyling store than are men. To test her hypothesis, she stands outside of a hairstyling store for 2 hours and counts the number of men and women who purchase haircuts. Her results are listed below:
Sex Sex Sex Sex Sex Sex
Man Woman Man Man Man Woman
Man Woman Man Woman Man Man
Woman Man Man Man Woman Man
Man Man Woman Man Man Man
Woman Man Man Man Woman
Man Man Man Woman Man
a. Construct a frequency distribution table for these data and a bar chart of these data (put the percentage [%] of men and women on the Y-axis).
3. A pollster asks a sample of voters to indicate whether they are in favor, undecided, or against a proposed law. Here are their answers:
Opinion Opinion Opinion Opinion Opinion
991
Undecided In favor In favor Against Undecided
In favor Against In favor In favor In favor
Against In favor Against In favor In favor
Undecided Against Undecided Undecided Against
Against Undecided In favor Against Against
Against In favor In favor Against In favor
In favor In favor Against In favor Undecided
In favor Against In favor In favor
Against In favor Against Against
a. Construct a frequency distribution table for these data and a bar chart of these data (put the percentage [%] of the three opinions on the Y-axis).
4. A Vermont university (www.middlebury.edu) surveyed its graduating seniors regarding their plans immediately following graduation. The majority of their responses fell into one of four categories: work full-time, go to graduate school, travel, or no plans. Below are the choices of a hypothetical sample of graduating seniors (which mirrors the actual findings):
Plans After Plans After Plans After Plans After
Graduation Graduation Graduation Graduation
Graduate school Work full-time Work full-time Work full-time
Work full-time Work full-time Work full-time Graduate school
Work full-time No plans Travel Work full-time
Travel Work full-time Work full-time Work full-time
Work full-time Graduate school Work full-time Graduate school
No plans Work full-time Work full-time Work full-time
Work full-time Work full-time No plans Work full-time
Work full-time Graduate school Work full-time Travel
a. Construct a frequency distribution table for these data and a bar chart of these data (put the percentage [%] of each group of seniors on the Y-axis).
5. An instructor wants to examine the distribution of grades students receive in her classes:
Grade Grade Grade Grade Grade Grade Grade
C F D D B C A
B D B A C B B
992
F A C C D A C
B C F B A F B
C B A C B B
a. Construct a frequency distribution table for these data and a bar chart of these data (put the percentage [%] of each grade on the Y-axis).
6. Given the importance of advertising and products related to the Super Bowl, one company (www.itsasurvey.com) commissioned a survey asking people to indicate their favorite snack to eat during the game. You're interested in seeing whether your friends expressed the same preferences so you visit several parties and see what people are choosing to eat. Here are your findings:
Favorite Snack Favorite Snack Favorite Snack Favorite Snack
Chicken fingers Pizza Chicken fingers Chicken fingers
Pizza Nachos Pizza Veggies/dip
Nachos Veggies/dip Veggies/dip Pizza
Potato chips Pizza Potato chips Veggies/dip
Pizza Chicken fingers Pizza Pizza
Nachos Potato chips Potato chips Potato chips
Veggies/dip Veggies/dip Veggies/dip Pizza
Nachos Chicken fingers Pizza Potato chips
Chicken fingers Pizza Veggies/dip Veggies/dip
Pizza Chicken fingers Nachos Nachos
Potato chips Veggies/dip Potato chips Chicken fingers
Chicken fingers Pizza Pizza Pizza
a. Construct a frequency distribution table for these data (put the percentage [%] of each snack on the Y-axis).
7. For each of the following variables, calculate the degrees of freedom (df) and identify the critical value (α = .05).
a. Drink (Coffee, Tea) b. Season of year (Winter, Spring, Summer, Fall) c. Movie (Comedy, Drama, Romance, Adventure, Horror) d. Instrument (Guitar, Drums, Trumpet, Saxophone, Flute, Violin, Cello)
8. For each of the following variables, calculate the degrees of freedom (df) and identify the critical value (α = .05).
a. Direction (Left, Right, Up, Down) b. Shape (Circle, Triangle, Square, Rectangle, Oval, Pentagon, Star)
993
c. Location (Front, Middle, Back) d. Flavor (Apple, Peach, Strawberry, Pumpkin, Pecan, Cherry)
9. Calculate the expected frequencies (fe) and chi-square statistic (χ2) for the goodness- of-fit test for the data in Exercises 1 to 3 (assume equal hypothesized proportions).
a. Exercise 1 (type of automobile) b. Exercise 2 (hairstyling) c. Exercise 3 (proposed law)
10. Calculate the expected frequencies (fe) and chi-square statistic (χ2) for the goodness- of-fit test for the data in Exercises 4 to 6 (assume equal hypothesized proportions).
a. Exercise 4 (plans after graduation) b. Exercise 5 (class grades) c. Exercise 6 (Super Bowl snack)
11. “Sex in advertising prompts much discussion and media attention, typically how advertisers blend sex with their brands to gain attention” (Reichert & Lambiase, 2003, p. 120). As part of their study, they asked the following question: “Are sexual ads more likely to appear in women's or men's magazines?” (p. 125). To address this question, they examined the full-page advertisements in a sample of women's and men's magazines and found 107 sexual ads that were divided into women's and men's magazines as follows:
Type of Magazine f %
Women's 59 55.1%
Men's 48 44.9%
Total 107 100.0%
To determine whether sexual ads were more likely to be found in women's or men's magazines, conduct the chi-square goodness-of-fit test (assume equal proportions):
a. State the null and alternative hypotheses (H0 and H1). b. Make a decision about the null hypothesis.
1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate a statistic: chi-square (χ2). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (Cramér's φ).
c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
12. An important principle in Sthapatya Veda, a holistic science of architecture, is that buildings should be designed in alignment with the earth's magnetic field and the movement of the sun. Consequently, “a southern orientation of the entrance or
994
sleeping with the head to the north is held to create negative influences” (Travis et al., 2005, p. 555). These researchers “tested the prediction that homes with south entrances would have higher incidents of burglaries” (p. 555) by recording the addresses of 95 burglaries in the crime section of a local newspaper. Next, they went to each address and recorded whether the main entrance of the house was oriented to the east, north, west, or south. The number (f) and percentage of homes in each of the four orientations are provided below:
Orientation f %
East 20 21.1%
North 20 21.1%
West 18 18.9%
South 37 38.9%
Total 95 100.0%
Conduct the chi-square goodness-of-fit test for these data (assume an equal proportion of burglaries for the four orientations):
a. State the null and alternative hypotheses (H0 and H1). b. Make a decision about the null hypothesis.
1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate a statistic: chi-square (χ2). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (Cramér's φ).
c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
13. Historians are interested in the accuracy of eyewitness accounts of traumatic events. One study examined survivors' recall of the sinking of the ship Titanic (Riniolo, Koledin, Drakulic, & Payne, 2003). The researchers reviewed the transcripts of survivors' testimony at governmental hearings (www.titanicinquiry.org) to see whether they testified that the ship was intact or breaking apart during the ship's final plunge (it was in fact breaking apart). Listed below is the testimony of 20 survivors:
Testimony Testimony Testimony Testimony
Breaking apart Intact Breaking apart Breaking apart
Intact Breaking apart Intact Breaking apart
Breaking apart Breaking apart Breaking apart Intact
Breaking apart Intact Breaking apart Breaking apart
995
Breaking apart Breaking apart Breaking apart Breaking apart
To test the hypothesis that survivors were able to accurately recall the state of the ship at the time of the final plunge, conduct the chi-square goodness-of-fit test for these data (assume there was an equal likelihood of saying the ship was intact or breaking apart).
a. State the null and alternative hypotheses (H0 and H1). b. Make a decision about the null hypothesis.
1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate a statistic: chi-square (χ2). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (Cramér's φ).
c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
14. One study investigated what children think about people who share their money (McCrink, Bloom, & Santos, 2010). In this study, sixteen 5-year-olds were shown two puppets: One had a large number of coins, and the other had a small number of coins. Next, the two puppets gave the child some of their coins: One puppet gave the child 3 of its 12 coins, and the other puppet gave 1 of its 4 coins; although the number of coins given by the two puppets differed, the percentage was the same (i.e., 3/12 = 1/4 = 25%). The child then selected the puppet he or she thought was nicer: The “rich” puppet that started with more and gave more or the “poor” puppet that started with less and ended with less. Each child was categorized into one of three groups based on which puppet was thought to be nicer: the rich puppet, the poor puppet, or unsure. The researchers hypothesized that, even though the percentages were the same, children will think a puppet who gives a greater number of coins is nicer than a puppet who gives fewer coins. The selections for the sixteen 5-year olds in this sample are presented below; conduct the chi-square goodness-of-fit test (assume equal hypothesized proportions):
Puppet Puppet Puppet Puppet Puppet Puppet
Rich Unsure Rich Rich Rich Rich
Rich Rich Rich Unsure Unsure
Unsure Rich Unsure Rich Rich
a. Create a frequency distribution table and bar chart of these data. b. State the null and alternative hypotheses (H0 and H1). c. Make a decision about the null hypothesis.
1. Calculate the degrees of freedom (df).
996
2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate a statistic: chi-square (χ2). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (Cramér's φ).
d. Draw a conclusion from the analysis. e. Relate the result of the analysis to the research hypothesis.
15. Exercise 2 tested the hypothesis that women are less likely to go to a chain hairstyling store than are men. However, rather than assume men and women are equally likely to go to chain hairstyle shops, you learn that the Supercuts chain of hairstyling shops reported that 65% of their customers are men while only 35% are women (www.answers.com). Calculate the chi-square goodness-of-fit using hypothesized unequal proportions of .65 for men and .35 for women.
a. State the null and alternative hypotheses (H0 and H1). b. Make a decision about the null hypothesis.
1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate a statistic: chi-square (χ2). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (Cramér's φ).
c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
16. Exercise 6 looked at people's preferences for snacks to eat during the Super Bowl. A survey commissioned by www.Itsasurvey.com found that 10% preferred chicken fingers, 27% nachos, 37% pizza, 13% potato chips, and 13% veggies and dip. Calculate the chi-square goodness-of-fit test comparing your sample's choices with those from itsasurvey.
a. State the null and alternative hypotheses (H0 and H1). b. Make a decision about the null hypothesis.
1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate a statistic: chi-square (χ2). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (Cramér's φ).
c. Draw a conclusion from the analysis. 17. Do people arrested for driving while intoxicated (DWI) believe they're problem
drinkers? Researchers in one study asked 199 people arrested for DWI, “Have you ever thought you might have a drinking problem?” (Adams & Dennis, 2011). Nineteen of the offenders (9.5%) answered “yes” and 180 (91.5%) answered “no.”
997
To test the hypothesis that people arrested for DWI do not believe they have a drinking problem, conduct a goodness-of-fit test using hypothesized proportions of 71.8% “yes” and 28.2% “no,” which are based on an experts opinion as to the percentage of offenders in DWI programs who do in fact have drinking problems.
a. State the null and alternative hypotheses (H0 and H1). b. Make a decision about the null hypothesis.
1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate a statistic: chi-square (χ2). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (Cramér's φ).
c. Draw a conclusion from the analysis. d. Relate the result of the analysis to the research hypothesis.
18. Each of the following situations consists of two variables; for each situation, calculate the degrees of freedom (df) and determine the critical value of χ2 (assume α = .05).
a. Condition (Drug, Placebo) and Age (Young, Old) b. Grade (Third, Sixth) and Score (Below average, Average, Above average) c. Time of day (Morning, Afternoon, Evening) and Activity (Running, Biking,
Swimming) d. Type of car (Luxury, Sports, Sedan, Station wagon) and Major (Business,
Engineering, Social sciences, Humanities) 19.
a. Flavor (Chocolate, Vanilla, Strawberry) and Attitude (Like, Dislike) b. Beverage (Water, Tea, Soda, Beer) and Food (American, Asian, Italian,
Mexican) c. Gender (Male, Female) and Education (Middle school, High school, College,
Grad school) d. Grade (Freshman, Sophomore, Junior, Senior) and Residence (On-campus,
Off-campus, Home) 20. According to one study, deception is an important strategy in dating relationships. “If
one does not honestly possess the desired characteristics to attract a mate, one must portray the image that he/she … has these qualities. This … may require some skill in deception” (Benz, Anderson, & Miller, 2005, p. 306). Do people believe men use deception more in presenting their financial situation (i.e., how much money they make) or their physical appearance (i.e., type of clothing worn)? This study hypothesized that women are more likely than men to believe that men use deception more often regarding their financial situation than their physical appearance. The contingency table of the two types of deception for the male and female respondents is presented below:
998
Use these data to address the question of whether men and women differ in their beliefs regarding men's deception.
a. Draw a bar chart of the data. b. State the null and alternative hypotheses (H0 and H1). c. Make a decision about the null hypothesis.
1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate a statistic: chi-square (χ2). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (Cramér's φ).
d. Draw a conclusion from the analysis. e. Relate the result of the analysis to the research hypothesis.
21. In terms of beliefs regarding women's deception in dating situations, the Benz et al. (2005) study hypothesized that men, more so than women, believe women use deception more often regarding their physical appearance than their financial situation. Below is the contingency table for this analysis:
Use these data to address the question of whether men and women differ in their beliefs regarding women's deception.
a. Draw a bar chart of the data. b. State the null and alternative hypotheses (H0 and H1). c. Make a decision about the null hypothesis.
1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate a statistic: chi-square (χ2). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance.
999
6. Calculate a measure of effect size (Cramér's φ). d. Draw a conclusion from the analysis. e. Relate the result of the analysis to the research hypothesis.
22. Exercise 11 discussed a study looking at sexual ads in women's and men's magazines. Although the researchers found sexual ads were equally likely to appear in women's and men's magazines, they wondered whether different types of sexual ads appeared in the two magazines. To study this question, they coded sexual ads into one of three categories: sexual attractiveness (using the product will make the person appear more attractive), sexual behavior (using the product will increase the likelihood of sexual activity), or sex esteem (using the product will make one feel sexier). The contingency table of the three types of ads in their sample of women's and men's magazines is presented below:
Use these data to address the question of whether the distribution of the three types of sexual ads is different for women's and men's magazines.
a. Draw a bar chart of the data. b. State the null and alternative hypotheses (H0 and H1). c. Make a decision about the null hypothesis.
1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate a statistic: chi-square (χ2). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (Cramér's φ).
d. Draw a conclusion from the analysis. 23. The personality study discussed earlier in this chapter not only compared the
distribution of four types (Alpha, Beta, Gamma, Delta) in a sample of 588 college students with a sample of college students classified during the 1980s but also compared it with a sample of 147 graduate students, saying that “we reasoned that graduate students would differ from typical undergraduates on personality characteristics that support academic achievement” (Stewart & Bernhardt, 2010, p. 582). The following contingency table shows the distribution of the types for the college students and the graduate students:
1000
Conduct the chi-square test of independence to address whether the distribution of the four types depends on the type of student (college student vs. graduate student).
a. Draw a bar chart of the data. b. State the null and alternative hypotheses (H0 and H1). c. Make a decision about the null hypothesis.
1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate a statistic: chi-square (χ2). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (Cramér's φ).
d. Draw a conclusion from the analysis. 24. Students and teachers sometimes disagree regarding what factors and information
teachers should take into account in assigning grades to their students. “Although faculty and students agree that grades should reflect achievement performance, they do not agree on the relative impact effort should have on grades” (J. B. Adams, 2005, pp. 21–22). If a student's performance in a general elective course is failing but at the same time he or she puts in a lot of effort, what grade should that student receive? This study asked this question to a sample of faculty and students; their responses are organized below:
Conduct the chi-square test of independence to address whether there is a difference in the distribution of grades assigned by faculty versus students.
a. Draw a bar chart of the data. b. State the null and alternative hypotheses (H0 and H1). c. Make a decision about the null hypothesis.
1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate a statistic: chi-square (χ2).
1001
4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (Cramér's φ).
d. Draw a conclusion from the analysis. e. Relate the result of the analysis to the research hypothesis.
25. Do children think bedtime stories are real? One study examined whether children's beliefs regarding the reality of people or events in stories depended on the type of book (Woolley & Cox, 2007). In the study, 3-year-old children were read four books that were one of three types: realistic (people interacting with family and friends), fantastical (people interacting with monsters), or religious (people interacting with religious figures). After listening to each book, the children indicated whether they believed the people or events could exist or happen in real life. The children can be grouped into one of two categories based on the number of books with people or events they believed could exist in real life: none of the books versus one or more books. Imagine the researchers hypothesized that children were more likely to believe realistic books were real than fantastical or religious books. Listed below are the data from this study:
Realistic Fantastical Religious
# Books # Books # Books # Books # Books # Books
None 1 or more None None None None
1 or more None None 1 or more None
None 1 or more None None 1 or more
None 1 or more 1 or more None None
None None None None 1 or more
1 or more 1 or more None 1 or more None
1 or more None None None
None 1 or more None 1 or more
1 or more None None None
None 1 or more None None
None None None
Conduct the chi-square test of independence to test whether children's beliefs regarding the reality of people or events in books depend on the type of book they are read.
a. Construct a contingency table and bar chart for these data. b. State the null and alternative hypotheses (H0 and H1). c. Make a decision about the null hypothesis.
1002
1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate a statistic: chi-square (χ2). 4. Make a decision whether to reject the null hypothesis. 5. Determine the level of significance. 6. Calculate a measure of effect size (Cramér's φ).
d. Draw a conclusion from the analysis. e. Relate the result of the analysis to the research hypothesis.
26. Another aspect of the study in Exercise 25 examined age-related changes in children's beliefs; more specifically the study hypothesized that older children were more likely to believe that the people or events in books could really exist or happen than younger children. To test this hypothesis, children who were 3, 4, or 5 years old were read four books in which people interacted with religious figures such as God. Listed below are the number of books (none, 1, 2 or more) that each child indicated he or she believed the people or events in the book could exist or happen in real life:
3 Years Old 4 Years Old 5 Years Old
# Books # Books # Books # Books # Books # Books
None None 1 None 2 or more 2 or more
2 or more None None 2 or more 2 or more None
None 2 or more 1 1 1
2 or more None None 2 or more 2 or more
None None 1 None 2 or more
None 1 None 2 or more 1
None None 2 or more 2 or more
2 or more 2 or more 1 None
None None 2 or more 2 or more
None None 2 or more
Conduct the chi-square test of independence to address whether the beliefs regarding the reality of people or events in religious books depend on the age of the child.
a. Construct a contingency table and bar chart for these data. b. State the null and alternative hypotheses (H0 and H1). c. Make a decision about the null hypothesis.
1. Calculate the degrees of freedom (df). 2. Set alpha (α), identify the critical value, and state a decision rule. 3. Calculate a statistic: chi-square (χ2). 4. Make a decision whether to reject the null hypothesis.
1003
5. Determine the level of significance. 6. Calculate a measure of effect size (Cramér's φ).
d. Draw a conclusion from the analysis. e. Relate the result of the analysis to the research hypothesis.
1004
Answers to Learning Checks
Learning Check 1
2.
Type of Phone f %
Keyboard 4 25.0%
Touch screen 12 75.0%
Total 16 100.0%
a.
Type of Milk f %
Whole 7 35.0%
Low fat 10 50.0%
Non-fat 3 15.0%
Total 20 100.0%
b.
Learning Check 2
2. a. df = 1, critical value = 3.84 b. df = 2, critical value = 5.99
1005
c. df = 4, critical value = 9.49 3.
a. CCC: fe = 22.00; MCCC: fe = 22.00 b. χ2 = 13.10
Learning Check 3
2.
Color f %
Purple 16 45.7%
Aqua 7 20.0%
Pink 12 34.3%
Total 35 100.0%
a.
b. H0: The distribution of observed frequencies fits the distribution of expected frequencies. H1: The distribution of observed frequencies does not fit the distribution of expected frequencies.
c. 1. df = 2 2. If χ2 > 5.99, reject H0; otherwise, do not reject H0 3. Purple: fe = 11.67; Aqua: fe = 11.67; Pink: fe = 11.67; χ2 = 3.49 4. χ2 = 3.49 < 5.99 ∴ do not reject H0 (p > .05) 5. Not necessary (H0 not rejected) 6. φ = .22
d. In the sample of 35 people, votes for three proposed M&M colors (Purple [f = 16 (45.7%)], Aqua [f = 7 (20.0%)], and Pink [f = 12 (34.3%)]) were not significantly different from each other, χ2(2, N = 35) = 3.49, p > .05, φ = .22.
Learning Check 4
2.
1006
a. H0: The distribution of observed frequencies fits the distribution of expected frequencies. H1: The distribution of observed frequencies does not fit the distribution of expected frequencies.
b. 1. df = 2 2. If χ2 > 5.99, reject H0; otherwise, do not reject H0 3. Purple: fe = 14.30; Aqua: fe = 13.30; Pink: fe = 7.00; χ2 = 6.66 4. χ2 = 6.66 > 5.99 ∴ reject H0 (p < .05) 5. χ2 = 6.66 < 9.21 ∴ p < .05 (but not < .01) 6. φ = .31
c. In the sample of 35 people, votes for three proposed M&M colors (Purple [f = 16 (45.7%)], Aqua [f = 7 (20.0%)], and Pink [f = 12 (34.3%)]) were significantly different from frequencies based on unequal hypothesized proportions (Purple [42%], Aqua [38%], Pink [20%]), χ2(2, N = 35) = 6.66, p > .05, φ = .31.
Learning Check 5
a.
b.
1007
c.
Learning Check 6
2. a. df = 1, critical value = 3.84 b. df = 3, critical value = 7.81 c. df = 6, critical value = 12.59
3.
a.
b. H0: The two variables are independent (i.e., there is no relationship between the two variables). H1: The two variables are not independent (i.e., there is a relationship between the two variables).
1008
c. 1. df = 3 2. If χ2 > 7.81, reject H0; otherwise, do not reject H0 3. χ2 = 2.79
4. χ2 = 2.79 < 7.81 ∴ do not reject H0 (p > .05) 5. Not necessary (H0 not rejected) 6. φ = .15
d. In a sample of 120 five- and nine-year-olds, the distribution of preferred colors for the 60 boys and 60 girls in this sample was not different, χ2(3) = 2.79, p > .05, φ = .15.
1009
Answers to Odd-Numbered Exercises
1.
Type of Automobile f %
Gasoline 10 38.5%
Hybrid 16 61.5%
Total 26 100.0%
a.
3.
Opinion f %
In Favor 20 46.5%
Undecided 7 16.3%
Against 16 37.2%
Total 43 100.0%
a.
5.
Grade f %
A 6 17.6%
B 11 32.4%
C 9 26.5%
1010
D 4 11.8%
F 4 11.8%
Total 34 100.0%
a.
7. a. df = 1, critical value = 3.84 b. df = 3, critical value = 7.81 c. df = 4, critical value = 9.49 d. df = 6, critical value = 12.59
9. a. Gasoline: fe = 13.00; Hybrid: fe = 13.00; χ2 = 1.38 b. Man: fe = 17.00; Woman: fe = 17.00; χ2 = 5.76 c. In favor: fe = 14.33; Undecided: fe = 14.33; Against: fe = 14.33 χ2 = 6.18
11. a. H0: The distribution of observed frequencies fits the distribution of expected
frequencies. H1: The distribution of observed frequencies does not fit the distribution of expected frequencies.
b. 1. df = 1 2. If χ2 > 3.84, reject H0; otherwise, do not reject H0 3. Women's: fe = 53.50; Men's: fe = 53.50; χ2 = 1.13 4. χ2 = 1.13 < 3.84 ∴ do not reject H0 (p > .05) 5. Not necessary (H0 not rejected) 6. φ = .10
c. “Although more sexual ads appeared in women's (55%, N = 59) than in men's (45%, N = 48) magazines, the overall difference was not significant, χ2 (1, N = 107) = 1.13, p > .05” (Reichert & Lambiase, 2003, p. 128).
d. This analysis does not support the research hypothesis that sexual ads are more likely to appear in women's versus men's magazines.
13.
Testimony f %
1011
Intact 5 25.0%
Breaking Up 15 75.0%
Total 20 100.0%
a.
b. H0: The distribution of observed frequencies fits the distribution of expected frequencies. H1: The distribution of observed frequencies does not fit the distribution of expected frequencies.
c. 1. df = 1 2. If χ2 > 3.84, reject H0; otherwise, do not reject H0 3. Intact: fe = 10.00; Breaking up: fe = 10.00; χ2 = 5.00 4. χ2 = 5.00 > 3.84 ∴ reject H0 (p < .05) 5. χ2 = 5.00 < 6.63 ∴ p < .05 (but not < .01) 6. φ = .50
d. In the sample of 20 Titanic survivors, a significantly greater number of survivors testified that the ship was breaking apart (f = 15 [75.0%]) than intact (f = 5 [25.0%]) during the ships final plunge, χ2(1, N = 20) = 5.00, p < .05, φ = .50.
e. This analysis supports the research hypothesis that hypothesis that Titanic survivors were able to accurately recall the state of the ship at the time of the final plunge.
15. a. H0: The distribution of observed frequencies fits the distribution of expected
frequencies. H1: The distribution of observed frequencies does not fit the distribution of expected frequencies.
b. 1. df = 1 2. If χ2 > 3.84, reject H0; otherwise, do not reject H0 3. Men: fe = 22.10; Women: fe = 11.90; χ2 = .46 4. χ2 = .46 < 3.84 ∴ do not reject H0 (p > .05) 5. Not necessary (H0 not rejected) 6. φ = .12
1012
c. The number of men (24 [70.6%]) and women (10 [29.4%]) shopping at the hairstyling chain store in this sample did not differ significantly from frequencies based on hypothesized proportions of 65% for men and 35% for women, χ2 (1, N = 34) = .46, p > .05, φ = .12.
d. This analysis does not support the research hypothesis that women are less likely to go to a chain hairstyling store than are men.
17. a. H0: The distribution of observed frequencies fits the distribution of expected
frequencies. H1: The distribution of observed frequencies does not fit the distribution of expected frequencies.
b. 1. df = 1 2. If χ2 > 3.84, reject H0; otherwise, do not reject H0 3. Yes: fe = 142.88; No: fe = 56.12; χ2 = 380.88 4. χ2 = 380.88 > 3.84 ∴ reject H0 (p < .05) 5. χ2 = 380.88 > 6.63 ∴ p < .01 6. φ = 1.38
c. In the sample of 199 people arrested for DWI, the percentage of people who answered “yes” (f = 19 [9.5%]) and “no” (f = 180 [91.5%]) to the question, “Have you ever thought you might have a drinking problem?” were significantly different from the experts opinion of 71.8% “yes” and 28.2% “no,” χ2(1, N = 199) = 380.88, p < .01, φ = 1.38.
d. This analysis supports the research hypothesis that people arrested for DWI do not believe they have a drinking problem.
19 a. df = 2, critical value = 5.99 b. df = 9, critical value = 16.92 c. df = 3, critical value = 7.81 d. df = 6, critical value = 12.59
21.
a.
1013
b. H0: The two variables are independent (i.e., there is no relationship between the two variables). H1: The two variables are not independent (i.e., there is a relationship between the two variables).
c. 1. df = 1 2. If χ2 > 3.84, reject H0; otherwise, do not reject H0 3. χ2 = 4.75
4. χ2 = 4.75 > 3.84 ∴ reject H0 (p < .05) 5. χ2 = 4.75 < 6.63 ∴ p < .05 (but not <.01) 6. φ = .23
d. This analysis found a significant relationship between Sex and Type of deception used by women such that a higher percentage of female respondents (100.0%) than male respondents (85.2%) believed women were more likely to use physical deception than financial deception, χ2 (1, N = 90) = 4.75, p < .05, φ = .23.
e. The result of this analysis supports the hypothesis that men, more so than women, believe women use deception more often regarding their physical appearance than their financial situation.
23.
a.
b. H0: The two variables are independent (i.e., there is no relationship between the two variables). H1: The two variables are not independent (i.e., there is a relationship between the two variables).
c. 1. df = 3
1014
2. If χ2 > 7.81, reject H0; otherwise, do not reject H0 3. χ2 = 16.86
4. χ2 = 16.86 > 7.81 ∴ reject H0 (p < .05) 5. χ2 = 16.86 > 11.34 ∴ p < .01 6. φ = .15
d. “A Chi-square test of independence was conducted to determine if the distribution of the four Lifestyle types differed depending on graduate status (comparing 2004-08 undergraduates to 2004-08 graduates). The test was significant, showing the distribution did depend on graduate status, χ2 (3, N =735) = 16.86, p = .001” (Stewart & Bernhardt, 2010, p. 591).
e. The result of this analysis supports the research hypothesis that the distribution of the four types depends on the type of student (college student vs. graduate student).
25.
a.
b. H0: The two variables are independent (i.e., there is no relationship between the two variables). H1: The two variables are not independent (i.e., there is a relationship between the two variables).
c. 1. df = 2
1015
2. If χ2 > 5.99, reject H0; otherwise, do not reject H0 3. χ2 = 2.69
4. χ2 = 2.69 < 5.99 ∴ do not reject H0 (p > .05) 5. Not necessary (H0 not rejected) 6. φ = .23
d. In a sample of fifty 3-year-old children, a chi-square test of independence indicated a nonsignificant relationship between the type of book read to them (realistic, fantastical, or religious) and the number of books (none, 1 or more) they believed contained people or events that could exist or happen in real life, χ2 (2, N = 50) = 2.69, p > .05, φ = .23.
e. The result of this analysis does not support the research hypothesis that children are more likely to believe that realistic books are real than that fantastical or religious books are real.
1016
Sharpen your skills with SAGE edge at edge.sagepub.com/tokunaga
SAGE edge for students provides a personalized approach to help you accomplish your coursework goals in an easy-to-use learning environment. Log on to access:
eFlashcards Web Quizzes Chapter Outlines Learning Objectives Media Links SPSS Data Files
1017
Tables
Table 1. Proportions of Area Under the Standard Normal Distribution
Table 2. Binomial Probabilities
Adapted from Burington, R. S., & May, D. C. (1970). Handbook of probability and statistics with tables (2nd ed.). New York: McGraw-Hill.
Table 3. Critical Values of t
Adapted from Pearson, E. S., & Hartley, H. O. (Eds.). (1966). Biometrika tables for statisticians (3rd ed., Vol. 1). New York: Cambridge University Press.
Table 4. Critical Values of F
Adapted from Pearson, E. S., & Hartley, H. O. (Eds.). (1966). Biometrika tables for statisticians (3rd ed., Vol. 1). New York: Cambridge University Press.
Table 5. The Studentized Range Statistic (qT)
Adapted from Pearson, E. S., & Hartley, H. O. (Eds.). (1966). Biometrika tables for statisticians (3rd ed., Vol. 1). New York: Cambridge University Press. Table 6. Critical Values of the Pearson Correlation (r)
Table 7. Critical Values of the Spearman Rank-Order Correlation (rS)
Adapted from Ramsey, P. H. (1989). Critical values for Spearman's rank order correlation. Journal of Educational Statistics, 14, 245–253.
Table 8. Critical Values of Chi-Square (χ2)
Adapted from Pearson, E. S., & Hartley, H. O. (Eds.). (1966). Biometrika tables for statisticians (3rd ed., Vol. 1). New York: Cambridge University Press.
Table 1 Proportions of Area Under the Standard Normal Distribution Table 1 Proportions of Area Under the Standard Normal Distribution
Instructions: The ‘z’ column lists z-score values. The ‘Area between mean and z’ column lists the proportion of the area between thc mean and the z-score value. The ‘Area beyond z’ column lists the proportion of the area beyond the z-score value in the tail of the distribution. Note: Because the distribution is symmetrical, areas for negative z-scores are the same as for positive z-scores.
1018
.00 .0000 .5000
.01 .0040 .4960
.02 .0080 .4920
.03 .0120 .4880
.04 .0160 .4840
.05 .0199 .4801
.06 .0239 .4761
.07 .0279 .4721
.08 .0319 .4681
.09 .0359 .4641
.10 .0398 .4602
.11 .0438 .4562
.12 .0478 .4522
.13 .0517 .4483
.14 .0557 .4443
.15 .0596 .4404
.16 .0636 .4364
.17 .0675 .4325
.18 .0714 .4286
.19 .0753 .4247
.20 .0793 .4207
.21 .0832 .4168
.22 .0871 .4129
.23 .0910 .4090
.24 .0948 .4052
1019
.25 .0987 .4013
.26 .1026 .3974
.27 .1064 .3936
.28 .1103 .3897
.29 .1141 .3859
.30 .1179 .3821
.31 .1217 .3783
.32 .1255 .3745
.33 .1293 .3707
.34 .1331 .3669
.35 .1368 .3632
.36 .1406 .3594
.37 .1443 .3557
.38 .1480 .3520
.39 .1517 .3483
.40 .1554 .3446
.41 .1591 .3409
.42 .1628 .3372
.43 .1664 .3336
.44 .1700 .3300
.45 .1736 .3264
.46 .1772 .3228
.47 .1808 .3192
.48 .1844 .3156
.49 .1879 .3121
.50 .1915 .3085
.51 .1950 .3050
.52 .1985 .3015
.53 .2019 .2981
1020
.54 .2054 .2946
.55 .2088 .2912
.56 .2123 .2877
.57 .2157 .2843
.58 .2190 .2810
.59 .2224 .2776
.60 .2257 .2743
.61 .2291 .2709
.62 .2324 .2676
.63 .2357 .2643
.64 .2389 .2611
.65 .2422 .2578
.66 .2454 .2546
.67 .2486 .2514
.68 .2517 .2483
.69 .2549 .2451
.70 .2580 .2420
.71 .2611 .2389
.72 .2642 .2358
.73 .2673 .2327
.74 .2704 .2296
.75 .2734 .2266
.76 .2764 .2236
.77 .2794 .2206
.78 .2823 .2177
.79 .2852 .2148
.80 .2881 .2119
.81 .2910 .2090
1021
.82 .2939 .2061
.83 .2967 .2033
.84 .2995 .2005
.85 .3023 .1977
.86 .3051 .1949
.87 .3078 .1922
.88 .3106 .1894
.89 .3133 .1867
.90 .3159 .1841
.91 .3186 .1814
.92 .3212 .1788
.93 .3238 .1762
.94 .3264 .1736
.95 .3289 .1711
.96 .3315 .1685
.97 .3340 .1660
.98 .3365 .1635
.99 .3389 .1611
1.00 .3413 .1587
1.01 .3438 .1562
1.02 .3461 .1539
1.03 .3485 .1515
1.04 .3508 .1492
1.05 .3531 .1469
1.06 .3554 .1446
1.07 .3577 .1423
1.08 .3599 .1401
1.09 .3621 .1379
1022
1.10 .3643 .1357
1.11 .3665 .1335
1.12 .3686 .1314
1.13 .3708 .1292
1.14 .3729 .1271
1.15 .3749 .1251
1.16 .3770 .1230
1.17 .3790 .1210
1.18 .3810 .1190
1.19 .3830 .1170
1.20 .3849 .1151
1.21 .3869 .1131
1.22 .3888 .1112
1.23 .3907 .1093
1.24 .3925 .1075
1.25 .3944 .1056
1.26 .3962 .1038
1.27 .3980 .1020
1.28 .3997 .1003
1.29 .4015 .0985
1.30 .4032 .0968
1.31 .4049 .0951
1.32 .4066 .0934
1.33 .4082 .0918
1.34 .4099 .0901
1.35 .4115 .0885
1.36 .4131 .0869
1.37 .4147 .0853
1.38 .4162 .0838
1023
1.39 .4177 .0823
1.40 .4192 .0808
1.41 .4207 .0793
1.42 .4222 .0778
1.43 .4236 .0764
1.44 .4251 .0749
1.45 .4265 .0735
1.46 .4279 .0721
1.47 .4292 .0708
1.48 .4306 .0694
1.49 .4319 .0681
1.50 .4332 .0668
1.51 .4345 .0655
1.52 .4357 .0643
1.53 .4370 .0630
1.54 .4382 .0618
1.55 .4394 .0606
1.56 .4406 .0594
1.57 .4418 .0582
1.58 .4429 .0571
1.59 .4441 .0559
1.60 .4452 .0548
1.61 .4463 .0537
1.62 .4474 .0526
1.63 .4484 .0516
1.64 .4495 .0505
1.65 .4505 .0495
1.66 .4515 .0485
1024
1.67 .4525 .0475
1.68 .4535 .0465
1.69 .4545 .0455
1.70 .4554 .0446
1.71 .4564 .0436
1.72 .4573 .0427
1.73 .4582 .0418
1.74 .4591 .0409
1.75 .4599 .0401
1.76 .4608 .0392
1.77 .4616 .0384
1.78 .4625 .0375
1.79 .4633 .0367
1.80 .4641 .0359
1.81 .4649 .0351
1.82 .4656 .0344
1.83 .4664 .0336
1.84 .4671 .0329
1.85 .4678 .0322
1.86 .4686 .0314
1.87 .4693 .0307
1.88 .4699 .0301
1.89 .4706 .0294
1.90 .4713 .0287
1.91 .4719 .0281
1.92 .4726 .0274
1.93 .4732 .0268
1.94 .4738 .0262
1025
1.95 .4744 .0256
1.96 .4750 .0250
1.97 .4756 .0244
1.98 .4761 .0239
1.99 .4767 .0233
2.00 .4772 .0228
2.01 .4778 .0222
2.02 .4783 .0217
2.03 .4788 .0212
2.04 .4793 .0207
2.05 .4798 .0202
2.06 .4803 .0197
2.07 .4808 .0192
2.08 .4812 .0188
2.09 .4817 .0183
2.10 .4821 .0179
2.11 .4826 .0174
2.12 .4830 .0170
2.13 .4834 .0166
2.14 .4838 .0162
2.15 .4842 .0158
2.16 .4846 .0154
2.17 .4850 .0150
2.18 .4854 .0146
2.19 .4857 .0143
2.20 .4861 .0139
2.21 .4864 .0136
2.22 .4868 .0132
2.23 .4871 .0129
1026
2.24 .4875 .0125
2.25 .4878 .0122
2.26 .4881 .0119
2.27 .4884 .0116
2.28 .4887 .0113
2.29 .4890 .0110
2.30 .4893 .0107
2.31 .4896 .0104
2.32 .4898 .0102
2.33 .4901 .0099
2.34 .4904 .0096
2.35 .4906 .0094
2.36 .4909 .0091
2.37 .4911 .0089
2.38 .4913 .0087
2.39 .4916 .0084
2.40 .4918 .0082
2.41 .4920 .0080
2.42 .4922 .0078
2.43 .4925 .0075
2.44 .4927 .0073
2.45 .4929 .0071
2.46 .4931 .0069
2.47 .4932 .0068
2.48 .4934 .0066
2.49 .4936 .0064
2.50 .4938 .0062
2.51 .4940 .0060
1027
2.52 .4941 .0059
2.53 .4943 .0057
2.54 .4945 .0055
2.55 .4946 .0054
2.56 .4948 .0052
2.57 .4949 .0051
2.58 .4951 .0049
2.59 .4952 .0048
2.60 .4953 .0047
2.61 .4955 .0045
2.62 .4956 .0044
2.63 .4957 .0043
2.64 .4959 .0041
2.65 .4960 .0040
2.66 .4961 .0039
2.67 .4962 .0038
2.68 .4963 .0037
2.69 .4964 .0036
2.70 .4965 .0035
2.71 .4966 .0034
2.72 .4967 .0033
2.73 .4968 .0032
2.74 .4969 .0031
2.75 .4970 .0030
2.76 .4971 .0029
2.77 .4972 .0028
2.78 .4973 .0027
2.79 .4974 .0026
1028
2.80 .4974 .0026
2.81 .4975 .0025
2.82 .4976 .0024
2.83 .4977 .0023
2.84 .4977 .0023
2.85 .4978 .0022
2.86 .4979 .0021
2.87 .4979 .0021
2.88 .4980 .0020
2.89 .4981 .0019
2.90 .4981 .0019
2.91 .4982 .0018
2.92 .4982 .0018
2.93 .4983 .0017
2.94 .4984 .0016
2.95 .4984 .0016
2.96 .4985 .0015
2.97 .4985 .0015
2.98 .4986 .0014
2.99 .4986 .0014
3.00 .4987 .0013
3.01 .4987 .0013
3.02 .4987 .0013
3.03 .4988 .0012
3.04 .4988 .0012
3.05 .4989 .0011
3.06 .4989 .0011
3.07 .4989 .0011
3.08 .4990 .0010
1029
3.09 .4990 .0010
3.10 .4990 .0010
3.11 .4991 .0009
3.12 .4991 .0009
3.13 .4991 .0009
3.14 .4992 .0008
3.15 .4992 .0008
3.20 .4993 .0007
3.25 .4994 .0006
3.30 .4995 .0005
3.35 .4996 .0004
3.40 .4997 .0003
3.45 .4997 .0003
3.50 .4998 .0002
3.60 .4998 .0002
3.70 .4999 .0001
3.80 .4999 .0001
3.90 .49995 .00005
4.00 .49997 .00003
Table 2 Binomial Distribution Table 2 Binomial Distribution
p
x .01 .05 .10 .15 .20 .25 .30 .35 .40
n=1 0 .9900 .9500 .9000 .8500 .8000 .7500 .7000 .6500 .6000
1 .0100 .0500 .1000 .1500 .2000 .2500 .3000 .3500 .4000
n=2
0 .9801 .9025 .8100 .7225 .6400 .5625 .4900 .4225 .3600
1 .0198 .0950 .1800 .2550 .3200 .3750 .4200 .4550 .4800
2 .0001 .0025 .0100 .0225 .0400 .0625 .0900 .1225 .1600
1030
n=3
0 .9703 .8574 .7290 .6141 .5120 .4219 .3430 .2746 .2160
1 .0294 .1354 .2430 .3251 .3840 .4219 .4410 .4436 .4320
2 .0003 .0071 .0270 .0574 .0960 .1406 .1890 .2389 .2880
3 .0001 .0010 .0034 .0080 .0156 .0270 .0429 .0640
n=4
0 .9606 .8145 .6561 .5220 .4096 .3164 .2401 .1785 .1296
1 .0388 .1715 .2916 .3685 .4096 .4219 .4116 .3845 .3456
2 .0006 .0135 .0486 .0975 .1536 .2109 .2646 .3105 .3456
3 .0005 .0036 .0115 .0256 .0469 .0756 .1115 .1536
4 .0001 .0005 .0016 .0039 .0081 .0150 .0256
n=5
0 .9510 .7738 .5905 .4437 .3277 .2373 .1681 .1160 .0778
1 .0480 .2036 .3281 .3915 .4096 .3955 .3602 .3124 .2592
2 .0010 .0214 .0729 .1382 .2048 .2637 .3087 .3364 .3456
3 .0011 .0081 .0244 .0512 .0879 .1323 .1811 .2304
4 .0005 .0022 .0064 .0146 .0284 .0488 .0768
5 .0001 .0003 .0010 .0024 .0053 .0102
.99 .95 .90 .85 .80 .75 .70 .65 .60
p
n=6
0 .9415 .7351 .5314 .3771 .2621 .1780 .1176 .0754 .0467
1 .0571 .2321 .3543 .3993 .3932 .3560 .3025 .2437 .1866
2 .0014 .0305 .0984 .1762 .2458 .2966 .3241 .3280 .3110
3 .0021 .0146 .0415 .0819 .1318 .1852 .2355 .2765
4 .0001 .0012 .0055 .0154 .0330 .0595 .0951 .1382
5 .0001 .0004 .0015 .0044 .0102 .0205 .0369
6 .0001 .0002 .0007 .0018 .0041
n=7
0 .9321 .6983 .4783 .3206 .2097 .1335 .0824 .0490 .0280
1 .0659 .2573 .3720 .3960 .3670 .3115 .2471 .1848 .1306
2 .0020 .0406 .1240 .2097 .2753 .3115 .3177 .2985 .2613
3 .0036 .0230 .0617 .1147 .1730 .2269 .2679 .2903
4 .0002 .0026 .0109 .0287 .0577 .0972 .1442 .1935
5 .0002 .0012 .0043 .0115 .0250 .0466 .0774
6 .0001 .0004 .0013 .0036 .0084 .0172
7 .0001 .0002 .0006 .0016
0 .9227 .6634 .4305 .2725 .1678 .1001 .0576 .0319 .0168
1031
n=8
1 .0746 .2793 .3826 .3847 .3355 .2670 .1977 .1373 .0896
2 .0026 .0515 .1488 .2376 .2936 .3115 .2965 .2587 .2090
3 .0001 .0054 .0331 .0839 .1468 .2076 .2541 .2786 .2787
4 .0004 .0046 .0185 .0459 .0865 .1361 .1875 .2322
5 .0004 .0026 .0092 .0231 .0467 .0808 .1239
6 .0002 .0011 .0038 .0100 .0217 .0413
7 .0001 .0004 .0012 .0033 .0079
8 .0001 .0002 .0007
.99 .95 .90 .85 .80 .75 .70 .65 .60
p
n=9
0 .9135 .6302 .3874 .2316 .1342 .0751 .0404 .0207 .0101
1 .0830 .2985 .3874 .3679 .3020 .2253 .1556 .1004 .0605
2 .0034 .0629 .1722 .2597 .3020 .3003 .2668 .2162 .1612
3 .0001 .0077 .0446 .1069 .1762 .2336 .2668 .2716 .2508
4 .0006 .0074 .0283 .0661 .1168 .1715 .2194 .2508
5 .0008 .0050 .0165 .0389 .0735 .1181 .1672
6 .0001 .0006 .0028 .0087 .0210 .0424 .0743
7 .0003 .0012 .0039 .0098 .0212
8 .0001 .0004 .0013 .0035
9 .0001 .0003
n=10
0 .9044 .5987 .3487 .1969 .1074 .0563 .0282 .0135 .0060
1 .0914 .3151 .3874 .3474 .2684 .1877 .1211 .0725 .0403
2 .0042 .0746 .1937 .2759 .3020 .2816 .2335 .1757 .1209
3 .0001 .0105 .0574 .1298 .2013 .2503 .2668 .2522 .2150
4 .0010 .0112 .0401 .0881 .1460 .2001 .2377 .2508
5 .0001 .0015 .0085 .0264 .0584 .1029 .1536 .2007
6 .0001 .0012 .0055 .0162 .0368 .0689 .1115
7 .0001 .0008 .0031 .0090 .0212 .0425
8 .0001 .0004 .0014 .0043 .0106
9 .0001 .0005 .0016
10 .0001
.99 .95 .90 .85 .80 .75 .70 .65 .60
p
1032
n=11
0 .8953 .5688 .3138 .1673 .0859 .0422 .0198 .0088 .0036
1 .0995 .3293 .3835 .3248 .2362 .1549 .0932 .0518 .0266
2 .0050 .0867 .2131 .2866 .2953 .2581 .1998 .1395 .0887
3 .0002 .0137 .0710 .1517 .2215 .2581 .2568 .2254 .1774
4 .0014 .0158 .0536 .1107 .1721 .2201 .2428 .2365
5 .0001 .0025 .0132 .0388 .0803 .1321 .1830 .2207
6 .0003 .0023 .0097 .0268 .0566 .0985 .1471
7 .0003 .0017 .0064 .0173 .0379 .0701
8 .0002 .0011 .0037 .0102 .0234
9 .0001 .0005 .0018 .0052
10 .0002 .0007
11
n=12
0 .8864 5404 .2824 .1422 .0687 .0317 .0138 .0057 .0022
1 .1074 .3413 .3766 .3012 .2062 .1267 .0712 .0368 .0174
2 .0060 .0988 .2301 .2924 .2835 .2323 .1678 .1088 .0639
3 .0002 .0173 .0852 .1720 .2362 .2581 .2397 .1954 .1419
4 .0021 .0213 .0683 .1329 .1936 .2311 .2367 .2128
5 .0002 .0038 .0193 .0532 .1032 .1585 .2039 .2270
6 .0005 .0040 .0155 .0401 .0792 .1281 .1766
7 .0006 .0033 .0115 .0291 .0591 .1009
8 .0001 .0005 .0024 .0078 .0199 .0420
9 .0001 .0004 .0015 .0048 .0125
10 .0002 .0008 .0025
11 .0001 .0003
12
.99 .95 .90 .85 .80 .75 .70 .65 .60
p
0 .8775 .5133 .2542 .1209 .0550 .0238 .0097 .0037 .0013
1 .1152 .3512 .3672 .2774 .1787 .1029 .0540 .0259 .0113
2 .0070 .1109 .2448 .2937 .2680 .2059 .1388 .0836 .0453
3 .0003 .0214 .0997 .1900 .2457 .2517 .2181 .1651 .1107
4 .0028 .0277 .0838 .1535 .2097 .2337 .2222 .1845
5 .0003 .0055 .0266 .0691 .1258 .1803 .2154 .2214
1033
n=13 6 .0008 .0063 .0230 .0559 .1030 .1546 .1968
7 .0001 .0011 .0058 .0186 .0442 .0833 .1312
8 .0001 .0011 .0047 .0142 .0336 .0656
9 .0001 .0009 .0034 .0101 .0243
10 .0001 .0006 .0022 .0065
11 .0001 .0003 .0012
12 .0001
13
.99 .95 .90 .85 .80 .75 .70 .65 .60
p
n=14
0 .8687 .4877 .2288 .1028 .0440 .0178 .0068 .0024 .0008
1 .1229 .3593 .3559 .2539 .1539 .0832 .0407 .0181 .0073
2 .0081 .1229 .2570 .2912 .2501 .1802 .1134 .0634 .0317
3 .0003 .0259 .1142 .2056 .2501 .2402 .1943 .1366 .0845
4 .0037 .0349 .0998 .1720 .2202 .2290 .2022 .1549
5 .0004 .0078 .0352 .0860 .1468 .1963 .2178 .2066
6 .0013 .0093 .0322 .0734 .1262 .1759 .2066
7 .0002 .0019 .0092 .0280 .0618 .1082 .1574
8 .0003 .0020 .0082 .0232 .0510 .0918
9 .0003 .0018 .0066 .0183 .0408
10 .0003 .0014 .0049 .0136
11 .0002 .0010 .0033
12 .0001 .0005
13 .0001
14
.99 .95 .90 .85 .80 .75 .70 .65 .60
p
0 .8601 .4633 .2059 .0874 .0352 .0134 .0047 .0016 .0005
1 .1303 .3658 .3432 .2312 .1319 .0668 .0305 .0126 .0047
2 .0092 .1348 .2669 .2856 .2309 .1559 .0916 .0476 .0219
3 .0004 .0307 .1285 .2184 .2501 .2252 .1700 .1110 .0634
4 .0049 .0428 .1156 .1876 .2252 .2186 .1792 .1268
5 .0006 .0105 .0449 .1032 .1651 .2061 .2123 .1859
1034
n=15
6 .0019 .0132 .0430 .0917 .1472 .1906 .2066
7 .0003 .0030 .0138 .0393 .0811 .1319 .1771
8 .0005 .0035 .0131 .0348 .0710 .1181
9 .0001 .0007 .0034 .0116 .0298 .0612
10 .0001 .0007 .0030 .0096 .0245
11 .0001 .0006 .0024 .0074
12 .0001 .0004 .0016
13 .0001 .0003
14
15
.99 .95 .90 .85 .80 .75 .70 .65 .60
p
n=16
0 .8515 .4401 .1853 .0743 .0281 .0100 .0033 .0010 .0003 .0001
1 .1376 .3706 .3294 .2097 .1126 .0535 .0228 .0087 .0030 .0009
2 .0104 .1463 .2745 .2775 .2111 .1336 .0732 .0353 .0150 .0056
3 .0005 .0359 .1423 .2285 .2463 .2079 .1465 .0888 .0468 .0215
4 .0061 .0514 .1311 .2001 .2252 .2040 .1553 .1014 .0572
5 .0008 .0137 .0555 .1201 .1802 .2099 .2008 .1623 .1123
6 .0001 .0028 .0180 .0550 .1101 .1649 .1982 .1983 .1684
7 .0004 .0045 .0197 .0524 .1010 .1524 .1889 .1969
8 .0001 .0009 .0055 .0197 .0487 .0923 .1417 .1812
9 .0001 .0012 .0058 .0185 .0442 .0840 .1318
10 .0002 .0014 .0056 .0167 .0392 .0755
11 .0002 .0013 .0049 .0142 .0337
12 .0002 .0011 .0040 .0115
13 .0000 .0002 .0008 .0029
14 .0001 .0005
15 .0001
16
.99 .95 .90 .85 .80 .75 .70 .65 .60
p
0 .8429 .4181 .1668 .0631 .0225 .0075 .0023 .0007 .0002
1035
n=17
1 .1447 .3741 .3150 .1893 .0957 .0426 .0169 .0060 .0019
2 .0117 .1575 .2800 .2673 .1914 .1136 .0581 .0260 .0102
3 .0006 .0415 .1556 .2359 .2393 .1893 .1245 .0701 .0341
4 .0076 .0605 .1457 .2093 .2209 .1868 .1320 .0796
5 .0010 .0175 .0668 .1361 .1914 .2081 .1849 .1379
6 .0001 .0039 .0236 .0680 .1276 .1784 .1991 .1839
7 .0007 .0065 .0267 .0668 .1201 .1685 .1927
8 .0001 .0014 .0084 .0279 .0644 .1134 .1606
9 .0003 .0021 .0093 .0276 .0611 .1070
10 .0004 .0025 .0095 .0263 .0571
11 .0001 .0005 .0026 .0090 .0242
12 .0001 .0006 .0024 .0081
13 .0001 .0005 .0021
14 .0001 .0004
15 .0001
16
17
.99 .95 .90 .85 .80 .75 .70 .65 .60
p
n=18
0 .8345 .3972 .1501 .0536 .0180 .0056 .0016 .0004 .0001
1 .1517 .3763 .3002 .1704 .0811 .0338 .0126 .0042 .0012
2 .0130 .1683 .2835 .2556 .1723 .0958 .0458 .0190 .0069
3 .0007 .0473 .1680 .2406 .2297 .1704 .1046 .0547 .0246
4 .0093 .0700 .1592 .2153 .2130 .1681 .1104 .0614
5 .0014 .0218 .0787 .1507 .1988 .2017 .1664 .1146
6 .0002 .0052 .0301 .0816 .1436 .1873 .1941 .1655
7 .0010 .0091 .0350 .0820 .1376 .1792 .1892
8 .0002 .0022 .0120 .0376 .0811 .1327 .1734
9 .0004 .0033 .0139 .0386 .0794 .1284
10 .0001 .0008 .0042 .0149 .0385 .0771
11 .0001 .0010 .0046 .0151 .0374
12 .0002 .0012 .0047 .0145
13 .0002 .0012 .0045
1036
14 .0002 .0011
15 .0002
16
17
18
.99 .95 .90 .85 .80 .75 .70 .65 .60
p
n=19
0 .8262 .3774 .1351 .0456 .0144 .0042 .0011 .0003 .0001
1 .1586 .3774 .2852 .1529 .0685 .0268 .0093 .0029 .0008
2 .0144 .1787 .2852 .2428 .1540 .0803 .0358 .0138 .0046
3 .0008 .0533 .1796 .2428 .2182 .1517 .0869 .0422 .0175
4 .0112 .0798 .1714 .2182 .2023 .1491 .0909 .0467
5 .0018 .0266 .0907 .1636 .2023 .1916 .1468 .0933
6 .0002 .0069 .0374 .0955 .1574 .1916 .1844 .1451
7 .0014 .0122 .0443 .0974 .1525 .1844 .1797
8 .0002 .0032 .0166 .0487 .0981 .1489 .1797
9 .0007 .0051 .0198 .0514 .0980 .1464
10 .0001 .0013 .0066 .0220 .0528 .0976
11 .0003 .0018 .0077 .0233 .0532
12 .0004 .0022 .0083 .0237
13 .0001 .0005 .0024 .0085
14 .0001 .0006 .0024
15 .0001 .0005
16 .0001
17
18
19
.99 .95 .90 .85 .80 .75 .70 .65 .60
p
0 .8179 .3585 .1216 .0388 .0115 .0032 .0008 .0002
1 .1652 .3774 .2702 .1368 .0576 .0211 .0068 .0020 .0005
2 .0159 .1887 .2852 .2293 .1369 .0669 .0278 .0100 .0031
3 .0010 .0596 .1901 .2428 .2054 .1339 .0716 .0323 .0123
1037
n=20
4 .0133 .0898 .1821 .2182 .1897 .1304 .0738 .0350
5 .0022 .0319 .1028 .1746 .2023 .1789 .1272 .0746
6 .0003 .0089 .0454 .1091 .1686 .1916 .1712 .1244
7 .0020 .0160 .0545 .1124 .1643 .1844 .1659
8 .0004 .0046 .0222 .0609 .1144 .1614 .1797
9 .0001 .0011 .0074 .0271 .0654 .1158 .1597
10 .0002 .0020 .0099 .0308 .0686 .1171
11 .0005 .0030 .0120 .0336 .0710
12 .0001 .0008 .0039 .0136 .0355
13 .0002 .0010 .0045 .0146
14 .0002 .0012 .0049
15 .0003 .0013
16 .0003
17
18
19
20
.99 .95 .90 .85 .80 .75 .70 .65 .60
p
Table 3 Critical Values of t Table 3 Critical Values of t
Instructions: The columns of this table are associated with different values of α for a one-tailed and two-tailed test: the rows are associated with the degrees of freedom for the t-statistic.
df
Level of significance for one-tailed test
Level of significance for two-tailed test
.10 .05 .01 .001 .10 .05 .01 .001
1 3.078 6.314 31.821 318.310 6.314 12.706 63.657 636.619
2 1.886 2.920 6.965 22.326 2.920 4.303 9.925 31.598
3 1.638 2.353 4.541 1.213 2.353 3.182 5.841 12.941
4 1.533 2.132 3.747 7.173 2.132 2.776 4.604 8.610
5 1.476 2.015 3.365 5.893 2.015 2.571 4.032 6.859
1038
6 1.440 1.943 3.143 5.208 1.943 2.447 3.707 5.959
7 1.415 1.895 2.998 4.785 1.895 2.365 3.499 5.405
8 1.397 1.860 2.896 4.501 1.860 2.306 3.355 5.041
9 1.383 1.833 2.821 4.297 1.833 2.262 3.250 4.781
10 1.372 1.812 2.764 4.144 1.812 2.228 3.169 4.587
11 1.363 1.796 2.718 4.025 1.796 2.201 3.106 4.437
12 1.356 1.782 2.681 3.930 1.782 2.179 3.055 4.318
13 1.350 1.771 2.650 3.852 1.771 2.160 3.012 4.221
14 1.345 1.761 2.624 3.787 1.761 2.145 2.977 4.140
15 1.341 1.753 2.602 3.733 1.753 2.131 2.947 4.073
16 1.337 1.746 2.583 3.686 1.746 2.120 2.921 4.015
17 1.333 1.740 2.567 3.646 1.740 2.110 2.898 3.965
18 1.330 1.734 2.552 3.610 1.734 2.101 2.878 3.922
19 1.328 1.729 2.639 3.579 1.729 2.093 2.861 3.883
20 1.325 1.725 2.528 3.552 1.725 2.086 2.845 3.850
21 1.323 1.721 2.518 3.527 1.721 2.080 2.831 3.819
22 1.321 1.717 2.508 3.505 1.717 2.074 2.819 3.792
23 1.319 1.714 2.500 3.485 1.714 2.069 2.807 3.767
24 1.318 1.711 2.492 3.467 1.711 2.064 2.797 3.745
25 1.316 1.708 2.485 3.450 1.708 2.060 2.787 3.725
26 1.315 1.706 2.479 3.435 1.706 2.056 2.779 3.707
27 1.314 1.703 2.473 3.421 1.703 2.052 2.771 3.690
28 1.313 1.701 2.467 3.408 1.701 2.048 2.763 3.674
29 1.311 1.699 2.462 3.396 1.699 2.045 2.756 3.659
30 1.310 1.697 2.457 3.385 1.697 2.042 2.750 3.646
40 1.303 1.684 2.423 3.307 1.684 2.021 2.704 3.551
50 1.299 1.676 2.403 3.261 1.676 2.009 2.678 3.496
60 1.296 1.671 2.390 3.232 1.671 2.000 2.660 3.460
1039
90 1.291 1.662 2.368 3.183 1.662 1.987 2.632 3.403
120 1.289 1.658 2.358 3.232 1.658 1.980 2.617 3.373
∞ 1.282 1.645 2.326 3.183 1.645 1.960 2.576 3.291
Table 4 Critical Values of F Table 4 Critical Values of F
Instructions: The columns of this table are associated with the degrees of freedom for the numerator of the the rows are associated with the degrees of freedom for the denominator of the F-ratio. The numbers listed in boldface type are critical vaues for α = .05; the numbers listed in lightface roman type are critical
Degrees of Freedom for Numerator
1 2 3 4 5 6 7 8
1
161.4 199.5 215.7 224.6 230.2 234.0 236.8 238.9
4052 4999 5403 5625 5764 5159 5928 5981
2
18.51 19.00 19.16 19.25 19.30 19.33 19.36 19.37
98.49 99.00 99.17 99.25 99.30 99.33 99.36 99.37
3
10.13 9.55 9.28 9.12 9.01 8.94 8.88 8.84
34.12 30.82 29.46 28.71 28.24 27.91 27.67 27.49
4
7.71 6.94 6.59 6.39 6.26 6.16 6.09 6.04
21.20 18.00 16.69 15.98 15.52 15.21 14.98 14.80
5
6.61 5.79 5.41 5.19 5.05 4.95 4.88 4.82
16.26 13.27 12.06 11.39 10.97 10.67 10.45 10.29
6
5.99 5.14 4.76 4.53 4.39 4.28 4.21 4.15
13.74 10.92 9.78 9.15 8.75 8.47 8.26 8.10
7
5.59 4.74 4.35 4.12 3.97 3.87 3.79 3.73
12.25 9.55 8.45 7.85 7.46 7.19 7.00 6.84
5.32 4.46 4.07 3.84 3.69 3.58 3.50 3.44
1040
Degrees of Freedom for Denominator
8 11.26 8.65 7.59 7.01 6.63 6.37 6.19 6.03
9
5.12 4.26 3.86 3.63 3.48 3.37 3.29 3.23
10.56 8.02 6.99 6.42 6.06 5.80 5.62 5.47
10
4.96 4.10 3.71 3.48 3.33 3.22 3.14 3.07
10.04 7.56 6.55 5.99 5.64 5.39 5.21 5.06
11
4.84 3.98 3.59 3.36 3.20 3.09 3.01 2.95
9.65 7.20 6.22 5.67 5.32 5.07 4.88 4.74
12
4.75 3.89 3.49 3.26 3.11 3.00 2.92 2.85
9.33 6.93 5.95 5.41 5.06 4.82 4.65 4.50
13
4.67 3.80 3.41 3.18 3.02 2.92 2.84 2.77
9.07 6.70 5.74 5.20 4.86 4.62 4.44 4.30
14
4.60 3.74 3.34 3.11 2.96 2.85 2.77 2.70
8.86 6.51 5.56 5.03 4.69 4.46 4.28 4.14
15
4.54 3.68 3.29 3.06 2.90 2.79 2.70 2.64
8.68 6.36 5.42 4.89 4.56 4.32 4.14 4.00
16
4.49 3.63 3.24 3.01 2.85 2.74 2.66 2.59
8.53 6.23 5.29 4.77 4.44 4.2 4.03 3.89
17
4.45 3.59 3.20 2.96 2.81 2.70 2.62 2.55 2.50
8.40 6.11 5.18 4.67 4.34 4.10 3.93 3.79 3.68
18
4.41 3.55 3.16 2.93 2.77 2.66 2.58 2.51 2.46
8.28 6.01 5.09 4.58 4.25 4.01 3.85 3.71 3.60
4.38 3.52 3.13 2.90 2.74 2.63 2.55 2.48 2.43
1041
19 4.38 3.52 3.13 2.90 2.74 2.63 2.55 2.48 2.43
8.18 5.93 5.01 4.50 4.17 3.94 3.77 3.63 3.52
20
4.35 3.49 3.10 2.87 2.71 2.60 2.52 2.45 2.40
8.10 5.85 4.94 4.43 4.10 3.87 3.71 3.56 3.45
21
4.32 3.47 3.07 2.84 2.68 2.57 2.49 2.42 2.37
8.02 5.78 4.87 4.37 4.04 3.81 3.65 3.51 3.40
22
4.30 3.44 3.05 2.82 2.66 2.55 2.47 2.40 2.35
7.94 5.72 4.82 4.31 3.99 3.76 3.59 3.45 3.35
23
4.28 3.42 3.03 2.80 2.64 2.53 2.45 2.38 2.32
7.88 5.66 4.76 4.26 3.94 3.71 3.54 3.41 3.30
24
4.26 3.40 3.01 2.78 2.62 2.51 2.43 2.36 2.30
7.82 5.61 4.72 4.22 3.90 3.67 3.50 3.36 3.25
25
4.24 3.38 2.99 2.76 2.60 2.49 2.41 2.34 2.28
7.77 5.57 4.68 4.18 3.86 3.63 3.46 3.32 3.21
26
4.22 3.37 2.98 2.74 2.59 2.47 2.39 2.32 2.27
7.72 5.53 4.64 4.14 3.82 3.59 3.42 3.29 3.17
27
4.21 3.35 2.96 2.73 2.57 2.46 2.37 2.30 2.25
7.68 5.49 4.60 4.11 3.79 3.56 3.39 3.26 3.14
28
4.20 3.34 2.95 2.71 2.56 2.44 2.36 2.29 2.24
7.64 5.45 4.57 4.07 3.76 3.53 3.36 3.23 3.11
29
4.18 3.33 2.93 2.70 2.54 2.43 2.35 2.28 2.22
7.60 5.42 4.54 4.04 3.73 3.50 3.33 3.20 3.08
1042
30 7.56 5.39 4.51 4.02 3.70 3.47 3.30 3.17 3.06
40
4.08 3.23 2.84 2.61 2.45 2.34 2.25 2.18 2.12
7.31 5.18 4.31 3.83 3.51 3.29 3.12 2.99 2.88
50
4.03 3.18 2.79 2.56 2.40 2.29 2.20 2.13 2.07
7.17 5.06 4.20 3.72 3.41 3.18 3.02 2.88 2.78
60
4.00 3.15 2.76 2.52 2.37 2.25 2.17 2.10 2.04
7.08 4.98 4.13 3.65 3.34 3.12 2.95 2.82 2.72
80
3.96 3.11 2.72 2.48 2.33 2.21 2.12 2.05 1.99
6.96 4.88 4.04 3.56 3.25 3.04 2.87 2.74 2.64
120
3.92 3.07 2.68 2.45 2.29 2.17 2.09 2.02 1.96
6.85 4.79 3.95 3.48 3.17 2.96 2.79 2.66 2.56
∞ 3.84 2.99 2.60 2.37 2.21 2.09 2.01 1.94 1.88
6.64 4.60 3.78 3.32 3.02 2.80 2.64 2.51 2.41
Table 5 The Studentized Range Statistic (qT) Table 5 The Studentized Range Statistic (qT)
Instructions: The columns of this table are associated with number of groups in the analysis ( rows are associated with the degrees of freedom for the error term (denominator) of the The numbers listed in boldface type are critical vaues for α = .05; the numbers listed in lightface roman type are critical values for α = .01.
dferror k = number of groups
2 3 4 5 6 7 8 9 10
1
17.97 26.98 32.82 37.08 40.41 43.12 45.40 47.36 49.07
90.03 135.00 164.30 185.60 202.20 215.80 227.20 237.00 245.60
2
6.08 8.33 9.80 10.88 11.74 12.44 13.03 13.54 13.99
14.04 19.02 22.29 24.72 26.63 28.20 29.53 30.68 31.69
1043
3 4.50 5.91 6.82 7.50 8.04 8.48 8.85 9.18 9.46
8.26 10.62 12.17 13.33 14.24 15.00 15.64 16.20 16.69
4
3.93 5.04 5.76 6.29 6.71 7.05 7.35 7.60 7.83
6.51 8.12 9.17 9.96 10.58 11.10 11.55 11.93 12.27
5
3.64 4.60 5.22 5.67 6.03 6.33 6.58 6.80 6.99
5.70 6.98 7.80 8.42 8.91 9.32 9.67 9.97 10.24
6
3.46 4.34 4.90 5.30 5.63 5.90 6.12 6.32 6.49
5.24 6.33 7.03 7.56 7.97 8.32 8.61 8.87 9.10
7
3.34 4.16 4.68 5.06 5.36 5.61 5.82 6.00 6.16
4.95 5.92 6.54 7.01 7.37 7.68 7.94 8.17 8.37
8
3.26 4.04 4.53 4.89 5.17 5.40 5.60 5.77 5.92
4.75 5.64 6.20 6.62 6.96 7.24 7.47 7.68 7.86
9
3.20 3.95 4.41 4.76 5.02 5.24 5.43 5.59 5.74
4.60 5.43 5.96 6.35 6.66 6.91 7.13 7.33 7.49
10
3.15 3.88 4.33 4.65 4.91 5.12 5.30 5.46 5.60
4.48 5.27 5.77 6.14 6.43 6.67 6.87 7.05 7.21
11
3.11 3.82 4.26 4.57 4.82 5.03 5.20 5.35 5.49
4.39 5.15 5.62 5.97 6.25 6.48 6.67 6.84 6.99
12
3.08 3.77 4.20 4.51 4.75 4.95 5.12 5.27 5.39
4.32 5.05 5.50 5.84 6.10 6.32 6.51 6.67 6.81
13
3.06 3.73 4.15 4.45 4.69 4.88 5.05 5.19 5.32
4.26 4.96 5.40 5.73 5.98 6.19 6.37 6.53 6.67
1044
13 4.26 4.96 5.40 5.73 5.98 6.19 6.37 6.53 6.67
14
3.03 3.70 4.11 4.41 4.64 4.83 4.99 5.13 5.25
4.21 4.89 5.32 5.63 5.88 6.08 6.26 6.41 6.54
15
3.01 3.67 4.08 4.37 4.59 4.78 4.94 5.08 5.20
4.17 4.84 5.25 5.56 5.80 5.99 6.16 6.31 6.44
16
3.00 3.65 4.05 4.33 4.56 4.74 4.90 5.03 5.15
4.13 4.79 5.19 5.49 5.72 5.92 6.08 6.22 6.35
17
2.98 3.63 4.02 4.30 4.52 4.70 4.86 4.99 5.11
4.10 4.74 5.14 5.43 5.66 5.85 6.01 6.15 6.27
18
2.97 3.61 4.00 4.28 4.49 4.67 4.82 4.96 5.07
4.07 4.70 5.09 5.38 5.60 5.79 5.94 6.08 6.20
19
2.96 3.59 3.98 4.25 4.47 4.65 4.79 4.92 5.04
4.05 4.67 5.05 5.33 5.55 5.73 5.89 6.02 6.14
20
2.95 3.58 3.96 4.23 4.45 4.62 4.77 4.90 5.01
4.02 4.64 5.02 5.29 5.51 5.69 5.84 5.97 6.09
22
2.93 3.55 3.93 4.20 4.41 4.58 4.72 4.85 4.96
3.99 4.59 4.96 5.22 5.43 5.61 5.75 5.88 5.99
24
2.92 3.53 3.90 4.17 4.37 4.54 4.68 4.81 4.92
3.96 4.55 4.91 5.17 5.37 5.54 5.69 5.81 5.92
26
2.91 3.51 3.88 4.14 4.35 4.51 4.65 4.77 4.88
3.93 4.51 4.87 5.12 5.32 5.49 5.63 5.75 5.86
2.90 3.50 3.86 4.12 4.32 4.49 4.62 4.74 4.85
1045
28 2.90 3.50 3.86 4.12 4.32 4.49 4.62 4.74 4.85
3.91 4.48 4.83 5.08 5.28 5.44 5.58 5.70 5.80
30
2.89 3.49 3.85 4.10 4.30 4.46 4.60 4.72 4.82
3.89 4.45 4.80 5.05 5.24 5.40 5.54 5.65 5.76
40
2.86 3.44 3.79 4.04 4.23 4.39 4.52 4.63 4.73
3.82 4.37 4.70 4.93 5.11 5.26 5.39 5.50 5.60
50
2.84 3.42 3.76 4.00 4.19 4.34 4.47 4.58 4.68
3.79 4.32 4.63 4.86 5.04 5.19 5.31 5.41 5.51
60
2.83 3.40 3.74 3.98 4.16 4.31 4.44 4.55 4.65
3.76 4.28 4.59 4.82 4.99 5.13 5.25 5.36 5.45
90
2.81 3.37 3.70 3.94 4.12 4.27 4.39 4.50 4.59
3.72 4.23 4.53 4.74 4.91 5.05 5.16 5.26 5.35
120
2.80 3.36 3.68 3.92 4.10 4.24 4.36 4.47 4.56
3.70 4.20 4.50 4.71 4.87 5.01 5.12 5.21 5.30
∞ 2.77 3.31 3.63 3.86 4.03 4.17 4.28 4.39 4.47
3.64 4.12 4.40 4.60 4.76 4.88 4.99 5.08 5.16
Table 6 Critical Values of the Pearson Correlation Coefficient (r)
1046
Table 7 Critical Values of the Spearman Rank-Order Correlation (rs
1047
Table 8 Critical Values of Chi-Square (χ2) Table 8 Critical Values of Chi-Square (χ2)
Instructions: The columns of this table are associated with different values of α; the rows are associated with the degrees of freedom for the chi-square statistic (χ2).
df Level of significance
.10 .05 .01 .001
1 2.71 3.84 6.63 10.83
2 4.61 5.99 9.21 13.82
3 6.25 7.81 11.34 16.27
4 7.78 9.49 13.28 18.46
5 9.24 11.07 15.09 20.52
6 10.64 12.59 16.81 22.46
1048
7 12.02 14.07 18.48 24.32
8 13.36 15.51 20.09 26.12
9 14.68 16.92 21.67 27.88
10 15.99 18.31 23.21 29.59
11 17.28 19.68 24.72 31.26
12 18.55 21.03 26.22 32.91
13 19.81 22.36 27.69 34.53
14 21.06 23.68 29.14 36.12
15 22.31 25.00 30.58 37.70
16 23.54 26.30 32.00 39.29
17 24.77 27.59 33.41 40.75
18 25.99 28.87 34.81 42.31
19 27.20 30.14 36.19 43.82
20 28.41 31.41 37.57 45.32
21 29.62 32.67 38.93 46.80
22 30.81 33.92 40.29 48.27
23 32.01 35.17 41.64 49.73
24 33.20 36.42 42.98 51.18
25 34.38 37.65 44.31 52.62
26 35.56 38.88 45.64 54.05
27 36.74 40.11 46.96 55.48
28 37.92 41.34 48.28 56.89
29 39.09 42.69 49.59 58.30
30 40.26 43.77 50.89 59.70
40 51.81 55.76 63.69 73.40
50 63.17 67.50 76.15 86.66
60 74.40 79.08 88.38 99.61
80 96.58 101.88 112.33 124.84
1049
100 118.50 124.34 135.81 149.45
1050
Appendix: Review of Basic Mathematics
The purpose of this appendix is to provide a review of some of the symbols, terms, and basic math skills needed to perform calculations presented in this book. Many students will know some or all of the material covered in this appendix; however, others may need a more extensive review. It's important for you to know that none of the statistical procedures conducted in this book require higher (i.e., geometry, trigonometry, calculus) math skills.
This appendix covers seven aspects of a basic math review: (1) symbols; (2) terms; (3) addition, subtraction, multiplication, and division; (4) order of operations; (5) fractions, decimals, and percentages; (6) solving simple equations with one unknown; and (7) solving more complex equations with one unknown. Rules, steps, and examples are provided within each of these sections, as well as practice problems designed to help you assess your understanding; the answers to the problems are provided at the end of this appendix. If you have difficulty with this review, we encourage you to refer to your previous math books and get assistance from your instructor or perhaps a tutor.
1051
I. Symbols
Symbol Meaning Example Read
+ addition 4 + 7 4 plus 7
– subtraction 14 – 12 14 minus 12
± addition and subtraction
5 ± 2 5 plus or minus 2
× multiplication 3 × 9 3 times 9
() () (2)(16) 2 times 16
7(8) 7 times 8
/ division 30/5 30 divided by 5
÷ 121 ÷ 12 121 divided by 12
X Y 476 147 476 divided by 147
= equals X = 9 X is equal to 9
≠ not equals X ≠ 14 X is not equal to 14
< less than X < 310 X is less than 310
> greater than X > 6 X is greater than 6
≤ less than or equal to X ≤ 3 X is less than or equal to 3
≥ greater than or equal to X ≥ 5 X is greater than or equal to 5
()2 square (16)2 or 162 16 squared (16 times 16)
√ square root 175 the square root of 175
ǁ absolute value |13| the absolute value of 13
… continuation of pattern 1, 2, 3 … 7, 8, 9
1, 2, 3, 4, 5, 6, 7, 8, 9
1052
II. Terms
Term Meaning Example Read
Sum Result of addition 4 + 7 = 11
The sum of 4 and 7 equals 11
Difference Result of subtraction 14 – 12 = 2
The difference between 14 and 12 equals 2
Product Result of multiplication 3 × 9 = 27
The product of 3 and 9 equals 27
Quotient Result of division 24/12 = 2
The quotient of 24 divided by 12 equals 2
Numerator Top expression in division 84/7 The numerator of 84/7 is 84
Denominator Bottom expression in division 84/7 The denominator of 84/7 is 7
Fraction Number represented with numerator and denominator
2/5, 1/4 2 divided by 5, 1 divided by 4
Percent Part per 100 33% 33 percent
Square Result of multiplying a number by itself 42 = 4 × 4
4 squared equals 4 times 4
Square root Square root of a value is equal to a number that, when multiplied by itself, results in the original value
25 = 5 The square root of 25 equals 5
Equation Mathematical statement consisting of two equal quantities
11 = 8 + 3
11 equals 8 plus 3
1053
III. Addition, Subtraction, Multiplication, Division
1054
A. All Positive Numbers
RULE: Adding, multiplying, or dividing numbers that are all positive results in a positive number.
Examples: 3 + 6 = 9
6 × 4 = 24
30/6 = 5
RULE: Subtracting a smaller positive number from a larger positive number results in a positive number.
Examples: 6 – 4 = 2
24 − 9 = 13
RULE: Subtracting a larger positive number from a smaller positive number results in a negative number.
Examples: 7 – 10 = −3
14 – 22 = −8
1055
B. Adding All Negative Numbers
STEPS: (1) Add the numbers as though they were positive.
(2) Place a negative sign (–) before the sum.
Examples: −1 + −4 + −7 = –(1 + 4 + 7) = –12
–5 + −16 + −9 + −3 = –(5 + 16 + 9 + 3) = –33
1056
C. Adding a Combination of Positive and Negative Numbers
STEPS: (1) Add the positive and negative numbers separately.
(2) Subtract the smaller sum from the larger sum.
(3) If the larger sum is negative, assign the minus sign to the difference.
Examples: 3 + −2 + −5 + 8 = (3 + 8) and –(2 + 5) = 11 – 7 = 4
–10 + −1 + 9 + 7 + −12 = (9 + 7) and –(10 + 1 + 12) = 23 – 16 = −7
16 + −11 + −6 = 16 and –(11 + 6) = 17 – 16 = –1
1057
D. Subtracting Negative Numbers
STEPS: (1) Change the sign of the negative number to positive.
(2) Add the numbers together.
Examples: 2 – (–5) = 2 + 5 = 7
–43 – (–16) – (–7) = −43 + 16 + 7 = –20
12 – (–8) – (–4) – (–17) = 12 + 8 + 4 +17 = 41
1058
E. Multiplying Two Negative Numbers
RULE: The product of two negative numbers is a positive number.
Examples: −4 × −9 = 36
(–7)(–3) = 21
–2(–5) = 10
1059
F. Multiplying a Combination of Positive and Negative Numbers
RULE: A product involving an odd number of negative numbers is a negative number.
Examples: 5 × −3 = –15
8(–9) = –72
(4)(–5)(–3)(–2) = –120
RULE: A product involving an even number of negative numbers is a positive number.
Examples: 4 × −5 × −2 = 40
(–3)(6)(4)(–1) = 72
–6(3)(–7) = 126
1060
G. Dividing Two Negative Numbers
RULE: The quotient of two negative numbers is a positive number.
Examples: − 6 − 2 = 3
–48 ÷ −12 = 4
–36/–4 = 9
1061
H. Dividing One Positive Number and One Negative Number
RULE: The quotient of one positive number and one negative number is a negative number.
Examples: 30/–5 = −6
− 8 − 2 = − 4
–80 ÷ 10 = –8
1062
Practice Problems: Addition, Subtraction, Multiplication, and Division
1. All positive numbers a. 5 + 8 b. 13 + 9 c. 6 × 7 d. (3)(11) e. 3(24) f. 27/3 g. 72 ÷ 9 h. 87 3 i. 9 – 2 j. 16 – 6 k. 6 – 7 l. 15 – 19
2. Adding all negative numbers a. –8 + −3 b. –17 + −5 + −14 c. –23 + −56 + −3 d. –11 + −5 + −2 + −39
3. Adding a combination of positive and negative numbers a. 6 + 3 + −4 + −7 b. 11 + −5 + 10 + −1 c. –9 + 4 + −5 + −17 + 25 d. 8 + 32 + −13 + 19 + 7 + −14
4. Subtracting negative numbers a. 21 – (–7) b. 43 – (–2) – (–19) c. –15 – (–7) – (–5) d. –24 – (–13) – (–16) – (–6)
5. Multiplying two negative numbers a. –2 × −6 b. (–5)(–10) c. (–9)(–7) d. –6(–15)
6. Multiplying combination of positive and negative numbers a. –6 × 9 b. (–5)(–19) c. 13 × −11 d. –7 × −8 e. 2(–4)(–12) f. –5 × −3 × −7 g. (–4)(3)(–12)(–2) h. (2)(2)(–4)(–9)
7. Dividing two negative numbers a. –32/–8 b. –63/–3 c. –57 ÷ −3 d. − 112 − 4
8. Dividing one positive number and one negative number
1063
a. 49/–7 b. –90 ÷ 10 c. − 85 5 d. 154 − 14
1064
IV. Order of Operations
The order of operations is a rule used within algebra to specify the order in which operations should be performed in a given mathematical expression.
1065
First: Parentheses ( ) and Brackets [ ]
Simplify the inside of parentheses and brackets and then remove the parentheses before continuing.
1066
Second: Squares and/or Square Roots
Simplify the square or square root of a number before multiplying, dividing, adding, or subtracting it.
1067
Third: Multiplication and/or Division
Simplify multiplication and division in the order they appear from left to right.
1068
Fourth: Addition and/or Subtraction
Simplify addition and subtraction in the order they appear from left to right.
1069
a. Parentheses () and Brackets [ ]
Examples: ( 1 + 3 ) 2 = ( 4 ) 2 = 16 [ ( 4 − 1 ) 2 − ( 3 − 1 ) 2 ] = [ ( 3 ) 2 − ( 2 ) 2 ] = [ 9 − 4 ] = 5 ( 3 + 2 ) × 10 = 5 × 10 = 50 3 + ( 2 × 10 ) = 3 + 20 = 23 12 / ( 6 − 2 ) − 2 = 2 − 2 = 0 ( 3 + 9 + 12 + 8 ) / 8 = 32 / 8 = 4 [ ( 12 × 4 ) / ( 6 × 2 ) ] / 2 = [ 48 / 12 ] / 2 = 4 / 2 = 2 [ ( 6 − 4 ) ( 9 − 4 ) ] + 11 = [ ( 2 ) ( 5 ) ] + 11 = 10 + 11 = 21 ( 4 − 1 ) 2 + ( 5 − 1 ) 2 = ( 3 ) 2 + ( 4 ) 2 = 9 + 16 = 25 = 5 [ ( 5 − 3 ) + ( 6 − 1 ) ] 7 2 = [ 2 + 5 ] 2 7 = ( 7 ) 2 7 = 49 7 = 7
1070
b. Squares and/or Square Roots
Examples: 2 2 + 5 2 = 4 + 25 = 29 8 − 3 2 = 8 + 9 = 17 26 + 38 = 64 = 8 4 + 49 = 2 + 7 = 9 774 86 = 9 = 3 31 − 88 + 12 = 31 − 100 = 31 − 10 = 21 ( 9 ) ( 3 ) 2 + ( 4 ) 2 = ( 9 ) 9 + 16 = ( 9 ) 5 = 45
c. Multiplication and/or Division
Examples: 12/6 × 2 = 2 × 2 = 4
3 × 6/3 = 18/3 = 6
16/8 + 7 = 2 + 7 = 9
2 × 8 − 10 = 16 − 10 = 6
19 + 21 ÷ 3 = 19 + 7 = 26
14 − 3 × 3 = 14 − 9 = 5
5 + 11 − 8 + 18/6 = 5 + 11 − 8 + 3 = 11
1071
Practice Problems: Order of Operations 1. Parentheses () and brackets [ ]
a. (9 – 2)2
b. (24 ÷ 3)2
c. (4 × 3)2
d. [(2 + 4) (3 + 2)] e. ( 8 − 36 ) ( 6 − 1 ) f. (6 – 2) × 2 g. 6 – (2 × 2) h. (16 – 4) ÷ 2 i. 16 – (4 ÷ 2) j. (8 + 4 + 7 + 5)/6 k. 4(9 + 2 + 6 + 1) l. 3[(2 × 6) + (4 × 2)]]
m. [ ( 3 × 6 ) + ( 4 + 8 ) ] 10 n. 64 16 o. ( 7 − 2 ) [ ( 5 − 3 ) + ( 6 − 1 ) ] 2 7 p. 11 × 2 − 3 × 2 q. ( 11 × 2 ) − ( 3 × 2 )
r. (3 − 1) [(4 − 1)2 + (5 − 1)2] 2. Squares and/or square roots
a. 42 + 62
b. 72 – 13
c. (3)52
d. 92/3 e. 8 16 f. 36 + 121 g. 81 − 7 h. 4 2 − 81 i. 6 − 2 2 + 5 − 2 2 j. 6 − 1 5 + 1 − 3 + 1 2 4
3. Multiplication and/or division a. 3 × 7 × 4 b. 3 × 6 ÷ 9 c. 7 × 12 ÷ 3 d. 33 ÷ 11 × 8 e. 6 × 12 + 3 f. 9 × 7 – 14 g. 36/9 − 10 h. 13 + 3 × 4 i. 17 – 39 ÷ 13 j. 3 + 2 + 4 + 6 ÷ 2 k. 6 – 2 + 4 × 7 l. 16 ÷ 2 + 7 × 3 – 11
1072
V. Fractions, Decimals, Percentages
1073
A. Fraction
DEFINITION: A fraction is a numerical quantity expressed with a numerator and denominator indicating that one number is being divided by another. Examples : 1 4 , 2 3 , − 11 14 , 4 2 , 15 4 , 20 / 8
1074
B. Decimal
DEFINITION: A decimal is the result of dividing a numerator by a denominator represented numerically with decimal places. Examples : 1 4 = .25 , 2 3 = .67 , − 11 14 = − .79 , 4 2 = 2.00 , 15 4 = 3.75 , 20 / 8 = 2.50
1075
C. Percentage
RULE: A percentage is calculated by multiplying a decimal by 100 and placing a percent sign (%) after the result.
Examples: .25 × 100 = 25%, .67 × 100 = 67%, –.79 × 100 = −79%, 2.00 × 100 = 200%
1076
Practice Problems: Fractions, Decimals, and Percentages
1. Convert the following fractions to decimals. a. 1 5 b. 3 4 c. 12 20 d. − 9 27 e. 18 6 f. 9 2 g. 14 8 h. 36/40 i. 75/12
2. Convert the following decimals to percentages. a. .05 b. .68 c. .16 d. –.52 e. 1.36 f. 2.33
3. Convert the following percentages to decimals. a. 1% b. 95% c. 47% d. 99%
RULE: A percentage is converted to a decimal by removing the percent sign and dividing the percentage by 100. Examples: 50%= 50 100 = .50 , 72 % = 72 100 = .72 , 150 % = 150 100 = 1.50
1077
VI. Solving Simple Equations with One Unknown (X)
RULE: Solve for the unknown (X) by isolating X on one side of the equation and then simplifying the steps on the other side of the equation.
1078
A. X has a Value Subtracted from it (Solve for X Using Addition)
Steps: Isolate X by adding a number to both sides of the equation. Example X − 6 = 4 Check: X − 6 = 4 X − 6 + 6 = 4 + 6 10 − 6 = 4 X = 10 4 = 4
1079
B. X has a Value Added to it (Solve for X Using Subtraction)
Steps: Isolate X by subtracting a number from both sides of the equation. E x a m p l e X − 12 = 19 C h e c k : X + 12 = 19 X + 12 − 12 = 19 7 + 12 = 19 X = 7 19 = 19
1080
C. X is Multiplied by a Value (Solve for X Using Division)
Steps: Isolate X by dividing both sides of the equation by a number. E x a m p l e 4 ( X ) = 36 C h e c k : 4 ( X ) = 36 4 ( X ) 4 = 36 4 4 ( 9 ) = 36 X = 9 36 = 36
1081
D. X is Divided by a Value (Solve for X Using Multiplication)
Steps: Isolate X by multiplying both sides of the equation by a number.
1082
Practice Problems: Simple Equations with One Unknown
1. X has a value subtracted from it (solve for X using addition). a. X – 5 = 3 b. X – 13 = 8 c. X – 37 = 79 d. X – 231 = 590 e. X – 1.54 = 3.21 f. X – 29.64 = 105.31
2. X has a value added to it (solve for X using subtraction). a. X + 5 = 9 b. X + 14 = 23 c. X + 53 = 104 d. X + 167 = 331 e. X + .83 = 6.72 f. X + 13.06 = 73.29
3. X is multiplied by a value (solve for X using division). a. 3(X) = 15 b. 6(X) = 42 c. 1.50(X) = 10.50 d. 2.25(X) = 8.46 e. –4(X) = −36 f. – 5(X) = 25
4. X is divided by a value (solve for X using multiplication). a. X 3 = 6 b. X 6 = 10 c. X 11 = 9 d. X 13 = 12 e. X 2.50 = 6.00 f. X 3.88 = 9.50
Examples: X 7 = 4 Check: X 7 = 4 ( 7 ) X 7 = 4 ( 7 ) 28 7 = 4 X = 28 4 = 4
1083
VII. Solving more Complex Equations with One Unknown (X)
RULE: Solve for the unknown (X) by using a combination of simple operations. Below are some possible complex equations.
1084
A. X is Multiplied by a Value and then has a Value Either Added to it or Subtracted from it (Solve for X Using Either Subtraction or Addition and then Division) Example: 2 ( X ) + 7 = 19 2 ( X ) + 7 − 7 = 19 − 7 ( subtract 7from both sides of the equation ) 2 ( X ) = 12 2 ( X ) 2 = 12 2 ( divide both sides of theequation by 2 ) X = 6 Check 2 ( X ) + 7 = 19 2 ( 6 ) + 7 = 19 12 + 7 = 19 19 = 19
1085
B. X is Divided by a Value and then has a Value Either Added to it or Subtracted from it (Solve for X Using Either Addition or Subtraction and then Multiplication) Example: X 6 − 5 = 2 X 6 − 5 + 5 = 2 + 5 ( add 5 to both sides of the equation ) X 6 = 7 6 ( X ) 6 = 7 ( 6 ) ( multiply both sides of the equation by 6 ) X = 42 Check: 42 6 − 5 = 2 7 − 5 = 2 2 = 2
1086
C. X has a Value Added to it or Subtracted from it and is then Multiplied by a Value (Solve for X Using Division and then Either Subtraction or Addition) Example : 4 ( X + 3 ) = 48 4 ( x + 3 ) 4 = 48 4 ( divide both sides of the equation by 4 ) X + 3 = 12 X + 3 − 3 = 12 − 3 ( subtract 3 from both sides of the equation ) X = 9 Check: 4 ( 9 + 3 ) = 48 4 ( 12 ) = 48 48 = 48
1087
D. X has a Value Either Added to it or Subtracted from it and is then Divided by a Value (Solve for X Using Multiplication and then Either Subtraction or Addition) Example: X − 11 9 = 3 ( 9 ) X − 11 9 = 3 ( 9 ) ( multiply both sides of the equation by 4 ) X − 11 = 27 X − 11 + 11 = 27 + 11 ( add 11 to both sides of the equation ) X = 38 Check: 38 − 11 9 = 3 27 9 = 3 3 = 3
1088
Practice Problems: More Complex Equations with One Unknown
1. X is multiplied by a value and then has a value either added to it or subtracted from it. a. 3(X) + 4 = 25 b. 5(X) – 3 = 17 c. 4(X) + 2.50 = 16.50 d. 6 + 2(X) = 36 e. 29 – 4(X) = 5 f. 7.25 – 1.50(X) = 3.47
2. X is divided by a value and then has a value either added to it or subtracted from it. a. X 2 − 4 = 1 b. X 4 + 5 = 11 c. X 3 + 1.75 = 8.25 d. 9 + X 8 = 13 e. 27 − X 5 = 16 f. 2.53 + X 2 = 3.60
3. X has a value added to it or subtracted from it and is then multiplied by a value. a. 3(X + 2) = 21 b. 10(X – 7) = 40 c. 6(X – 6) = 78 d. .50(X + 2) = 14.50 e. 1.34(X – 2.75) = 21.24 f. 4.00(X + 7.44) = 56.52
4. X has a value added to it or subtracted from it and is then divided by a value. a. ( X − 5 ) 2 = 4 b. ( X − 7 ) 3 = 9 c. ( X − 11 ) 14 = 3 d. ( X − 1.75 ) 4 = 11.00 e. ( X + 5.85 ) 3.70 = 12.50 f. ( X + 25.88 ) 7.20 = 34.90
1089
Answers to Problems
1090
Addition, Subtraction, Multiplication, Division
1. All positive numbers a. 13 b. 22 c. 42 d. 33 e. 72 f. 9 g. 8 h. 27 i. 7 j. 10 k. −1 l. −4
2. Adding all negative numbers a. −11 b. −36 c. −82 d. −67
3. Adding a combination of positive and negative numbers a. −2 b. 15 c. −2 d. 39
4. Subtracting negative numbers a. 28 b. 64 c. −3 d. 11
5. Multiplying two negative numbers a. 12 b. 50 c. 63 d. 90
6. Multiplying combination of positive and negative numbers a. −54 b. 95 c. −143 d. −56 e. 96
1091
f. 75 g. -288 h. 144
7. Dividing two negative numbers a. 4 b. 21 c. 19 d. 28
8. Dividing one positive number and one negative number a. −7 b. −9 c. −17 d. −11
1092
Order of Operations
1. Parentheses () and brackets [ ] a. (7)2 = 49 b. (8)2 = 64 c. (12)2 =144 d. (6)(5) = 30 e. ( 5 ) ( 5 ) = 25 = 5 f. 4 × 2 = 8 g. 6 – 4 = 2 h. 12 ÷ 2 = 6 i. 16 – 2 = 14 j. 24/6 = 4 k. 4(18) = 72 l. 3[(12)+(8)] = 3(20) = 60
m. [ 18 + 32 ] 10 = 50 10 = 5 n. 22 − 6 = 16 = 4 o. 16 4 = 4 p. (2)[9 + 16] = (2)[25] = 50 q. ( 5 ) [ 2 + 5 ] 2 7 = ( 5 ) 49 7 = ( 5 ) ( 7 ) = 35
2. Squares and/or square roots a. 16 + 36 = 52 b. 49 – 13 = 36 c. (3)25 = 75 d. 81/3 = 27 e. 8 4 = 2 f. 6 + 11 = 17 g. 9 – 7 = 2 h. 16 – 9 = 7 i. 16 + 9 = 25 = 5 j. ( 5 ) 16 4 = ( 5 ) ( 2 ) = 10
3. Multiplication and/or division a. 84 b. 28 c. 84 d. 24 e. 75 f. 49 g. −6 h. 25
1093
i. 14 j. 12 k. 32 l. 18
1094
Fractions, Decimals, and Percentages
1. Convert the following fractions to decimals. a. .20 b. .75 c. .60 d. –.33 e. 3.00 f. 4.50 g. 1.75 h. .90 i. 6.25
2. Convert the following decimals to percentages. a. 5% b. 68% c. 16% d. −52% e. 136% f. 233%
3. Convert the following percentages to decimals. a. .01 b. .95 c. .47 d. .99
1095
Simple Equations with One Unknown
1. X has a value subtracted from it (solve for X using addition). a. 8 b. 21 c. 116 d. 821 e. 4.75 f. 134.95
2. X has a value added to it (solve for X using subtraction). a. 4 b. 9 c. 51 d. 164 e. 5.89 f. 60.23
3. X is multiplied by a value (solve for X using division). a. 5 b. 7 c. 7.00 d. 3.76 e. 9 f. −5
4. X is divided by a value (solve for X using multiplication). a. 18 b. 60 c. 99 d. 156 e. 15.00 f. 36.86
1096
More Complex Equations with One Unknown
1. X is multiplied by a value and then has a value either added to it or subtracted from it.
a. 7 b. 4 c. 3.50 d. 15 e. 6 f. 2.52
2. X is divided by a value and then has a value either added to it or subtracted from it. a. 10 b. 24 c. 19.50 d. 32 e. 55 f. 12.26
3. X has a value added to it or subtracted from it and is then multiplied by a value. a. 5 b. 11 c. 19 d. 27 e. 18.60 f. 6.69
4. X has a value added to it or subtracted from it and is then divided by a value. a. 13 b. 20 c. 31 d. 45.75 e. 52.10 f. 277.16
1097
G-1
1098
Glossary
Addition rule: combined probability of mutually exclusive outcomes is the sum of the individual probabilities.
Alpha (α): probability of a statistic used to make the decision whether to reject the null hypothesis.
Alternative hypothesis (H1): statistical hypothesis that a hypothesized change, difference, or relationship among groups or variables does exist in the population.
Analysis of variance (ANOVA): family of statistical procedure used to test differences between two or more group means.
Analytical comparisons: comparisons between groups that are part of a larger research design.
ANOVA summary table: table summarizing the calculations of an analysis of variance.
Archival research: research methods involving the use of records or documents of the activities of individuals, groups, or organizations.
Assumption of interval or ratio scale of measurement: assumption that variables are measured at the interval or ratio scale of measurement.
Assumption of Independence of Observations. Assumption of normality:
assumption that scores for a variable are approximately normally distributed in the population.
Assumption of random sampling: assumption that a sample is randomly selected from the population.
Asymmetric distribution: distribution in which the frequencies change in a different manner, moving away in both directions from the most frequently occurring values.
Bar chart: figure that uses bars to represent the frequency or percentage of a sample corresponding to each value of a variable.
Bar graph: figure in which bars are used to represent the mean of the dependent variable for each level of the independent variable.
Between-group variability: differences among the means of groups that comprise an independent variable.
1099
Between-subjects research design: research design in which each research participant appears in only one level or category of the independent variable.
Biased estimate: statistic based on a sample that systematically underestimates or overestimates the population from which the sample was drawn.
Bimodal distribution: distribution where two values occur with the greatest frequency.
Binomial distribution: distribution of probabilities for a binomial variable.
Binomial variable: variable consisting of exactly two categories.
Cell: combination of independent variables within a factorial research design.
Cell mean: mean of the dependent variable for a combination of independent variables.
Central limit theorem: theorem stating that sample means are approximately normally distributed for an infinite number of random samples drawn from a population.
Chi-square goodness-of-fit test: statistic that tests the difference between the distribution of observed frequencies and expected frequencies for a single variable.
Chi-square statistic: statistic that tests the difference between the distribution of observed frequencies and expected frequencies.
Chi-square test of independence: statistic that tests the difference between the distribution of observed frequencies and expected frequencies for two categorical variables.
Cohen's d: estimate of the magnitude of the difference between the means of two groups measured in standard deviation units.
Complex comparisons: analytical comparisons involving more than two groups.
Computational formula: formula that is an algebraic manipulation of a definitional formula designed to minimize the complexity of mathematical calculations.
Confidence interval: range or interval of values with a stated probability of containing an unknown population parameter.
Confidence interval for the mean: interval of values for a variable with a stated probability of containing an unknown population mean.
1100
Confidence limits: lower and upper boundaries of a confidence interval.
Confounding variable: variable related to an independent variable that provides an alternative explanation for the relationship between the independent and dependent variables.
Contingency table: table in which the rows and columns represent the values of categorical variables and the cells contain observed frequencies for combinations of the variables.
Control group: group of participants in an experiment not exposed to the independent variable.
Correlation: mutual or reciprocal relationship between two variables such that systematic changes in the values of one variable are accompanied by systematic changes in the values of another variable.
Correlational statistics: statistics designed to measure relationships between variables.
Covariance: extent to which two variables vary together such that they have shared variance.
Cramér's φ: measure of effect size for the chi-square statistic.
Criterion variable: variable in a correlational research study predicted by a predictor variable.
Critical value: value of a statistic that separates the regions of rejection and nonrejection.
Cumulative percentage: percentage of a sample at or below a particular value of a variable.
Decision rule: rule specifying the values of the statistic that result in the decision to reject or not reject the null hypothesis.
Definitional formula: formula based on the actual or literal definition of a concept.
Degrees of freedom (df): number of values or quantities free to vary when a statistic is used to estimate a parameter.
Dependent variable: variable measured by the researcher.
Descriptive statistics: statistics used to summarize and describe a set of data for a variable.
Directional alternative hypothesis: alternative hypothesis that indicates the direction of the change, difference, or relationship.
Dunnett test:
1101
statistical procedure used to control familywise error when comparing each group with a single reference group.
Error variance: variability that cannot be explained or accounted for.
Expected frequency (fe): frequency of a group expected to occur in a sample of data under the assumption the null hypothesis is true.
Experimental research methods: research methods designed to test causal relationships between variables.
Factorial research design: research design consisting of all possible combinations of two or more independent variables.
Familywise error: probability of making at least one Type I error across a set of comparisons.
Fisher's exact test: chi-square statistic used if any of the expected frequencies in a sample of data are less than 5.
Flat distribution: distribution where the data are spread evenly across the values of a variable.
F-ratio: statistic used to test differences between group means.
Frequency: number of participants in a sample corresponding to a value of a variable.
Frequency distribution table: table summarizing the number and percentage of participants for the different values of a variable.
Frequency polygon: line graph that uses data points to represent the frequency of each value of a variable.
Grouped frequency distribution table: table that groups values of a variable into intervals and provides the frequency and percentage within each interval.
Histogram: figure that uses connected bars to represent the frequency of each value of a variable.
Homogeneity of variance: assumption that the variance of scores for a variable is the same for different populations.
Independent variable: variable manipulated by the researcher.
Inferential statistics: statistical procedures used to test hypotheses and draw conclusions from data collected during research studies.
Interaction effect:
1102
effect of one independent variable on the dependent variable at the different levels of another independent variable.
Interquartile range: range of the middle 50% of the scores in a set of data.
Interval estimation: estimation of a population parameter with a range, or interval, of values within which one has a certain degree of confidence the population parameter falls.
Interval level of measurement: values of variables equally spaced along a numeric continuum.
Linear regression: statistical procedure in which a straight line is fitted to a set of data to best represent the relationship between two variables.
Linear regression equation: mathematical equation that predicts a score on one variable from a score on another variable based on the relationship between two variables.
Linear relationship: relationship between variables appropriately represented by a straight line.
Linear transformation: mathematical transformation of a variable comprised of addition, subtraction, multiplication, or division.
Longitudinal research: research design in which the same information is collected from a research participant over two or more administrations.
Main effect: effect of an independent variable on the dependent variable within a factorial research design.
Marginal mean: mean representing a main effect within a factorial research design.
Mean: the arithmetic average of a set of scores.
Measure of central tendency: a descriptive statistic that is the most typical, common, or frequently occurring value for a variable.
Measure of effect size: statistic that measures the magnitude of the relationship between variables.
Measure of variability: a descriptive statistic of the amount of differences in a set of data for a variable.
Measurement: assignment of categories or numbers to objects or events according to rules.
Median: value of a variable that splits a distribution of scores in half, with the same number of scores above the median as below.
1103
Modality: value or values of a variable that have the highest frequency in a set of data.
Mode: most frequently occurring score or value of a variable in a set of data.
Multimodal distribution: distribution where more than two values have the greatest frequency.
Negative relationship: relationship in which increases in scores for one variable are associated with decreases in scores of another variable.
Negatively skewed distribution: distribution in which the higher frequencies are at the upper end of the distribution, with the tail on the lower (left) end of the distribution.
Nominal level of measurement: values of variables differing in category or type.
Nondirectional alternative hypothesis: alternative hypothesis that does not indicate the direction of the change, difference, or relationship between groups or variables.
Non-experimental research methods: research methods designed to measure naturally occurring relationships between variables.
Nonlinear relationship: relationship between variables not appropriately represented by a straight line.
Nonparametric statistical tests: statistical tests that do not involve the estimation of population parameters or the assumption that scores for a variable are normally distributed in the larger population.
Normal curve table: table containing the percentage of the standard normal distribution associated with different z-scores.
Normal distribution: distribution based on a population of an infinite number of scores calculated from a mathematical formula.
Normally distributed variable: distribution for a variable considered to be unimodal, symmetric, and neither peaked nor flat.
Null hypothesis (H0): statistical hypothesis that a hypothesized change, difference, or relationship among groups or variables does not exist in the population.
Observational research: research methods involving the systematic and objective observation of naturally occurring behavior or events.
Observed frequency (f0):
1104
frequency of a group in the sample of data. One-way ANOVA:
statistical procedure used to test differences between the means of three or more groups comprising a single independent variable.
Ordinal level of measurement: values of variables can be placed in an order relative to the other values.
Outlier: rare, extreme score that lies outside of the range of the majority of scores in a set of data.
Parameter: numeric characteristic of a population.
Parametric statistical tests: statistical tests based on assumption that data from a sample are being used to estimate population parameters and that the distribution of scores for a variable is normally distributed in the larger population.
Peaked distribution: distribution where much of the data are in a small number of values of a variable.
Pearson correlation coefficient: statistic measuring the linear relationship between two variables measured at the interval and/or ratio level of measurement.
Pie chart: figure that uses a circle divided into proportions to represent the percentage of the sample corresponding to each value of a variable.
Planned comparisons: comparisons built into a research design prior to data collection.
Point estimate: single value used to estimate an unknown population parameter.
Population: total number of possible units or elements that could potentially be included in a study.
Population mean: mean associated with an entire population.
Population standard deviation: square root of the population variance that represents the average deviation of a score from the population mean.
Population standard error of the mean σ x ¯ : standard deviation of the sampling distribution of the mean when the population standard deviation (σ) is known.
Population variance: average squared deviation of a score from the population mean.
Positive relationship: relationship in which increases in scores for one variable are associated with increases
1105
in scores of another variable. Positively skewed distribution:
distribution in which the higher frequencies are at the lower end of the distribution, with the tail on the upper (right) end of the distribution.
Predictor variable: variable in a correlational research study used to predict a criterion variable.
Pretest-posttest research design: research design in which data are collected from participants before and after the introduction of an intervention or experimental manipulation.
Probability: likelihood of occurrence of a particular outcome of an event given all possible outcomes.
Quasi-experimental research: research methods comparing naturally formed or preexisting groups.
r squared (r2): percentage of variance in one variable accounted for by another variable.
R2: percentage of variance in the dependent variable associated with differences between groups comprising an independent variable.
Random assignment: the assignment of participants to categories of an independent variable such that each participant has an equal chance of being assigned to each category.
Range: mathematical difference between the lowest and highest scores in a set of data.
Rank: relative position in a graded or evaluated group.
Ratio level of measurement: values of variables equally spaced along a numeric continuum with a true zero point.
Real limits: values of a variable located halfway between the top of one interval and the bottom of the next interval.
Real lower limit: smallest value of a variable in a particular interval of values.
Real upper limit: largest value of a variable in a particular interval of values.
Region of non-rejection: values of a statistic that result in the decision to not reject the null hypothesis.
Region of rejection: values of a statistic that result in the decision to reject the null hypothesis.
Regression: use of a relationship between two or more correlated variables to predict values of one variable from values of other variables.
1106
Research hypothesis: statement regarding an expected or predicted relationship between variables.
Robust: ability of a statistical procedure to withstand moderate violations of the assumption of normality.
Sample: subset or portion of a population.
Sample mean: mean calculated from data for a sample.
Sampling distribution: distribution of statistics for samples randomly drawn from populations.
Sampling distribution of the difference: distribution of all possible values of the difference between two sample means when an infinite number of pairs of samples of size N are randomly selected from two populations.
Sampling distribution of the mean: distribution of sample means for an infinite number of samples of size N randomly selected from the population.
Sampling error: differences between statistics calculated from a sample and statistics pertaining to the population from which the sample is drawn due to random, chance factors.
Scatterplot: graphical display of paired scores on two variables measured at the interval or ratio level of measurement.
Scheffé test: statistical procedure used to control familywise error when conducting all possible simple and complex comparisons in a set of data.
Scientific method: method of investigation that uses the objective and systematic collection and analysis of empirical data to test theories and hypotheses.
Simple comparisons: analytical comparisons between two groups.
Simple effect: effect of one independent variable at one of the levels of another independent variable.
Single-factor research design: research design consisting of one independent variable.
Spearman rank order correlation (rs): statistic measuring the relationship between two variables measured at the ordinal level of measurement.
Standard deviation: square root of the variance that represents the average deviation of a score from the
1107
mean. Standard error of the difference:
standard deviation of the sampling distribution of the difference. Standard error of the mean s x ¯ :
standard deviation of the sampling distribution of the mean when the population standard deviation (s) is not known.
Standard error of the mean: average deviation of a sample mean from the population mean.
Standard normal distribution: normal distribution measured in standard deviation units with a mean equal to 0 and a standard deviation equal to 1.
Standardized distribution: distribution of z-scores for scores in a frequency distribution.
Standardized score: z-score within a standardized distribution.
Statistic: numeric characteristic of a sample.
Statistical hypothesis: statement about an expected outcome or relationship involving population parameters.
Statistical power: probability of rejecting the null hypothesis when the alternative hypothesis is true.
Statistics: branch of mathematics dealing with the collection, analysis, interpretation, and presentation of masses of numerical data.
Student t-distribution: distribution of the t-statistic based on an infinite number of samples of size N randomly drawn from the population.
Survey research: research methods obtaining information directly from a group of people regarding their opinions, beliefs, or behavior.
Symmetric distribution: distribution in which the frequencies change in a similar manner moving away in both directions from the most frequently occurring values.
Symmetry: how the frequencies of values of a variable change in relation to the most common or frequently occurring values.
Theory: set of propositions used to describe or explain a phenomenon.
t-test for dependent means: inferential statistic testing the difference between two means that are based on the same participant or paired participants.
1108
t-test for independent means: inferential statistic that tests the difference between the means of two samples drawn from two populations.
t-test for one mean: statistical procedure testing the difference between a sample mean and a hypothesized population mean when the population standard deviation is not known.
Tukey test: statistical procedure used to control familywise error when conducting all possible simple comparisons between groups.
Two-factor (A × B) research design: factorial research design consisting of two independent variables.
Type I error: error that occurs when the null hypothesis is true but the decision is made to reject the null hypothesis.
Type II error: error that occurs when the alternative hypothesis is true but the decision is made to not reject the null hypothesis.
Unbiased estimate: statistic based on a sample that is equally likely to underestimate or overestimate the population from which the sample was drawn.
Unimodal distribution: distribution where one value occurs with the greatest frequency.
Unplanned comparisons: comparisons not been built into a research design prior to data collection.
Variability: the amount of differences in a distribution of data for a variable.
Variable: property or characteristic of an object, event, or person that can take on different values.
Variance: average squared deviation from the mean.
Within-group variability: variability of scores within groups that comprise an independent variable.
Within-subjects research design: research design in which each participant appears in all levels or categories of the independent variable.
Yates' correction for continuity: chi-square statistic used when any of the expected frequencies in a sample of data are between 5 and 10.
z-score: scores for the standard normal distribution measured in standard deviation units.
z-test for one mean:
1109
statistical procedure testing the difference between a sample mean and a population mean when σ is known.
1110
References
Abar C. C. (2012). Examining the relationship between parenting types and patterns of student alcohol-related behavior during the transition to college. Psychology of Addictive Behaviors, 26, 20–29.
Abeles H. (2009). Are musical instrument gender associations changing? Journal of Research in Music Education, 57, 127–139.
Adams A. E., & Dennis M. (2011). A comparison of actual and perceived problem drinking among driving while intoxicated (DWI) offenders. Journal of Alcohol & Drug Education, 55, 53–69.
Adams J. B. (2005). What makes the grade? Faculty and student perceptions. Teaching of Psychology, 32, 21–24.
Adams R. J. (1987). An evaluation of color preferences in early infancy. Infant Behavior and Development, 10, 143–150.
Ahmadi M., Raiszadeh F., & Helms M. (1997). An examination of the admission criteria for the MBA programs: A case study. Education, 117, 540–546.
Alterovitz S. S., & Mendelsohn G. A. (2009). Partner preferences across the life span: Online dating by older adults. Psychology and Aging, 24, 513–517.
Ambady N., & Rosenthal R. (1993). Half a minute: Predicting teacher evaluations from thin slices of nonverbal behavior and physical attractiveness. Journal of Personality and Social Psychology, 64, 431–441.
American Psychological Association. (2010). Publication manual of the American Psychological Association (6th ed.) Washington, DC: Author.
Anderson C. A., & Dill K. E. (2000). Video games and aggressive thoughts, feelings, and
1111
behavior in the laboratory and in life. Journal of Personality and Social Psychology, 78, 772–790.
Arano K. G. (2006). Credit card usage among students: Evidence from a survey of Fort Hays State University Students. Kansas Policy Review, 28(1), 31–38.
Auwarter A. E., & Aruguete M. S. (2008). Effects of student gender and socioeconomic status on teacher perceptions. Journal of Educational Research, 101, 243–246.
Banerjee S. C., Campo S., & Green K. (2008). Fact or wishful thinking? Biased expectations in “I think I look better when I'm tanned.” American Journal of Health Behavior, 32, 243–252.
Bar M., Neta M., & Linz H. (2006). Very first impressions. Emotion, 6, 269–278.
Barlett C. P., Harris R. J., & Baldassaro R. (2007). Longer you play, the more hostile you feel: Examination of first person shooter video games and aggression during video game play. Aggressive Behavior, 33, 458–466.
Barnett M. A., Quackenbush S. W., & Sinisi C. S. (1996). Factors affecting children's, adolescents', and young adults' perceptions of parental discipline. Journal of Genetic Psychology: Research and Theory on Human Development, 157, 411–424.
Benz J. J., Anderson M. K., & Miller R. L. (2005). Attributions of deception in dating situations. The Psychological Record, 55, 305–314.
Berg E. M., & Lippman L. G. (2001). Does humor in radio advertising affect recognition of novel product brand names? Journal of General Psychology, 128, 194–205.
Beyerstein B. L. (1999). Whence cometh the myth that we only use 10% of our brains? In Sala S. Della, (Ed.), Mind-myths: Exploring popular assumptions about the mind and brain (pp. 3–24). New York: John Wiley.
Bhat R. A., & Kushtagi P. (2006). A re-look at the duration of human pregnancy.
1112
Singapore Medical Journal, 47, 1044–1048.
Borofsky L. A., Kellerman I., Baucom B., Oliver P. H., & Margolin G. (2013). Community violence exposure and adolescents' school engagement and academic achievement over time. Psychology of Violence, 3, 381–395.
Brenner V., & Fox R. A. (1998). Parental discipline and behavior problems in young children. Journal of Genetic Psychology: Research and Theory on Human Development, 159, 251–256.
Busato V. V., Prins F. J., Elshout J. J., & Hamaker C. (2000). Intellectual ability, learning style, personality, achievement motivation and academic success of psychology students in higher education. Personality and Individual Differences, 29, 1057–1068.
Carlton P., & Pyle R. (2007). A program for parents of teens with anorexia nervosa and eating disorder not otherwise specified. International Journal of Psychiatry in Clinical Practice, 11, 9–15.
Cohen J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Hillsdale, NJ: Lawrence Erlbaum.
Cohen J., & Cohen P. (1983). Applied multiple regression/correlation analysis for the behavioral sciences (2nd ed.). Hillsdale, NJ: Lawrence Erlbaum.
Cumming G. (2014). The new statistics: Why and how. Psychological Science, 25, 7–29.
de Bruin A. B., Rikers R. M., & Schmidt H. G. (2007). The effect of self-explanation and prediction on the development of principled understanding of chess in novices. Contemporary Educational Psychology, 32, 188–205.
Detweiler J. B., Bedell B. T., Salovey P., Pronin E., & Rothman A. J. (1999). Message framing and sunscreen use: Gain-framed messages motivate beach-goers. Health Psychology, 18, 189–196.
1113
Devlin J., & Peterson R. (1994). Student perceptions of entry-level employment goals: An international comparison. Journal of Education for Business, 69, 154–158.
Diener E., & Seligman M. E. P. (2002). Very happy people. Psychological Science, 13, 81–84.
Dietz D. K., Cook R. F., Billings D. W., & Hendrickson A. (2009). A web-based mental health program: Reaching parents at work. Journal of Pediatric Psychology, 34, 488–494.
Dietz T. L. (1998). An examination of violence and gender role portrayals in video games: Implications for gender socialization and aggressive behavior. Sex Roles, 38, 425–442.
Doll G. A. (2009). An exploratory study of resistance training and functional ability in older adults. Activities, Adaptation & Aging, 33, 179–190.
Dunn M. J., & Searle J. (2010). Effect of manipulated prestige-car ownership on both sex attractiveness ratings. British Journal of Psychology, 101, 69–80.
Dunnett C. W. (1955). A multiple comparisons procedure for comparing several treatments with a control. Journal of the American Statistical Association, 50, 1096–1121.
DuRant R. H., Treiber F., Getts A., McCloud K., Linder C. W., & Woods E. R. (1996). Comparison of two violence prevention curricula for middle school adolescents. Journal of Adolescent Health, 19, 111–117.
Edmonds E., O'Donoghue C., Spano S., & Algozzine R. F. (2009). Learning when school is out. Journal of Educational Research, 102, 213–221.
Eisenman D. P., Glik D., Ong M., Zhou Q., Tseng C., Long A., … Asch S. (2009). Terrorism-related fear and avoidance behavior in a multiethnic urban population. American Journal of Public Health, 99, 168–174.
1114
Elaad E., Ginton A., & Jungman N. (1992). Detection measures in real-life criminal Guilty Knowledge Test. Journal of Applied Psychology, 77, 757–767.
Elliot K., & Shelley K. (2006). Effects of drugs and alcohol on behavior, job performance, and workplace safety. Journal of Employment Counseling, 43, 130–134.
El-Sheikh M., & Elmore-Staton L. (2007). The alcohol–aggression link: Children's aggression expectancies in marital arguments as a function of the sobriety or intoxication of the arguing couple. Aggressive Behavior, 33, 458–466.
Erdodi L. A. (2012). What makes a test difficult? Exploring the effect of item content on student's performance. Journal of Instructional Psychology, 39, 171–176.
Eysenck H. J. (1982). The biological basis of cross-cultural differences in personality: Blood group antigens. Psychological Reports, 51, 531–540.
Fontaine K. R., Cheskin L. J., & Barofsky I. (1996). Health-related quality of life in obese persons seeking treatment. Journal of Family Practice, 43, 265–270.
Frohlich C. (1994). Baseball: Pitching no-hitters. Chance, 7(3), 24–30.
Gallucci N. T. (1997). An evaluation of the characteristics of undergraduate psychology majors. Psychological Reports, 81, 879–889.
Gardner D. G., Van Dyne L., & Pierce J. L. (2004). The effects of pay level on organization-based self-esteem and performance: A field study. Journal of Occupational Psychology, 77, 307–322.
Gough H. G., & Bradley P. (2005). CPI 260™manual. Mountain View, CA: CPP, Inc.
Grafman J. M., & Cates G. L. (2010). The differential effects of two self-managed math instruction procedures: Cover, copy, and compare versus copy, cover, and compare. Psychology in the Schools, 47, 153–165.
1115
Gupta S. (1992). Season of birth in relation to personality and blood groups. Personality and Individual Differences, 13, 631–633.
Gurm H., & Litaker D. G. (2000). Framing procedural risks to patients: Is 99% safe the same as a risk of 1 in 100? Academic Medicine, 75, 840–842.
Hard S. F., Conway J. M., & Moran A. C. (2006). Faculty and college student beliefs about the frequency of student academic misconduct. Journal of Higher Education, 77, 1058–1080.
Hardoon K. K., Baboushkin H. R., Derevensky J. L., & Gupta R. (2001). Underlying cognitions in the selection of lottery tickets. Journal of Clinical Psychology, 57, 749–763.
Hays W. L. (1988). Statistics (4th ed.). New York: Holt, Rinehart & Winston.
Herman W. E. (1997). The relationship between time to completion and achievement on multiple choice exams. Journal of Research and Development in Education, 30, 113–117.
Herrick R. (2001). The effects of political ambition on legislative behavior: A replication. The Social Science Journal, 38, 469–474.
Hetsroni A. (2000). Choosing a mate in television dating games: The influence of setting, culture, and gender. Sex Roles, 42, 83–106.
Higbee K. L., & Clay S. L. (1998). College students' beliefs in the ten-percent myth. Journal of Psychology, 132, 469–476.
Ho J., & Lindquist M. (2001). Time saved with the use of emergency warning lights and siren while responding to requests for emergency medical aid in a rural environment. Prehospital Emergency Care, 5, 159–162.
Hong J., & Sun Y. (2012). Warm it up with love: The effect of physical coldness on liking
1116
of romance movies. Journal of Consumer Research, 39, 293–306.
Jaccard J., Hand D., Ku L., Richardson K., & Abella R. (1981). Attitudes toward male oral contraceptives: Implications for models of the relationship between beliefs and attitudes. Journal of Applied Social Psychology, 11, 181–191.
James W. (1911). The energies of men. New York: Henry Holt and Co.
John D., Bassett D., Thompson D., Fairbrother J., & Baldwin D. (2009). Effect of using a treadmill workstation on performance of simulated office work tasks. Journal of Physical Activity and Health, 6, 617–624.
Katkowski D. A., & Medsker G. J. (2001). Income and employment of SIOP members in 2000. The Industrial-Organizational Psychologist, 39(1), 21–36.
Kebbell M. R., & Milne R. (1998). Police officers' perceptions of eyewitness performance in forensic investigations. Journal of Social Psychology, 138, 323–330.
Keppel G., Saufley W. H., & Tokunaga H. (1992). Introduction to design and analysis: A student's handbook (2nd ed.). New York: W. H. Freeman.
Keppel G., & Wickens T. D. (2004). Design and analysis: A researcher's handbook (4th ed.). Upper Saddle River, NJ: Pearson Prentice Hall.
Keppel G., & Zedeck S. (1989). Data analysis for research designs: Analysis of variance and multiple regression/correlation approaches. New York: W. H. Freeman.
Khanna C., & Medsker G. J. (2007). 2006 income and employment survey results for the society for industrial and organizational psychology. The Industrial-Organizational Psychologist, 45(1), 17–32.
Kirsh S. J., & Olczak P. V. (2001). Rating comic book violence: Contributions of gender and trait hostility. Social Behavior and Personality, 29, 833–836.
1117
Larimer M. E., Turner A. P., Anderson B. K., Fader J. S., Kilmer J. R., Palmer R. S., & Cronce J. M. (2001). Evaluating a brief alcohol intervention with fraternities. Journal of Studies on Alcohol, 62, 370–380.
Latané B., & Rodin J. (1969). A lady in distress: Inhibiting effects of friends and strangers on bystander intervention. Journal of Experimental Social Psychology, 5, 189–202.
Lavine H., Burgess D., Snyder M., Transue J., Sullivan J. L., Haney B., & Wagner S. H. (1999). Threat, authoritarianism, and voting: An investigation of personality and persuasion. Personality & Social Psychology Bulletin, 25, 337–347.
Lewis B. A., & O'Neill H. K. (2000). Alcohol expectancies and social deficits relating to problem drinking among college students. Addictive Behaviors, 25, 295–299.
Lo S., Wang C., & Fang W. (2005). Physical interpersonal relationships and social anxiety among online game players. Cyberpsychology & Behavior, 8, 15–20.
Loftus G. R. (1996). Psychology will be a much better science when we change the way we analyze data. Current Directions in Psychological Science, 5, 161–171.
Lohse G. L., & Rosen D. L. (2001). Signaling quality and credibility in Yellow Pages advertising: The influence of color and graphics on choice. Journal of Advertising, 30, 73–85.
Ma S. (2008). Paternal race/ethnicity and birth outcomes. American Journal of Public Health, 98, 2285–2292.
Maier M. A., Barchfeld P., Elliot A. J., & Peckrun R. (2009). Context specificity of implicit preferences: The case of human preference for red. Emotion, 9, 734–738.
Matvienko O., & Ahrabi-Fard I. (2010). The effects of a 4-week after-school program on motor skills and fitness of kindergarten and first-grade students. American Journal of Health Promotion, 24, 299–303.
1118
McCrink K., Bloom P., & Santos L. R. (2010). Children's and adults' judgments of equitable resource distributions. Developmental Science, 13, 37–45.
McKee L., Roland E., Coffelt N., Olson A. L., Forehand R., Massari C., … Zens M. S. (2007). Harsh discipline and child problem behaviors: The roles of positive parenting and gender. Journal of Family Violence, 22, 187–196.
McLeish A. C., & Del Ben K. S. (2008). Symptoms of depression and posttraumatic stress disorder in an outpatient population before and after Hurricane Katrina. Depression and Anxiety, 25, 416–421.
McNemar Q. (1969). Psychological statistics (4th ed.). New York: John Wiley.
Michalski D., Kohout J., Wicherski M., & Hart B. (2011). 2009 Doctorate employment survey. Washington, DC: American Psychological Association.
Miranda A., Villaescusa M. I., & Vidal-Abarca E. (1997). Is attribution retraining necessary? Use of self-regulation procedures for enhancing the reading comprehension strategies of children with learning disabilities. Journal of Learning Disabilities, 30, 503–512.
Miyazaki A. D., Langenderfer J., & Sprott D. E. (1999). Government-sponsored lotteries: Exploring purchase and nonpurchase motivations. Psychology & Marketing, 16, 1–20.
Mondin G. W., Morgan W. P., Piering P. N., Stegner A. J., Stotesbery C. L., Trine M. R., & Wu M. (1996). Psychological consequences of exercise deprivation in habitual exercisers. Medicine & Science in Sports & Exercise, 28, 1199–1203.
Nickerson R. S. (2000). Null hypothesis significance testing: A review of an old and continuing controversy. Psychological Methods, 5, 241–301.
O'Hagan J., & Kelly E. (2005). Identifying the most important artists in a historical context: Methods used and initial results. Historical Methods, 38, 118–125.
1119
Osberg J. S., & Di Scala C. (1992). Morbidity among pediatric motor vehicle crash victims: The effectiveness of seat belts. American Journal of Public Health, 82, 422–425.
Pellegrini R. J. (1973). The astrological ‘theory’ of personality: An unbiased test by a biased observer. Journal of Psychology: Interdisciplinary and Applied, 85, 21–28.
Pryor T., & Wiederman M. W. (1998). Personality features and expressed concerns of adolescents with eating disorders. Adolescence, 33, 291–300.
Psunder M. (2009). Future teachers' knowledge and awareness of their role in student misbehaviour. The New Educational Review, 19, 247–262.
Read M. A., & Upington D. (2009). Young children's color preferences in the interior environment. Early Childhood Education Journal, 36, 491–496.
Reed J. M., Marchand-Martella N. E., Martella R. C., & Kolts R. L. (2007). Assessing the effects of the Reading Success Level A Program with fourth-grade students at a Title I elementary school. Education and Treatment of Children, 30, 45–68.
Reichert T., & Lambiase J. (2003). How to get “kissably close”: Examining how advertisers appeal to consumers' sexual needs and desires. Sexuality and Culture, 7, 120–136.
Ridge R. D., & Reber J. S. (2002). ‘I think she's attracted to me’: The effect of men's beliefs on women's behavior in a job interview scenario. Basic and Applied Social Psychology, 24, 1–14.
Riniolo T. C., Koledin M., Drakulic G. M., & Payne R. A. (2003). An archival study of eyewitness memory of the Titanic's final plunge. Journal of General Psychology, 130, 89–95.
Rist C. (2001, May). The physics of … baseball: Unraveling the mystery of why it's so easy to hit a home run. Discover Magazine, pp. 33–34.
Rosenthal R., & Rosnow R. L. (1991). Essentials of behavioral research: Methods and data
1120
analysis. New York: McGraw-Hill.
Rossi J. S. (1997). A case study in the failure of psychology as a cumulative science: The spontaneous recovery of verbal learning. In Harlow L. L. Mulaik S. A., & Steiger J. H. (Eds.), What if there were no significance tests? (pp. 175–197). Mahwah, NJ: Lawrence Erlbaum.
Ruback R. B., & Juieng D. (1997). Territorial defense in parking lots: Retaliation against waiting drivers. Journal of Applied Social Psychology, 27, 821–834.
Rudski J. M., & Edwards A. (2007). Malinowski goes to college: Factors influencing students' use of ritual and superstition. Journal of General Psychology, 134, 389–403.
Scheffé H. (1953). A method for judging all contrasts in the analysis of variance. Biometrika, 40, 87–104.
Schmidt F. L. (1996). Statistical significance testing and cumulative knowledge in psychology: Implications for training of researchers. Psychological Methods, 1, 115–129.
Schuetze P., Lewis A., & DiMartino D. (1999). Relation between time spent in daycare and exploratory behaviors in 9-month-old infants. Infant Behavior & Development, 22, 267–276.
Sebastian R. J., & Bristow D. (2008). Formal or informal? The impact of style of dress and forms of address on business students' perceptions of professors. Journal of Education for Business, 83, 196–201.
Shaw T., & Duys D. K. (2005). Work values of mortuary science students. Career Development Quarterly, 53, 348–352.
Signorielli N., & Staples J. (1997). Television and children's conceptions of nutrition. Health Communication, 9, 289–301.
Skevington S. M., & Tucker C. (1999). Designing response scales for cross-cultural use in
1121
health care: Data from the development of the UK WHOQOL. British Journal of Medical Psychology, 72, 51–61.
Snyder C. R. (1997). Unique invulnerability: A classroom demonstration in estimating personal mortality. Teaching of Psychology, 24, 197–199.
Steffensmeier D., Ulmer J., & Kramer J. (1998). The interaction of race, gender, and age in criminal sentencing: The punishment cost of being young, black, and male. Criminology, 36, 763–798.
Stewart K. D., & Bernhardt P. C. (2010). Comparing millennials to pre-1987 students and with one another. North American Journal of Psychology, 12, 579–602.
Tokunaga H. T. (1985). The effect of bereavement upon death-related attitudes and fears. Omega, 16, 267–280.
Tom G., & Ruiz S. (1997). Everyday low price or sale price. Journal of Psychology: Interdisciplinary and Applied, 131, 401–406.
Travis F., Bonshek A., Butler V., Rainforth M., Alexander C. N., Khare R., & Lipman J. (2005). Can a building's orientation affect the quality of life of the people within? Testing principles of Maharishi Sthapatya Veda. Journal of Social Behavior and Personality, 17, 553–564.
Tuckman B. W. (1996). The relative effectiveness of incentive motivation and prescribed learning strategy in improving college students' course performance. Journal of Experimental Education, 64, 197–210.
Tukey J. W. (1953). The problem of multiple comparisons. Unpublished manuscript, Princeton University.
Walsh L. M., Toma R. B., Tuveson R. V., & Sondhi L. (1990). Color preference and food choice among children. Journal of Psychology, 124, 645–653.
1122
Warner B., & Rutledge J. (1999). Checking the Chips Ahoy! guarantee. Chance, 12(1), 10–14.
Watson J. C., & Bedard D. L. (2006). Clients' emotional processing in psychotherapy: A comparison between cognitive-behavioral and process-experiential therapies. Journal of Consulting and Clinical Psychology, 74, 152–159.
Welch B. L. (1938). The significance of the difference between two means when the population variances are unequal. Biometrika, 29, 350–362.
Williams V. A., Johnson C. E., & Danhauer J. L. (2009). Hearing aid outcomes: Effects of gender and experience on patients' use and satisfaction. Journal of the American Academy of Audiology, 20, 422–432.
Woolley J. D., & Cox V. (2007). Development of beliefs about storybook reality. Developmental Science, 10, 681–693.
Zentner M. R. (2001). Preferences for colors and color-emotion combinations in early childhood. Developmental Science, 4, 389–398.
Zimmerman D.W. (2004). A note on preliminary test of equality of variances. British Journal of Mathematical and Statistical Psychology, 57, 173–181.
1123
Index
Accounting for variance, 404 Addition rule, 176, 204, G-1 Ahrabi-Fard, Iradge, 332 Alcohol-related behavior study, 16 Alpha
and decision about null hypothesis in testing one sample mean, 250 defined, 186, 205, G-1 setting, for null hypothesis, 186–187, 197–199 and Type I error, 383, 389–390 and Type II error, 393–394
Alternative hypothesis correlation between two variables, 582–583 defined, 185, 205, G-1 directionality, and decision about null hypothesis in testing one sample mean, 250–253 directionality of, 199–201 stating, in hypothesis testing, 184–186 See also Directional alternative hypothesis
American Psychological Association (APA) abbreviation style of measures of central tendency, 74 Cohen lifetime achievement award, 405 communicating findings of research, 17 guidelines for creating figures, 40 (table)
Analysis of simple effects, 531–533 Analysis of variance (ANOVA)
between-group and within-group variability, 426–428 correlational statistics and correlational research, 621 defined, 429, 465, G-1 and F-ratio, 429–430 one-way. See One-way ANOVA review of testing the difference between two sample means, 425–428 for two-factor research design, 506–508 two-way. See Two-way ANOVA
Analytical comparisons within one-way ANOVA, 425–464 between-group and within-group variance in, 453–454 conducting planned comparisons, 454–458. See also Planned comparisons
1124
conducting unplanned comparisons, 458–463. See also Unplanned comparisons defined, 453, 465, G-1 planned vs. unplanned comparisons, 453 using SPSS, 470–471
ANOVA. See Analysis of variance (ANOVA) ANOVA summary table
defined, 439, 465, G-1 one-way ANOVA, 439, 440 (table) two-way ANOVA, 520, 521 (table)
“a priori” comparisons, 453 Archer, Dane, xvi Archival research, 14–15, 19, G-1 Assumption of homogeneity of variance, of t-test for independent means, 329–330 Assumption of interval or ratio scale of measurement, of z-test for one mean, 231, G- 1 Assumption of normality
defined, of z-test for one mean, 231, G-1 of t-test for independent means, 329–330
Assumption of random sampling, of z-test for one mean, 231, G-1 Assumptions
of chi-square statistic, 654–655 of t-test for dependent means, 350 of t-test for independent means, 327–330, 338 of z-test for one mean, 231, 246
Asymmetric distributions, 43–44 defined, 43, 47, G-1 and measures of central tendency, 88–90, 92, 94 (figure)
Authoritarianism, 487, 621 Average deviation from the mean, 113 A x B research design, 502
Bar charts, 35–36, 36 (figure), 46, G-1 Bar graph, 314, G-1 Bar, Moshe, 571 Bell-shaped distribution, 44, 140 Bernhardt, Paul, 649 Best-fit line, 603 Beta, and Type II error, 385 Between-group variability
defined, 396, G-1 increasing to control Type II error, 396, 397 (table)
1125
in two-factor research design, 507–508 and within-group variability in ANOVA, 426–428 and within-group variance in analytical comparisons, 453–454
Between-subjects research designs, 338, 353, G-1 Beyerstein, Barry, 70 Biased estimate, in variance, 120, G-1 Bimodal distribution
defined, 42, 43 (figure), 47, G-1 and measures of central tendency, 87, 91, 92, 94 (figure)
Binomial distributions applying probability to, 178–180 defined, 178, 204, G-1 probabilities table, 187, T-7-T-18
Binomial variable, 178, G-1
Calculators, xix California Psychological Inventory (CPI), 649–651 Causality, and correlation, 621–622 Cause-and-effect relationships, 620 Cell, 490, G-1 Cell means, 490, G-1 Central limit theorem
and assumptions of z-test for one mean, 231 defined, 219, 254, G-1 in sampling distribution of the mean, 219
Central tendency understanding, 73–74 See also Measures of central tendency
Chi-square, 648–720 example: color preference study, 669–682, 683–684 (table) example: personality study, 648–654, 657–669 parametric and nonparametric statistical tests, 682–686 statistic. See Chi-square statistic using SPSS, 688–695
Chi-square goodness-of-fit test (personality study), 655–669 alpha set, 658 alternative hypotheses statement, 657 conclusion drawn, 663–664 critical value of chi-square statistic identified, 658, T-30-T-31 decision rule stated, 659 defined, 655, 687, G-1 degrees of freedom calculated, 657–658, 688
1126
expected frequencies calculated, 659, 688 level of significance determined, 662 measures of effect size calculated, 662–663 null hypotheses rejection decision, 661 null hypotheses statement and decision, 657–663 result of analysis related to research hypothesis, 664 statistic calculated: chi-square, 659–660 summary, 665–666 (table) using SPSS, 688–690
Chi-square goodness-of-fit test with unequal hypothesized proportions (personality study), 666–669
expected frequencies calculated, 667–668 statistic calculated: chi-square, 668–669 using SPSS, 691–692
Chi-square statistic, 651–655 assumptions underlying, 654–655 characteristics of distribution of, 654 defined, 653, 687, G-2 formula for, 659, 688 as nonparametric statistical test, 685
Chi-square test of independence (color preference study), 674–682 alpha set, 676 alternative hypotheses statement, 675 conclusion drawn, 679–680 critical value of chi-square statistic identified, 676, T-30-T-31 decision rule stated, 676 defined, 674, 687, G-2 degrees of freedom calculated, 675–676, 688 expected frequencies calculated, 676–677, 688 level of significance determined, 678–679 measures of effect size calculated, 679 null hypotheses rejection decision, 678 null hypotheses statement and decision, 675–679 result of analysis related to research hypothesis, 680–681 statistic calculated: chi-square, 676–678 summary, 682, 683–684 (table) using SPSS, 692–695
Clay, Samuel, 71 Cohen, Jacob, 405, 579 Cohen's d
calculation of, 407–408, 411 defined, 407, 411, G-2
1127
guidelines to interpret, 408, 443 Communication of findings, 17 Comparisons
analytical (planned and unplanned). See Analytical comparisons within one- way ANOVA simple and complex, 461
Complex comparisons, 461, G-2 Computational formula
defined, 116, 127, G-2 for standard deviation, 122, 123 (table) for variance, 116–119
Conclusions, drawing of, 16 based on null hypothesis rejection, 192 inappropriate, from figures, 39–41 and measures of central tendency, 93–95 and measures of variability, 126 transforming scores into z-scores, 156
Confidence interval 90%, 293 95%, 277–280, 293 99%, 293 defined, 275, 299, G-2
Confidence interval for the mean defined, 275, 299, G-2 example: salary survey study, 271–273, 300–302 factors affecting width of, 286–295 interval estimation and hypothesis testing, 295–298 introduction to, 271, 273–276 using sampling distribution of the mean to create, 275–276
Confidence interval for the mean, population standard deviation known (SAT scores), 284–286, 300
calculate confidence interval and confidence limits, 284–286 conclusion drawn, 286 critical value of z identified, 285–286 level of confidence stated, 284 population standard error of the mean calculated, 285 summary, 286, 287 (table)
Confidence interval for the mean, population standard deviation not known (salary survey study), 276–283, 299
calculate confidence interval and confidence limits, 277–279 conclusion drawn, 279–281
1128
critical value of t identified, 278–279 degrees of freedom calculated, 278 level of confidence stated, 276–277 probability of confidence interval or probability of population mean, 280–281 standard error of the mean calculated, 277–278 summary, 282–283, 282 (table) using SPSS, 300–302
Confidence limits calculation of, 277–279, 284–286 defined, 279, 299, G-2
Confounding variable, 13, 19, G-2 Contingency table, 671–672
defined, 671, 687, G-2 Control group, 13, 19, G-2 Correlation, 570–647
and causality, 621–622 correlational statistics vs. correlational research, 617, 620–621 defined, 574, 622, G-2 example: first impression study, 570–572 introduction to, 574–581 linear regression and prediction, 600–610 Pearson correlation coefficient. See Pearson correlation coefficient relationship between variables described, 574–578 relationship between variables measured, 578–580 Spearman rank-order correlation. See Spearman rank-order correlation
Correlational research methods, 14 Correlational statistics
compared to correlational research, 617, 620–621 defined, 578, 623, G-2
Covariance defined, 580, 623, G-2 in Pearson correlation coefficient, 580, 585, 596
Cramer's V, 662 Cramer's Φ (phi)
defined, 662, G-2 formula for chi-square goodness-of-fit test, 662–663, 688 formula for chi-square test of independence, 679, 688
Criterion variable, 622, 623, G-2 Critical values
of chi-square statistic, T-30–T-31 defined, 188, 205, G-2 example, 188, 197
1129
of F-ratio, 433, 434 (table), 435, T-21–T-23 of f-statistic, T-21–T-23 of Pearson correlation coefficient, 584, T-27–T-28 of Spearman rank-order correlation, 614, T-29–T-30 of t-statistic, 242, 278–279, T-19–T-20
Cumming, Geoff, 296 Cumulative percentages, 32–34
defined, 34, G-2
Data analysis, 15–16 Data coding errors, 27 Data collection, 8–15
determination of levels of measurement, 9–12 methods of collection, 12–15 samples from populations, 9
Data entry errors, 27 Data, examination of, 24–69
distribution descriptions, 42–45 reasons for, 25–28 using figures, 34–41 using SPSS, 47–50 using tables, 28–34
de Bruin, Anique, 421 Decision making errors, 195–196 Decision rule, 190–191, 205, G-2 Definitional formula
defined, 116, 127, G-2 for population standard deviation, 125 for standard deviation, 120–122 for variance, 114–116
Degrees of freedom, 240, 254, G-2 Dependent means, 347 Dependent variable
defined, 5, 19, G-2 example of, 6t and measurement of variables, 11t renamed in correlational research, 622
Descriptive statistics defined, 16, 19, G-2 measures of central tendency as, 70 in reading skills study, 216–218 two-factor research design analysis, 502–506
1130
Devlin, J. Stuart, 611 Dietz, Diane, 341 Difference between two sample means. See Testing the difference between two sample means Directional alternative hypothesis
defined, 201, 202 (figure), 205, G-2 in testing one sample mean, 222, 241, 250–253 using to control Type II error, 394–396
Directional research hypothesis, 7t Direction of the relationship between variables, 576–577, 579 Distribution-free tests, 685 Distributions
of chi-square statistics, characteristics of, 654 of F-ratio, characteristics of, 429–430 of frequencies of groups, and chi-square statistic, 651–652 mean as a balancing point, 83–84 See also Normal distributions; Sampling distribution; Standard normal distribution
Distributions, in examining data, 42–45, 46–47 modality, 42 symmetry, 42–44 variability, 44, 45 (figure)
Drawing conclusions. See Conclusions, drawing of Dunnett test, 461, 466, G-2
Effect size. See Measures of effect size The Energies of Men (James), 71 Errors in hypothesis testing, 377–400, 410–420
example: criminal trials, 377–381 Type I error. See Type I error Type II error. See Type II error
Error variance, 398, G-2 Estimation of population mean, 271–311
example: salary survey study, 271–273 interval estimation and hypothesis testing, 295–298 See also Confidence interval for the mean
Events, 176 Examples
business student study, 611–617, 618 (table) chess study, 421–424, 431–446, 452–464, 459 (table), 620 color preference study, 669–682, 683–684 (table) criminal trials and GKT study, 377–381
1131
exercise deprivation study, 489–501 first impression study, 570–574, 582–600, 603–606 frequency rating study, 104–106 grade point averages, 32, 33 (table) income and skewed distribution, 89–90 kids' motor skills and fitness, 331–332 legislative behavior of politicians, 158–164 lottery study, 24–25, 28–30, 51–52 parking lot study, 313–317, 319–331, 389–390, 393–398, 401–408, 447–452 personality study, 648–654, 657–669 reading skills study, 214–218, 221–231 salary survey study, 271–282, 286–295 SAT scores, 139, 152–156, 284–286 Super Bowl, 181–184, 185–192 ten-percent myth study, 70–73 unique invulnerability study, 234–236, 239–247, 255–257 video games and aggressive behavior, 15 voting message study, 486–488, 502–531, 621 web-based intervention study, 341–342
Expected frequencies, 652–653 defined, 652, 687, G-3 minimum, 655
Experimental research methods compared to correlational research, 620, 621 defined, 12, 19, G-3 example of combining with non-experimental research, 15
Factorial research designs, 488–501 compared to single-factor designs, 498–500 defined, 489, 534, G-3 relationship between main effects and interaction effects, 495–498 strategy for analyzing, 498, 499 (figure) testing interaction effects, 489–493 testing main effects, 494–495
Familywise error defined, 460, 466, G-3 methods of controlling, 461–462 Tukey test, 462–463 in unplanned comparisons, 460–461
Figures, examining data using, 34–41 bar charts, 35–36, 36 (figure) frequency polygons, 37–39, 38 (figure)
1132
guidelines for creating figures, 40 (table) histograms, 37, 38 (figure) pie charts, 36, 37 (figure)
Fisher's exact test, 655, 687, G-3 Flat distributions, 44, 47, G-3 F-ratio
for analytical comparison, 453–454, 467 defined, 429, 465, G-3 distribution characteristics for ANOVA, 429–430 relationship with t-test, 452, 467 in two-factor research design, 507–508
Frequency, 28, G-3 Frequency distributions, standardizing, 157–166 Frequency distribution tables, 28–30
calculating mean from, 82–83 defined, 28, 46, G-3 using SPSS, 51–52
Frequency polygons, 37–39, 38 (figure) defined, 37, 46, G-3
Gender, as example of nominal variable, 9 Goodness-of-fit. See Chi-square Gosset, W. S., 236 Gough, Harrison, 649 Grouped frequency distribution tables, 31–34
defined, 32, 46, G-3 guidelines for creation of, 34 (table)
Guilty knowledge test (GKT), 379–381
Hardoon, Karen, 25 Harrington, David, xvi Herman, William, 607 Herrick, Rebekah, 158 Higbee, Kenneth, 71 “Highly significant,” inappropriateness of, 402–404 Histograms, 37, 38 (figure), 46, G-3 Homogeneity of variance, 329, G-3 Hypothesis of research. See Research hypothesis Hypothesis testing, 184–213
concerns about using, 296–297 deciding about null hypothesis, 186–191 drawing conclusion, 192
1133
errors in. See Errors in hypothesis testing errors in decision making, 195–196 factors influencing decision about null hypothesis, 196–201 and interval estimation, 295–298 “proof” in, 195 result of analysis related to research hypothesis, 192 stating null and alternative hypotheses, 184–186 steps in, 184, 192, 193 (table) vs. criminal trials, 377–381
IBM SPSS software. See SPSS software Independence of observations, 654 Independent means, 322–323 Independent variable
analysis of studies consisting of two, 486 defined, 5, 19, G-3 example of, 6t and measurement of variables, 11t renamed in correlational research, 622
Inferential statistics chi-square tests, 655, 674 defined, 16, 19, G-3 interpreting, and influence of sample size, 401 one-way ANOVA, 431 Pearson correlation coefficient, 582 Spearman rank-order correlation coefficient, 612 testing one sample mean, 218, 221, 239 testing the difference between two sample means, 319, 331, 340 two-way ANOVA, 508
Infinity, 140 Interaction effects
defined, 489, 534, G-3 relationship between main effects and, 495–498 significant A x B, 531–533 testing: example of interaction effect, 489–492 testing: example of no interaction effect, 492–493
Interquartile range, 110–112 calculation of, 111 defined, 110, 127, G-3 strengths and weaknesses of, 112
Interval estimation alternative to hypothesis testing, 297–298
1134
defined, 271, 299, G-3 and hypothesis testing, 295–298
Interval level of measurement assumption of interval or ratio scale of measurement, 231 defined, 10, 19, G-3 displaying variables with histograms and frequency polygons, 36–39 example of, 12 Pearson correlation coefficient and, 578
James, William, 71 Juieng, Daniel, 312
Keppel, Geoff, xvi Khanna, Charu, 271 Kruskal-Wallis test, 686 (table)
Lavine, Howard, 486 Level of confidence
probability vs. precision, 293–295 relation to sample size, 295 relation to width of confidence interval, 291–293
Levels of measurement, 9–12 Linear regression, 600–610
defined, 603, 623, G-3 equation, 603 prediction and, 600–601 using SPSS, 628–629
Linear regression equation calculating, 603–605, 606 (table), 607–608 defined, 603, 623, 625, G-3 drawing, 605–607, 608–609
Linear relationship, 574, 622, G-3 Linear transformation, 164, G-3 Line, formula for, 603 Loftus, Geoffrey, 296 Longitudinal research, 341, 353, G-3
Maier, Markus, 669 Main effects
defined, 494, 534, G-3 relationship between interaction effects and, 495–498 testing, 494–495
1135
Mann-Whitney U, 330, 686 (table) Marginal means, 494, G-3 Mathematics review, A-1–A-18
addition, subtraction, multiplication, division, A-2–A-5 answers to problems, A-15–A-18 fractions, decimals, percentages, A-9–A-10 order of operations, A-6–A-8 solving complex equations with one unknown, A-13–A-14 solving simple equations with one unknown, A-11–A-12 symbols, A-1 terms, A-2
Matvienko, Oksana, 332 Mean
as balancing point in a distribution, 83–84 calculating, 80–82, 97 calculating from frequency distribution tables, 82–83, 97 compared to median and mode, 86, 93 (table) defined, 80, 96, G-4 example of descriptive statistic, 16 notational system in two-factor research design, 503–504 sample mean vs. population mean, 86 strengths and weaknesses of, 91–92, 93 (table) and z-scores: distance from, in normal distributions, 143–144
Mean squares, in two-way ANOVA, 513, 516–519 Measurement, 9, G-4 Measure of variability of sample means, in calculation of confidence interval, 277–278 Measures of central tendency, 70–103
comparison of mode, median, and mean, 86–93 defined, 73, 96, G-4 drawing conclusions, 93–95 mean, 80–86 median, 74–80 mode, 74 ten-percent myth, 70–73 understanding central tendency, 73–74 using SPSS, 129–131
Measures of effect size, 401–420 Cohen's d, for difference between two sample means, 407–408 Cramer's Φ (phi), 662–663, 679 defined, 405, 410, G-4 effect size as variance accounted for, 404
1136
“highly significant” inappropriateness, 402–404 interpreting, 405–406 presenting, 406 reasons for calculating, 406–407 r squared, for difference between two sample means, 405, 411 R squared, for one-way ANOVA, 442–443, 451 sample size influence, 401
Measures of variability, 104–138 defined, 107, 127, G-4 drawing conclusions, 126 interquartile range, 110–112 population standard deviation, 124–125 population variance, 124 range, 108–110 for samples vs. populations, 125 “sometimes” in frequency rating study, 104–106 standard deviation, 120–123 understanding variability, 106–108 using SPSS, 129–131 variance, 112–120
Median calculating when multiple scores have median value, 78–80 calculating, with even number of scores, 76–78, 96 calculating, with odd number of scores, 74–76, 96 compared to mode and mean, 86, 93 (table) defined, 74, 96, G-4 strengths and weaknesses of, 88–91, 93 (table)
Medsker, Gina, 271 Minimum expected frequencies, 655 Modality, 42, 43 (figure), 47, G-4 Modal score (mode), 74 Mode
compared to median and mean, 86, 93 (table) defined, 74, 96, G-4 determination of, 74 strengths and weaknesses of, 87–88, 93 (table)
Mondlin, Gregory, 489 Multimodal distribution, 42, 43 (figure), 47, G-4 Mutually exclusive outcomes, 176
Nature of the relationship between variables, 574–576 Negatively skewed distribution, 44, G-4
1137
Negative relationship, 576, 623, G-4 Nickerson, Raymond, 298 Nominal level of measurement
chi-square statistic and, 648 defined, 9, 19, G-4 displaying variables with bar charts and pie charts, 35–36 example of, 12
Nondirectional alternative hypothesis, 201, 202 (figure), 205, G-4 Nondirectional research hypothesis, 7t Non-experimental research methods, 14–15
defined, 14, 19, G-4 example of combining with experimental research, 15
Nonlinear relationship, 575, 623, G-4 Nonparametric statistical tests, 682–686
compared to parametric statistical tests, 682–686 defined, 684, 687, G-4 examples of, 685–686, 686 (table) reasons for using, 685
Nonsignificance, 227, 297 Normal curve table
defined, 144, 167, G-4 in standard normal distribution, 144, 145 (table), T-2–T-6
Normal distributions, 139–174 applying probability to, 177–178 applying z-scores to, 150–157 defined, 140, 166, G-4 features of, 140 importance of, 140–141 interpreting standardized scores, 161–162 interpreting z-scores, 143–150, 153–156 standardizing frequency distributions, 157–166 standard normal distribution, 142–150. See also Standard normal distribution transforming into standard normal distribution, 151 See also z-scores
Normally distributed variable, 44, 47, G-4 Null hypothesis
alpha relationship to, 197–199 correlation between two variables, 582–583 defined, 185, 205, G-4 directionality of alternative hypothesis, 199–201 factors affecting decision about, in testing one sample mean, 248–253
1138
making decision about, 186–191 risk in not rejecting (Type II error), 385–387 risk in rejecting (Type I error), 381–385 sample size as factor influencing, 196–197 stating, in hypothesis testing, 184–185
Observational research defined, 14, 19, G-4 example, parking lot study, 313
Observations, independence of, 654 Observed frequencies, 652–653
defined, 652, 687, G-4 independence of observations, 654
One-tailed alternative hypothesis, 201, 205 vs. two-tailed, in testing one sample mean, 251–253
One-way ANOVA, 421–485 analytical comparisons within, 425–464. See also Analytical comparisons within one-way ANOVA defined, 429, 465, G-4 example: chess study, 421–424, 431–446, 452–464 example: parking lot study, 447–452 introduction to ANOVA, and F-ratio, 429–430 introduction to ANOVA, review of testing the difference between two sample means, 425–428
One-way ANOVA (chess study), 431–446 alpha set, 433 alternative hypotheses statement, 431 between-group variance calculation, 435–438, 467 conclusion drawn, 443–444 create ANOVA summary table, 439, 440 (table) critical value of F-ratio identified, 433, 434 (table), 435, T-21–T-23 decision rule stated, 435 degrees of freedom calculated, 432–433, 466 level of significance determined, 441–442 measure of effect size calculated, 442–443, 467 null hypotheses rejection decision, 441 null hypotheses statement and decision, 431–443 result of analysis related to research hypothesis, 444 statistic calculated: F-ratio for one-way ANOVA, 435–439, 467 summary, 444, 445–446 (table) total sample mean calculation, 437, 467 using SPSS, 468–469
1139
within-group variance calculation, 438–439, 467 One-way ANOVA (parking lot study), 447–452
alpha set, 448 alternative hypotheses statement, 448 calculate between-group variance, 449 calculate within-group variance, 450 conclusion drawn, 452 critical value of F-ratio identified, 448–449 decision rule stated, 449 degrees of freedom calculated, 448 level of significance determined, 451 measure of effect size calculated, 451 null hypotheses rejection decision, 450 null hypotheses statement and decision, 448–451 relationship between t-test and F-ratio, 452, 467 statistic calculated: F-ratio for one-way ANOVA, 449–450 summary, (table)
Ordinal level of measurement defined, 9, 19, G-4 displaying variables with bar charts and pie charts, 35–36 example of, 12 Spearman rank-order correlation and, 610
Outcomes, 176 Outliers
defined, identification of, 27, 46, G-5 range and, 110, 112
Paired means. See Testing the difference between paired means Parameter
defined, sample mean vs. population mean, 86, 96, G-5 and measures of variability for samples vs. populations, 125
Parametric statistical tests, 682–686 compared to nonparametric statistical tests, 682–686 defined, 684, 687, G-5 examples of, 685, 686 (table)
Peaked distributions, 44, 47, G-5 Pearson correlation coefficient
compared to Spearman rank-order correlation, 613 defined, 578, 623, G-5 example: first impression study, 570–574, 582–600, 603–606 example of inferential statistic, 16, 582 formulas, computational, 595–600, 601 (table), 625
1140
formulas, conceptual, 579–580 formulas, definitional, 589, 592, 593–594 (table), 595, 600, 624 guidelines for interpreting values of, 579, 591–592 introduction to, 578–579
Pearson correlation coefficient (first impression study), 582–600 alpha set, 583–584 alternative hypotheses statement, 582–583 conclusion drawn, 592 critical value r identified, 584, T-27–T-28 decision rule stated, 584 degrees of freedom calculated, 583, 624 level of significance determined, 590–591 measure of effect size calculated, 591–592, 624 null hypotheses rejection decision, 590 null hypotheses statement and decision, 582–592 result of analysis related to research hypothesis, 592 statistic calculated: r, 585–589 summary, 593–594 (table), 601 (table) using SPSS, 626–627
Pearson, Karl, 578, 653 Pearson's chi-square statistic, 653 Percentage, calculation of, 29, 47 Perfect relationship, 577–578, 579 Peterson, Robin, 611 Pie charts, 36, 37 (figure), 46, G-5 Planned comparisons (chess study), 454–458
alpha set, 455 alternative hypotheses statement, 455 conclusion drawn, 458 critical value of F-ratio identified, 455 decision rule stated, 456 defined, 453, 466, G-5 degrees of freedom calculated, 455 level of significance determined, 457 measure of effect size calculated, 457–458, 467 null hypotheses rejection decision, 457 null hypotheses statement and decision, 455–458 result of analysis related to research hypothesis, 458 statistic calculated: F-ratio for analytical comparisons, 456, 467 summary, 459 (table)
Point estimate, 273, 299, G-5 Population
1141
defined, 9, 19, G-5 measures of variability for, 124–126 vs. samples, measures of variability for, 125
Population mean compared to sample mean, 86 defined, 86, 96, G-5 estimating. See Estimation of population mean in normal distributions, 141–142
Population standard deviation, 124–125 defined, 124, 127, G-5 definitional formula for, 125 known, in testing one sample mean, 218, 221–233 not known, in testing one sample mean, 239–248
Population standard error of the mean calculation of, testing one sample mean, 225, 255 defined, 225, 254, G-5
Population variance defined, 124, 127, G-5 definitional formula for, 124
Positively skewed distribution, 44, G-5 Positive relationship, 576, 623, G-5 “post hoc” comparisons, 453 Prediction, 600 Predictor variable, 622, 623, G-5 Pretest-posttest research design, 341, 353, G-5 Principled understanding, 421 Probability, 175–184
applied to binomial distributions, 178–180 applied to normal distributions, 177–178 binomial distributions, 187, T-7–T-18 of confidence interval, 280 defined, 176, 204, G-5 events and outcomes, 176–177 example, Super Bowl, 181–184 formula, 176, 206 importance of, 177 of making Type I error, 382–383, 411 of making Type II error, 385–386, 411 of not making Type I error, 383 of not making Type II error, 386 of population mean, 280
“Proof” in hypothesis testing, 195, 327
1142
Quasi-experimental research defined, 14, G-5 example, kids' motor skills and fitness, 332
Random assignment, 14, 19, G-5 Random sampling, assumption of, 231 Range, 108–110
calculation of, 109–110 defined, 108, 127, G-5 interquartile. See Interquartile range strengths and weaknesses of, 110
Rank, 611, G-5 Ranked variables, 611 Rankings, as example of ordinal scale, 9 Ratio level of measurement
assumption of interval or ratio scale of measurement, 231 defined, 10, 19, G-5 displaying variables with histograms and frequency polygons, 36–39 example of, 12 Pearson correlation coefficient and, 578
Ratio scales, 10, 36 Real limits, 32, 46, G-6 Real lower limit, 32, G-6 Real upper limit, 32, G-6 Reed, Jaclyn, 214 Region of non-rejection
defined, 188, 205, G-6 example, 188–190, 197
Region of rejection defined, 187, 205, G-6 example, 187–190, 196–197
Regression, 601, G-6 Research hypothesis
compared to statistical hypothesis, 222 defined, 3, 18, G-6 question identification, 4–5 review and evaluation of relevant theories, 5 statement of research hypothesis, 5–8 testing. See Hypothesis testing
Research methodology, evaluating, 27 Research process, stages of, 2–17
communication of findings, 17
1143
data analysis, 15–16 data collection, 8–15 drawing of conclusion, 16 five steps of, 3 research hypothesis development, 3–8
Robust defined, z-statistic, 231, G-6 t-statistic, 330, 338
Rossi, Joseph, 297 Rounding numbers, when calculating the mean, 81–82 r squared (r2) statistic, 405, 411, G-5 R squared (R2) statistic, 442, G-5 Ruback, Barry, 312
Sample mean compared to population mean, 86 defined, 86, 96, G-6 hypothesis testing. See Testing one sample mean; Testing the difference between two sample means
Samples defined, 9, 19, G-6 vs. populations, measures of variability for, 125
Sample size and decision about null hypothesis, 196–197 and decision about null hypothesis in testing one sample mean, 249–250 estimated, for confidence interval for the mean, 300 increasing to control Type II error, 392–393 influence on interpreting inferential statistics, 401 influence on statistical analyses, 297 relation to level of confidence, 295 relation to width of confidence interval, 288–291
Sampling distribution, 218, 254, G-6 Sampling distribution of the difference, 317–318
characteristics of, 318 defined, 317, 353, G-6
Sampling distribution of the mean, 218–220 characteristics of, 219–220 defined, 218, 254, G-6 using to create confidence interval for the mean, 275–276
Sampling error, in probability, 177, 204, G-6 Scatterplot
in correlation example, 572–578
1144
defined, 572, 622, G-6 drawing linear regression equation, 605, 606 (figure), 608–609
Scheffé test, 461, 466, G-6 Schmidt, Frank, 296 Scientific method, 3, G-6 Significance, 227, 297
inappropriateness of “highly significant,” 402–404 statistical vs. practical, 407
Significant interaction effect, 531–533 Simple comparisons, 461, G-6 Simple effects
analysis of, 531–533 defined, 531, 535, G-6
Single-factor research designs compared to factorial designs, 498–500 defined, 488, 534, G-6 introduction to factorial research designs, 488–501
Size, as example of ordinal scale, 9 Skevington, Suzanne, 105 Skewed distribution, 43–44
of chi-square statistic, 654 F-ratio distribution, 429 and measures of central tendency, 88–90, 92, 94 (figure)
Slope of a line, 603–604, 625 Snyder, C. R., 234, 246 “Sometimes” in frequency rating study, 104–106 Spearman, Charles, 612 Spearman rank-order correlation
compared to Pearson correlation coefficient, 613 defined, 612, 623, G-6 example: business student study, 611–617, 618 (table) introduction to ranked variables, 611
Spearman rank-order correlation (business student study), 612–617 alpha set, 614 alternative hypotheses statement, 613–614 conclusion drawn, 617 critical values identified, 614, T-29–T-30 decision rule stated, 614 degrees of freedom calculated, 614 level of significance determined, 616 measure of effect size calculated, 617, 625 as nonparametric statistical test, 685, 686 (table)
1145
null hypotheses rejection decision, 616 null hypotheses statement and decision, 613–617 result of analysis related to research hypothesis, 617 statistic calculated, 615–616, 625 summary, 618 (table) using SPSS, 629–630
Spontaneous recovery, 297 SPSS software
analytical comparisons, planned comparisons, 470–471 calculating confidence interval for the mean, salary survey study, 300–302 calculating measures of central tendency and variability, frequency rating study, 129–131 chi-square goodness-of-fit test, personality study, 688–690 chi-square goodness-of-fit test with unequal hypothesized proportions, personality study, 691–692 chi-square test of independence, color preference study, 692–695 creation of data files, defining variables and entering data, 47–50 data examination, frequency distribution tables and figures, 51–52 linear regression, first impression study, 628–629 one-way ANOVA, chess study, 468–469 pearson correlation coefficient, first impression study, 626–627 spearman rank-order correlation, business student study, 629–630 testing one sample mean, unique invulnerability study, 255–257 testing the difference between paired means, web intervention study, 357–358 testing the difference between two sample means, parking lot study, 355–356 two-way ANOVA, voting message study, 537–540
Standard deviation (s), 120–123 computational formula for, 122, 123 (table) defined, 120, 127, G-6 definitional formula for, 120–122 in normal distributions, 142
Standard error of the difference calculation of, testing the difference between two sample means, unequal sample sizes, 335–336, 354 defined, testing the difference between two sample means, 318, 353, 354, G-6
Standard error of the difference scores, in testing the difference between paired means, 348, 354 Standard error of the mean
calculation of, for confidence interval, 277–278 calculation of, testing one sample mean, 243–244, 255 calculation of, testing the difference between two sample means, 315–316 defined, 219, 243, 254, G-6
1146
Standardized distribution characteristics of, 162–164 defined, 160, 167, G-7
Standardized scores defined, 160, 167, G-7 formula for, 160, 167 interpreting, 161–162 using to compare or combine variables, 162
Standardizing frequency distributions, 157–166 example, politicians' behavior, 158–164
Standard normal distribution, 142–150 applying probability to, 177–178 area between mean and z-score, 146–147 area between two z-scores, 148–150 area less or greater than z-score, 147–148 defined, 142, 167, G-7 proportions of area under curve, 144, 145 (table), T-2–T-6 rules for finding areas under, 150 (table) transforming normal distributions into, 151 z-scores, 143 z-scores: distance from the mean, 143–144 z-scores: relative position within distribution, 144–150
Statistic (numeric characteristic of sample) defined, sample mean vs. population mean, 86, 96, G-7 low probability of, 186–190 and measures of variability for samples vs. populations, 125
Statistical hypothesis compared to research hypothesis, 222 defined, 184, 205, G-7
Statistical power, Type II error, 386, 410, G-7 Statistical significance vs. nonsignificance, 227
interpretation of, in hypothesis testing, 297 Statistical significance vs. practical significance, 407 Statistics
defined, 1, 18, G-7 reasons to learn, 1–2
Stewart, Kenneth, 649 Strength of the relationship between variables, 577–578, 579 Studentized Range statistic, 462, T-24–T-26 Student t-distribution, 236–238
characteristics of, 237–238 defined, 236, 254, G-7
1147
Sum of products, 585, 596, 624 Sum of squared deviations, 587 Sum of Squares (SS), 435, 587, 598, 624, 625 Support, importance of term, 16, 195 Survey research, 14, 19, G-7 Symmetric distributions
defined, 43, 47, G-7 and measures of central tendency, 92, 94 (figure)
Symmetry, 42–44 defined, 42, 47, G-7
Tables, examining data using, 28–34 cumulative percentages, 32–34 frequency distribution tables, 28–30 grouped frequency distribution tables, 31–34
t-distribution introduction to, 236–238 See also Student t-distribution
Temperature conversion between Fahrenheit and Celsius, 164 as example of interval scale, 10
Testing one sample mean, 214–270 example: reading skills study, 214–218, 221–231 example: unique invulnerability bias, 234–236, 239–247, 255–257 factors affecting decision about null hypothesis, 248–253 sampling distribution of the mean, 218–220 Student t-distribution, 236–238
Testing one sample mean, population standard deviation known (reading skills study), 218, 221–233
alpha set, 223 alternative hypothesis statement, 221–222 assumptions of z-test for one mean, 231 conclusion drawn, 229–230 critical values identified, 223 decision rule stated, 223–224 level of significance determined, 227–229 null hypotheses rejection decision, 226–227 null hypothesis statement and decision, 221–229 result of analysis related to research hypothesis, 230–231 statistic calculated: z-test for one mean, 224–226 summary, 232–233, 232 (table)
Testing one sample mean, population standard deviation not known (unique
1148
invulnerability study), 239–248 alpha set, 241 alternative hypothesis statement, 239 assumptions of t-test for one mean, 246 conclusion drawn, 246 critical values indentified, 241 decision rule stated, 241–242 degrees of freedom calculated, 240–241, 255 level of significance determined, 245 null hypotheses rejection decision, 245 null hypothesis statement and decision, 239–245 result of analysis related to research hypothesis, 246 statistic calculated: t-test for one mean, 242–244 summary, 247–248, 247 (table) using SPSS, 255–257
Testing the difference between paired means (web-based intervention study), 340–352
alpha set, 347 alternative hypotheses statement, 345–346 assumptions of t-test for dependent means, 350 calculate difference between paired scores, 342–345 conclusion drawn, 349 critical values identified, 347 decision rule stated, 347 degrees of freedom calculated, 347, 354 level of significance determined, 348–349 null hypotheses rejection decision, 348 null hypotheses statement and decision, 345–349 research situations for within-subjects research designs, 341 result of analysis related to research hypothesis, 349–350 statistic calculated: t-test for dependent means, 347–348 summary, 350, 350–351 (table) using SPSS, 357–358
Testing the difference between two sample means, 312–376 example: kids' motor skills and fitness, 331–332 example: parking lot study, 313–317, 319–331 example: web-based intervention study, 341–342 review for ANOVA, 425–428 sampling distribution of the difference, 317–318
Testing the difference between two sample means, equal sample sizes (parking lot study), 319–331
alpha set, 321
1149
alternative hypotheses statement, 320 assumptions of t-test for independent means, 327–330 conclusion drawn, 326–327 critical values identified, 321–322 decision rule stated, 321–322 degrees of freedom calculated, 321, 354 level of significance determined, 326 null hypotheses rejection decision, 325 null hypotheses statement and decision, 320–326 result of analysis related to research hypothesis, 327 statistic calculated: t-test for independent means, 322–324 summary, 328 (table), 331 using SPSS, 355–356
Testing the difference between two sample means, unequal sample sizes (kids' motor skills and fitness), 331–340
alpha set, 335 alternative hypotheses statement, 332–334 assumptions of t-test for independent means, 338 conclusion drawn, 337 critical values identified, 335 decision rule stated, 335 degrees of freedom calculated, 334 level of significance determined, 337 null hypotheses rejection decision, 336 null hypotheses statement and decision, 332–337 result of analysis related to research hypothesis, 337–338 statistic calculated: t-test for independent means, 335–336 summary, 338, 339 (table)
Theory, 5, 19, G-7 Theory of territorial behavior, 312–313 3 × 4 research design, 513, 514 (table) Tick mark method, 28–29 Total degrees of freedom, 439 Total sum of squares, 439 True zero point, 10, 36 t-statistic, 236, 241 t-test for dependent means
calculation of, testing the difference between paired means, 347–348, 354 defined, 347, 353, G-7
t-test for independent means and analysis of variance (ANOVA), 425–429, 447 assumptions of, 327–330
1150
calculation of, testing the difference between two sample means, 322–324, 354 defined, 322, 353, G-7 increasing between-group variability to control Type II error, 396 relationship with F-ratio, 452, 467
t-test for one mean assumptions of, 246 calculation of, 242–244, 255 defined, 242, 253, G-7
Tucker, Christine, 105 Tukey Honestly Significant Difference (HSD) Test, 462 Tukey test
controlling Familywise error, 462–463 defined, 462, 466, G-7
2 × 2 research design, 502 Two-factor (A × B) research design
defined, 502, 534, G-7 descriptive statistics, 502–506 F-ratios in, 507–508 introduction to ANOVA, 506–508 notational system for means, 503–504
Two-tailed alternative hypothesis, 201, 205 vs. one-tailed, in testing one sample mean, 251–253
Two-way ANOVA, 486–569 example: exercise deprivation study, 489–501 example: voting message study, 486–488, 502–531, 621 F-ratios in two-factor research design, 507–508 introduction to ANOVA for two-factor research design, 506–508 introduction to factorial research designs, 488–501. See also Factorial research designs significant interaction effect, 531–533 testing interaction effects, 489–493 testing main effects, 494–495 two-factor research design, 502–506
Two-way ANOVA (voting message study), 508–531 alpha set, 511–512 alternative hypotheses statements, 508–509 ANOVA summary table, 520, 521 (table) conclusions drawn, 526 critical values for F-ratio identified, 511–513 decision rules stated, 511–512 degrees of freedom calculated, 510–511, 535–536
1151
levels of significance determined, 523 mean squares calculations, 513, 516–519, 536 measures of effect size calculated, 523–526, 537 null hypotheses rejection decisions, 520–523 null hypotheses statements and decisions, 508–526 results of analysis related to research hypothesis, 526, 530–531 statistics calculated: F-ratios for two-way ANOVA, 513–520, 536–537 summary, 527–530 (table) using SPSS, 537–540
Type I error defined, 382, 410, G-7 example: criminal trials and GKT study, 377–381 and Familywise error, 460–461 introduction to, 381–385 probability of making, 382–383, 411 probability of not making, 383 reasons for concern, 383 reasons for occurrence, 384–385 summary, 387–389, 388 (table)
Type I error, control of, 389–392 concerns about, 390–392 lowering alpha, 389–390, 391 (table) summary, 399, 399 (table)
Type II error defined, 385, 410, G-7 example: criminal trials and GKT study, 377–381 introduction to, 385–387 nonparametric tests and, 686 probability of making, 385–386, 411 probability of not making, 386 reasons for concern, 386–387 reasons for occurrence, 387 summary, 387–389, 388 (table)
Type II error, control of, 392–399 decreasing within-group variability, 396–399 increasing between-group variability, 396 increasing sample size, 392–393 raising alpha, 393–394 summary, 399, 399 (table) using directional alternative hypothesis, 394–396
Unbiased estimate, in variance, 120, G-7
1152
Unimodal distribution, 42, 43 (figure), 47, G-7 Unplanned comparisons, 458, 460–463
concerns regarding, 460–461 critical value of FT identified, 462–463 decision rule stated, 463 defined, 453, 466, G-7 desired probability of familywise error set, 462 Familywise error, 460–461 Familywise error, methods of controlling, 461 Familywise error, Tukey test, 462–463 null hypotheses rejection decision, 463 null hypotheses statement and decision, 463 statistic calculated: F-ratio for analytical comparison, 463 summary, 463–464
Variability defined, 44, 47, G-7 types of distributions, 44, 45 (figure) understanding, 106–108 See also Measures of variability
Variables defined, 3, 19, G-7 determination of measurement of, 9–12 relationship between, described, 574–578 relationship between, measured, 578–580 using SPSS, 47–50 values of, examining data in tables, 28–29
Variance (s2), 112–120 calculation considerations, 119–120 computational formula for, 116–119 defined, 114, 127, G-7 definitional formula for, 114–116 effect size as accounted for, 404 in Pearson correlation coefficient, 579–580, 587–589, 596–598
Wilcoxon matched-pairs signed-rank test, 686 (table) Within-group variability
and between-group variability in ANOVA, 426–428 and between-group variance in analytical comparisons, 453–454 decreasing to control Type II error, 396–399, 398 (table) defined, 397, G-8 in two-factor research design, 507–508
1153
Within-subjects research design defined, 341, 353, G-8 research situations appropriate for, 341
Yates's correction for continuity, 655, 687, G-8 Y-intercept, 603, 604, 625
z-scores applying to normal distributions, 150–157 area between mean and z-score, 146–147, 154 area between two z-scores, 148–150, 155–156 area less or greater than z-score, 147–148, 154–155 defined, 143, 167, G-8 distance from the mean, 143–144, 153 drawing conclusions, transforming scores into, 156 formula to transform scores in normal distribution, 152, 167 relative position within distribution, 144–150, 153–156 in standardized distribution, 160
z-statistic, 223 z-test for one mean
assumptions of, 231 calculation of, 224–226, 255 defined, 225, 253, G-8
1154
- Preface
- Acknowledgments
- About the Author
- Chapter 1. Introduction to Statistics
- Chapter 2. Examining Data: Tables and Figures
- Chapter 3. Measures of Central Tendency
- Chapter 4. Measures of Variability
- Chapter 5. Normal Distributions
- Chapter 6. Probability and Introduction to Hypothesis Testing
- Chapter 7. Testing One Sample Mean
- Chapter 8. Estimating the Mean of a Population
- Chapter 9. Testing the Difference between Two Means
- Chapter 10. Errors in Hypothesis Testing, Statistical Power, and Effect Size
- Chapter 11. One-Way Analysis of Variance (ANOVA)
- Chapter 12. Two-Way Analysis of Variance (ANOVA)
- Chapter 13. Correlation and Linear Regression
- Chapter 14. Chi-Square
- Tables
- Appendix: Review of Basic Mathematics
- Glossary
- References
- Index