due 10/27/2018. I have attached the chapter 2 textbook and data set for this assignment.
CHAPTER 2
QUANTITATIVE RESEARCH METHODS
Data Coding, Entry, and Checking
This chapter begins with a very brief overview of the initial steps in a research project. After this introduction, the chapter focuses on: (a) getting your data ready to enter into the data editor or a spreadsheet, (b) defining and labeling variables, (c) entering the data appropriately, and (d) checking to be sure that data entry was done correctly without errors.
Plan the Study, Pilot Test, and Collect Data
Plan the study. As discussed in Chapter 1, the research starts with identification of a research problem and research questions or hypotheses. It is also necessary to plan the research design before you select the data collection instrument(s) and begin to collect data. Most research methods books discuss this part of the research process extensively (e.g., see Gliner, Morgan, & Leech, 2009).
Select or develop the instrument(s). If there is an appropriate, available instrument that provides reliable and valid data and it has been used with a population similar to yours, it is usually desirable to use it. However, sometimes it is necessary to modify an existing instrument or develop your own. For this chapter, we have developed a short questionnaire to be given to students at the end of a course. Remember that questionnaires or surveys are only one way to collect quantitative data. You could also use structured interviews, observations, tests, standardized inventories, or some other type of data collection method. Research methods and measurement books have one or more chapters devoted to the selection and development of data collection instruments. A useful book on the development of questionnaires is Fink (2009).
Pilot test and refine instruments. It is always desirable to try out your instrument and directions with, at the very least, a few colleagues or friends. When possible, you also should conduct a pilot study with a sample similar to the one you plan to use later. This is especially important if you developed the instrument or if it is going to be used with a population different from the one(s) for which it was developed and on which it was previously used.
Pilot participants should be asked about the clarity of the items and whether they think any items should be added or deleted. Then, use the feedback to make modifications in the instrument before beginning the actual data collection. If the instrument is changed, the pilot data should not be added to the data collected for the study. Content validity can also be checked by asking experts to judge whether your items cover all aspects of the domain you intended to measure and whether they are in appropriate proportions relative to that domain.
Collect the data. The next step in the research process is to collect the data. There are several ways to collect questionnaire or survey data (such as telephone, mail, or e-mail). We do not discuss them here because that is not the purpose of this book. The Fink (2009) book, How to Conduct Surveys: A Step by Step Guide, provides information on the various methods for collecting survey data.
You should check your raw data after you collect it even before it is entered into the computer. Make sure that the participants marked their score sheets or questionnaires appropriately; check to see if there are double answers to a question (when only one is expected) or answers that are marked between two rating points. If this happens, you need to have a rule (e.g., “use the average”) that you can apply consistently. Thus, you should “clean your data, making sure they are clear, consistent, and readable, before entering them into a data file.
Let’s assume that the completed questionnaires shown in Figs. 2.1 and 2.2 were given to a small class of 12 students and that they filled them out and turned them in at the end of the class. The researcher numbered the forms from 1 to 12, as shown opposite ID.
After the questionnaires were turned in and numbered (i.e., given an ID number in the top right corner), the researcher was ready to begin the coding process, which we describe in the next section.
FIG 2.1 Completed questionnaires for participants 1 through 6.
FIG 2.2 Completed questionnaires for participants 7 through 12.
Code Data for Data Entry
Guidelines for Data Coding
Coding is the process of assigning numbers to the values or levels of each variable. Before starting the coding process, we want to present some broad suggestions or rules to keep in mind as you proceed. These suggestions are adapted from rules proposed in Newton and Rudestam’s (1999) useful book entitled Your Statistical Consultant. We believe that our suggestions are appropriate, but some researchers might propose alternatives, especially for guidelines 1, 2, 4, 5, and 7.
All data should be numeric. Even though it is possible to use letters or words (string variables) as data, it is not desirable when using SPSS to do so. For example, we could code gender as M for male and F for female, but in order to do most statistics you would have to convert the letters or words to numbers. It is easier to do this conversion before entering the data into the computer as we have done with the HSB data set (see Fig. 1.3). You will see in Fig. 2.3 that we decided to code females as 1 and males as 0. This is called dummy coding. In essence, the 0 means “not female.” Dummy coding is useful if you will want to use the data in some types of analyses and for obtaining descriptive statistics. For example, the mean of data coded this way will tell you the percentage of participants who fall in the category coded as “1.” We could, of course, code males as 1 and females as 0, or we could code one gender as 1 and the other as 2. However, it is crucial that you be consistent in your coding (e.g., for this study, all males are coded 0 and females 1) and that you have a way to remind yourself and others of how you did the coding. Later in this chapter, we show how you can provide such a record, called a codebook or dictionary.
Each variable for each case or participant must occupy the same column in the Data Editor. It is important that data from each participant occupy only one line (row), and each column must contain data on the same variable for all the participants. The data editor, into which you will enter data, facilitates this by putting the short variable names that you choose at the top of each column, as you saw in Chapter 1, Fig. 1.3. If a variable is measured more than once (e.g., pretest and posttest), it will be entered in two columns with somewhat different names, such as mathpre and mathpost.1
All values (codes) for a variable must be mutually exclusive. That is, only one value or number can be recorded for each variable. Some items, like our item 6 in Fig. 2.3, allow for participants to check more than one response. In that case, the item should be divided into a separate variable for each possible response choice, with one value of each variable (usually 1) corresponding to yes (i.e., checked) and the other to no (usually 0, for not checked). For example, item 6 becomes variables 6, 7, and 8 (see Fig. 2.3). Items should be phrased so that persons would logically choose only one of the provided options, and all possible options should be provided. A final category labeled “other” may be provided in cases where all possible options cannot be listed, but these “other” responses are usually quite diverse and thus may not be very useful for statistical purposes.
Each variable should be coded to obtain maximum information. Do not collapse categories or values when you set up the codes for them. If needed, let the computer do it later. In general, it is desirable to code and enter data in as detailed a form as available. Thus, enter actual test scores, ages, GPAs, and so forth, if you know them. It is good practice to ask participants to provide information that is quite specific. However, you should be careful not to ask questions that are so specific that the respondent may not know the answer or may not feel comfortable providing it. For example, you will obtain more information by asking participants to state their GPA to two decimals (as in Figs. 2.1 and 2.2) than if you asked them to select from a few broad categories (e.g., less than 2.0, 2.0–2.49, 2.50–2.99, etc). However, if students don’t know their GPA or don’t want to reveal it precisely, they may leave the question blank or write in a difficult to interpret answer, as discussed later.These issues might lead you to provide a number of categories, each with a relatively narrow range of values, for variables such as age, weight, and income. Never collapse such categories before you enter the data into the data editor. For example, if you have age categories for university undergraduates 15–17, 18–20, 21–23, and so forth, and you realize that there are only a few students younger than 18, keep the codes as is for now. Later you can make a new category of 20 or younger by using a function, Transform => Recode. If you collapse categories before you enter the data, the extra information will no longer be available.
For each participant, there must be a code or value for each variable. These codes should be numbers, except for variables for which the data are missing. We recommend using blanks when data are missing or unusable because SPSS is designed to handle blanks as missing values. However, sometimes you may have more than one type of missing data, such as items left blank and those that had an answer that was not appropriate or usable. In this case you may assign numeric codes such as 98 and 99 to them, but you must tell the program that these codes are for missing values, or it will treat them as actual data. (See Fig. 1.2 and Table 1.2.)
Apply any coding rules consistently for all participants. This means that if you decide to treat a certain type of response as, say, missing for one person, you must do the same for all other participants.
Use high numbers (values or codes) for the “agree,” “good,” or “positive” end of a variable that is ordered. Sometimes you will see questionnaires that use 1 for “strongly agree,” and 5 for “strongly disagree.” This is not wrong as long as you are clear and consistent. However, you are less likely to get confused when interpreting your results if high values have a positive meaning.
Make a Coding Form
Now you need to make some decisions about how to code the data provided in Figs. 2.1 and 2.2, especially data that are not already in numerical form. When the responses provided by participants are numbers, the variable is said to be “self-coding.” You can just enter the number that was circled or checked. On the other hand, variables such as gender or college have no intrinsic value associated with them. See Fig. 2.3 for the decisions we made about how to number the variables, code the values, and name the eight variables. Don’t forget to number each of the questionnaires so that you can later check the entered data against the questionnaires.
Fig. 2.3 A blank survey showing how to code the data.
Problem 2.1: Check the Completed Questionnaires
Now, examine Figs. 2.1 and 2.2 for incomplete, unclear, or double answers. Stop and do this now, before proceeding. What issues did you see? The researcher needs to make rules about how to handle these problems and note them on the questionnaires or on a master “coding instructions” sheet so that the same rules are used for all cases.
We have identified at least 11 responses on 6 of the 12 questionnaires that need to be clarified. Can you find them all? How would you resolve them? Write on Figs. 2.1 and 2.2 how you would handle each issue that you see.
Make Rules About How to Handle These Problems
For each type of incomplete, blank, unclear, or double answer, you need to make a rule for what to do. As much as possible, you should make these rules before data collection, but there may well be some unanticipated issues. It is important that you apply the rules consistently for all similar problems so as not to bias your results.
Interpretation of Problem 2.1 and Fig. 2.4
Now, we will discuss each of the issues and how we decided to handle them. Of course, some reasonable choices could have been different from ours. We think that the data for Participants 1–6 are quite clear and ready to enter with the help of Fig. 2.3. However, the questionnaires for participants 7–12 pose a number of minor and more serious problems for the person entering the data. We discuss next and have written our decisions in numbered callout boxes on Fig. 2.4, which are the surveys and responses for Subjects 7–12.
For Participant 7, the GPA appears to be written as 250. It seems reasonable to assume that he meant to include a decimal after the 2, and so we would enter 2.50. We could instead have said that this was an invalid response and coded it as missing. However, missing data create problems in later data analysis, especially for complex statistics. Thus, we want to use as much of the data provided as is reasonable. The important thing here is that you must treat all other similar problems the same way.
For Subject 8, two colleges were checked. We could have developed a new legitimate response value (4 = other). Because this fictitious university requires that students be identified with one and only one of its three colleges, we have developed two missing value codes (as we did for ethnic group and religion in the HSB data set). Thus, for this variable only, we used 98 for multiple checked colleges or other written-in responses that did not fit clearly into one of the colleges (e.g., business engineering or history and business). We treated such responses as missing because they seemed to be invalid and/or because we would not have had enough of any given response to form a reasonable size group for analysis. We used 99 as the code for cases where nothing was checked or written on the form. Having two codes enabled us to distinguish between these two types of missing data, if we ever wanted to later. Other researchers (e.g., Newton & Rudestam, 1999) recommend using 8 and 9 in this case, but we think that it is best to use a code that is very different from the “valid” codes so that they stand out visually in the Data View and will lead to noticeable differences in the Descriptives if you forget to code them as missing values.
Also, Subject 8 wrote 2.2 for his GPA. It seems reasonable to enter 2.20 as the GPA. Actually, in this case, if we enter 2.2, the program will treat it as 2.20 because we will tell it to use two decimal places for this variable.
We decided to enter 3.00 for Participant 9’s GPA. Of course, the actual GPA could be higher or, more likely, lower, but 3.00 seems to be the best choice given the information provided by the student (i.e., “about 3 pt”).
Participant 10 only answered the first two questions, so there were lots of missing data. It appears that he or she decided not to complete the questionnaire. We made a rule that if three out of the first five items were blank or invalid, we would throw out that whole questionnaire as invalid. In your research report, you should state how many questionnaires were thrown out and for what reason(s). Usually you would not enter any data from that questionnaire, so you would only have 11 subjects or cases to enterTo show you how you would code someone’s college if they left it blank, we did not delete this subject at this time.
For Subject 11, there are several problems. First, she circled both 3 and 4 for the first item; a reasonable decision is to enter the average or midpoint, 3.50.
Participant 11 has written in “biology” for college. Although there is no biology college at this university, it seems reasonable to enter 1 = arts and sciences in this case and in other cases (e.g., history = 1, marketing = 2, civil = 3) where the actual college is clear. See the discussion of Issue 2 for how to handle unclear examples.
Participant 11 also entered 9.67 for the GPA, which is an invalid response because this university has a 4-point grading system (4.00 is the maximum possible GPA). To show you one method of checking the entered data for errors, we will go ahead and enter 9.67. If you examine the completed questionnaires carefully, you should be able to spot errors like this in the data and enter a blank for missing/invalid data.
Enter 1 for reading and homework for Participant 11 (even though they were circled rather than checked). Also enter 0 for extra credit (not checked) as you would for all the boxes left unchecked by other participants (except Subject 10, who, as stated in number 5 above, did not complete the questionnaire). Even though this person circled the boxes rather than putting X’s or checks in them, her intent is clear.
As in Point 6, we decided to enter 2.5 for Participant 12’s X between 2 and 3.
Participant 12 also left GPA blank so, using the general (system) missing value code, we left it blank.
Clean up Completed Questionnaires
Now that you have made your rules and decided how to handle each problem, you need to make these rules clear to whoever will enter the data. As mentioned earlier, we put our decisions in callout boxes on Fig. 2.4; a common procedure would be to write your decisions on the questionnaires, perhaps in a different color.
FIG 2.4 Completed survey with callout boxes showing how we handled problem responses.
Problem 2.2: Define and Label the Variables
The next step is to create a data file into which you will enter the data. If you do not have the program open, you need to log on. When you see the startup window, click the Type in data button; then you should see a blank Data Editor that will look something like Fig. 2.5. Also be sure that Display Commands in the Log is checked (see Appendix A). You should also examine Appendix A if you need more help getting started.
This section helps you name and label the variables. In the next section, we show you how to enter data. First, let’s define and label the first two variables, which are two 5-point Likert ratings. To do this we need to use the Variable View screen. Look at the bottom left corner of the Data Editor to see whether you are in the Data View or Variable View screen by noting which tab is white. If you are in Data View, to get to Variable View do the following:
Click on the Variable View tab at the bottom left of your screen. This will bring up a screen similar to Fig. 2.5. (Or, double click on var above the blank column to the far left side of the Data View.)
Fig. 2.5 Blank variable view screen in the data editor.
Blank variable view screen in the data editor.
In this window, you will see 11 columns that will allow you to input the variable name, type of variable, width, number of decimals, variable label, value labels, missing values other than blanks, columns, align data left or right, measurement type, and variable role.
Define and Label Two Likert-Type Variables
We now begin to enter information to name, label, and define the characteristics of the variables used in this chapter.
Click in the blank box directly under Name in Fig. 2.5.
Type recommend in this box. Notice the number 1 to the left of this box. This indicates that you are entering your first variable.2
Press enter. This will insert the program’s default values for variables. You need to check to be sure these are correct for each of your variables and make changes if needed.
Note that the Type is numeric, Width = 8, Decimals = 2, Label = (blank), Values = None, Missing = None, Columns = 8, Align = right, Measure = scale, Role = input.
For this assignment, we will keep the default values for Type, Width, Columns, and Align. On the Variable View screen, you will notice that the default for Type is Numeric. This refers to the type of variable you are entering. Usually, you will only use the Numeric option. Numeric means the data are numbers. String would be used if you input words or letters such as “M” for males and “F” for females. However, it is best not to enter words or letters because you wouldn’t be able to do many statistics without recoding them as numbers. In this book, we will always keep the Type as Numeric.
We recommend keeping the Width at eight, and keeping the Columns at eight. We will always Align the numbers to the right. Sometimes, we will change the settings for the other columns.
Now let’s continue with defining and labeling the recommend variable.
For this variable, leave the decimals at 2.
Click on the box under “Label” and type I recommend course in the Label box. This longer label will show in appropriate windows and on your printouts. The labels can be up to 40 characters but it is best to keep them about 20 or less or your outputs may be difficult to read.
In the Values column of Fig. 2.5, do the following:
Click on the word “None” and you will see a small blue box with three dots.
Click on the three dots. You will then see a screen like Fig. 2.6. We decided to add value labels for the lower and upper end of the Likert scale to help us interpret the data, but it is not as important to add labels for Likert or other ordered data as it is when the data are nominal or unordered.
Value labels window.
Type 1 in the Value box in Fig. 2.6.
Type strongly disagree in the Value Label box. Press Add.
Type 5 and strongly agree in the Values and Value Labels boxes. Your window should look like Fig. 2.6 just before you click on Add for the second time.
Click on Add.
Then click OK.
Leave the cells for the Missing to Measure columns in Fig. 2.5 as they currently appear.
Change Role to Target because recommend could be used as a Target (dependent or outcome) variable. See Figure 2.7. Different researchers might code these variables differently. For example, if they planned to use recommend only as an independent variable in their study, they would code Role as Input.
Role selection.
Now let’s define and label the next variable.
Click on the next blank box under Name (in Row 2) to enter the name of the next variable. Note spaces are not allowed in variable names. Spaces are allowed in labels.
Type workhard in the Name column and press Enter.
Click on the box in Row 2 under Label and type I worked hard in the Label column.
Insert the highest and lowest Values for this variable the same way you did for recommend (1 = strongly disagree and 5 = strongly agree).
Change Role to Both because it could be used as either an Input (independent) or Target (dependent) variable.
Define and Label College and Gender
Now, select the cell under Name in Row 3.
Call this third variable college by typing that in the box.
Click on the third box under Decimals. For this variable, there is no reason to have any decimal places because people were asked to choose only one of the three colleges. You will notice that when you select the box under Decimals, up and down arrows appear on the right side of the box. You can either click the arrows to raise or lower the number of decimals, or you can double click on the box and manually type in the desired number.
For the purposes of this variable, select or type 0 as the number of decimals.
Next, click the box under Label to type in the variable label college.
Under Values, click on None and then click on the small blue box with three dots.
In the Value Labels window, type 1 in the Value box, type arts and sciences in the Value Label box.
Then click Add. Do the same for 2 = business, 3 = engineering, 98 = other, multiple ans., 99 = blank.
The Value Labels window should resemble Fig. 2.8 just before you click Add for the last time.
Value labels window.
Then click OK.
Under Measure, click the box that reads Scale.
Click the down arrow and choose Nominal because for this variable the categories are unordered or nominal.
Your screen should look like Fig. 2.9 just after you click on nominal.
FIG 2.9 Measurement selection.
Change Role to Input because college will only be used as an independent variable.
Under Missing, click on None and then on the three dots. Click on Discrete Missing Values and enter 98 and 99 in the first two boxes. (See Fig. 2.10.) This step is essential if you have one or more specific values that you want to use as missing value code(s). If you leave the Missing cell at None, the program will not know that 98 and 99 should be considered missing. None in this column is somewhat misleading. None means no special missing values (i.e., only blanks are considered missing).
Fig. 2.10 Missing values.
Missing values.
Then click on OK.
Your Data Editor should now look like Fig. 2.11.
FIG 2.11 Completed variable view for the first three variables.
Now define and label gender similarly to how you did this for college.
First, type the variable Name gender in the next blank row in Fig. 2.11.
Click on Decimals to change the decimal places to 0 (zero).
Now click on Labels and label the variable gender.
Next you must label the values or levels of the gender variable. You need to be sure your coding matches your labels. We arbitrarily decided to code male as zero and female as 1. We could have coded female as zero and male as 1. There are some advantages to using 0 and 1 for the codes (“dummy coding”), as indicated below.
Click on the Values cell.
Then, click on the blue three-dot box to get a window like Fig. 2.6 again. Remember, this is the same process you conducted when entering the labels for the values of the first three variables.
Now, type 0 to the right of Value.
To the right of Label type male. Click on Add.
Repeat this process for 1 = female. Click on Add.
Click OK.
Click on Scale under Measure to change the level of measurement to Nominal because this is an unordered, dichotomous variable.
Finally, click on Input under Role because gender will be an independent variable.
Once again, realize that the researcher has made a series of decisions that another researcher could have done differently, as we noted earlier with the Role of the recommend variable. For example, you could have used 1 and 2 as the values for gender, and you might have given males the higher number. We have chosen, in this case, to do what is called dummy coding. In essence, 1 is female and 0 is not female. This type of coding is useful for interpreting gender when used in statistical analysis. Similarly, we could have decided to consider the level of measurement ordinal, since dummy coded dichotomous variables can be used in analyses that require ordered data, as we will discuss in later chapters.
Define and Label Grade Point Average
You should now have enough practice to define and label the gpa variable. After naming the variable gpa, do the following:
For Decimals leave the decimals at 2.
Now click on Label and label it grade point average.
Click on Values. Type 0 = All Fs and 4 = All As. (Note that for this variable, we have used actual GPA to 2 decimals, rather than dividing it into ordered groups such as a C average, B average, A average.)
Under Measure, leave it as Scale because this variable has many ordered values and is likely to be normally distributed.
Under Role, click on Both.
Define and Label the Last Three Variables
Now you should define the three variables related to the parts of the class that a student completed. Remember we said the Names of these variables would be: reading, homework, and extracrd. The variable Labels will be I did the reading, I did the homework, I did extra credit. The Value labels are: 0 = not checked/blank and 1 = checked. These variables should have no decimals, and the Measure should be changed to Nominal. Role should be changed to Both because these could be used as other independent or dependent variables. Your complete Variable View should look like Fig. 2.12.
Problem 2.3: Display Your Dictionary or Codebook
Now that you have defined and labeled your variables, you can print a codebook or dictionary of your variables. It is a very useful record of what you have done. Notice that the information in the codebook is essentially the same as that in the variable view (Fig. 2.12) so you do not really have to have both, but the codebook makes a more complete printed record of your labels and values.
Select File → Display Data File Information → Working File. Your codebook should look like Output 2.1, without the callout boxes. The codebook is divided into parts: the Variable Information (which is very similar to the Variable View in Fig. 2.12) and the Variable Values (which are partially hidden in the Variable View).
To rescale (shrink) the wide Variable Information table and print the whole codebook, see the Appendix A section about Resize/Rescale to Print. You may not be able to see all of the file information/codebook on your computer screen. However, you should be able to print the entire codebook.
OUTPUT 2.1-CODEBOOK
Problem 2.4: Enter Data
Close the codebook, and then click on the Data View tab on the bottom of the screen to give you the data editor. Note that the spreadsheet has numbers down the left-hand side (see Fig. 2.13). These numbers represent each participant in the study. The data for each participant’s questionnaire go on one and only one line across the page with each column representing a variable from our questionnaire. Therefore, the first column will be recommend, the second will be workhard, the third will be college, and so forth.
After defining and labeling the variables, your next task is to enter the data directly from the questionnaires or from a data entry form.
Occasionally a researcher will transfer the data from the questionnaires to a data entry form (like Table 2.1) by hand before entering the data into SPSS. This may be helpful if the questionnaires or answer sheet are not easily readable by the data entry person, if the responses are to be entered from several different sources, or if additional coding or recoding is required before data entry. In these situations, you could make mistakes entering the data directly from the questionnaires. However, if you use a data entry form, you could make copying mistakes, and it takes time to transfer the data from questionnaires to the data entry form. Thus, few researchers use a data entry form as an intermediate step between the questionnaire and the data editor. Our cleaned up questionnaires should be easy enough to use so that you could enter the data directly from Fig. 2.1 and Fig. 2.4.into the data editor. Try to do that using the directions below. If you have difficulty, you may use Table 2.1, but remember that it took an extra step to produce.
In Table 2.1, the data are shown as they would look if we copied the cleaned up data from the questionnaires to a data entry sheet, except that the data entry form could be handwritten on ruled paper.
Table 2.1. A Data Entry Form: Responses Copied From the Questionnaires
o enter the data, ensure that your Data Editor is showing.
If it is not already highlighted, click on the far left column, which should say recommend.
To enter the data into this highlighted column, simply type the number and press the right arrow. For example, first type 3 (the number will show up in the blank space above the row of variable names) and then press the right arrow; the number will be entered into the highlighted box. Next, type 5 in the workhard column and so forth.
In Fig. 2.13, all the data for the participants have been entered.
Data Editor participants entered.
Now enter from your cleaned up questionnaires the rest of the data in Fig. 2.1 and Fig. 2.4. If you make a mistake when entering data, correct it by clicking on the cell (the cell will be highlighted), type the correct score, and press enter or the arrow key.
Before you do any analysis, compare the data on your questionnaires with the data in the Data Editor. If you have lots of data, a sample can be checked, but it is preferable to check all of the data. If you find errors in your sample, you should check all the entries.
Problem 2.5: Run Descriptives and Check the Data
In order to get a better “feel” for the data and to check for other types of errors or problems on the questionnaires, we recommend that you run the statistics program called Descriptives. To compute basic descriptive statistics for all your subjects, you will need to do these steps:
Select Analyze → Descriptive Statistics → Descriptives… (see Fig. 2.14).3
After selecting Descriptives, you will be ready to compute the mean, minimum, and maximum values for all participants or cases on all variables in order to examine the data.
Analyze menu.
Now highlight all of the variables. To highlight, click on the first variable, then hold down the “shift” key and click on the last variable so that all of the variables listed are highlighted (see Fig. 2.15a). Note that in SPSS 14 and later versions, there is a symbol to the left of each variable name; it indicates whether you have labeled the measurement level as nominal , ordinal , or scale . Measurement levels are discussed in detail in Chapter 3 of this book.
Fig. 2.15a Descriptives— before moving variables.
Fig 2.15b Descriptives— after moving variables.
Be sure that all of the variables have moved out of the left window. If your screen looks like Fig. 2.15b, then click on Options. You will get Fig. 2.16.
Fig. 2.16 Descriptives: options.
Follow these steps:
Notice that the Mean, Std. deviation, Minimum, and Maximum were already checked. Click off Std. deviation. At this time, we will not request more descriptive statistics. We will do them in Chapter 4.
Ensure that the Variable list bubble is checked in the Display Order section. Note: You can also click on Ascending or Descending means if you want your variables listed in order of the means. If you wanted the variables listed alphabetically, you would check Alphabetic, but we won’t check any of these alternatives so the variables will be printed in the same order as on the Variable View.
Click on Continue, which will bring you back to the main Descriptives dialog box (Fig. 2.15b).
Then click on OK to run the program.
You should get an output like Fig. 2.17. If it looks similar, you have done the steps correctly
Output viewer for descriptives.
The left side of Fig. 2.17 lists the various parts of your output. You can click on any item on the left (e.g., Title, Notes, or Descriptive Statistics) to activate the output for that item, and then you can edit it. For example, you can click on Title and then expand the title or add information such as your name and the date. (See Appendix A for more on editing outputs.)
Double click on the large, bold word Descriptives in Fig. 2.17. Type your name in the box that appears so it will appear on your output when you print it later. Also type “Output 2.2” at the top so you and/or your instructor will know what it is later.
For each variable, compare the minimum and maximum scores in Fig. 2.17 with the highest and lowest appropriate values in the codebook (Output 2.1). This checking of data before doing any more statistics is important to further ensure that data entry errors have not been made and that the missing data codes are being used properly.
Note that after each output we have provided a brief interpretation in a box. On the output itself, we have pointed out some of the key things by circling them and making some comments in boxes, which are known as callout boxes. Of course, these circles and information boxes will not show up on your printout.
Output 2.2: Descriptives
Interpretation of Output 2.2
This output shows, for each of the eight variables, the number (N) of participants with no missing data on that variable. The Valid N (listwise) is the number (9) who have no missing data on any variable. The table also shows the Minimum and Maximum score that any participants had on that variable. For example, no one circled a 1, but one or more persons circled a 2 for the I recommend course variable, and at least one person circled 5. Notice that for I worked hard, 5 is both the minimum and maximum, everyone said they worked hard. This item is, therefore, really a constant and not a variable; it will not be useful in statistical analyses.
The table also provides the Mean or average score for each variable. Notice the mean for I worked hard is 5 because everyone circled 5. The mean of 1.80 for college, a nominal (unordered) variable with 3 or more categories, is nonsense, so ignore it. However, the means of .55 for the dichotomous variables gender, I did the reading, and I did the homework indicate that in each case 55% chose the answers that corresponded to 1 (female gender and “yes” for doing the reading and homework). The mean grade point average was 3.58, which is probably an error because it is too high for the overall GPA for most groups of undergrads. Note also that there has to be an error in GPA because the maximum GPA of 9.67 is not possible at this university, which has a 4.00 maximum (see codebook). Thus the 9.67 for participant 11 is an invalid response. The questionnaires should be checked again to be sure there wasn’t a data entry error. If, as in this case, the survey says 9.67, it should be changed to blank, the missing value code.