Unit VI discussion board

profileMalkta
CH24.docx

4 Analysing and Presenting Quantitative Data

Chapter outline

· Categorizing data

· Data entry, layout and quality

· Presenting data using descriptive statistics

· Analysing data using descriptive statistics

· The process of hypothesis testing: inferential statistics

· Statistical analysis: comparing variables

· Statistical analysis: associations between variables

Keywords https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img60.jpg

· Categorizing data

· Data entry

· Descriptive statistics

· Distributions

· Hypotheses

· Inferential statistics

· Significance

· Correlation analysis

· Regression

Icon Key

Read

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img3.jpg

Explore

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img4.jpg

Define

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img5.jpg

Apply

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img10.jpg

Watch

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img7.jpg

Build

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img8.jpg

Practise

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img9.jpg

Discover

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img11.jpg

Author video

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img62.jpg

Chapter objectives

After reading this chapter you will be able to:

· Prepare quantitative data for analysis.

· Select appropriate formats for the presentation of quantitative data.

· Choose the most appropriate techniques for describing data (descriptive statistics).

· Choose and apply the most appropriate statistical techniques for exploring relationships and trends in data (correlation and inferential statistics).

As we have seen in previous chapters, the distinction between quantitative and qualitative research methods is often blurred. Take, for example, survey methods. These can be purely descriptive in design, but, on the other hand, the gathering of respondent profile data provides an opportunity for finding associations between classifications of respondents and their attitudes or behaviour, providing the potential for quantitative analysis.

One of the essential features of quantitative analysis is that, if you have planned your research tool, collected your data and now you are thinking of how to analyse them – you are too late! The process of selecting statistical tests should take place at the planning stage of research, not at implementation. This is because it is so easy to end up with data for which there is no meaningful statistical test. Robson (2002) also provides an astute warning that, particularly with the aid of the modern computer, it becomes much easier to generate elegantly presented rubbish, reminding us of GIGO – Garbage In, Garbage Out (Robson, 2002).

The aim of this chapter is to introduce you to some of the basic statistical techniques. It does not pretend to provide you with an in-depth analysis of more complex statistics, since there are specialized textbooks for this purpose. It is assumed that you will have access to a computer and an appropriate software application for statistical analysis, particularly IBM SPSS Statistics. Note that in this chapter, rather than offer you Activities, Worked Examples using statistical formulae will be provided. In some cases, these will be supported with data that you can access by clicking on the icons in the margin or by visiting the book’s online resources. These datasets allow you to apply some statistical tests to ‘real’ data.

Watch: Basic statistics

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img7.jpg

Top Tip 24.1

If you are relatively new to statistics, try to get access to someone more experienced than yourself to act as a guide or mentor. Also, of course, if you have an academic supervisor, ensure that you maintain regular contact and ask for advice. As suggested in  Chapter 23 , there are also many useful online tutorials on statistics on YouTube. If you are new to statistics, you might find it helpful if you add the word ‘basic’ to ‘statistics’ in the YouTube search engine.

Categorizing data

The process of categorizing data is important because, as was noted in  Chapter 6 , the statistical tests that are used for data analysis will depend on the type of data being collected. Hence, the first step is to classify your data into one of two categories, categorical or quantifiable (see  Figure 24.1 ).  Categorical data  cannot be quantified numerically but are either placed into sets or categories (nominal data) or ranked in some way (ordinal data). Quantifiable data can be measured numerically, which means that they are more precise. Within the quantifiable classification there are two additional categories of interval and ratio data. All of these categories are described in more detail below. Saunders et al. (2012) warn that if you are not sure about the level of detail you need in your research study, it is safest to collect data at the highest level of precision possible.

Read: Categorical data analysis

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img3.jpg

In simple terms, these data are used for different analysis purposes.  Table 24.1  suggests some typical uses and the kinds of statistical tests that are appropriate.

As Diamantopoulos and Schlegelmilch (1997) point out, the four kinds of measurement scale are nested within one another: as we move from a lower level of measurement to a higher one, the properties of the lower type are retained. Thus, all the statistical tests appropriate to the lower type of data can be used with the higher types as well as additional, more powerful tests. But this does not work in reverse: as we move from, say, interval data to ordinal, the tests appropriate for the former cannot be applied to the latter. For categorical data only, non-parametric statistical tests can be used, but for quantifiable data (see  Figure 24.1 ), more powerful parametric tests need to be applied. Hence, in planning data collection it is better to design data gathering instruments that yield interval and ratio data, if this is appropriate to the research objectives. Let us look at each of the four data categories in turn.

Figure 24.1 Types of categorical and quantifiable data

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-fig100.jpg

Table 24.1 https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table134.jpg

Nominal data

Define: Nominal scale

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img5.jpg

Nominal data constitute a name value or category with no order or ranking implied (for example, sales departments, occupational descriptors of employees, etc.). A typical question that yields nominal data is presented in  Figure 24.2 , with a set of data that results from this presented in  Table 24.2 . Thus, we can see that with nominal data, we build up a simple frequency count of how often the nominal category occurs.

Figure 24.2 Types of questions that yield nominal data

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-fig101.jpg

Table 24.2 https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table135.jpg

Ordinal data

Define: Ordinal measure

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img5.jpg

Ordinal data comprise an ordering or ranking of values, although the intervals between the ranks are not intended to be equal (for example, an attitude questionnaire). A type of question that yields ordinal data is presented in  Figure 24.3 . Here there is a ranking of views (Sometimes, Never, etc.) where the order of such views is important but there is no suggestion that the differences between each scale are identical. Ordinal scales are also used for questions that rate the quality of something (for example, very good, good, fair, poor, etc.) and agreements (for example, Strongly Agree, Agree, Disagree, etc.). The typical results of gathering ordinal data are taken from  Figure 24.3  and presented in  Table 24.3 .

Interval data

Define: Interval measure

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img5.jpg

With quantifiable measures such as interval data, numerical values are assigned along an interval scale with equal intervals, but there is no zero point where the trait being measured does not exist. For example, a score of zero on a traditional IQ test would have no meaning. This is because the traditional IQ score is the raw (actual) score converted into a mental age divided by chronological age. Another characteristic of interval data is that the difference between a score of 14 and 15 would be the same as the difference between a score of 91 and 92. Hence, in contrast to ordinal data, the differences between categories are identical. The kinds of results from interval data are illustrated in  Table 24.4 , delivered as part of a company’s aptitude assessment of staff.

Figure 24.3 Types of questions that yield ordinal data

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-fig102.jpg

Table 24.3 https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table136.jpg

Table 24.4 https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table137.jpg

Ratio Data

Define: Ratio scale

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img5.jpg

Ratio data are a sub-set of interval data, and the scale is again interval, but there is an absolute zero that represents some meaning – for example, scores on an achievement test. If an employee, for example, undertakes a work-related test and scores zero, this would indicate a complete lack of knowledge or ability in this subject! An example of ratio data is presented in  Table 24.5 .

This sort of classification scheme is important because it influences the ways in which data are analysed and what kind of statistical tests can be applied. Having incorporated variables into a classification scheme, the next stage is to look at how data should be captured and laid out, prior to analysis and presentation.

Table 24.5 https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table138.jpg

Data entry, layout and quality

Data entry involves a number of stages, beginning with ‘cleaning’ the data, planning and implementing the actual input of the data, and dealing with the thorny problem of missing data. Ways of avoiding the degradation of data will also be discussed.

Cleaning the data

Watch: Cleaning data

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img7.jpg

Data analysis will only be reliable if it is built upon the foundations of ‘clean’ data: that is, data that have been entered into the computer accurately. When entering data containing a large number of variables and many individual records, it is easy to enter a wrong figure or to miss an entry. One solution is for two people to enter data separately and to compare the results, but this is expensive. Another approach is to use frequency analysis on a column of data that will throw up any spurious figures that have been entered. For example, if you are using numbers 1 to 5 to represent individual codes for each of five variables, the frequency analysis might show that you had also entered the number 8 – clearly a mistake. Where there are branching or skip questions (recall  Chapter 14 ) it may also be necessary to check that respondents are going through the questions carefully. For example, they may be completing sections that do not apply to them or missing other sections.

Data coding and layout

Coding usually involves allocating an identification number (Id) to data. Take care, however, not to make the mistake of subsequently analysing the codes as raw data! The codes are merely shorthand ways of describing the data. Once the coding is completed, it is possible to collate the data into groups of less detailed categories. So, in  Case Study 24.1  the categories could be recoded to form the groups Legal and Financial and then Health and Safety.

The most obvious approach to data layout is the use of tables in the form of a data matrix. Within each data matrix, columns will represent a single variable while each row presents a case or profile. Hence,  Table 24.6  illustrates an example of data from a survey of employee attitudes. The second column, labelled ‘Id’, is the survey form identifier, allowing the researcher to check back to the original survey form when checking for errors. The next column contains numbers, each of which signifies a particular department. Length of service is quantifiable data representing actual years spent in the organization, while seniority is again coded data signifying different scales of seniority. Thus, the numerical values have different meanings for different variables. Note that  Table 24.6  is typical of the kind of data matrix that can be set up in a computer program such as SPSS, ready for the application of statistical formulae.

Case Study 24.1  illustrates the kind of survey layout and structure that yields data suitable for a data matrix (presented at the end of the case study). Hence, we have a range of variables and structured responses, each of which can be coded.

Table 24.6 https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table139.jpg

Case Study 24.1 From survey instrument to data matrix

A voluntary association that provides free advice to the public seeks to discover which of its services are most utilized. A survey form is designed dealing with four potential areas, namely the law, finance, health and safety in the home.

Question: Please look at the following services and indicate whether you have used any of them in the last 12 months.

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table140.jpg

The data are collected from 100 respondents and input into the following data matrix using the numerical codes: 1 = Yes; 2 = No. For no data or non-response the cell should be left blank.

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table141.jpg

Note that Respondent 3 has ticked the box for ‘Legal advice’ but has failed to complete any of the others – hence, a ‘0’ for no data has to be put in the matrix.

Dealing with missing data

Oppenheim (1992) notes that the best approach to dealing with missing data is not to have any! Hence, steps should be taken to ensure that data are collected from the entire intended sample and that non-response is kept to a minimum. But in practice, we know that there will be cases where a respondent either has not replied or has not answered all the questions. The issue here is one of potential bias – has the respondent omitted those questions s/he feels uneasy about or hostile to answering? For example, in answering a staff survey on working practices, are those with the worst records on absenteeism more likely to omit the questions on this (hence, potentially biasing the analysis)?

It might be useful to distinguish between four different types of missing values: ‘Not applicable’ (NA), ‘Refused’ (RF), ‘Did not know’ (DK) and ‘Forgot to answer’ (FA). Making this distinction may help you to adopt strategies for coping with this data loss.  Table 24.7  illustrates examples of these responses.

Table 24.7 https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table142.jpg

You may note that the categories for non-response chosen may depend largely on the researcher’s inferences or guesswork. How do we know that someone forgot to answer or simply did not know how to respond? Of course, if many people fail to answer the same question, this might suggest there is something about the question they do not like – in which case, this could be construed as ‘Refusal’. You may decide to ignore these separate categories and just use one ‘No answer’ label. Alternatively, you might put in a value if this is possible by taking the average of other people’s responses. There are dangers, however, in this approach, particularly for single item questions. Note that some statisticians have spent almost a lifetime pondering issues of this kind! It would be safer if missing data were entered for a sub-question that comprised just one of a number of sub-questions (for which data were available). Note, also, that this becomes unfeasible if there are many non-responses to the same question, since it would leave the calculation based on a small sample.

Avoiding the degradation of data

It is fairly clear when non-response has occurred, but it is also possible to compromise the quality of data by the process of degradation. Say we were interested in measuring the age profile of the workforce and drew up a questionnaire, as illustrated in  Figure 24.4 . One problem here is that the age categories are unequal (for example, 18–24 compared with 25–34). But a further difficulty is the loss of information that comes with collecting the data in this way. We have ended up with an ordinal measure of what should be ratio data and cannot even calculate the average age of the workforce. Far better would have been simply to ask for each person’s exact age (for example, by requesting their date of birth) and the date the questionnaire was completed. After this, we could calculate the average age (mean), the modal (most frequently occurring) age and identify both the oldest and youngest worker, etc.

Figure 24.4 Section of questionnaire comprising an age profile

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-fig103.jpg

Presenting data using descriptive statistics

One of the aims of descriptive statistics is to describe the basic features of a study, often through the use of graphical analysis. Descriptive statistics are distinguished from inferential statistics in that they attempt to show what the data are, while inferential statistics try to draw conclusions beyond the data – for example, inferring what a population may think on the basis of sample data.

Descriptive statistics, and in particular the use of charts or graphs, certainly provide the potential for the communication of data in readily accessible formats, but the kinds of graphics used will depend on the types of data being presented. This is why the start of this chapter focused on classifying data into nominal, ordinal, interval and ratio categories, since not all types of graph are appropriate for all kinds of data. Black (1999) provides a neat summary of what is appropriate (see  Table 24.8 ).

Table 24.8 https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table143.jpg

Source: Adapted from Black, 1999: 306

Build: Choosing graph types

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img8.jpg

Employability Skill 24.1 Selecting appropriate graphs and tables for the presentation of information

Selecting appropriate charts and graphs is the key to making your data meaningful and communicating your message to your audience. The aim is to summarize and organize your material in the most easily understood format for the type of data that you have.

Nominal and ordinal data – single groups

As we saw earlier, nominal data are a record of categories or names, with no intended order or ranking, while ordinal data do assume some intended ordering of categories. Taking the nominal data in  Table 24.2 , we can present a bar chart ( Figure 24.5 ) for the frequency count of staff in different departments.

Figure 24.5 Bar chart for the nominal data in  Table 24.2

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-fig104.jpg

Figure 24.6  shows that this same set of data can also be presented in the form of a pie chart. Note that pie charts are suitable for illustrating nominal data but are not appropriate for ordinal data – because a pie chart presents proportions of a total, not the ordering of categories.

Figure 24.6 Pie chart of the nominal data in  Figure 24.5

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-fig105.jpg

Interval and ratio data – single groups

Interval and ratio data describe scores on tests, age, weight, annual income, etc., for a group of individuals. These numbers are then, usually, translated into a frequency table, such as in  Table 24.3 . The first stage is to decide on the number of intervals in the data. Black (1999) recommends between 10 and 20 as acceptable, since going outside this range would tend to distort the shape of the histogram or frequency polygon. Take a look at the data on an age profile of the entire workforce in an e-commerce development organization, presented in  Table 24.9 . The age range is from 23 to 43, a difference of 21. If we selected an interval range of three, this would only give us a set of seven age ranges and conflict with Black’s (1999) recommendation that only a minimum of 10 ranges is acceptable. If, however, we took two as the interval range, we would end up with 11 sets of intervals, as in  Table 24.10 , which is acceptable. We then take these data for graphical presentation in the form of a histogram, as in  Figure 24.7 .

Table 24.9 https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table144.jpg

Table 24.10 https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table145.jpg

Figure 24.7 Histogram illustrating interval data in  Table 24.10

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-fig106.jpg

Nominal data – comparing groups

So far, we have looked at presenting single sets of data. But often research will require us to gather data on a number of related characteristics and it is useful to be able to compare these graphically. For example, returning to  Table 24.2  and the number of employees per department, these may be aggregate frequencies, based on the spread of both male and female workers per department, as in  Figure 24.8 .

Another way of presenting these kinds of data is where it is useful to show not only the distribution between groups, but also the total size of each group, as in  Figure 24.9 .

Figure 24.8 Bar chart for nominal data with comparison between groups

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-fig107.jpg

Figure 24.9 Stacked bar chart for nominal data with comparison between groups

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-fig108.jpg

Interval and ratio data – comparing groups

It is sometimes necessary to compare two groups for traits that are measured as continuous data. While this exercise is, as we have seen, relatively easy for nominal data that are discrete, for interval and ratio data the two sets of data may overlap and one hide the other. The solution is to use a frequency polygon. As we can see in  Figure 24.10 , we have two sets of continuous data of test scores, one set for a group of employees who have received training and another for those who have not. The frequency polygon enables us to see both sets of results simultaneously and to compare the trends.

Figure 24.10 Frequency polygons for two sets of continuous data showing test scores

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-fig109.jpg

Two variables for a single group

You may also want to compare two variables for a single group. Returning once more to our example of departments, we might look at the age profiles of the workers in each of them.  Figure 24.11  shows the result.

Figure 24.11 Solid polygon showing data for two variables: department and age

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-fig110.jpg

Analysing data using descriptive statistics

Watch: Analysing and presenting data

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img7.jpg

A descriptive focus involves the creation of a summary picture of a sample or population in terms of key variables being researched. This may involve the presentation of data in graphical form (as in the previous section) or the use of descriptive statistics, as discussed here.

Frequency distribution and central tendency

Frequency distribution is one of the most common methods of data analysis, particularly for analysing survey data. Frequency simply means the number of instances in a class, and in surveys it is often associated with the use of Likert scales. So, for example, a survey might measure customer satisfaction for a particular product over a two-year period.  Table 24.11  presents a typical set of results, showing what percentage of customers answered for each attitude category to the statement: ‘We think that the Squeezy floor cleaner is good value for money.’

Table 24.11 https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table146.jpg

Comparing the data between the two years, it appears that there has been a 7 per cent rise in the number of customers who ‘Strongly agree’ that the floor cleaner is good value for money. Unfortunately, just to report this result would be misleading because, as we can see, there has also been a 6 per cent rise in those who ‘Strongly disagree’ with the statement. So what are we to make of the results? Given that the ‘Agree’ category has fallen by 7 per cent and the ‘Disagree’ category by 6 per cent, have attitudes moved for or against the product? To make sense of the data, two approaches need to be adopted.

· The use of all the data, not just selected figures that meet the researcher’s agendas.

· A way of quantifying the results using a single, representative figure.

This scoring method involves the calculation of a mean score for each set of data. Hence the categories could be given a score, as illustrated in  Table 24.12 .

All respondents’ scores can then be added up, yielding the set of scores presented in  Table 24.13 , and the mean, showing that, overall, attitudes have moved very slightly in favour of the product.

Table 24.12 https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table147.jpg

Since the data can be described by the mean, a single figure, it becomes possible to make comparisons between different parts of the data or, if, say, two surveys are carried out at different periods, across time. Of course, there are also dangers in this approach. There is an assumption (possibly a mistaken one) that the differences between these ordinal categories are identical. Furthermore, the mean is only one  measure of central tendency , others include the  median  and the  mode . The median is the central value when all the scores are arranged in order. The mode is simply the most frequently occurring value. If the median and mode scores are less than the mean, the distribution of scores will be skewed to the left (positive skew); if they are greater than the mean, the scores are said to be skewed to the right (negative skew). So, while two mean scores could be identical, this need not imply that two sets of scores were the same, since each might have a different distribution of scores.

Table 24.13 https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table148.jpg

Having made these qualifications, this scoring method can still be used, but is probably best utilized over a multiple set of scores rather than just a single set. It is also safest used for descriptive rather than for inferential statistics.

Measuring dispersion

In addition to measuring central tendency, it may also be important to measure the spread of responses around the mean to show whether the mean is representative of the responses or not.

There are a number of ways of calculating  measures of dispersion :

· The  range : the difference between the highest and the lowest scores.

· The inter-quartile range: the difference between the score that has a quarter of the scores below it (often known as the first quartile or the 25th percentile) and the score that has three-quarters of the scores below it (the 75th percentile).

· The variance: a measure of the average of the squared deviations of individual scores from the mean.

· The standard deviation: a measure of the extent to which responses vary from the mean, and is derived by calculating the variation from the mean, squaring them, adding them and calculating the square root. Like the mean, because you are able to calculate a single figure, it allows comparisons to be made between different parts of a survey and across time periods.

Normal and skewed distributions

The  normal distribution  curve is bell-shaped, that is symmetrical around the mean, which means that there are an equal number of subjects above and below the mean (x–). The shape of the curve also indicates the proportion of subjects at each of the standard deviations (SD, 1SD, etc.) above and below the mean. Thus in  Figure 24.12 , 34.13 per cent of the subjects are one  standard deviation  above the mean and another 34.13 per cent below it.

In the real world, however, it is often the case that distributions are not normal, but skewed, and this will have implications for the relationship between the mean, the mode and the median. A distribution is said to be skewed if one of its tails is longer than the other. Where the distribution is positively skewed, it has a long tail in a positive direction (to the right) and the majority of the subjects are below, to the left of the mean in terms of the trait or attitude being measured. With a negative skew, the tail is in a negative direction (to the left) and the majority of subjects are above the mean (to the right).

Figure 24.12 The theoretical ‘normal’ distribution with mean = 0

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-fig111.jpg

The process of hypothesis testing: inferential statistics

Explore: Descriptive vs inferential statistics

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img4.jpg

We saw in  Chapter 3  that the research process may involve the formulation of a hypothesis or hypotheses that describe the relationship between two variables. In this section we will re-examine hypothesis testing in a number of stages, which comprise:

· Hypothesis formulation.

· Specification of significance level (to see how safe it is to accept or reject the hypothesis).

· Identification of the probability distribution and definition of the region of rejection.

· Selection of appropriate statistical tests.

· Calculation of the test statistic and acceptance or rejection of the hypothesis.

Hypothesis formulation

As we saw in  Chapter 3 , a hypothesis is a statement concerning a population (or populations) that may or may not be true, and constitutes an inference or inferences about a population, drawn from sample information.

Let us say that we are interested in the relationship between corporate entrepreneurship and strategic management. According to Schumpeter (1950) entrepreneurship involves the introduction of new products, new methods of production and other innovations. Barrington and Bluedorn (1999) suggest that strategic management involves five dimensions, namely: scanning intensity, locus of planning, planning flexibility, planning horizon and control attributes. Taking just the first dimension, scanning intensity is the managerial activity of learning about events and trends in an organization’s environment, which should yield new business opportunities. Hence, Barrington and Bluedorn (1999) formulate a hypothesis in the following manner:

Hypothesis 1: A positive relationship exists between scanning intensity and corporate entrepreneurship intensity.

However, we can never ‘prove’ something to be true, because there always remains a finite possibility that one day someone will emerge with a refutation. Hence, for research purposes, we usually phrase a hypothesis in its null (negative) form. So, we would state the hypothesis as:

Hypothesis 1: There is no relationship between scanning intensity and corporate entrepreneurship intensity.

Then, if we find that a statistically significant relationship exists, we can reject the  null hypothesis .

Hypotheses come in essentially three forms. Those that:

· Examine the characteristics of a single population (and may involve calculating the mean, median and standard deviation and the shape of the distribution).

· Explore contrasts and comparisons between groups.

· Examine associations and relationships between groups.

For one research study, it may be necessary to formulate a number of null hypotheses incorporating statements about distributions, scores, frequencies, associations and correlations.

Specification of significance level

Having formulated the null hypothesis, we must next decide on the circumstances in which it will be accepted or rejected. Since we do not know with absolute certainty whether the hypothesis is true or false, ideally we would want to reject the null hypothesis when it is false, and to accept it when it is true. However, since there is no such thing as an absolute certainty (especially in the real world!), there is always a chance of rejecting the null hypothesis when in fact it is true (called a  Type I error ) and accepting it when it is in fact false (a  Type II error ).  Table 24.14  presents a summary of possible outcomes.

Table 24.14 https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table149.jpg

What is the potential impact of these errors? Say, for example, we measure whether a new training programme improves staff attitudes to customers, and we express this in null terms (the training will have no effect). If we made a Type I error then we are rejecting the null hypothesis, and therefore claim that the training does have an effect when, in fact, this is not true. You will, no doubt, recognize that we do not want to make claims for the impact of independent variables that are actually false. Think of the implications if we made a Type I error when testing a new drug! We also want to avoid Type II errors, since here we would be accepting the null hypothesis and therefore failing to notice the impact that an independent variable was having.

Type I and Type II errors are the converse of each other. As Fielding and Gilbert (2006) observe, anything we do to reduce a Type I error will increase the likelihood of a Type II error, and vice versa. Whichever error is the most likely depends on how we set the significance level (see following section).

Identification of the probability distribution

What are the chances of making a Type I error? This is measured by what is called the  significance level , which measures the probability of making a mistake. The significance level is always set before a test is carried out, and is traditionally set at either 0.05, 0.01 or 0.001. Thus, if we set our significance level at 5 per cent (p = .05), we are willing to take the risk of rejecting the null hypothesis when in fact it is correct 5 times out of 100.

All statistical tests are based on an  area of acceptance  and an  area of rejection . For what is termed a  one-tailed test , the rejection area is either the upper or lower tail of the distribution. A one-tailed test is used when the hypothesis is directional: that is, it predicts an outcome at either the higher or lower end of the distribution. But there may be cases when it is not possible to make such a prediction. In these circumstances, a  two-tailed test  is used, for which there are two areas of rejection – both the upper and lower tails. For example, for the z distribution where p = .05 and a two-tailed test, statistical tables show that the area of acceptance for the null hypothesis is the central 95 per cent of the distribution and the areas of rejection are the 2.5 per cent of each tail (see  Figure 24.13 ). Hence, if the test statistic is less than -1.96 or greater than 1.96 the null hypothesis will be rejected.

Figure 24.13 Areas of acceptance and rejection in a standard normal distribution with an α of .05

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-fig112.jpg

Selection of appropriate statistical tests

The selection of statistical tests appropriate for each hypothesis is perhaps the most challenging feature of using statistics but also the most necessary. It is all too easy to formulate a valid hypothesis only to choose an inappropriate test, with the result – statistical nonsense! The type of statistical test used will depend on quite a broad range of factors.

Watch: Selecting statistical tests

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img7.jpg

Firstly, the type of hypothesis – for example, hypotheses concerned with the characteristics of groups, compared with relationships between variables. Even within these broad groups of hypotheses different tests may be needed. So a test for comparing differences between group means will be different to one comparing differences between medians. Even for the same sample, different tests may be used depending on the size of the sample. Secondly, assumptions about the distribution of populations will affect the type of statistical test used. For example, different tests will be used for populations for which the data are evenly distributed compared with those that are not. A third consideration is the level of measurement of the variables in the hypothesis. As we saw earlier, different tests are appropriate for nominal, ordinal, interval and ratio data, and only  non-parametric tests  are suitable for nominal and ordinal data, but  parametric tests  can be used with interval and ratio data. Parametric tests also work best with larger sample sizes (that is, at least 30 observations per variable or group) and are more powerful than non-parametric tests. This simply means that they are more likely to reject the null hypothesis when it should be rejected, avoiding Type I errors. Motulsky (1995) advises that parametric tests should usually be selected if you are sure that the population is normally distributed.  Table 24.15  provides a summary of the kinds of statistical test available in the variety of circumstances just described.

Table 24.15 https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table150.jpg

Source: Adapted from Fink, 2003

In the sections that follow, we will take some examples from  Table 24.15  and apply them for the purpose of illustration.

Statistical analysis: comparing variables

In this section and the one that follows, we will be performing a number of statistical tests. It will be assumed that readers will have access to SPSS.

Nominal data – one sample

In the following section we will look at comparing relationships between variables, but here we will confine ourselves to exploring the distribution of a variable. Firstly, if we assume a pre-specified distribution (such as a normal distribution) we can compare the observed (actual data) frequencies against  expected  (theoretical) frequencies , to measure what is termed the  goodness-of-fit .

Let us say that a company is interested in comparing disciplinary records across its four production sites by measuring the number of written warnings issued in the past two years. We might assume that, since the sites are of broadly equal size in terms of people employed, the warnings might be evenly spread across these sites, that is 25 per cent for each. Since the total number of recorded written warnings is 116 (see  Table 24.16 ), this represents 29 expected warnings per site. Data are gathered ( observed frequencies ) to see if they match the expected frequencies. The null hypothesis is that there will be no difference between the observed and expected frequencies. Following our earlier advice, we set the level of significance in advance. In this case let us say that we set it at p = .05. If any significant difference is found, then the null hypothesis will be rejected.  Table 24.16  presents the data in what is called a  contingency table .

Table 24.16 https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table151.jpg

The appropriate test here is the  chi-square distribution . For each case we deduct the expected frequency from the observed frequency, square the result and divide by the expected frequency; the chi-square statistic is the sum of the totals (see  Table 24.17 ). Is the chi-square statistic of 71.86 significant? To find out, we look up the figure in an appropriate statistical table for the chi-square statistic. The value to use will be in the column for p = .05 and for 3  degrees of freedom  (the number of categories minus one). This figure turns out to be 7.81, which is far exceeded by our chi-square figure. Hence, we can say that the difference is significant and we can reject the null hypothesis that there is no difference between the issue of written warnings between the sites.

Table 24.17 Table 24.16  https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table152.jpg

Note, however, that the expected frequencies do not have to be equal. Say, we know through some prior research that site B is three times as likely to issue warnings as the other sites.  Table 24.18  presents the new data.

Table 24.18 https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table153.jpg

Here we find that the new chi-square statistic is only 6.34, which is not significant. Diamantopoulos and Schlegelmilch (1997) warn that when the number of categories in the variable is greater than two, the  chi-square test  should not be used where:

· More than 20 per cent of the expected frequencies are smaller than five.

· Any expected frequency is less than one.

If the numbers within  cells  are small, and it is possible to combine adjacent categories, then it is advisable to do so. For example, if some of our expected frequencies in  Table 24.14  were rather small but sites A and B were in England and site C and D in Germany, we might sensibly combine A with B and C with D in order to make an international comparison study.

Nominal groups and quantifiable data (normally distributed)

Let us say that you want to compare the performance of two groups, or to compare the performance of one group over a period of time using quantifiable variables such as scores. In these circumstances we can use a  paired t-test . If we were to have two different samples of people for which we wish to compare scores, then we would use an independent  t -test . It is assumed that in t-tests the data are normally distributed, and that the two groups have the same variance (the standard deviation squared). If the data are not normally distributed then usually a non-parametric test, the Wilcoxon signed-rank test, can be used – although, as we shall see, t-tests can be used even when the distribution is not perfectly normal. The t-test compares the means of the two groups to see if any differences between them are statistically significant. If the  p -value  associated with t is low (< .05), then there is evidence to accept the alternate hypothesis (and reject the null hypothesis): that is, the means of the two groups are statistically different.

Say that we want to examine the effectiveness of a workplace stress counselling programme. Taking a simple before and after design (recall  Chapter 6  for some of the limitations of this design), we get respondents to complete a stress assessment questionnaire before the counselling and then after it. We can see from the data set provided (see the book’s website and the link to Data sets: t-test data) that in a number of cases the levels of stress have actually increased! But in most cases stress levels have fallen, in some cases quite sharply. Worked Example 24.1 shows how we can use SPSS to see if this is statistically significant.

Practise: t-tests

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img9.jpg

Worked Example 24.1

Type the gain scores for both the experimental and control groups into an SPSS data file. Before we begin any data analysis, we need to determine the normality of the data distribution, since this will influence whether we should use parametric or non-parametric statistical tests. Remember that parametric tests are the more powerful, but can only be used if the data are relatively normally distributed.

For this worked example, use the t-test dataset (available as part of the online resources). Save the data and open it in SPSS.

1. Click on [Analyze], then on [Descriptive statistics] followed by [Explore].

2. Click on Experimental A and Experimental B and move them into the [Dependent List] box by clicking on the arrow.

3. In the [Display] section make sure that [Both] is ticked.

4. Click on [Statistics] and then on [Descriptives] and [Outliers]. Click on [Continue].

5. Click on the [Plots] button. Then under [Descriptive] click on [Histogram]. Select [Normality plots with tests] and [Continue].

6. Click on the [Options] button and in the [Missing values section] select [Exclude cases pairwise]. To complete the process click on [Continue] followed by [OK].

7. You should then see the data as presented in the outputs below.

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table154.jpg

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table155.jpg

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table156.jpg

Only a partial list of cases with the value 16.00 are shown in the table of upper extremes.

Only a partial list of cases with the value 13.00 are shown in the table of upper extremes.

Only a partial list of cases with the value 3.00 are shown in the table of lower extremes.

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table157.jpg

Lilliefors Significance Correction

In the Descriptives output, note the statistic for 5% Trimmed Mean. SPSS removes the top and bottom 5 per cent of cases and recalculates this new mean, to see if extreme scores (outliers) have much impact. In our example above, the mean and trimmed means are very similar so we should not be concerned about outliers distorting the results. The output also provides values for skewness and kurtosis. Skewness provides an indication of the symmetry of the distribution and (as discussed above) can be reported as positive (if scores are clustered to the left) and negative (if clustered to the right). Kurtosis refers to the peakness or otherwise of the distribution. Values of less than 0 indicate a relatively flat distribution; that is, too many cases at the extremes (as in Experimental A example).

Watch: Testing for normality

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img7.jpg

The table labelled Test for Normality contains the Shapiro–Wilk statistic, which is generally used for samples ranging from 3 to 2,000. Above 2,000 the Kolmogorov–Smirnov statistic is generally used to test for the normality of the distribution. A result where the Sig. value is more than 0.05 indicates normality, while a result that is less than 0.05 violates the assumption of normality. Given that the sample size in this study is below 100 we will use the Shapiro–Wilk statistic. In the above table we can see that the statistic for Experimental A is above 0.05, indicating normality, whereas the statistic for Experimental B is below 0.05, violating the assumption of normality. Does this mean that we must use a non-parametric test? Not necessarily. For sample sizes over 30, Pallant (2013) suggests that violation of the normality assumption should not lead the researcher to panic, with use of parametric tests being permissible. The next step is to take a look at the results for Skewness and Kurtosis in the Descriptives table. As long as these are between -1.0 and +1.0, we can assume that the distribution is sufficiently normal for the use of parametric tests.

Hence, we apply the procedure for a  paired sample  t-test as follows:

1. Click on [Analyze] then on [Compare Means] and then on [Paired Samples T-test].

2. Click on the variables Experimental A and Experimental B and on the arrow to move them into the [Paired Variables] box.

3. Click on [OK]. You should see the output as presented below.

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table158.jpg

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table158a.jpg

The procedure for interpreting these results is as follows.

1. Look at the Paired Samples Tests, at the right-hand column labelled Sig. (2-tailed) which gives the probability value. If this is less than 0.05 then we can assume that the difference between the two scores is significant. In our case the Sig. = 0.00 so the differences in the stress scores is, indeed, significant.

2. Establish which set of scores is the higher (Experimental A or Experimental B). The box Paired Samples Statistics gives the mean for each set of scores. The mean for Experimental A was 10.3736 while that for Experimental B was 8.4176. We can therefore conclude that the workplace counselling programme did, indeed, help to reduce stress.

Now a note of caution. Although we obtained differences in the two sets of scores (and the Sig. result suggests that this did not occur by chance alone), we must be careful when it comes to attributing causation. We also need to take into account other factors that could explain the fall in stress levels – refer to ‘Design 3: One group, pre-test/post-test’ in  Chapter 6 . The researcher should try to anticipate the kinds of contaminating factors that could confound the results. One approach would be to improve the research design – for example, by introducing a control group that does not receive the intervention (in this case the stress counselling).

Read: Normality test

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img3.jpg

Nominal groups and quantifiable data (not normally distributed)

In the section above we looked at differences in normally (or near normally) distributed data. But what if the data do not satisfy the assumptions required for statistical tests based on a normal distribution? Let us say that we are exploring the attitudes of men and women towards the purchase of skin care products. Do women prefer these types of product more than men?  Figure 24.14  provides an example of part of a survey dealing in attitudes towards personal grooming. The resulting data from this imaginary survey are provided on the book’s website (see Data sets: Mann-Whitney U data).

The data are captured into an SPSS file, with each questionnaire being allocated its own Id number. Male respondents are allocated the code 1 and females 2. The response of each person is allotted a score by adding their responses. Note that in  Figure 24.14 , question 3 has been posed in a negative form to encourage respondents to think more carefully about their answers. This needs to be allocated a score of 1. Hence, the total score for this respondent would be coded as 6. Total scores for each respondent range from 4 to 20.

Figure 24.14 Example of a portion of a survey on skin care products

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-fig113.jpg

Practise: Mann- Whitney U tests

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img9.jpg

Watch: Doing a Mann-Whitney U test

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img7.jpg

Worked Example 24.2

For this worked example, use the Mann-Whitney dataset (available as part of the online resources). Save the data and open it in SPSS.

First of all, we test for whether the data are normally distributed (see Worked Example 24.1 for how to test for this). Note that as we have both a dependent variable (attitude) and an independent variable (sex), you can generate data for both male and female groups by moving the categorical variable (sex) into the [Factor List] box in the [Explore] dialogue box.

Looking at the Kolmogorov–Smirnov statistic in the Tests for Normality table below, we note that the figure for Sig. is 0.00, indicating that the assumption of normality has been violated. Rather than an independent t-test, we now need to make use of its non-parametric alternative, the  Mann-Whitney U test .

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table159.jpg

Lilliefors Significance Correction

The procedure for the Mann–Whitney U test is as follows:

1. Click on [Analyze], then on [Nonparametric Tests], followed by [2 Independent Samples].

2. Click on the dependent variable [Attitudes] and the arrow to move it into the [Test Variable List] box.

3. Click on the categorical (independent) variable [sex] and the arrow to move this into the [Group Variable] box.

4. Click on the [Define Groups] button. In the [Group 1] box input the number ‘1’, and in the [Group 2] box, input ‘2’ to match sex Id numbers in the data set. Click on [Continue].

5. Click on [Mann-Whitney U] box under the label [Text Type].

6. Click on [Options] and then [Descriptive]. Then click on [Continue] and finally, [OK].

You should see the output as presented below.

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table160.jpg

Grouping Variable: Sex

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table161.jpg

To analyse the data, look at the Test Statistics box for the value of Z and the significance level. The Z value has a significance level of 0.000. Given that this figure is lower than the probability value of 0.05, we can say that this result is significant. Since the result is significant we now need to make reference to the [Ranks] box and particularly the differences between the mean ranks, commenting on which is higher (in our example, it is females).

Note that the Mann-Whitney U test is also useful in other situations. Say, for example, we employ two different training programmes that teach the same topic and want to see which is the most effective. If it cannot be assumed that the data come from a normal distribution, we would use the Mann–Whitney U test to compare the test scores of the two sets of learners.

Statistical analysis: associations between variables

This section examines situations where the study contains two independent variables of the same type (nominal, ordinal, interval/ratio).  Table 24.19  illustrates the different kinds of measurement of association between two variables, depending on the type of variable involved.

Table 24.19 https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table162.jpg

Associations between two nominal variables

Sometimes we may want to investigate relationships between two nominal variables – for example:

· Educational attainment and choice of career.

· Type of recruit (graduate/non-graduate) and level of responsibility in an organization.

You will recall in the discussions about chi-square, above, that we used the statistic to see whether the distribution of a variable occurred by chance or not. Chi-square is appropriate when you have two or more variables each of which contains at least two or more categories.

Let us say that a research team is studying a coaching programme and that a set of interviews with coachees (the recipients of coaching) has indicated that, when it came to a choice of coach, many (both males and females) expressed positive preferences for female coaches. Given that these comments were made by several respondents, the researchers turned to the quantitative data to see whether this was true.  Table 24.20  illustrates the observed values, that is the dataset that shows the gender of coach selected by both female and male coachees. We can see that in both cases, both male and female coachees did, indeed, choose more female than male coaches. But is this difference significant? To find out, we need to use the chi-square statistic. Worked Example 24.3 shows how SPSS can be used for this data analysis.

Table 24.20 https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table163.jpg

Practise: Chi square test

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-img9.jpg

Worked Example 24.3

For this worked example, use the Chi square dataset (available as part of the online resources). Save the data and open it in SPSS.

1. Click on [Analyze] and then on [Descriptive Statistics] followed by [Crosstabs].

2. Click on one of the variables, for example [GenderCoachee], and then click on the arrow to move this variable to the [Rows] box. Then click on [GenderCoach] and then the arrow to move this to the [Columns] box.

3. Click on the [Statistics] button, followed by [Chi-square] and [Phi and Cramer’s V]. Then click on [Continue].

4. Having clicked on the [Cells] button, click on [Observed] in the [Counts] box. In the [Percentages] box, click on [Row], [Column] and [Total].

5. Click on [Continue], followed by [OK]. You should see the output as illustrated below.

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table164.jpg

In analysing the above output, the first step is to ensure that one of the assumptions of the chi-square test has not been violated: that is, that the expected cell frequency should never be less than five. We can see from the footnote (b) under the Chi-Square Tests table that 0 per cent of cells have an expected count of less than five – so we have not violated the assumption. In the study we are discussing, the minimum expected count is, in fact, 33.08.

https://jigsaw.vitalsource.com/books/9781529700527/epub/OEBPS/images/10.4135_9781529700404-table165.jpg

a. Computed only for a 2 × 2 table

b. 0 cells (.0%) have expected count less than 5. The minimum expected count is 33.08