Statistical Discussion questions

profiledjownson33ah
opening_lecture__chi_square.doc

Chi-Square Tests in Marketing

 Preamble:

 In this article, an attempt is made to bring into sharp focus the use of  2  in marketing function. By no means, the coverage is exhaustive. The aim is to make the reader appreciate the conceptual  framework of Chi-Square analysis through problem illustrations in marketing. The ideas presented in this article certainly can be extended to many decision situations in marketing that can fruitfully employ chi-square tests.

Contents:

1. Chi-Square Analysis-Introduction

2. Chi-Square Test-Goodness of Fit

3. Chi-Square Test of Independence

1. Chi-Square (2) Analysis- Introduction

 Consider the following decision situations:

· Are all package designs equally preferred?

· Are all brands equally preferred?

· Is their any association between income level and brand preference?

· Is their any association between family size and size of washing machine bought?

· Are the attributes educational background and type of job chosen independent?

The answer to these questions require the help of Chi-Square (2) analysis. The first two questions can be unfolded using Chi-Square test of goodness of fit for a single variable while solution to questions 3, 4, and 5 need the help of Chi-Square test of independence (Contingency Table Analysis).

Please note that the variables involved in Chi-Square analysis are nominally scaled. Nominal data are also known by two names: categorical data and attribute data. The symbol 2 used here is to denote the chi-square distribution whose value depends upon the number of degrees of freedom (d.f.). As we know, chi-square distribution is a skewed distribution particularly with smaller d.f. As the sample size and associated the d.f. increase, the 2 distribution approaches normality.

 

2 tests are nonparametric or distribution-free in nature. This means that no assumption needs to be made about the form of the original population distribution from which the samples are drawn. Please note that all parametric tests (tests like the one- and two-sample t-tests we conducted the past two weeks make the assumption that the samples are drawn from a normally distributed set of sample means or proportions.

 

2. Chi-Square Test-Goodness of Fit

A number of marketing problems (and others as well) involve decision situations in which it is important for a marketing manager to know whether a pattern of frequencies that are observed (from a sample) fit well what we will call an “expected” distribution of frequencies.  The expected frequencies are drawn from past experience or from a “hoped for” situation – it represents what the researcher hypothesizes the population distribution to be. The appropriate test is the 2 test of Goodness of Fit Test. The observed frequencies will be compared statistically to the expected frequencies to see how well the observed “fit” the expected. A probability is calculated which is the likelihood that the observed set of frequencies came from population like the “expected.”

The illustration given below will clarify the role of 2 in which only one categorical variable is involved.

Problem: In consumer marketing, a common problem that any marketing manager faces is the selection of appropriate colors for package design.  Assume that a marketing manager wishes to compare five different colors of package design. He is interested in knowing which of the five is the most preferred one so that it can be introduced in the market. A random sample of 400 consumers reveals the following:  

Package Color

Preference by Consumers

Red

70

Blue

106

Green

80

Pink

70

Orange

74

Total

400

Do the consumer preferences for package colors show any significant difference? (Note that in the sample we do observe differences. However, remember that this is just a sample drawn from a population where many, many, many samples of 400 are possible. The question is, could this sample have been drawn from a population where preferences are evenly distributed???

So, if you look at the data, you may be tempted to infer that Blue is the most preferred color. Statistically, you have to find out whether this preference could have arisen due to chance. The appropriate test statistic is the 2 test of goodness of fit.

Null Hypothesis (H0): All colors are equally preferred (observed came from “expected” population).

Alternative Hypothesis (H1): They are not equally preferred  

Package Color

Observed

Frequencies (O)

Expected

Frequencies (E)

image1.png

image2.png

Red

70

80

100

1.250

Blue

106

80

676

8.450

Green

80

80

0

0.000

Pink

70

80

100

1.250

Orange

74

80

36

0.450

Total

400

400

 

11.400

Please note that under the null hypothesis of equal preference for all colors being true, the expected frequencies for all the colors will be equal to 80. Applying the formula

image3.png which = [(70 – 80)2]/80 + [(106 – 80)2]/80 + [(80 – 80)2]/80 +

[(70 – 80)2]/80 + [(74 – 80)2]/80 =

the computed value of chi-square ( 2 ) = 11.40

To find the critical value – the value at which we reject the null – you first need to calculate degrees of freedom. The formula for this is k – 1, meaning the number of categories in the problem – 1 (don’t ask why they use “k” instead of “c” – I don’t know! ( If you find out, let me know.)

This problem assumes the level of significance = .05

Go to Appendix I in your statistical tables provided in the rEsource page. The critical value of   2  at 5% level of significance for 5 – 1 or 4 degrees of freedom is 9.488.

The decision rule, then, is “reject the null if computed 2 is > 9.488.

The null hypothesis is rejected. This means that it is unlikely that the observed sample came from a population where the distribution of preferences is equal – no preferences. So, the inference is that all colors are not equally preferred by the consumers. It seems that blue is the most preferred one – but to say that definitively takes a little more testing which we will not go into now. However, if I am this marketing manager I would take this an indication that I could introduce the blue color package in the market.

3. Chi-Square Test of Independence

The goodness-of-fit test discussed above is appropriate for situations that involve one categorical variable. If there are two categorical variables, and our interest is to examine whether these two variables are associated with each other, the chi-square( 2 )  test of independence is the correct tool to use. This is test is like correlation analysis but correlation analysis requires interval and ratio level data. The Chi-Square Test of Independence (Contingency Table Analysis) is meant for nominal or ordinal level data. . This test is very popular in analyzing cross-tabulations in which an investigator is keen to find out whether the two attributes of interest have any relationship with each other.

Problem: A marketing firm producing detergents is interested in studying consumer behavior in the context of purchase decisions regarding detergents in a specific market. This company is a major player in a detergent market that is characterized by intense competition.  It would like to know whether the income level of the consumers influence their choice of the brand. Currently there are four brands in the market. Brand 1 and Brand 2 are the premium brands while Brand 3 and Brand 4 are the economy brands.

A representative stratified random sampling procedure was adopted covering the entire market using income as the basis of selection. The categories that were used in classifying income level are: Lower, Middle, Upper Middle and High. A sample of 600 consumers participated in this study. The following data emerged from the study.

Cross Tabulation of Income versus Brand chosen (Figures in the cells represent number of consumers)

 

Brands

 

Brand1

Brand2

Brand3

Brand4

Total

Income

 

 

 

 

 

Lower

25

15

55

65

160

Middle

30

25

35

30

120

Upper Middle

50

55

20

22

147

Upper

60

80

15

18

173

Total

165

175

125

135

600

 Analyze the cross-tabulation data above using chi-square test of independence and draw your conclusions.

Solution:

 Null Hypothesis:  There is no association between the brand preference and income level (These two attributes are independent).

 Alternative Hypothesis: There is association between brand preference and income level  (These two attributes are dependent).

 Level of significance = 5%.

In order to calculate the 2  value, you need to work out the expected frequency in each cell in the contingency table. In our example, there are 4 rows and 4 columns amounting to 16 elements. There will be 16 expected frequencies. These are calculated by taking the row total of the cell you are interested in, times the column total of the same cell, divided by the grand total – in this case, 600. If you have MegaStat this is done for you automatically.

Observed Frequencies (These are actual frequencies observed in the survey)

 

Brands

 

Brand1

Brand2

Brand3

Brand4

Total

Income

 

 

 

 

 

Lower

25

15

55

65

160

Middle

30

25

35

30

120

Upper Middle

50

55

20

22

147

Upper

60

80

15

18

173

Total

165

175

125

135

600

Expected Frequencies (These are calculated on the assumption of the null hypothesis being true: That is, income level and brand preference are independent)

 

Brands

 

Brand1

Brand2

Brand3

Brand4

Total

Income

 

 

 

 

 

Lower

44.000

46.667

33.333

36.000

160.000

Middle

33.000

35.000

25.000  

27.000

120.000

Upper Middle

40.425

42.875

30.625

33.075

147.000

Upper

47.575  

50.458

36.042

38.925

173.000

Total

165.000

175.000

125.000

135.000

600.000

Note: The fractional expected frequencies are retained for the purpose of accuracy. Do not round them.

Calculation:

 Compute

image4.png.

There are 16 observed frequencies (O) and 16 expected frequencies (E). As in the case of the goodness of fit, calculate this 2  value. In our case, the computed 2  =131.76 as shown below: Each cell in the table below shows (O-E)2/(E)

 

Brand1

Brand2

Brand3

Brand4

Income

 

 

 

 

Lower

8.20

21.49

14.08

23.36

Middle

0.27

2.86

4.00

0.33

Upper Middle

2.27

3.43

3.69

3.71

Upper

3.24

17.30

12.28

11.25

and there are 16 such cells. Adding all these 16 values, we get 2  =131.76

The critical value of 2  depends on the degrees of freedom. The degrees of freedom = (row total -1) x (column total -1). In our case, there are 4 rows and 4 columns. So the degrees of freedom = (4-1) x (4-1) = 3 x 3 = 9. At  5% level of significance, critical 2  for 9 d.f = 16.92 (Appendix I).

The decision rule is: reject the null hypothesis and accept the alternative hypothesis if computed 2 is > 16.92.

Since 131.76 is > 16.92 we will reject the null. It is unlikely – very improbable that – that the observed sample came from a population like the “expected” distribution.

The inference is that brand preference is highly associated with income level. Thus, the choice of the brand depends on the income strata. Consumers in different income strata prefer different brands.  From this sample distribution it seems that consumers in upper middle and upper income group prefer premium brands while consumers in lower income and middle-income category prefer economy brands. The company should develop suitable strategies to position its detergent products. In the marketplace, it should position economy brands to lower and middle-income category and premium brands to upper middle and upper income category.

Adapted from paper by P.K. Viswanathan

Adjunct Professor and Management Consultant

Chennai-India