FOR BUSINESS INTELLIGENCE
8. Discriminant Analysis
*
Basic Research Question
- the primary goal is to find a dimension(s) that groups differ on and create classification functions
- i.e., Can group membership be accurately predicted by a set of predictors?
Similarities and Differences between ANOVA, Regression, and Discriminant Analysis
ANOVA REGRESSION DISCRIMINANT
Similarities
Number of One One One
dependent
variables
Number of
independent Multiple Multiple Multiple
variables
Differences
Nature of the
dependent Metric Metric Categorical
variables
Nature of the
independent Categorical Metric Metric
variables
Discriminant Analysis
Discriminant analysis is a technique for analyzing data when the criterion variable (DV) is categorical and the predictor variables (IVs) are interval in nature.
The objectives of discriminant analysis are as follows:
Development of discriminant functions (n groups, n-1 discrininant functions), or linear combinations of the predictor or IVs which will best discriminate between the categories of the DVs(groups).
Examination of whether significant differences exist among the groups, in terms of the predictor variables. (Tests of Equality of Group Means)
Determination of which predictor variables contribute to most of the intergroup differences. (The smaller the variable Wilks' lambda for an independent variable, the more that variable contributes to the discriminant function.)
Evaluation of the accuracy of classification. (classification result table)
Discriminant Analysis Model
The discriminant analysis model involves linear combinations of
the following form:
D = b0 + b1X1 + b2X2 + b3X3 + . . . + bkXk
Where:
D = discriminant score
b 's = discriminant coefficient or weight
X 's = predictor (independent variable)
The coefficients, or weights (b), are estimated so that the groups differ as much as possible on the values of the discriminant function.
This occurs when the ratio of between-group sum of squares to within-group sum of squares for the discriminant scores is at a maximum.
Key terms and concepts
- Discriminating variables = IV, also called predictors.
- Criterion variable =DV, also called the grouping variable in SPSS.
- Discriminant function: A discriminant function, also called a canonical root, is a latent variable (e.g., somebody’s credit) which is created as a linear combination of discriminating (independent) variables, such that L (or D) = b1x1 + b2x2 + ... + bnxn + c, where the b's are discriminant coefficients, the x's are discriminating variables, and c is a constant.
Discriminant analysis
- Discriminant analysis has two steps:
- (1) Wilks' lambda with Chi-Square transformations is used to test if the discriminant model as a whole is significant
- (2) if the F test shows significance, then the individual independent variables are assessed to see which differ significantly in mean by group and these are used to classify the dependent variable.
Assumptions
Discriminant analysis requires following assumptions:
- linear relationships
- homoscedastic
- proper model specification (inclusion of all important independents and exclusion of extraneous variables)
Canonical correlation. is a measure of the association between the groups and the given discriminant function. When it is zero, there is no relation between the groups and the function.
Squared Canonical correlation, Rc2: Squared canonical correlation, Rc2, is the percent of variation in the dependent discriminated by the set of independents in DA or MDA. A canonical correlation square close to 1 means that nearly all the variance in the discriminant scores can be attributed to group differences.
Centroid. The centroid is the mean values of the discriminant scores for a particular group. There are as many centroids as there are groups, as there is one for each group. The means for a group on all the functions are the group centroids.
Statistics Associated with Discriminant Analysis
Unstandardized discriminant function coefficients. are used in the formula for making the classifications in DA, much as b coefficients are used in regression in making predictions. The constant plus the sum of products of the unstandardized coefficients with the observations yields the discriminant scores.
Standardized discriminant function coefficients. The standardized discriminant function coefficients are the discriminant function coefficients and are used as the multipliers when the variables have been standardized to a mean of 0 and a variance of 1.
Discriminant scores. also called the DA score, is the value resulting from applying a discriminant function formula to the data for a given case. The Z score is the discriminant score for standardized data. To get discriminant scores in SPSS, select Analyze, Classify, Discriminant; click the Save button; check "Discriminant scores".
Statistics Associated with Discriminant Analysis
Eigenvalue. For each discriminant function, the Eigenvalue is the ratio of between-group to within-group sums of squares. It reflects the importance of the discriminant function. There is one eigenvalue for each discriminant function. For two-group DA, there is one discriminant function and one eigenvalue. If there is more than one discriminant function, the first will be the largest and most important, the second next most important in explanatory power, and so on.
F values and their significance. F values are calculated from ANOVA, with the grouping variable serving as the categorical independent variable. Each predictor (IV), in turn, serves as the metric dependent variable in the ANOVA.
Statistics Associated with Discriminant Analysis
Structure correlations. Also referred to as discriminant loadings, the structure correlations represent the simple correlations between the predictors and the discriminant function. The correlations then serve like factor loadings in factor analysis.
(model)Wilks' Lambda ( ) . is used to test the significance of the discriminant function as a whole.
(variable) Wilks’ Lambda. Sometimes also called the U statistic, Wilks' lambda for each predictor is the ratio of the within-group sum of squares to the total sum of squares. Its value varies between 0 and 1. Large values
of Wilks’ Lambda (near 1) indicate that group means do not seem to be different. Small values of Wilk’s Lambda (near 0) indicate that the group means seem to be different. Thus, the smaller the variable Wilks' lambda for an independent variable, the more that variable contributes to the discriminant function.
l
Statistics Associated with Discriminant Analysis
Conducting Discriminant Analysis
Formulate the Problem
The criterion variable (DV) must consist of two or more mutually exclusive and collectively exhaustive categories.
The predictor variables (IV) should be selected based on a theoretical model or previous research, or the experience of the researcher.
One part of the sample, called the estimation or analysis sample, is used for estimation of the discriminant function.
The other part, called the holdout or validation sample, is reserved for validating the discriminant function.
Information on Resort Visits: Analysis Sample
- A total sample contains 42 households
- 30 households are included in the analysis sample and the remaining 12 were part of the validation sample.
- Variables:
- Families visited (VISIT) a resort during the last two years: 1, Families didn’t: 2
- Annual income (INCOME)
- Attitude toward travel(TRAVEL): 9 point likert scale
- Important attached to family vacation (VACATION): 9 point scale
- Household size (HSIZE)
- Age of the head of the household (AGE)
Information on Resort Visits: Analysis Sample
Annual Attitude Importance Household Age of Amount
Resort Family Toward Attached Size Head of Spent on No. Visit Income Travel to Family Household Family
($000) Vacation Vacation
1 1 50.2 5 8 3 43 M (2)
2 1 70.3 6 7 4 61 H (3)
3 1 62.9 7 5 6 52 H (3)
4 1 48.5 7 5 5 36 L (1)
5 1 52.7 6 6 4 55 H (3)
6 1 75.0 8 7 5 68 H (3)
7 1 46.2 5 3 3 62 M (2)
8 1 57.0 2 4 6 51 M (2)
9 1 64.1 7 5 4 57 H (3)
10 1 68.1 7 6 5 45 H (3)
11 1 73.4 6 7 5 44 H (3)
12 1 71.9 5 8 4 64 H (3)
13 1 56.2 1 8 6 54 M (2)
14 1 49.3 4 2 3 56 H (3)
15 1 62.0 5 6 2 58 H (3)
Information on Resort Visits: Analysis Sample
Table 18.2, cont.
Annual Attitude Importance Household Age of Amount
Resort Family Toward Attached Size Head of Spent on No. Visit Income Travel to Family Household Family
($000) Vacation Vacation
16 2 32.1 5 4 3 58 L (1)
17 2 36.2 4 3 2 55 L (1)
18 2 43.2 2 5 2 57 M (2)
19 2 50.4 5 2 4 37 M (2)
20 2 44.1 6 6 3 42 M (2)
21 2 38.3 6 6 2 45 L (1)
22 2 55.0 1 2 2 57 M (2)
23 2 46.1 3 5 3 51 L (1)
24 2 35.0 6 4 5 64 L (1)
25 2 37.3 2 7 4 54 L (1)
26 2 41.8 5 1 3 56 M (2)
27 2 57.0 8 3 2 36 M (2)
28 2 33.4 6 8 2 50 L (1)
29 2 37.5 3 2 3 48 L (1)
30 2 41.3 3 3 2 42 L (1)
Information on Resort Visits:
Holdout Sample
Annual Attitude Importance Household Age of Amount
Resort Family Toward Attached Size Head of Spent on No. Visit Income Travel to Family Household Family
($000) Vacation Vacation
1 1 50.8 4 7 3 45 M(2)
2 1 63.6 7 4 7 55 H (3)
3 1 54.0 6 7 4 58 M(2)
4 1 45.0 5 4 3 60 M(2)
5 1 68.0 6 6 6 46 H (3)
6 1 62.1 5 6 3 56 H (3)
7 2 35.0 4 3 4 54 L (1)
8 2 49.6 5 3 5 39 L (1)
9 2 39.4 6 5 3 44 H (3)
10 2 37.0 2 6 5 51 L (1)
11 2 54.5 7 3 3 37 M(2)
12 2 38.2 2 2 3 49 L (1)
Results
In the testing for significance in the vacation resort analysis, we found the Wilks’ lambda is 0.359, which transforms to a chi-square of 26.13 with 5 degrees of freedom. This model is significant (p<0.001).
Conducting Discriminant Analysis
Determine the Significance of Discriminant Function
The null hypothesis that, in the population, the means of all discriminant functions in all groups are equal.
In SPSS this test is based on Wilks' . If several functions are tested simultaneously (as in the case of multiple discriminant analysis), the Wilks' statistic is the product of the univariate for each function. The significance level is estimated based on a chi-square transformation of the statistic.
If the null hypothesis is rejected, indicating significant discrimination.
l
l
Results
The pooled within-groups correlation matrix indicates low correlations between the predictors. Multicollinearity is unlikely to be problem.
The significance of the univariate F ratio indicate that when the predictors are considered individually, only income, importance of vacation, and household size significantly differentiate between these those who visited a resort and those who did not.
Results
Predictors with relatively large standardized coefficients contribute more to the discriminating power of the function.
The relative importance of the predictors can also be obtained by examining the structure correlations between each predictors and the discriminant function represent the variance that the predictor shares with the function. Thus, income, household size, importance attached to vacation, attitudes toward travel, and age of the household are in the more important to less important sequence. The signs of the coefficients associated with all the predictors are positive. This suggests that higher family income, household size, importance attached to family vacation, attitude toward travel, and age are more likely to result in the family visiting the resort.
The group centriod is also given. Group 1, those who have visited a resort, has a positive value (1.291), whereas group 2 has an equal negative value.
Validation of Discriminant Analysis
Conducting Discriminant Analysis
Assess Validity of Discriminant Analysis
Many computer programs, such as SPSS, offer a leave-one-out cross-validation option.
The hit ratio, or the percentage of cases correctly classified, can then be determined by summing the diagonal elements and dividing by the total number of cases. Here it is (12+15)/30=90%. Leave-one-out cross validation correctly classifies only (11+13)/30=80% of the cases. Conduction classification analysis on an independent holdout set of data, we have (4+6)/12=83% hit ratio.
Classification accuracy achieved by discriminant analysis should be at least 25% greater than that obtained by chance.
Given two groups of equal size, by chance one could expect a hit raio of ½=0.5, or 50%. Hence, the improvement over chance is more than 25%, and the validity of the discriminant analysis is judged as satisfactory.
Results for three group DA
Because there are 3 groups, a maximum of 2 functions can be extracted. The eigenvalue associated with the first function is 3.819, and this function accounts for 93.93% of the explained variance. The second function has a small eigenvalue of 0.247 and accounts for only 6.1% of the explained variance.
To test the null hypothesis of equal group centroids, both the functions must be considered simultaneously.
The value of wilk’s lambda is 0.166. this transforms to a chi-square of 44.831 with 10 df, which is significant at 0.05 level (p<0.001).
Thus, the two functions together significantly discriminant among the three groups. However, when the first function is removed, the wilk’s lambda associated with the second function is 0.8020, which is not significant at the 0.05 level. Therefore, the second function does not contribute significantly to group differences.
Results of Three-Group Discriminant Analysis
Income and attitudes toward travel are significant to separate three groups of amount spend on family vacation (high, medium, and low groups) at 0.05 level (p<0.05). The other predictors are not significant at 0.05 level (p>0.05)
Results of Three-Group Discriminant Analysis
Predictors with relatively large standardized coefficients contribute more to the discriminating power of the function. Income is more important predictor than attitude towards travel on function 1, whereas function 2 has relatively larger coefficients for travel, vacation, and age. A similar conclusion is reached by an examination of the structure matrix.
To help interpret the functions, variables with large coefficients for a particular function are grouped together. Those groups are show with asterisks.
Assess validity of discriminant analysis
The classification results indicate that (9+9+8)/30 =86.7% of the cases are correctly classified. Leave-one-out cross-validation correctly classified only (7+5+8)/30=66.7% of the case. When the classification analysis is conducted on the independent holdout sample, a hit ratio of (3+3+3)/12=75% is obtained.
Given three groups of equal size, by chance alone one could expect a hit ratio of 1/3=33%. Thus, the improvement over chance is greater than 25%, indicating a satisfactory validity.
SPSS Windows
The DISCRIMINANT program performs both two-group and multiple discriminant analysis. To select this procedure using SPSS for Windows click:
Analyze>Classify>Discriminant …
- http://www.utexas.edu/courses/schwab/sw388r7/Tutorials/TwoGroupHatcoDiscriminantAnalysis_doc_html/