BUS 308 Assignment
1
Introduction to Statistics for Managers
Learning Objectives
After reading this chapter, you should be able to:
• Explain the importance of statistics for managers.
• Demonstrate an understanding of basic statistical symbols and terminology.
• Distinguish between populations and samples.
• Compare and contrast nominal, ordinal, interval, and ratio data scales.
• Distinguish between descriptive and inferential statistical analyses.
• Calculate statistics describing what is most typical, and how much variety there is in a data set.
iStockphoto/Thinkstock
tan81004_01_c01_001-024.indd 1 2/22/13 3:30 PM
CHAPTER 1Section 1.1 Regarding the Math Involved . . .
Chapter Outline
1.1 Regarding the Math Involved . . .
1.2 The Notation Populations and Parameters Samples and Statistics
1.3 Types of Data Scales Nominal Data Ordinal Data Interval Data Ratio Data
1.4 Describing or Inferring?
1.5 Descriptive Statistics Measures of Central Tendency Measures of Variability Populations Versus Samples and a Correction for a Biased Estimator Degrees of Freedom About Being Reasonable
1.6 Calculating Descriptive Statistics With Excel Navigating Excel Entering the Command for the Mean Using the Descriptive Statistics Option
1.7 Dependent and Independent Variables
Chapter Summary
Introduction
The “information age” has transformed society and with it, the roles and responsibili-ties of managers. In particular, computer technology, with all of its elements including the World Wide Web, has made an astounding array of data available. The sheer amount of data available makes its management, organization, and analysis a much larger task than it used to be. Managers need to learn to efficiently and effectively access, filter, sum- marize, interpret, and report data. Happily, the greater volume of data available because of computers is also much easier to manage because of those same computers and easy- to-use data analysis software.
1.1 Regarding the Math Involved . . .
Indeed, the business field embraced computer technology very early. It may seem ironic, therefore, that in a book written for students of a discipline that pioneered computer
tan81004_01_c01_001-024.indd 2 2/22/13 3:30 PM
CHAPTER 1Section 1.2 The Notation
use, a good deal of hand calculation is expected. The rationale for hand calculations is this: the logic involved in statistical analysis is never clearer than it is when some of the calcula- tions are completed by hand. The hand calculations will be limited to small data sets, but that work will be an important component of your progress.
It’s likely that management students study statistical analysis less out of a passion for sta- tistical analysis per se than for the direction it can provide to decision-makers. Generally speaking, statistical analysis is simply a means to an end. The same is true of the math. Its value is in the answers it provides to important questions. The derivations of the differ- ent formulas and the nuances of the statistical theory that support them are less relevant than the business-related questions that the formulas let us address. While this book will contain some discussion of statistical theory, it is elementary in nature and limited to what you, as a management student, need in order to understand why the related procedures operate the way they do. Understanding the formulas will involve nothing more complex than the math from that first secondary school algebra course. It will be helpful to review issues related to the order of mathematical operations since some of the formulas involve several calculations, but if you can remember whether to multiply or add first, and where parentheses and exponents fit into the mix, all will be well. Recall that the “please excuse my dear Aunt Sally,” mnemonic device taught that when there are multiple mathematical calculations involved, first:
• do what’s in the parentheses; • then deal with any exponents (any squaring, for example); • then do any multiplication and division, working from left to right, and • last, complete any addition and subtraction problems, working from left to
right.
Dedicated statistical packages are generally quite accessible and easy to use, but this book specifically emphasizes the Excel® spreadsheet program in Microsoft Office because it is more readily available to managers. Although you will be performing the calculations by hand, each chapter will contain a section on how Excel can be used to complete some of the same calculations and procedures.
1.2 The Notation
By tradition, Greek letters are used to signify statistical procedures that are used repeat-edly, or values that are frequently part of a formula. The benefit of symbols in statistics is in the economy they provide. Because the same procedures and values tend to be used repeatedly, it is much simpler to insert a symbol for a value or procedure than it is to pro- vide a narrative description of what is occurring each time. For example, the upper-case Greek letter sigma () indicates “summation.” That symbol suggests that several values are to be added. When paired with “x,” the x expression (said “summation ex”) indicates that the several values in a data set, each designated “x,” are to be added. If x refers to monthly sales figures for the year, x means that those figures for each month are added to produce the annual total. The sigma symbol means the same thing in Excel.
tan81004_01_c01_001-024.indd 3 2/22/13 3:30 PM
CHAPTER 1Section 1.3 Types of Data Scales
Populations and Parameters
Greek letters also sometimes indicate the characteristics of populations. A population includes all members of a defined group. All members of a particular year in school at a university are a population, as are all blue-collar workers or all retail outlets in the county.
As that last example indicates, in its statistical usage, the word “population” need not refer to people.
When the reference is to a population’s average, referred to as the “mean” in statistics, the symbol is the lower-case Greek letter μ. It is spelled “mu” and pronounced “mew” (like the cat, not the cow). A statement that the assistant managers of local banks have a mean age of 38.750 would be indicated
μ 5 38.750. All population characteristics, including this mean value, are referred to as parameters.
Samples and Statistics
If a population is all members of a defined group, a sample is any group involving anything less than all of the population. Remove one member of a popu- lation, and what’s left is a sample. The language is used quite loosely, but technically the word statistic refers to a characteristic of a sample.
Although the symbols used for population param- eters are quite consistent, the symbols used for statis- tics reflect some variety. It is common, for example, to see the mean of a sample represented with x (“ex-bar”). Rather than that symbol, the practice here will be to indicate the sample mean as an upper-case M, a convention that has become common in scholarly journals reporting quantitative analyses.
1.3 Types of Data Scales
The pioneering American psychologist E. L. Thorndike is quoted as saying that any-thing that exists, exists in some quantity, and if it exists in some quantity it can be measured. Managers generally find themselves in sympathy with that expression in their efforts to measure everything from job satisfaction to achievement motivation to verbal ability, which do not lend themselves to quantification as readily as more “objective” data on production and sales. Many different characteristics need to be quantified so that they can guide business decisions.
The different types of measurement data are classified in terms of the kind of informa- tion they provide, a characteristic referring to the scale of the data. There are several approaches to doing classification. It is common to distinguish between categorical and continuous data, for example. Categorical data, also referred to as nominal data, reflect
Review Question A: What are the charac- teristics of populations called?
ˉ
Key Terms: A population includes all members of a defined group. Any group less than the population is a sample. Popula- tion characteristics are param- eters. Sample characteristics are statistics.
tan81004_01_c01_001-024.indd 4 2/22/13 3:30 PM
CHAPTER 1Section 1.3 Types of Data Scales
mutually exclusive categories such as gender, ethnicity, political party affiliation, or mari- tal status. Continuous data involve measures of qualities that differ by degree, such as level of income, or height, or ambition. There are three types of continuous data: ordinal data, interval data, and ratio data, and then there are the nominal scale data noted. Every- thing the manager measures fits one those four scales.
Nominal Data
Besides gender, ethnicity, and marital status, nominal classifications can include charac- teristics as diverse as state or city of residence, the university attended (although U.S. News and World Report, and the Princeton Review are inclined to rank them, which associ- ates them with the next category), and whether one voted in the last election. There are many nominal classifications that are important to business analyses. Determining the brand of car someone drives, the credit card a consumer prefers to use, and the website people rely on for their online shopping involves gathering nominal data.
Although nominal data can be assigned numbers, as when ethnic origin is labeled with “1” for African American, “2” for Asian American, and so on, the numbers serve only as identifiers. They have no mathematical meaning. Their function is purely to label the individual. As a result, the statistics that are calculated to summarize nominal data, and the procedures that are used to analyze them have to be different than they are for data for which the numbers are mathematically relevant. Correspondingly, most of the analy- ses where the data are nominal revolve around the frequency with which a particular value occurs.
Ordinal Data
Ordinal data values reflect a ranking of the measured characteristic. A higher number indicates more of whatever is measured. However, the precise amount of whatever is measured isn’t evident from the ranking. A higher number does not indicate how much more of the quality the individual possesses. For example, probably everyone has taken a Likert-type survey at some point. The respondent indicates her or his level of agreement or disagreement with a statement that is presented. As an example, someone doing mar- ket research for a supermarket might ask exiting customers to complete a survey includ- ing items such as
I find the prices at Evergreen Market to be very competitive, and
I find the customer service at Evergreen Market to be very helpful.
To which the responses to each item are choices such as
• Strongly Agree • Agree • Disagree • Strongly Disagree
tan81004_01_c01_001-024.indd 5 2/22/13 3:30 PM
CHAPTER 1Section 1.3 Types of Data Scales
Once respondents have indicated their choices, the person summarizing the data will prob- ably assign a numeric value to each response, perhaps 4 for strongly agree, 3 for agree, and so on. Clearly a “strongly agree” is a more positive response than “agree,” and “agree” is more positive than “disagree,” but ordinal data do not indicate how much more positive.
When there are several items, as there usually are on a survey, differences in the total scores for two respondents will indicate who was the more positively disposed toward Evergreen Market. But ordinal data cannot indicate how much difference there is between the two total scores. The only accurate interpretation will be that the higher score indicates the most positive responses.
Interval Data
With interval data the difference in the amount of whatever is measured is the same between any two consecutive whole numbers. The increase from 5 to 6 will be the same as the increase from 13 to 14. A one-point difference will be the same amount of increase or decrease, wherever it occurs along the range of values. For example, aptitude tests often used in employee selection decisions are considered interval scale. The difference between someone who scores 50 and someone who scores 60 on a test will be the same as the dif- ference between 2 others who score 70 and 80.
One of the hallmarks of interval data is that a measure of zero does not mean the absence of the measured characteristic. In the example above, if an applicant does not answer any of the questions on an aptitude test correctly, it probably does not mean that the individ- ual has no aptitude at all. By analogy, 0 degrees on the Fahrenheit temperature scale does not mean the absence of heat, as anyone who has lived in snow country would know. In either of those examples, a 0 is just a point on the scale midway between 11 and 21, or for that matter, between 110 and 210. This indicates that Fahrenheit temperature is an interval scale measure.
Ratio Data
Ratio scale data have all the characteristics that inter- val data have, but with two important additions. First, a zero in ratio scale measure indicates that the characteristic measured is entirely absent. Sec- ond, as suggested by its name, ratio data allow for ratio comparisons. Someone who generates $90,000 in sales has generated 3 times the sales of the per- son who is responsible for $30,000 in sales. Someone who has been with the company for 20 years has been there 10 times as long as someone who has been there only 2 years.
Such comparisons make sense with ratio data, but they are problematic with interval data. In the aptitude example, a human resources specialist would probably be reluctant to say that someone who scores 60 on an aptitude test has only half the aptitude of a person who scores 120. There are too many other factors that might explain the score difference to be comfortable with such an explanation.
Key Terms: Nominal data indicate an individual’s category. Ordinal data allow ranking. Interval data have consistent differences between consecutive measures. With ratio data, zero indicates the absence of the qual- ity measured.
tan81004_01_c01_001-024.indd 6 2/22/13 3:30 PM
CHAPTER 1Section 1.4 Describing or Inferring?
Ratio data are quite common in business. As long as the figures are derived the same way in each case, an unemployment rate of 7% means that twice as many people are unem- ployed as when unemployment is at 3.5%. Fast-food sales of $32,000 in an outlet are twice the sales of another outlet for which the sales for same the period total $16,000.
In psychology and other social sciences, by contrast, ratio data are quite rare. Charac- teristics like intelligence, anxiety, and so on, are measured on an interval scale at best. Although the difference between interval and ratio data is very important for what each can reveal about some measured characteristic, the difference is unimportant from the standpoint of analytical procedure. Statistics and procedures that are appropriate for interval scale data are also appropriate for ratio scale data. So as far as analysis is concerned, we’ll no longer distinguish between the two scales. From this point for- ward, the reference will be to nominal data, ordinal data, and interval/ratio data.
1.4 Describing or Inferring?
Sometimes the analytical task is to summarize the characteristics of a group. This is the domain of descriptive statistics. Rather than repeating every measure for what may be quite a large group, analysts find it helpful to calculate descriptors that represent the essence of the group without the burden of repeating each data point.
There are many kinds of descriptive statistics. In this first chapter, we will concentrate on calculating measures of central tendency and measures of variability. As their names suggest, measures of central tendency indicate what is most typical in a data set. Measures of variability gauge how much difference there is in a set of measures. Because there are many ways to define what is typical, as well as many ways to analyze differences, there are multiple indicators of central tendency and variability. Each measure provides differ- ent information.
For some analytical tasks, descriptive statistics are all that is needed. If the task is to summarize the produc- tivity of workers in a manufacturing plant, perhaps statistics that indicate how many components the average worker assembles and how much the per- formance of individual workers tends to vary from that mean will suffice. Often, however, the descrip- tive data are a means to a different end. When certain conditions prevail, the sample can be a mechanism for understanding the characteristics of the more- difficult-to-access population, which means that the
descriptive statistics for samples can provide a window into the characteristics of the pop- ulation. This is the domain of inferential statistical analysis.
When the population is defined to be the management team at a local sandwich shop, perhaps access to population data is not a problem. But for the business analyst, the infor- mation that the local team can provide about conditions throughout the industry is prob- ably quite limited. What will be needed is a more representative sample, one that more
Key Terms: Measures of cen- tral tendency and variability are descriptive statistics that refer to what is typical and to how much variety there is in a data set. Inferential statistics generalize the characteristics of samples to populations.
tan81004_01_c01_001-024.indd 7 2/22/13 3:30 PM
CHAPTER 1Section 1.5 Descriptive Statistics
probably represents the characteristics of all management teams. When such a sample is assembled, its characteristics can provide important information about how all manage- ment teams function.
1.5 Descriptive Statistics
This type of data informs the descriptive statistics that ought to be calculated and show which analytical procedures are appropriate. What can be done with nominal scale data, both by way of description and analysis, is different from what can be done with interval scale data. Earlier we noted that what is most typical in a data set is indicated by the central tendency measures, and the measures of variability or dispersion indicate how much variety there is in a set of measures. There are other descriptive statistics beyond those, but central tendency and variability are a good place to begin.
Measures of Central Tendency
There are three different measures of central tendency. Each measure takes a different approach to describing what is “typical.”
The Mean
As mentioned earlier, the mean (M) is the arithmetic average of a set of interval or ratio values. For example, suppose that a call center manager measured the number of minutes a customer service representative spends on the phone with each caller, and the results were: 3, 4, 4, 5, 5, 5, 6, 6, 7, 7, 8. The formula for calculating the mean is:
Formula 1.1 M 5 x/n
Where
M 5 the mean 5 summation x 5 each value in the number set n 5 the number of values or scores
For the minutes on the phone data:
x 5 60 This is the sum of the number of minutes spent on all the calls. n 5 11 This is the number of calls.
x/n 5 60/11 5 5.450 M 5 5.450
The average (mean) number of minutes spent on customer service calls is 5.450.
tan81004_01_c01_001-024.indd 8 2/22/13 3:30 PM
CHAPTER 1Section 1.5 Descriptive Statistics
The Median
When scores are arranged in order either from largest to smallest or from smallest to larg- est, the median (Mdn) is the middle score. The median requires data of at least ordinal scale and can also be calculated for interval/ratio data. The median is not calculated for nomi- nal data since the numbers are only category labels, as discussed earlier. Using the same data above, the median is the middle-most ranking. With 11 values, the median call length is the 6th, which in this case is 5; Mdn 5 5. If the 8-minute call were eliminated from the
data set above, the result would be calls of 3, 4, 4, 5, 5, 5, 6, 6, 7, 7 minutes. Here, because there is an even number of calls, the median is the average of the time for the two middle calls. The 5th and 6th calls were both for 5 minutes: 5 1 5 5 10 4 2 5 5. With the 8-minute call eliminated, the median remains 5.
The Mode
The mode (Mo) is the value that occurs most frequently in a data set. Using the same data above, the mode is 5 minutes per call; Mo 5 5, since that value occurs most frequently in this data set. Based on frequency alone, the mode is a relatively crude measure, but the discussion in Chapter 2 will indicate how the mode is related to data normality.
When data are nominal, the mode is the only measure of central tendency that ought to be calculated. As noted above, when nominal data are counted, they are often numerically coded according to category. However, these numerical labels have no mathematical meaning. They serve as category iden- tifiers only. It is for that reason that calculating the median or mean of nominal data makes little sense. The mode is often calculated for ordinal, interval, and ratio data as well, but it’s the only measure of central tendency that makes sense for nominal data. For ordinal or interval/ratio data, the mode is usually reported along with other measures of central tendency.
Measures of Variability
Variability measures typically go hand-in-hand with measures of central tendency. In addition to what is most typical, it is helpful to know how much individual measures tend to vary from what is most typical. Two vendors might sell the following number of cell phone contracts in a five-day week:
Vendor A: 6, 6, 6, 6, 6
Vendor B: 1, 0, 28, 1, 0
Review Question B: For ordinal data, which measure(s) of central tendency is/ are appropriate?
Key Terms: The mode is the most frequent value. The median is the middle-most value. The mean is the average.
tan81004_01_c01_001-024.indd 9 2/22/13 3:30 PM
CHAPTER 1Section 1.5 Descriptive Statistics
The standard deviation:
Both vendors have a mean for daily sales of M 5 6.0, but their day-to-day performances are very different. Those differences are indicated not by central tendency measures, but by a measure of data variability.
The more variable the measures are, of course, the larger the values are that indicate variability. Data that are homogeneous produce relatively small vari- ability values.
The Range
The range (R) is the easiest measure of data variabil- ity to calculate. It is the difference between the high- est and lowest values. For vendor A, R 5 6 2 6 5 0. For vendor B, R 5 28 2 0 5 28.
It’s common to hear people say something like “scores ranged from __ to __,” but technically the range is just one value. It indicates only the difference between the highest and lowest measures, which makes the range not very informative when it is reported alone. If a report to the company president indi- cated that vendor A’s sales have R 5 0, there is no way to know how many contracts were sold or what the lowest and highest values were. Another vendor that consistently sold 100 contracts every day of the week would have had the same R 5 0 despite performing much better than vendor A.
The Variance and the Standard Deviation
The variance (s2) and the standard deviation (s) are related statistics that provide a dif- ferent view of data variability from what the range provides. The variance and standard deviation are measures of how much individual values tend to differ from the mean (M) of the group. As can be inferred from their symbols, the formulas for these two statistics are similar:
The variance:
Key Terms: The range is the difference between the highest and lowest score. The standard deviation and variance both measure the typical amount of variability between the mean and individual measures in the group. The variance is the square of the standard deviation.
Review Question C: How many values in the data set contrib- ute to the range? How many contribute to the value of the stan- dard deviation?
Formula 1.2 s2 5 S1x 2 M2 2
n 2 1
Formula 1.3 s 5 Å S1x 2 M2 2
n 2 1
tan81004_01_c01_001-024.indd 10 2/22/13 3:30 PM
CHAPTER 1Section 1.5 Descriptive Statistics
In the case of either formula,
5 summation
x 5 each score in the sample
M 5 the mean of the data set
n 5 the number of scores in the data set
If the standard deviation (s) is squared (s2), the result is the variance. Or working from the other direction, taking the square root of the variance produces the standard deviation. The steps for calculating the variance are:
1. Determine the mean for the group, M. 2. Take the difference between each individual score in the group and the mean
(x 2 M for each score). 3. Square the difference from all of the x 2 M calculations (x 2 M)2. 4. Sum all of the squared differences, (x 2 M)2. 5. Divide the sum of the squared differences by the number of scores minus one,
n 2 1.
Using vendor B data (1, 0, 28, 1, 0) the procedure is:
1. M 5 x/n 5 30/5 5 6 2. Subtract M from each x.
1 2 6 5 25 0 2 6 5 26 28 2 6 5 22 1 2 6 5 25 0 2 6 5 26
3. Square each result of x – M, remembering that squaring a negative number makes it positive.
252 5 25 262 5 36 222 5 484 252 5 25 262 5 36
4. Sum the squared differences.
25 1 36 1 484 1 25 1 36 5 606
5. Divide by the number of scores (n), minus 1.
s2 5 606/4 5 151.50
The one additional step needed to derive the standard deviation is taking the square root of the variance:
s 5 "s2 5 "151.50 5 12.31
Following the same procedure for vendor A’s sales data (6, 6, 6, 6, 6) will indicate that the variance and the standard deviation values are both 0. This occurs because there is
tan81004_01_c01_001-024.indd 11 2/22/13 3:30 PM
CHAPTER 1Section 1.5 Descriptive Statistics
no variation in the sales data across the five days. At the heart of Formula 1.2 and 1.3 are those repeated x – M difference calculations. They should serve as a reminder that s and s2 measure the distance between individual measures in the data set (x) and the mean (M) of the data set.
Since subtracting each individual value (x) from the mean (M) indicates how far indi- vidual data points are from M, why is it necessary to square the result? The answer is that whenever when x . M the difference will be positive, and when x , M the difference will be negative. Summing the positive and negative differences without squaring them will result in a value near zero, which will be little help in understanding data variability. The sum of the squares of those differences will indicate the absolute amount of variability between individual measures and the mean of the data set, regardless of the direction of variability.
It can be difficult to know what constitutes a small or large standard deviation of variance value. Making that judgment will be simplified in Chapter 2, when the standard deviation is explained in terms of how it compares to the range.
The Impact of Different Score Values
For vendor B, the highest number of cell phone contracts sold was 28. If it had been 12 instead, note the impact on the variance and standard deviation values. Since 12 is a less extreme value than 28, including it diminishes the average variation between individual scores and the mean. The result is that:
s2 becomes 26.7, rather than 151.5, and
s becomes 5.167, rather than the original 12.310
Because both statistics are based on the square of the difference between individual scores and the mean, and because extreme scores produce the largest squared differences between x and M, extreme values in either direction have a disproportionate effect on the size of s2 and s.
The effect of lower values in the data set creates circumstances that provide an interesting contrast between variance and standard deviation statistics on the one hand, and the range on the other. Once the range for a set of values is established, no additional value can shrink it. The range can increase in size with addition of scores that are farther from the mean than those already included, but it cannot shrink.
In this chapter the standard deviation and variance are cal- culated only for their descriptive value. That will change later when these statistics become part of more involved procedures.
Review Question D: Why is it that either very high or very low scores have a dispro- portionate impact on the value of the standard deviation and the variance?
tan81004_01_c01_001-024.indd 12 2/22/13 3:30 PM
CHAPTER 1Section 1.5 Descriptive Statistics
A Comment About Hand Calculators
Hand calculators with a built-in standard deviation function make it possible to enter the values and produce the statistics without the several x 2 M steps. The directions vary somewhat depending upon the particular calculator, but most have a key marked some- thing like “sxn21” or “sn21” for the standard deviation. The s is the lowercase letter sigma, which is the Greek equivalent of s. It is often used as a symbol for the standard deviation. A s2 key for the variance isn’t as common since it’s the lesser-used of the two statistics, and in any case, using the x2 key on the calculator to square the standard deviation will produce the variance.
Populations Versus Samples and a Correction for a Biased Estimator
Earlier, the population was defined as every member of a particular group. The sample is any subset of the population. Every municipal employee in San Diego defines a popula- tion, as do all law enforcement officers in the state, or everyone in your family. Remove at least one individual from any of those populations, and the group becomes a sample.
A “representative” sample has descriptive characteristics that are very similar to those of the population. One of the limitations in samples is that their data tend to be less variable than population data. Because this underrepresentation of population variability tends to be consistent across samples, it represents bias. To counter the bias, an adjustment is made in the formulas for sample variances and sample standard deviations called a “correc- tion for a biased estimator.” The correction is the “2 1” in the denominators of Formulas 1.2 and 1.3. In any division problem, if the divisor is reduced, the resulting quotient gets larger. The influence that the “correction for a biased estimator” has is greatest when the samples are the smallest. If n 5 10, the “2 1” impact on the quotient is proportionately greater than when n 5 100:
• 10 4 9 5 1.111 • but 100 4 99 5 1.010.
That the impact of the correction diminishes as the sample size increases stands to reason since the potential for the sample to distort population characteristics will generally be greatest when sample sizes are smallest. If all the data are available for every possible member of a group, there won’t be any concern for bias, and the use of the sample stan- dard deviation and variance with their n 2 1 adjustments is unnecessary. The formulas of population variance (s2) and population standard deviation (s) are as follows:
Formula 1.4 s2 5 S1x 2 m2 2
n
Formula 1.5 s 5 Å S1x 2 m2 2
n
tan81004_01_c01_001-024.indd 13 2/22/13 3:30 PM
CHAPTER 1Section 1.5 Descriptive Statistics
Note three changes:
• The “2 1” is gone from the denominators. • Sigma (s), signifying the population standard deviation, has been substituted
for s. • Mu (m), signifying the population mean, has been substituted for M.
Differentiating Sample and Population Characteristics
Earlier we noted that the descriptive characteristics of popu- lations are referred to as parameters, and although the word is used more loosely, the term “statistic” refers to a sample characteristic. To summarize:
Characteristics
Population Sample
Parameter Statistic
Mean µ M
Standard Deviation s s
Unless the population is defined in very restrictive terms (all tire retailers in 1 square mile of the city), managers rarely have population data. The analyses that they must perform are much more likely to be based on sample data. As a result, it will be the sample standard deviation, using Formula 1.3, that will be calculated from this point on in the book.
Degrees of Freedom
The n 2 1 correction in the variance and standard deviation formulas for samples gives rise to another topic called degrees of freedom. Degrees of free- dom are one of those odd, statistical abstractions that are very difficult to explain briefly, but that affect the way the t-tests (Chapter 4), analysis of variance (Chapter 5), and other statistical procedures that will come up later in the book are computed and inter- preted. Degrees of freedom are the number of scores in a calculation that are free to vary when the final result of the calculation is known. If the sum of three integers is 6:
_ 1 _ 1 _ 5 6,
Then two of those three integers can have any value—they could be 2 1 2, or 3 1 1, or 23 1 10, or any other two values—as long as the value of the third integer makes the result come out to 6. The third value can’t vary; it must be 2, or 2, or 21 in the 3 examples. In this case, the problem has 2 degrees of freedom.
Review Question E: Why are the mean, standard deviation, and variance inappro- priate for ordinal data?
Key Terms: A problem’s degrees of freedom are the number of values free to vary when the value of the result is fixed.
tan81004_01_c01_001-024.indd 14 2/22/13 3:30 PM
CHAPTER 1Section 1.6 Calculating Descriptive Statistics With Excel
If the number of integers that make up the problem is the value of n, then degrees of freedom (df ) for a problem like this are n 2 1. That same expression, n 2 1 also defines degrees of freedom for the standard deviation and the variance. If the final value of either s2 or s is known, all of the scores except for one can have any value. But that final value must be whatever makes the value of s or s2. Other procedures will have df values that differ, but the sample standard deviation and the sample variance have n 2 1 degrees of freedom.
About Being Reasonable
In the course of calculating variances or standard deviations, it is important to be reminded of some boundaries. For any measure of variability the lowest possible value is zero; there is no such thing as negative variance. If the accountant at a construction firm examines the cost of material orders over time and finds that any of s2, s, or R equal zero, it means that all the orders were for the same amount. If a negative value emerges for any of the calculations, the accountant should look for an error.
1.6 Calculating Descriptive Statistics With Excel
Even with hand calculators that include statistical functions, determining the standard deviation, or variance, or even the mean for large data sets can be tedious. Excel, how- ever, deals with such tasks readily. Like all business spreadsheets, Excel is laid out so that data can be arranged in rows, columns, or both. Although the commands vary for different kinds of software, all spreadsheets will produce descriptive statistics and most programs, including Excel, will also complete some of the basic statistical tests.
As a point of information, the command structure in Excel changes modestly with the different versions of the software. The commands provided in this book are specific to the version of Excel in Office 2010. Earlier versions of the software may have slightly differ- ent commands, but the processes described are always accessible. A brief Web search will usually provide the alternate approach.
There are two ways to get descriptive statistics in Excel. They can be calculated directly by entering the commands for the standard deviation, or for whatever value is needed. The standard deviation can also be produced as part of the “Descriptive Statistics” command. It will be done here both ways, first with the individual commands.
An account manager at an advertising firm sends out a customer satisfaction survey to the 12 clients she developed advertising campaigns for over the last year. The 12 clients register the following scores on the survey the manager sent: 11, 14, 14, 15, 17, 17, 17, 19, 22, 22, 23, 27.
Navigating Excel
The individual boxes in an Excel spreadsheet are called “cells.” Each cell is identified by its column and row.
Review Question F: What is the small- est value a standard deviation can have?
tan81004_01_c01_001-024.indd 15 2/22/13 3:30 PM
CHAPTER 1Section 1.6 Calculating Descriptive Statistics With Excel
• The columns are labeled from left to right, alphabetically from column A. • The rows are numbered from 1 down the left side of the window. • Cell “A1” is the cell in column A and row 1—the upper left. The next cell
down is cell A2, and so on. • When identifying a cell, the column letter is first, and then the row number. • When cell locations are entered in Excel there is no space between letter and
number.
The version of the software for this text is from Office 2010. As Excel has been updated from version to version, the appearance and some of the procedures have changed, but changes in the command structure have been relatively minor.
The steps for entering the customer satisfaction data into the spreadsheet are as follows:
• Place the cursor in the cell where the data are to begin—cell A1, for example, either · by moving the cursor with the arrow keys in the key-pad, · by clicking the mouse on the particular cell, · or by using the touch-pad on a laptop.
• In cell A1, key in the number 11, followed by the Enter key. The Enter key will move the cursor to the next cell down.
• Enter each of the other 11 values so that the data are arranged vertically in all the cells from A1 to A12. That will make the spreadsheet look as it does in Figure 1.1.
Figure 1.1: A data set entered in Excel
tan81004_01_c01_001-024.indd 16 2/22/13 3:30 PM
CHAPTER 1Section 1.6 Calculating Descriptive Statistics With Excel
Entering the Command for the Mean
In Excel, when the initial entry in any cell is an equal sign, “5,” it indicates that a math- ematical operation will follow. The operation can be entered either as one of Excel’s pro- grammed commands, such as “average,” which is the command for the mean, or as a formula that the user can construct with mathematical operators. To calculate the mean for the 12 customer satisfaction scores and have that value appear in cell A13:
• Place the cursor in cell A13. • Enter the command, 5average(a1:a12) which will calculate the mean for the
data in cells A1 to A12. Note the Excel command is “average” rather than “mean.”
• Press Enter. • The value in cell A13 is the mean, 18.16667. • In the Home tab, click the arrow in the bottom right corner of the Number
tab—it’s in the middle near the top of the screen. • Under Category, click Number (the Format tab for Mac users) and then to
the right you can indicate the number of decimal places. We’ll round to three decimal places. That will make M 5 18.167.
Using the Descriptive Statistics Option
Entering the specific command works well when a particular statistic is needed, but some- times several statistics are needed. The mean is reported with other descriptive values as part of a descriptive statistics option that is available to PC users but, unfortunately, is not present on Mac versions of Excel. For the customer satisfaction data list, the commands for that package of statistics are the following:
• Click the Data tab, which is the sixth tab at the top of the page. (For Mac users, click the Formulas tab, and then fx. Selecting the Statistical option from the list will indicate the various statistical procedures that are available.)
• Click the Data Analysis window at the extreme right just below the tabs. This will open a small window in the page with a list of options (see Figure 1.2).
tan81004_01_c01_001-024.indd 17 2/22/13 3:30 PM
CHAPTER 1Section 1.6 Calculating Descriptive Statistics With Excel
Figure 1.2: The Excel data analysis, descriptive statistics option
• If the “Data Analysis” option doesn’t appear at the extreme right, PC users can “add it in.” The commands vary with the different versions of Microsoft Office, but for Excel 2010 the commands for adding in the Data Analysis option are as follows: 1. Click the File Tab in the upper left corner of the screen. 2. Click Options toward the bottom of the left column 3. In the Excel Options window that appears, select Add-Ins in the left
column. This will produce a list of Inactive Application Add-Ins, one of which is Analysis ToolPak.
4. Click on Analysis ToolPak. 5. Click Go toward the bottom of the View and manage Microsoft Office
Add-Ins window. 6. In the window there will now be a check beside Analysis ToolPak. Click OK.
At this point, clicking the Data Tab, sixth from right at the top of the page, will reveal the Data Analysis option at the extreme right.
• Click the Data Analysis option and then in the Data Analysis window that appears in the middle of the page.
• Click on the Descriptive Statistics option and then click OK. • In the small window labeled “Input Range” type in the cells for which we
wish the values to be included, A1:A12, just as we did when we entered the formula for the mean. When entering the letter for the column, it doesn’t mat- ter whether it’s upper- or lowercase.
tan81004_01_c01_001-024.indd 18 2/22/13 3:30 PM
CHAPTER 1Section 1.7 Dependent and Independent Variables
• Note that the default is that data are “Grouped by” columns. If the data were listed along a row, we would have to change the default.
• Click Output Range and indicate where the results display is to begin, perhaps cell C1, so that results are next to the original data but not over top of them.
• Finally, click the particular output we wish, which is, Summary Statistics. • Click OK.
Results are in Figure 1.3 with the values rounded to three decimals.
Figure 1.3: The output from the descriptive statistics option
The output includes more descriptive statistics than have yet been discussed. In addi- tion to the median (Mdn), the mode (M), the lowest and highest values, the range (R), the sample standard deviation (s), and the variance (s2) values, there are also skewness and kurtosis values. Those descriptive statistics will come up in Chapter 2 as part of the dis- cussion of data normality. The standard error will come up later in the book.
1.7 Dependent and Independent Variables
When one variable is thought to have an effect on another, the variable affected is referred to as the dependent variable, and the variable creating the effect is the independent variable. The relationship is examined by manipulating the independent variable to see if there is some corresponding impact on the dependent variable. Produce managers place produce in more prominent positions (the independent variable) in an
tan81004_01_c01_001-024.indd 19 2/22/13 3:30 PM
CHAPTER 1Chapter Summary
effort to increase sales of slow-moving produce (the dependent variable). The bank offers incentives (the independent variable) in an effort to entice customers to open certificates of deposit (the dependent variable).
Sometimes it is tempting to place the relationship in causal terms and say that the inde- pendent variable causes whatever happens to the dependent variable, but cause is often very difficult to verify. Suppose the manager of a company that installs solar panels on residential homes decides to use free installation as an incentive to potential customers. The assumption is that removing the cost of the installation (the independent variable) can cause an increase in the sales of solar energy systems (the dependent variable). If free installation is offered and sales rise compared to installations with the usual charges, is the increase in sales caused by the free installation? Perhaps the causal factor is the tax incentive that the government chose to offer to those who install solar energy systems, or perhaps the factor is declining interest rates on borrowed money. The point is that unless all other variables are strictly controlled, as they sometimes are in laboratory experiments, cause is very difficult to establish. The independent variable/dependent variable lan- guage helps describe how variables may be related, but it is ordinarily very difficult to establish a causal relationship.
Chapter Summary
Managers use statistics in making strategic decisions and finding answers to impor-tant business questions (Objective 1). Descriptive statistics provide indicators of data characteristics. Descriptive characteristics can provide a great economy when data sets are large. Inferential statistics are utilized when the sample’s characteristics are important for what they reveal about the entire population (Objective 5).
There are many descriptive statistics. The most common are those that describe central ten- dency, or “typicality,” in a data set, and those that describe data variability (Objective 6). The type of data scale also informs the particular descriptive statistics that can be calcu- lated (Objective 4). The mean and standard deviation, for example, assume either interval or ratio data. If a measure of central tendency is required for nominal data, the mode is used. Because descriptive statistics each provide a different view of what is most central, or how variable data are, the different values are complementary. It is not uncommon to see multiple measures of central tendency and variability reported in a data description.
Part of the transition in any new discipline is learning the terminology. It is important to have a common language. Some of the symbols that indicate descriptive statistics depend upon whether they describe samples or populations. A sample mean, for example, is indi- cated by M; but the symbol for a population mean is m (Objectives 2 and 3). Inexpensive but powerful hand calculators and computer software have made statistical analysis much more readily accessible. They ease the burden of describing and analyzing larger data sets.
Statistical analysis is incremental in its structure. The topics developed in each chapter become elements of more involved procedures later, something that will be apparent in Chapter 2. Virtually nothing is raised, discussed, and then permanently set aside, which makes it important to integrate each new concept as we continue. The relationships will become clearer with repeated review and with attention to the end-of-chapter problems.
tan81004_01_c01_001-024.indd 20 2/22/13 3:30 PM
CHAPTER 1Management Application Exercises
Answers to Review Questions
A. Population characteristics are called parameters. B. For ordinal data, both the mode and the median can be calculated. C. The range is based on just two values, the highest and lowest. The standard
deviation, on the other hand, is based on all the measures in the group, a char- acteristic that makes the standard deviation the more informative of the two.
D. Extreme scores have a disproportionate impact on the standard deviation and the variance because calculating these statistics involves squaring the differ- ence between each individual value and the mean of the sample. Depending upon how extreme the score is, squaring the difference can have an inordinate impact on the resulting value of the statistic.
E. The way the mean, standard deviation, and variance are calculated and the way these statistics are interpreted depend on data that have equal intervals.
F. The smallest value a standard deviation can have is zero. It occurs when all the values in the data set are the same, which should make sense. With no variabil- ity from measure to measure, s 5 0.
The Chapter Formulas
Formula 1.1 M 5 Sx/n The mean of a set of scores
Formula 1.2 s2 5 S1x 2 M2 2
n 2 1 The variance of a sample
Formula 1.3 s 5 Å S1x 2 M2 2
n 2 1 The standard deviation of a sample
Formula 1.4 s2 5 S1x 2 m2 2
n The variance of a population
Formula 1.5 s 5 Å S1x 2 m2 2
n The standard deviation of a population
Management Application Exercises
Unless otherwise stated, use p 5 .05 in all your answers.
1. A marketing agency is interested in the buying habits of those who shop online versus those who shop in person.
a. A database is created, and shoppers are classified as online or in-person shoppers. Such classifications represent data of which scale?
b. Online and in-person shoppers are to be compared on their relative incomes. Income data represent data of which scale?
c. If shoppers are also ranked from the most to the least frequent shoppers, those rankings represent data of which scale?
tan81004_01_c01_001-024.indd 21 2/22/13 3:30 PM
CHAPTER 1Management Application Exercises
d. Shoppers are asked to complete a customer satisfaction survey. In response to each question on the survey, the shoppers circle one of five answers: strongly agree, agree, neutral, disagree, or strongly disagree. Their responses are data of which scale?
e. What measure(s) of central tendency will be most informative for each of the above data?
f. For which one(s) would calculating a standard deviation be meaningful? Why or why not?
2. A shift supervisor is interested in the number of times employees take breaks from the assembly line that require someone else to step in.
a. Determining the number of breaks involves data of which scale? b. Which measure(s) of central tendency is/are appropriate for data of this
scale?
3. The performances of a group of interns are evaluated by their supervisors at the end of their internships. Their scores are: 55, 47, 62, 27, 50, 49, 66, 53, 50, 44, 63, 59. Calculate:
a. the mean b. the median c. the range d. the standard deviation
4. A home products company produces a new scented candle that is supposed to relax people and prompt a sense of well-being. The company recruits potential buyers to test the product. Testers are randomly assigned to one of two groups. One group is exposed to the new product. The other group is a control group that is exposed to a different product. All testers then rank how calm they feel on a scale of 1–10.
a. Categorizing participants in terms of whether they were exposed to the scented candles involves data of which scale?
b. Which measure(s) of central tendency is/are appropriate for summarizing the exposure versus no exposure results?
c. If the testers are ranked from the most to the least calm as a result of being exposed to the candles, those rankings represent data of which scale?
d. If “calmness” is equated with how many minutes participants are willing to sit, the results can be summarized by which measure(s) of central tendency?
5. The following are the dollars that 13 consecutive shoppers at a sports equipment outlet spend on purchases: 24, 25, 28, 28, 31, 33, 36, 36, 36, 39, 40, 53, 54.
a. Calculate the mean, the variance, and the range. b. If the 14th and 15th shoppers spend $8 and $30 respectively, which shopper’s
expenditure will have the greatest effect on the value of the variance? Why?
tan81004_01_c01_001-024.indd 22 2/22/13 3:30 PM
CHAPTER 1Key Terms
6. The human resources department is interested in predicting retirements based on the number of years employees have been employed with the organization. In a particular department, the years each employee has been employed are as follows: 2, 5, 8, 11, 13, 16, 22, 27.
a. What is the scale of this data? b. What are the values of the mean and the median? c. Without doing the calculations, what will be the effect on the standard
deviation of adding a 9th person who has been employed 13 years? d. What will be the effect of that 9th person on the value of the range? e. Which would be increased most dramatically by the addition of someone
who had been employed 28 years, the range or the standard deviation? f. Now perform the calculations necessary for c through e and check your
answers.
7. A manager wants to determine the impact of a new training program on employee performance. Employee performance is measured before the new program is implemented. For a full year, the manager then keeps track of who attends the training program and how many levels each employee has successfully completed. At the end of the year, employee performance is measured again.
a. What is the independent variable? b. What is the dependent variable?
8. The national manager of a chain of restaurants has data gathered on average monthly sales for all the restaurants in the chain. What symbol should be used to indicate the average monthly sales?
9. If the manager referenced in item 8 generalizes what is happening nationally from what is occurring in just a sample of restaurants, what type of statistical analysis is being performed?
10. The manager of a hardware store has received data from the regional office that indicate that sales for a particular month have R 5 $1,320.
a. Is there any way to determine what the mean was for daily sales? b. Can the lowest and highest daily sales be determined from what the region-
al office provided? c. If the manager happens to know that the day for which power tools sales
were lowest was a day where sales 5 $548, what were sales for the highest sales day?
Key Terms
• Data scale, or the scale of the data, refers to the kind of information that a mea- sure provides, and so the kinds of statistics that can be calculated and the types of analyses performed. Nominal data, for example, indicate only the presence or ab- sence of qualitative characteristics such as race or marital status. The only measure of central tendency that should be determined is the mode.
tan81004_01_c01_001-024.indd 23 2/22/13 3:30 PM
CHAPTER 1Key Terms
• A population includes all members of a defined group. Anything less than a popula- tion is a sample. Descriptive characteristics can be determined in either case, but for samples, those characteristics are statistics; for populations, they are parameters.
• Measures of central tendency and measures of variability are descriptive statistics. Measures of central tendency refer to what is typical, and measures of variability indicate how much variety there is in a data set. When these values are employed so that the sample is used to describe the population, inferential statistical analysis is being applied.
• Nominal data indicate the category to which one individual belongs. Although nominal data have numbers, they serve as identifiers only; they have no math- ematical significance. Ordinal data provide enough information that individuals can be ranked according to some quality. Consecutive integers for interval data involve consistent differences. With ratio data, zero indicates the absence of the quality measured. Ratio data also make possible ratio comparisons of the quality measured—twice as much, half as much, and so on.
• Among measures of central tendency, the mode is the most frequently occurring value in a data set. When measures are arranged in order, the median is the middle- most value. The mean is the average. Of the more common measures of variabil- ity, the range indicates the difference between the highest and lowest score. The standard deviation and variance both measure how much individual measures tend to vary from the mean of the data set. The variance is the square of the stan- dard deviation. In all measures of variability, larger values indicate less homoge- neous data sets.
• Some amount of random variability is inescapable in statistical analysis. However, when the same error occurs systematically, the problem is referred to as bias. Be- cause the population standard deviation formula consistently underestimates data variability when used with samples, that formula results in a bias. It is for that rea- son that the n – 1 correction is made in the sample standard deviation formula.
• The degrees of freedom in a statistics problem are indicated by the number of values free to vary when the value of the result has been determined. In an addition problem with 3 values and a sum of 10, the first 2 measures can have any value so long as the 3rd value makes the sum of the 3 values 10. The problem has 2 degrees of freedom. Degrees of freedom are relevant to inferential statistics problems that use statistics to estimate parameter values.
• Statistical analysis is frequently about how variables affect each other. When the influence is thought to be unidirectional, the independent variable is the one be- lieved to have the influence; the dependent variable is the one affected.
tan81004_01_c01_001-024.indd 24 2/22/13 3:30 PM