lOMoARcPSD|21919833
BUSI 600 TEST 4 Notes
International Business (Liberty
University)
lOMoARcPSD|21919833
Chapter 16: Exploring, displaying, and examining data
Exploratory data analysis (P406-419)
Exploratory data analysis (EDA) (P406)
Confirmatory data analysis (P406)
- Compared EDA to the role of police detective and other investigators.
- A major contribution of exploratory approach.
- Frequency tables, bar charts, and pie charts
Frequency chart (Exhibit16-2, P407)
Histograms (P408-410):
●Histograms are useful for two points
- Use of histograms
- When the variable of interest is measured on an interval-ratio scale and is one
with many potential values, these techniques are not particularly informative.
- Exhibit 16-4, shown in the slide, is a condensed frequency table of the average
annual purchases of Prime Sell’s top 50 customers. Only two values, 59.9 and
66, have a frequency greater than 1. Thus, the primary contribution of this table
is an ordered list of values. If the table were converted to a bar chart, it would
have 48 bars of equal length and two bars with two occurrences.
Stem and leaf displays (P411-415)
Advantages of stem-and-leaf displays
Develop stem-and-leaf displays
Using Tables to understanding data
(P412) Pareto diagrams (P415)
lOMoARcPSD|21919833
Box plots (P415-418)
- Five-number summary
- Resistant statistics
- Resistance
- Mean and standard deviation
- Nonresistant statistics
- The basic ingredients of the plot
- Interquartile range (IQR)
- Outliers
Boxplot components (Exhibit 16-8)
Diagnostics with Boxplots (Exhibit 16-9)
Boxplot comparison of customer sectors (exhibit 16-10)
Mapping (P418-419)
Geographic information system (GIS)
Radio frequency identification (RFID)
- The most common way to display such data is with a map
Cross-Tabulation (P419-424)
- The definition
An example of computer-generated cross tabulation (Exhibit 16-11)
Cell Marginal Contingency tables The use of percentages (P420-
423)
The purpose of percentages Comparison of several distributions of data (Exhibit 16-12)
- Two-dimension tables
The guidelines to prevent errors in reporting Other table-based analysis (P423-424)
lOMoARcPSD|21919833
- Control variable
- Statistics packages
Three variables are handled under the same banner (Exhibit 16-13)
N-way tables Automatic interaction detection (AID)
Tree diagram resulted from AID (Exhibit 16-14)
An advanced variation on n-way tables is automatic interaction detection (AID).
●AID is a computerized statistical process that requires that the researcher identify a
dependent variable and a set of predictors or independent variables. The computer then
searches among up to 300 variables for the best single division of the data according to
each predictor variable, chooses one, and splits the sample using a statistical test to verify
the appropriateness of this choice.
Resolution of the problem
●Exhibit 16-14 shows the tree diagram that resulted from an AID study of customer
satisfaction with Mind Writer’s Complete Care repair service. The initial dependent
variable is the overall impression of the repair service. The variable was measured on an
interval scale of 1 to 5. The variables that contribute to perceptions of repair
effectiveness were also measured on the same scale but were rescaled to ordinal data for
this example. The top box shows that 62% of the respondents rated the repair service as
excellent. The best predictor of repair effectiveness s <resolution of the problem.=
●Condition on arrival Researcher separately studied
Chapter 17: Hypothesis Testing Introduction (P430-
440) Inferential statistics
lOMoARcPSD|21919833
Inferential statistics includes the estimation of population values and the testing of statistical
hypotheses. Descriptive statistics simply describe the characteristics of the data by giving
frequencies, measures of central tendency, and dispersion.
Hypothesis testing and research process (Exhibit 17-1)
Two approached to hypothesis testing Classical statistics classical (sampling-theory) approach
is more established Bayesian statistics Statistical significance (P430-431)
The definition
Practical significance
The logic of Hypothesis testing (P432-435)
Null hypothesis
The null hypothesis is used for testing. It is a statement that no difference exists between the
parameter and the statistic being compared to it. The parameter is a measure taken by a census of
the population or a prior measurement of a sample of the population.
Alternative hypothesis
Two-tailed test, or non-directional
test One-tailed test, or directional test
Two- and one-tailed tests at the 6 percent level of significance (Exhibit 17-
2) Decision rule in testing hypotheses
Statistics testing give only chance to disprove or fail to reject the hypothesis
Analogy to the American legal system
Comparison of statistical decision to legal analogy (Exhibit 17-3)
Type I error (P436-436)
lOMoARcPSD|21919833
Exhibit 17-4
Regions of rejection
Region of acceptance
Critical values
The boundary between the regions of acceptance and rejection
Type II error (P436-438)
The probability of committing a type II error depends on five factors
Several ways to reduce a Type II error
Probability of making a Type II error (Exhibit 17-
5) Statistical testing procedures (P438)
Six-stage sequence
Level of significance
Probability values (p Values) (P438-
440) The definition of p value
The p value is determined using the standard normal table
Tests of significance (P440-462)
Parametric tests are significance tests for data from interval or ratio scales. They are more
powerful than non-parametric tests. Non-parametric tests are used to test hypotheses with
nominal and ordinal data. Parametric tests should be used if their assumptions are met. Types of
tests (P440-442) Parametric tests
Nonparametric tests
Assumption for parametric tests
lOMoARcPSD|21919833
Normal probability plot
Probability plots and tests of normality (Exhibit 17-6)
The normality of the distribution may be checked in several ways. One such tool is the normal
probability plot. This plot compares the observed values with those expected from a normal
distribution. If the data display the characteristics of normality, the points will fall within a
narrowband along a straight line
Assumption for non-parametric tests
How to select a test (P442-443)
Three questions before choose a test
Decision trees
An expert system
Selecting tests using the choice criteria (P443)
Recommended statistical techniques by measurement level and testing situation (Exhibit 177)
One-sample tests (P444)
The definition
Encounter questions
Parametric tests (P444)
Z test, or t-test
Z distribution and t distribution
The t has more tail area than that found in the normal distribution.
Real-world applications of the one-sample test
Example Non-parametric Tests
(P445) The definition
lOMoARcPSD|21919833
Chi-Square Test (P445-446)
The most widely used non-parametric tests
Differences between the observed distribution and the expected distribution
The use of Chi-Square tests
The Z and t-tests are frequently used parametric tests for independent samples, although the F
test can also be used. The Z test is used with large sample sizes (exceeding 30 for both
independent samples) or with smaller samples when the data are normally distributed and
population variances are known. The formula is shown in the slide. Example In a one-sample
situation, a variety of non-parametric tests may be used, depending on the measurement scale
and other conditions. If the measurement scale is nominal, it is possible to use either the binomial
test or the chi-square test. The binomial test is appropriate when the population is viewed as only
two classes such as male and female. It is also useful when the sample size is so small that the
chi-square test cannot be used.
Two-Independent-Samples Tests (P447-450)
Parametric tests (P447-448)
Non-parametric tests (P448-450)
Two-related-sample tests (P450-453)
The definition of Parametric tests and non-parametric tests
Observed significance level
McNeal test
K-Independent-Samples
Tests (P453-460)
lOMoARcPSD|21919833
The definition
Parametric test
Analysis of variances
(ANOVA) One-way analysis
The use of ANOVA
Within-groups variance
F ratio
Mean square
Illustrate one-way
ANOVA
A priori contrasts (P457)
The definition
The use of contrasts
Multiple Comparison Tests (P457-458)
Post hoc tests Selection of Multiple comparison procedure (Exhibit 17-13, P458)
Exploring the findings with two-way ANOVA (P458-460)
Summary of table for Two-way ANOVA examples (Exhibit 17-15)
Three questions may be considered with two-way analysis
Two-way analysis of variance plots with means and standard deviation listed (Exhibit 17-16)
Analysis of variance
The Kruskal-Wallis
test K-Related-Samples
Tests (P460-462)
lOMoARcPSD|21919833
A K-related-samples test is required for three situations.
ANOVA is a special type of n-way analysis of variance
Summary tables for repeated-measures
ANOVA (Exhibit 17-17)
Repeated-measures
ANOVA plot (Exhibit 17-18)
Cochran Q test
Friedman two-way analysis of variance
Chapter 18: Measures of Association
Introduction (P468-469)
Relational hypothesis
Management questions
Common measure and their uses in exhibit 18-1
Exhibit 18-1 in the text presents a list of commonly used measures of association. These
measures of association will be discussed further in this chapter. This slide shows those
measures of association relevant for interval and ratio scaled data. The measures used for ordinal
and nominal data are presented on the following slides. Bivariate Correlation Analysis (P469-
479) Differs from non-parametric measures
The person correlation coefficient
The magnitude
Scatter plots
Scatter plots of correlations between two variables (Exhibit 18-2, P471)
lOMoARcPSD|21919833
Exhibit 18-2 contains a series of scatter plots that depict some relationships across the range r.
The three plots on the left side of the figure have their points sloping from the upper left to the
lower right of each x-y plot. They represent different magnitudes of negative relationships.
On the right side of the figure, the three plots have opposite directional patterns and show
positive relationships.
The shape of linear relationships
The need for data visualization (Exhibit 18-3)
Exhibit 18-3 contains the data and the exhibit shown in the slide plots them.
●In Plot 1, the variables are positively related.
●In Plot 2, the data are curvilinear and r is an inappropriate measure of their relationship.
●In Plot 3, an influential point that changed the coefficient is shown.
●In Plot 4, the values of x are constant.
Different scatter plots for the same summary statistics (Exhibit 18-4)
The need for data visualization is illustrated with four small data sets possessing identical
summary statistics but displaying different patterns.
The assumptions of r R is linearity
Correlation is a bivariate normal distribution
Computation and testing of r (P473)
Common variance as an explanation
Coefficient of determination, or r^2
The amount of common variance in X and Y may be summarized by r 2 , the coefficient
of determination.
Testing the significance of r (P475-476)
lOMoARcPSD|21919833
Interpretation of correlations
A correlation coefficient of any magnitude or sign, regardless of its statistical significance, does
not imply causation.
Correlation provides no evidence of cause and effect.
There are several alternate explanations that may be provided for correlation results and these are
listed in the slide.
Note that causation may actually exist, but the correlation only shows that a relationship exists.
Therefore, x may cause y or y may cause x.
However, it may also be that z causes x and y or that y and x interact.
Artifact correlations (Exhibit 18-8)
Artifact correlations occur where distinct groups combine to give the impression of one.
Artifact correlations should be avoided because they appear as one group when, in fact, there are
distinct groups present.
The left panel shows data from two business sectors. If all the data points for x and y variables
are aggregated and a correlation is computed for a single group, a positive correlation results.
Separate calculations for each sector reveal no relationship. In the right panel, the companies in
the financial sector score high on assets and low in sales but they are all banks. When the data
for banks are removed, the correlation is nearly perfect. When banks are kept in, the overall
relationship drops considerably.
Practical significance
A coefficient is not remarkable simply because it is statistically significant
Comparison of bivariate linear correlation and regression (Exhibit 18-9)
Relationships also serve as a basis for estimation and prediction.
lOMoARcPSD|21919833
Simple and multiple predictions are made with a technique called regression analysis.
●When we take the observed values of X to estimate or predict corresponding Y
values, the process is called simple prediction.
●When more than one X variable is used, the outcome is a function of multiple predictors.
Simple Linear Regression (P479-490)
Simple prediction and regression analysis
The basic model Regression coefficient
The slope
The ratio of change
Examples of different slopes (Exhibit 18-
10) The intercept
Concept application
Plot of wine price by average growing temperature (Exhibit 18-11)
Distribution of Y for observations of X (Exhibit 18-12)
It is more likely that we will collect data where the values of Y vary for each X value.
Considering Exhibit 18-12, we should expect a distribution of price values for the temperature X
= 16, another for X = 20, and another for each value of X.
The means of these Y distributions will also vary in some systematic way with X.
These variations lead us to construct a probabilistic model that also uses a linear function.
Method of least squares (P482)
Data for least square (Exhibit 18-13)
Scatter plot and possible regression lines based on visual inspection:
lOMoARcPSD|21919833
Wine price study (Exhibit 18-14)
Exhibit 18-14 contains a new data set for the wine price example. Our prediction of Y from X
must now account for the fact that the X and Y pairs do not fall neatly along the line. Exhibit
18-14, shown in the slide, suggests two alternatives based on visual inspection. The method of
least squares allows us to find a regression line, or line of best fit, which will keep these errors to
a minimum. It uses the criterion of minimizing the total squared errors of estimate. When we
predict values of Y for each Xi, the difference between the actual Yi and the predicted Y is the
error. This error is squared and then summed. The line of best fit is the one that minimizes the
total squared errors of prediction.
Drawing the regression line
Residuals Plot of standardized residuals (Exhibit 18-16)
Prediction and confidence bands (P486)
Testing the Goodness of Fit (P486)
Zero slopes result from various conditions Prediction and confidence bands on proximity to X
(Exhibit 18-17)
Prediction and confidence bands are bow-tie shaped confidence interval around a predictor.
Predictors farther from the mean have larger bandwidths. If we wanted to predict the price of a
case of investment-grade red wine for a growing season that averages 21 degrees Celsius
The t-test
The F test
Components of variation (Exhibit 18-18)
Computer printouts generally contain an analysis of variance (ANOVA) table with an F test
of the regression model.
lOMoARcPSD|21919833
●In bivariate regression, t and F tests produce the same results since t 2 is equal to F.
●In multiple regression, the F test has an overall role for the model and each of the
independent variables is evaluated in a separate t-test.
●For regression, the ANOVA comprises explained deviations and unexplained deviations.
Together, they constitute total deviation. This is shown graphically in the exhibit in the
slide. These sources of deviation are squared for all observations and summed across the
data points.
Progressive application of partitioned variance concept (Exhibit 18-19)
Coefficient of determination is the ratio of the line of best fit’s error over that incurred by using
Y.
Non-parametric measures of association (P490-498)
Two characteristics with nominal measures
Chi-Square-based measures (Exhibit 18-20)
Nominal measures are used to assess the strength of relationships in cross-classification tables.
They are often used with chi-square.
Phi Cramer’s V
The contingency coefficient C
Proportional reduction in error (PRE)
Lambda
The computation of Lambda
lOMoARcPSD|21919833
Proportional reduction of error measures (Exhibit 18-21)
Proportional reduction in error (PRE) statistics are the second type used with contingency tables.
Lambda and tau are the examples discussed in the text. Lambda is a measure of how well the
frequencies of one nominal variable predict the frequencies of another variable.
Goodman and Kruskal’s tau Measures for Ordinal data (P494)
Several statistical alternatives
Concordant and discordant
Gamma
Tau b
Calculation of concordant, discordant, tied, and total paired observations (Exhibit 18-23)
Tau c Somers’s d
Spearman’s rho
KDL data for Spearman’s rho (Exhibit 18-24)
Spearman’s rho correlates ranks been two ordinal variables. Spearman’s rho has many
advantages. When data are transformed by logs or squaring, rho remains unaffected. Outliers or
extreme scores are not a threat. It is easy to compute. The relationship between the panel’s and
the psychologist’s rankings is moderately high, suggesting agreement between the two
measures. The test of the null hypothesis that there is no relationship between the measures is
rejected at the .05 level with n-2 degrees of freedom.
lOMoARcPSD|21919833
Chapter 19:
Presenting Insights and Findings:
Written Reports
The Written Research report (P504-
507) The definition
Reports may be defined by their degree of formality and design.
●The formal report follows a well-delineated and relatively long format.
●The short report is more informal.
●Short reports are appropriate when the problem is well-defined, of limited scope, and has
a simple and straightforward methodology. They are usually about 5 pages in length. A
letter of transmittal is a vehicle to convey short reports.
●The letter is a form of a short report. Its tone should be informal. The format follows
that of any good business letter and should not exceed a few pages. A letter report is
often written in a personal style. Short reports can also follow the style of a memo. The
suggestions in the slide are provided for writing short reports.
●Report Access: With managers who have an interest in research often located in
different locations, report access has become of increasing interest. Often paper based
reports are delivered to the primary sponsor, with electronic versions made available to a
wider audience.
Short reports
Suggestion for writing short reports
lOMoARcPSD|21919833
Written presentation and the research process (Exhibit 19-
1) Long Reports
- Long reports may be technical or management reports. Some projects require both forms.
- a management report is written for the non-technically oriented manager or client.
- The management report focuses on an introduction with conclusions and
recommendations. Individual findings follow to support the conclusions already made.
The appendices provide any required methodological details. It also makes liberal use of
visual displays.
- A technical report is written for an audience of researchers
- The technical report should include full documentation and detail. It has the full story of
what was done and how. A good guide is to provide sufficient information that would
enable others to replicate the study.
- The Technical report should also include a full presentation and analysis of significant
data with conclusions and recommendations. Technical report and management report
The long management report contains all prefatory information (letter of transmittal, title
page, authorization statement, executive summary, and table of contents, introduction
(including problem statement, research objectives, and background as well as a brief
statement of the methods and limitations), conclusions and recommendations, followed
by the findings, and relevant appendices. The short technical report contains all prefatory
information (letter of transmittal, title page, authorization statement, executive summary,
and table of contents, introduction (including problem statement, research objectives, and
background plus a brief statement on the methods and limitations of the study), findings,
lOMoARcPSD|21919833
conclusions and recommendations, and relevant appendices. The long technical report
contains all possible components in the order designated in Exhibit 21-2.
Research report components (P507-512)
Research report sections and their order of inclusion (Exhibit 19-2)
Prefatory items (P508)
Letter of transmittal
Title page
Three acceptable way to word report title
Authorization letter
Executive summary
Table of Contents
Introduction (P509)
Problem statement
Research objectives
Background Methodology (P510)
Sampling design
Research design
Data collection
Data analysis
Limitation Finding (P511)
Example of a finding page (Exhibit 19-3)
Each report needs to have a standard findings page style guide.
lOMoARcPSD|21919833
●This is especially true if the report pages are prepared by distinct individuals.
●This common style makes it easier for the reader to quickly grasp research results.
●Many organizations have a template for findings pages.
●The template used for this slide requires the summarization of findings to lead the page,
the question to appear as it appeared on the questionnaire, and the data table below.
Conclusions (P512)
Summary and conclusions
Recommendations
Appendices
Bibliography
Writing the report (P512-
517) Prewriting concerns
Before writing the report, one should ask and answer these questions to help frame the situation.
The outline
Before writing, the researcher should develop an outline.
This slide presents a useful organizational structure.
Topic outline and Sentence outline
In a topic outline, a key word or two is used.
The sentence outline expresses the essential thoughts associated with the specific topic.
The bibliography
Writing the
draft
Readability
lOMoARcPSD|21919833
Readability index
Grammar and style proofreader results (Exhibit 19-4)
Comprehensibility
Pace Service words are words that transition from one idea to another; examples include:
●On the other hand
●In summary
●In contrast
Variety of methods to adjust the
pace Tone Final proof
Presentation considerations
Overcrowded text may be avoided in following ways
Presentation of statistics (P517-536)
Text presentation
Semi-tabular presentation
Tabular presentation
Sample tabular finding (Exhibit 19-5)
Graphics
Line Graphs
Line graphs are chiefly used for time series and frequency distributions. There are
several guidelines for designing a line graph:
lOMoARcPSD|21919833
Put the time units or the independent variable on the horizontal axis.
When showing more than one line, use different line types.
Try not to put more than four lines on one chart.
Use a solid line for the primary data.
Guidelines for designing a line graph
Guide to chart for written reports (Exhibit 19-6)
Examples of area charts (Exhibit 19-9)
An area chart is also used for a time series.
Consisting of a line that has been divided into component parts, it is best used to show changes
in patterns over time.
Area chart:
Stratum or surface
Pie chart
A pie chart is another form of area chart. It is often used with business data. They can easily
be improperly prepared, though. Pie charts are useful for frequency data.
Consider the following suggestions when designing pie charts:
Show 100% of the subject being graphed
Label the slides with <call outs=
Put the largest slice at twelve o’clock and move clockwise in descending order
Use light colors for large slices
lOMoARcPSD|21919833
In a pie chart of black and white slices, a single red one will command the most
attention
Do not show evolution over time
Suggestion for designing pie chart Bar chart
A bar chart is a graphical presentation technique that represents frequency data as horizontal or
vertical bars. It can be very effective when properly constructed. Pictographs and geography
Pictographs are bar charts using pictorial symbols rather than bars to represent frequency data.
This one is from a study of pizza consumption. It was used in both the written report AND the
oral presentation. A geographic chart uses a map to show regional variations in data. This one
is for digital camera ownership.
3-D graphics
A 3-D graphic is a presentation technique that permits a graphical comparison of three or
more variables.
Exhibit 19-10 illustrates a 3-D column, 3-D Ribbon, 3-D Wireframe, and 3-D Surface Line. 3-
D charts (Exhibit 19-10)
Not all researchers are asked to prepare recommendations, but increasingly many are. The
researcher needs to clarify the extent to which the sponsor seeks recommendations before
preparing the report. Students need to clearly distinguish between data, the interpretation of data,
a conclusion drawn from the data, and a recommendation related to the manager’s dilemma that
stimulated the need for the research. Compiling the written report means preparing and gathering
the totality of all written materials which will be delivered to the sponsor and the format in which
these will be delivered. 3-ring binders, bound printed reports, and PDF reports are all fairly
lOMoARcPSD|21919833
common for research report compilations. All require a detailed table of contents. The PDF
report has the added value of being key-word searchable by the reader. Decisions at this stage
involve determining order of material within the report (usually determined by sponsor
preference or researcher template, and quantity of copies. Delivery of the report often is
determined by whether an oral presentation of data findings is planned. Written reports are
delivered before oral presentations or following oral presentations, depending on the preference
of the sponsor. Reports are delivered by courier or package delivery service or in person,
depending on the arrangements to address questions if no oral presentation is planned. Compiling
the written report means preparing and gathering the totality of all written materials which will
be delivered to the sponsor and the format in which these will be delivered.
This can include materials the sponsor has provided (prior research reports, promotional
materials, etc.)
3-ring binders, bound printed reports, and PDF reports are all fairly common for research report
compilations. Protecting the anonymity of respondents must be balanced against the sponsor's
need for data at this stage. Usually, actual completed questionnaires are not provided to the
sponsor to protect anonymity. Outside research suppliers often spend considerable time on the
appearance of the total compilation as it affects how professionally the report is perceived by the
sponsor.
The report access decision influences the decisions at this stage:
●use of color
●type of report holder/binder
●order of material within the report (usually determined by sponsor preference or
researcher template),
lOMoARcPSD|21919833
●quantity of report copies