Applied Business Research and Analysis (assessment 2)
Salford Business School
Applied Business Research and Analysis
Quantitative Research
Data Analysis and Presentation
Lecture
Professor David F. Percy
Contents
1. Summaries
2. Probability
3. Inference
4. Modelling
5. Forecasting
Bibliography
1. Rees D.G. (2000) Essential Statistics, Chapman & Hall
2. Weiss N.A. (2012) Introductory Statistics, Addison-Wesley
3. Freund J.E. & Perles B.M. (2006) Modern Elementary Statistics, Prentice Hall
Syllabus
1. Summaries
65
55
45
35
C o s t
Sample size
� � 24 Measures of location
mode (most likely value): £51.00
median (middle ordered value): £51.00
mean (average value): £50.50
1. Summaries Gas heating costs (£) per unit area in 24 UK factories
51 52 50 47 55 57 43 59
52 51 40 48 53 47 54 63
51 45 55 53 48 56 36 46
Measures of spread
range (difference between maximum and minimum): £27.00
interquartile range (difference between upper and lower quartiles): £7.75
standard deviation (root mean square distance about mean): £6.03
1. Summaries
Summary Statistics
Consider the current ratios (assets÷liabilities) of 8 market traders:
1.43, 1.02, 2.07, 2.35, 0.81, 1.73, 2.99, 1.26
Sample median:
median � �. ���.��� � 1.58
Sample mean:
�̅ � �� ∑ �� � ��� � �. ��⋯��.��� � 1.71
Sample range:
maximum � minimum � 2.99 � 0.81 � 2.18
Sample standard deviation:
s � ��#� ∑ �� � �̅ � � ��� �
�. �#�.�� $�⋯� �.��#�.�� $ � � 0.73
1. Summaries
Accuracy of Data
• Express different observations of a measurement, such as monthly sales figures, to
the same degree of precision.
• Most observations are rounded up or down in the last decimal place: heights of
1.341 metres and 1.417 metres become 1.34m and 1.42m respectively.
• Retain as much accuracy as possible in intermediate calculations: do not round the
sample mean when calculating a sample standard deviation.
• Avoid displaying results of calculations to more accuracy than is needed: present a
sample mean as 5.2 rather than 5.166666667.
• Enter data carefully into computer spread sheets and perform simple numerical and
graphical checks that the data are reasonable.
• Use a special symbol for missing data such as * rather than a space, zero or minus
number to ensure that these are excluded from your analysis.
• Avoid guessing missing data values without proper justification.
1. Summaries
Event % collection of outcomes Probability & % chance that % occurs Complement %′ & % ( & %′ � 1
impossible evens certain
0 ½ 1
2. Probability
Union % ∪ * % or * or both Intersection % ∩ * % and * Conditioning *|% * given %
% *
2. Probability
If
% = “motorist makes insurance claim this year” * = “motorist makes insurance claim next year”
& % � & * � -. and & *|% � -$ then
& %′ � 1 � & % � � & % ∩ * � & % / & * % � ��
& % ∪ * � & % ( & * � & % ∩ * � ��
2. Probability
θ
2. Probability
2. Probability
95% probability interval 479,521 has limits 3 4 1.966
2. Probability
2. Probability
3. Inference
7~9 3,6�
�̅ � 1� : �� �
���
;� � 1� � 1 : �� � �̅ �
�
���
95% confidence interval for 3 has limits
�̅ 4 < / ;�
0.1000 0.0500 0.0250 0.0100 0.0050 0.0025 0.0010 0.0005
1 3.078 6.314 12.71 31.82 63.66 127.3 318.3 636.6
2 1.886 2.920 4.303 6.965 9.925 14.09 22.33 31.60
3 1.638 2.353 3.182 4.541 5.841 7.453 10.21 12.92
4 1.533 2.132 2.776 3.747 4.604 5.598 7.173 8.610
5 1.476 2.015 2.571 3.365 4.032 4.773 5.893 6.869
6 1.440 1.943 2.447 3.143 3.707 4.317 5.208 5.959
7 1.415 1.895 2.365 2.998 3.499 4.029 4.785 5.408
8 1.397 1.860 2.306 2.896 3.355 3.833 4.501 5.041
9 1.383 1.833 2.262 2.821 3.250 3.690 4.297 4.781
10 1.372 1.812 2.228 2.764 3.169 3.581 4.144 4.587
11 1.363 1.796 2.201 2.718 3.106 3.497 4.025 4.437
12 1.356 1.782 2.179 2.681 3.055 3.428 3.930 4.318
13 1.350 1.771 2.160 2.650 3.012 3.372 3.852 4.221
14 1.345 1.761 2.145 2.624 2.977 3.326 3.787 4.140
15 1.341 1.753 2.131 2.602 2.947 3.286 3.733 4.073
16 1.337 1.746 2.120 2.583 2.921 3.252 3.686 4.015
17 1.333 1.740 2.110 2.567 2.898 3.222 3.646 3.965
18 1.330 1.734 2.101 2.552 2.878 3.197 3.610 3.922
19 1.328 1.729 2.093 2.539 2.861 3.174 3.579 3.883
20 1.325 1.725 2.086 2.528 2.845 3.153 3.552 3.850
21 1.323 1.721 2.080 2.518 2.831 3.135 3.527 3.819
22 1.321 1.717 2.074 2.508 2.819 3.119 3.505 3.792
23 1.319 1.714 2.069 2.500 2.807 3.104 3.485 3.768
24 1.318 1.711 2.064 2.492 2.797 3.091 3.467 3.745
25 1.316 1.708 2.060 2.485 2.787 3.078 3.450 3.725
26 1.315 1.706 2.056 2.479 2.779 3.067 3.435 3.707
27 1.314 1.703 2.052 2.473 2.771 3.057 3.421 3.690
28 1.313 1.701 2.048 2.467 2.763 3.047 3.408 3.674
29 1.311 1.699 2.045 2.462 2.756 3.038 3.396 3.659
30 1.310 1.697 2.042 2.457 2.750 3.030 3.385 3.646
40 1.303 1.684 2.021 2.423 2.704 2.971 3.307 3.551
60 1.296 1.671 2.000 2.390 2.660 2.915 3.232 3.460
120 1.289 1.658 1.980 2.358 2.617 2.860 3.160 3.373
∞∞∞∞ 1.282 1.645 1.960 2.326 2.576 2.807 3.090 3.291
4 3 2 1 0 1 2 3 4
0.1
0.2
0.3
−
critical value
<=.=�> � � 1
3. Inference
3. Inference
7~? @
@A � 1� :�� �
���
95% confidence interval for @ has limits
@A 4 1.96 / @ A 1 � @A
�
3. Inference
0.13
0.13
3. Inference
null hypothesis for:
mean B= ∶ 3 � 8.3 proportion B= ∶ @ � 0.1
two-sided alternative hypothesis for:
mean B� ∶ 3 E 8.3 proportion B� ∶ @ E 0.1
one-sided alternative hypothesis for:
mean B� ∶ 3 F 8.3 B� ∶ 3 G 8.3 proportion B� ∶ @ F 0.1 B� ∶ @ G 0.1
3. Inference
calculate test statistic and compare with critical value
reject B= or do not reject B= at 5% level of significance p<0.05 (reject) or p>0.05 (do not reject)
3. Inference
3. Inference
4. Modelling
4. Modelling
4. Modelling
4. Modelling
Chi-square Test for Association
Test for association between two factors with B=:"no association" and B�:"association".
For a table with r rows and c columns, compare the test statistic
N� � : obs � exp �
expQRR STRRU with the upper critical value from the N� distribution with V � 1 W � 1 degrees of freedom.
4. Modelling
4. Modelling
4. Modelling
4. Modelling
0.995 0.990 0.975 0.950 0.900 0.100 0.050 0.025 0.010 0.005
1 0.000 0.000 0.001 0.004 0.016 2.706 3.841 5.024 6.635 7.879
2 0.010 0.020 0.051 0.103 0.211 4.605 5.991 7.378 9.210 10.60
3 0.072 0.115 0.216 0.352 0.584 6.251 7.815 9.348 11.34 12.84
4 0.207 0.297 0.484 0.711 1.064 7.779 9.488 11.14 13.28 14.86
5 0.412 0.554 0.831 1.145 1.610 9.236 11.07 12.83 15.09 16.75
6 0.676 0.872 1.237 1.635 2.204 10.64 12.59 14.45 16.81 18.55
7 0.989 1.239 1.690 2.167 2.833 12.02 14.07 16.01 18.48 20.28
8 1.344 1.646 2.180 2.733 3.490 13.36 15.51 17.53 20.09 21.95
9 1.735 2.088 2.700 3.325 4.168 14.68 16.92 19.02 21.67 23.59
10 2.156 2.558 3.247 3.940 4.865 15.99 18.31 20.48 23.21 25.19
0 5 10 15 20 25 30
5. Forecasting
Business applications often involve data that are collected sequentially over
time (time series) with the aim of predicting (forecasting) future values.
Trend and Seasonality
If a time series 7X is not stationary, we remove trend and seasonality by transforming the original data. Trend is removed by taking lag-one differences,
also called integrating, to generate a new time series YX where
YX � 7X � 7X#� and seasonality is removed by taking seasonal differences to generate a new
time series YX where
YX � 7X � 7X#Z and ; typically takes the values 24 (hours per day), 7 (days per week), 12 (months per year) and 365 (days per year), though other values can arise.
5. Forecasting
Autoregressive AR(1) Model
For a stationary time series 7X , an autoregressive model of order 1 is defined by
7X � [7X#� ( \X where % \X � 0, var \X � 6� and cov \X-,\X$ � 0 for <� E <�, for weight [ that is estimated from data.
10 20 30 40 50 60 70 80 90 100
4
2
2
4 4
4−
xt
1001 t
5. Forecasting
Moving Average MA(1) Model
For a stationary time series 7X , a moving average model of order 1 is defined by
7X � _\X#� ( \X where % \X � 0, var \X � 6� and cov \X-,\X$ � 0 for <� E <�, for weight _ that is estimated from data.
10 20 30 40 50 60 70 80 90 100
4
2
2
4 4
4−
xt
1001 t
5. Forecasting
The Forecast
The one-step-ahead forecast for theAR(1) model is:
7AX��|X � [�X
The one-step-ahead forecast for the MA(1) model is:
7AX��|X � 0
5. Forecasting
10 20 30 40 50 60 70 80 90 100
4
2
2
4 4
4−
xt
1001 t
10 20 30 40 50 60 70 80 90 100
4
2
2
4 4
4−
xt
1001 t