NEED TO USE SPSS TO COMPLETE
Empirical Research Methods for Business
Q1a) Price
The total samples are 120 and SPSS REGARD 100% of the observations.
The mean amount of house in AUD (In $000s) is 886.575 based on descriptive statistics with a standard error of 29.6344 from the mean of the population. The confidence interval of the mean between the upper bound of 945.3116 and the lower bound of 827.8384 is 95%.
The mean value is 886.5750 which is very close to the trimmed mean value which is 876.7778. Through this, it shows that there is a very small impact of the extreme values. 852.000 is the median cost of the house which shows that the distribution score is skewed slightly towards the lower end of the distribution with 0.426 as the value as shown in the table on the left side. 1761.00 is the maximum value of a house in a population while 192.00 is the minimum value of a house in the population. 324.94666 is the value of the standard deviation which is slightly high in this phenomenon therefore, it means that the prices of the houses are spread out over a wide range of values.
The box plot here shows that the distribution is skewed towers the left of the distribution graph this is because the box is closer towards the lower end of the graph. Therefore, there is also an individual large outlier, which is at the top of the line.
The significance of Kolmogorov-Smirnov is shown at 0.200 and the significance of Shapiro-Wilk is at 0.116 as for the normality test. The data can be stated to be normal distributed since both of the values are above 0.05.
The histogram is not normally distributed and is skewed more towards the left and it is asymmetric.
In conclusion, the indicated value of each house which is plotted against the expected value of the normal distribution does not bring a reasonably straight line therefore, indicating that it’s not a normal distribution.
1b) Lot Size
100% of the observations are regarded by SPSS whereby the total samples are 120 as for the Lot size for the houses.
The mean size of the house in the data set is 1175.2250 which is suggested by the descriptive statistics for the Lot Size 34.041 being the standard error from the mean of the population. Between the upper bound of 1242.6296 and the lower bound of 1107.8204 there is a confidence interval of 95%.
1161.3796 is the trimmed mean value which is closer to 1174.2250 which is the mean value. It communicates that there is a small level of impact of the extreme values. Since the distribution score is skewed towards the lower end of the distribution, therefore the median size of the house is 980.00 having a value of 0.838 as indicated in the figure on the left. The maximum size of a house in the dataset is 1950.00 while its minimum size is 632.00.The sizes of the houses varies over a wide range of sizes since the standard deviation is rather higher in this phenomenon and that is 372.90043.
Since the box is closer towards the lower end of the graph therefore the box plot shows that the distribution is skewed towers the left of the distribution graph. This means that there are no outliers.
The significance of Kolmogorov-Smirnov is suggested to be at 0.000 while the significance of Shapiro-Wilk is at 0.000 since both of these values are below 0.05 hence this is for the normality test. Therefore, the data is regarded to be not normally distributed
The histogram shows that is asymmetric and skewed more towards the left as left side of the graph shows extremities at the left therefore it is not normally distributed.
Since the line is not straight, it shows that it is not a normal distribution because the QQ plot shows a large number datasets falling out of the line when plotted against the expected value of the normal distribution therefore it shows that it is not a normal distribution.
Q1c) Material
|
Material |
|||||
|
|
Frequency |
Percent |
Valid Percent |
Cumulative Percent |
|
|
Valid |
Timber |
45 |
37.5 |
37.5 |
37.5 |
|
|
Veneer |
36 |
30.0 |
30.0 |
67.5 |
|
|
Brick |
39 |
32.5 |
32.5 |
100.0 |
|
|
Total |
120 |
100.0 |
100.0 |
|
The data from the dataset cannot be measured since the material is considered nominal scale. Out the entire population from the descriptive stats since materials is considered to be nominal scale, the data from the dataset cannot be measure of 120houses, 45 houses, use timber which is equivalent to 37.5% of the population, 36houses use Veneer, which is 30% of the population. The rest of the 39 houses brick is the material used which is 32.5%.
|
Condition |
|||||
|
|
Frequency |
Percent |
Valid Percent |
Cumulative Percent |
|
|
Valid |
Very Poor |
15 |
12.5 |
12.5 |
12.5 |
|
|
Poor |
40 |
33.3 |
33.3 |
45.8 |
|
|
Good |
42 |
35.0 |
35.0 |
80.8 |
|
|
Excellent |
23 |
19.2 |
19.2 |
100.0 |
|
|
Total |
120 |
100.0 |
100.0 |
|
Q1d) Condition
The data set cannot be measured since the condition of the houses is considered ordinal. 15 houses are in poor condition which makes up 12.5% of the population of 120 houses. 40 houses are in poor condition making up 33.3% of the population, the majority of the houses are in good condition with a total up to 42 of them making up 35% of the population. Finally, 23 houses are in excellent condition making 19.2% of the total population.
Q2)
The value of significance of K Kolmogorov-Smirnov is 0.01 while the value of significance of Shapiro-Wilk is 0.002 according to results from the normality test for distance to train. This indicates that it is not normally distributed since both values are less than 0.05.
Since a large portion of the values goes to high extremes, then the histogram to distance to Train appears to be asymmetric. Therefore, it shows that it is not normally distributed.
|
Correlations |
|||
|
|
Price |
To Train |
|
|
Price |
Pearson Correlation |
1 |
.003 |
|
|
Sig. (2-tailed) |
|
.974 |
|
|
N |
120 |
120 |
|
To Train |
Pearson Correlation |
.003 |
1 |
|
|
Sig. (2-tailed) |
.974 |
|
|
|
N |
120 |
120 |
The Pearson Correlation shows that there is a very weak relationship between the distance to the train and also including the price of the houses.
The distance to the train station rarely can affect the price of the houses.
By using the scatter plot for further test the hypothesis, it is indicated that there is a weak linear relationship between the distance of the house to the train station and the prices of the house.
Since both the results show that the significance is zero then, as per the normality test, Kolmogorov-Smirnov and Shapiro-Wilk normality test, it is indicated that it is not normally distributed.
The Histogram for distance to bus station is not normally distributed since it appears to be asymmetric since there are large extremities at both ends of the graph.
|
Correlations |
|||
|
|
Price |
ToBus |
|
|
Price |
Pearson Correlation |
1 |
-.024 |
|
|
Sig. (2-tailed) |
|
.796 |
|
|
N |
120 |
120 |
|
ToBus |
Pearson Correlation |
-.024 |
1 |
|
|
Sig. (2-tailed) |
.796 |
|
|
|
N |
120 |
120 |
Per the Pearson Correlation results between the distances to the bus station to the price of the house shows that there are a very weak negative relationship at only negative 0.024.
The Scatter Plot results of the distance to the bus stations and the prices of the houses indicate that the data points are spread unevenly on the diagram with no patterns. Therefore, there is a very weak linear relationship and little correlation too.
Q3)
Cronbach’s Alpha from reliability statistic indicates 0.923 as a result. Since 0.923 values is above 0.9 then it is regarded to have an excellent level of internal consistency of scale. Moreover, by eliminating one or more variables the reliability of the statistic can be increased.
By through deleting Q8 which is the highest value at 0.928, will further increase the Cronbach’s Alpha value.
Through deleting off Q8 results in the Cronbach’s Alpha an increase to 0.926 from 0.923 is significant. Though it is being considered to be an excellent level of internal consistency of scale, therefore, the value can be increased by eliminating off another variable.
To increase the Cronbach’s Alpha value, while judging from the dataset, Q11 would be the highest value at 9.13, deleting this variable, will increase Cronbach’s Alpha value.
When you delete both variables Q8 and Q11 the Cronbach’s Alpha value will increase to 0.931 from 0.926.
Since the Cronbach’s Alpha is above 0.9, it is considered to have an excellent level of internal consistency of scale and eliminating off a variable to increase the Cronbach’s Alpha value is not allowed. If that happens then the dataset suggest that the next variable to eliminate off would be Q9.