#7208 Topic: Statistic Project
Introduction: In the major project student could select a date set with at least 25 cases. The data set must also have two categorical variables and two quantitative variables. The dataset that is used in this project is about nutrition study, a study that consistent of 315 cases. The cases have 17 variables. This variable reveals the result from some experiments, that is more like a simple given about people’s nutrition intake base of their ages and lifestyle. Below are the variable name, description and variable type that were used for this project.
Analysis: In this portion of the project, Statkey was use for all graphing and most calculating use. Analysis of each variable and the relationships such as a two ways table, correlation, correlation coefficient, and confidence interval was used for some of the variables.
A) Analysis of one Quantitative Variable:
calories
From the statistic above, it discusses the calories from the dataset called nutrition study. The value of the mean seems to be more than the median which is 1666.800. Also, this show that the dataset is a left-skewed which is 1.7. The minimum is 445.2 while the maximum net sales is 6662.2. Also, the outlier is 3711 the one in the middle ,6662.2 the one at the very end.
B) Analysis of one Categorical Variable:
Gender Variable with frequency table
|
Gender |
Frequency |
Relative frequency |
|
Female |
273 |
0.87 |
|
Male |
42 |
7.5 |
In the categorical Gender variable, it shows the frequency and the relative frequency. Sample taken from 273 females and 42 males which total up to 315 explains the relationship between the frequencies. Both male and female was asked about their nutrition habits, and it appear that the relative frequency shows a higher percentage of 87 female compare to only 42% of males when to come to changes in their diet.
C) Analysis of One Relationship between two Categorical Variables:
There seem to be some association between the two Gender and vitamin use. The grand total of female and male who use the vitamin occasional or answer no are both in the 20’s % castigate. Also, it shows that 75.24% of female used the vitamin while only 11.11% of male used the vitamin. Also, the graph show a higher percentage of people who people who uses the vitamin regularly.
D) Analysis of One Relationship between a Categorical Variable and a Quantitative Variable:
Relationship between calories and Gender.
The relationship between the calories and Gander shows that it has a correlation which is (-0.87156). The two variables seem to have a linear correlation, because it in the negative which mean that the magnitude of the correlation is close to zero. So therefore, as the age increase, the amount of age will decrease. The mean however, look larger for male than female who has more sample size.
E) Analysis of One Relationship between Two Quantitative Variables
Relationship between Alcohol and Cholesterol
Graph showing the confidence interval (another relationship) between alcohol and cholesterol
Summer Statistics
The statistic above analyses the relation and correlation between two quantitative variables. The two variable that was being used are alcohol and cholesterol. The cholesterol seems to be a larger data set than the alcohol. For example, the mean for the alcohol is 3.279, while the mean for the cholesterol is 242. 461.This also clarify that more than half of both the data set is worth 206.3. The maximum is 900.7 and 203, while the minimum is 37.7 and 0.3. The slope is 1.952 and intercept is 236.059
For the relationship between the two datasets, the correlation is (0.182263984). The Cholesterol and alcohol does not seem to have a linear correlation, because it in the positive. This means that the magnitude of the correlation is a bit far from zero. So therefore, both dataset have no effect on each other.
Also for the graph that shows the confidence interval relationship between alcohol and cholesterol, explains that we are 95% sure that there is a 0.031 to 0.298 that people who drink alcohol have high cholesterol.
Conclusion: overall, this major project has been very helpful in helping me understand how to properly analysis the relationship between variables. With the help of Statkey graphing and summary statistics which points out the key findings, such as five number summaries, left or right skewed, outliers, and pattern of correlation, which makes it easy to visualize the data set. Also, next time we could use confounding variable because it has extra variables that has unusual outcome on experiment results; or continuous variable because of the infinite number of values that it has. Those variables will better help me understand all the different ways that data sets can be analysis.
Worked cited:
http://higheredbcs.wiley.com/legacy/college/lock/0470601876/data_sets/datapage.html
Statkey.com