Using the same data set and variables for your selected topic, add the following information to your analysis:
Running Head: INFECTIOUS DISEASE STATISTICS 1
INFECTIOUS DISEASE STATISTICS 10
Infectious Disease Statistics
Author Note
Infectious Disease Statistics
Scenario
You are currently working at NCLEX Memorial Hospital in the Infectious Diseases Unit. Over the past few days, you have noticed an increase in patients admitted with a particular infectious disease. You believe that the ages of these patients play a critical role in the method used to treat the patients. You decide to speak to your manager and together you work to use statistical analysis to look more closely at the ages of these patients. You do some research and put together a spreadsheet of the data that contains the Client number, Infection Disease Status, and the Age of the patient. The data set consists of 60 patients that have the infectious disease with ages ranging from 35 years of age to 76 years of age for NCLEX Memorial Hospital. It is as shown below
|
Patient # |
Infectious Disease |
Age |
|
Patient # |
Infectious Disease |
Age |
|
Patient # |
Infectious Disease |
Age |
|
1 |
Yes |
69 |
|
21 |
Yes |
72 |
|
41 |
Yes |
60 |
|
2 |
Yes |
35 |
|
22 |
Yes |
70 |
|
42 |
Yes |
62 |
|
3 |
Yes |
60 |
|
23 |
Yes |
76 |
|
43 |
Yes |
63 |
|
4 |
Yes |
55 |
|
24 |
Yes |
56 |
|
44 |
Yes |
53 |
|
5 |
Yes |
49 |
|
25 |
Yes |
59 |
|
45 |
Yes |
64 |
|
6 |
Yes |
60 |
|
26 |
Yes |
64 |
|
46 |
Yes |
50 |
|
7 |
Yes |
72 |
|
27 |
Yes |
71 |
|
47 |
Yes |
69 |
|
8 |
Yes |
70 |
|
28 |
Yes |
69 |
|
48 |
Yes |
52 |
|
9 |
Yes |
70 |
|
29 |
Yes |
55 |
|
49 |
Yes |
68 |
|
10 |
Yes |
73 |
|
30 |
Yes |
61 |
|
50 |
Yes |
70 |
|
11 |
Yes |
68 |
|
31 |
Yes |
70 |
|
51 |
Yes |
69 |
|
12 |
Yes |
72 |
|
32 |
Yes |
55 |
|
52 |
Yes |
59 |
|
13 |
Yes |
74 |
|
33 |
Yes |
45 |
|
53 |
Yes |
58 |
|
14 |
Yes |
69 |
|
34 |
Yes |
69 |
|
54 |
Yes |
69 |
|
15 |
Yes |
46 |
|
35 |
Yes |
54 |
|
55 |
Yes |
65 |
|
16 |
Yes |
48 |
|
36 |
Yes |
48 |
|
56 |
Yes |
61 |
|
17 |
Yes |
70 |
|
37 |
Yes |
60 |
|
57 |
Yes |
59 |
|
18 |
Yes |
55 |
|
38 |
Yes |
61 |
|
58 |
Yes |
71 |
|
19 |
Yes |
49 |
|
39 |
Yes |
50 |
|
59 |
Yes |
71 |
|
20 |
Yes |
60 |
|
40 |
Yes |
59 |
|
60 |
Yes |
68 |
Classification of the Variables in the Data Set.
Which variables are quantitative/qualitative?
The ages are quantitative. The column on the ages provide information in which the count, or quantity of the characteristic is most important. For example, we are interested in the total number of years in each .This type of variable is called numerical (or quantitative). Which variables are discrete/continuous?
The ages can be viewed as discrete. A discrete numerical variable can only have values at specific values. For example, the number of ages must be a whole number. (How would you introduce half of a YEAR?!) But don’t get the wrong idea! It is possible for a variable to have fractional values and still be discrete. The ages can also be viewed as continuous data because they vary for every particular person measured. The patient numbers are discrete that is they increment uniformly.
Measures of Center
The Mode
The mode is defined as the most frequently occurring number in a data set. The Importance of the mode is its most useful in situations that involve categorical (qualitative) data that is measured at the nominal level.
The Median
The median is simply the middle number in a set of data. When there is an even number of numbers, there is no true value in the middle. In this case we take the two middle numbers and find their mean. The importance is it helps us understand more about a data set
The Mean/Average
The average is a measure of center that statisticians call the mean. The importance of the mean is actually the numerical “balancing point” of the data set. Because is the numerical balancing point for the data, is in an extremely important measure of center that is the basis for many other calculations and processes necessary for making useful conclusions about a set of data.
The Midrange
The midrange (sometimes called the midextreme), is found by taking the mean of the maximum and minimum values of the data set. One of the reasons that the midrange is not commonly used is that it is only based on two values of the data set, and not just any two, but the values that are most likely to be outliers! It would be like basing your class grade on only two assessments and ignoring all the other work you may have done. Even if it works out as a higher grade for you, much of your accomplishments would be meaningless!
Weighted mean
A weighted mean, involves multiplying individual data values by their frequencies or percentages before adding them and then dividing by the total of the weights.
Measures of Variation
The Range
The range is simply the difference between the smallest value (minimum) and the largest value (maximum) in the data. The range is useful because it requires very little calculation and therefore gives a quick and easy “snapshot” of how the data is spread, but it is limited because it only involves two values in the data set and it is not resistant to outliers.
The disease data is here arranged in ascending order:
35 45 46 48 48 49 49 50 50 52 53 54 55 55 55 55 56 58 59 59 59 59 60 60 60 60 60 61 61
61 62 63 64 64 65 68 68 68 69 69 69 69 69 69 69 70 70 70 70 70 70 71 71 71 72 72 72 73 74 76
35 and 76 are the extreme values (76-35=41)
Interquartile Range
Similar to the range, the interquartile range is the difference between the quartiles. If the range tells us how widely spread the entire data set is, the interquartile range (abbreviated IQR) gives information about how the middle 50% of the data is spread.
Standard Deviation
The standard deviation is an extremely important measure of spread that is based on
the mean. Recalling that the mean is the numerical balancing point of the data. One way to measure how the data is spread is to look at how far away the values are from the mean. The difference between the actual value and the mean is called the deviation. Written symbolically it would be Deviation = x – x
Variance
The Square of the standard deviation.
Calculate the measures of center and measures of variation. Interpret your results in context of the selected topic. From the table above the following was obtained taking n=60 ages to be xi and mean = and the above guidelines used.
|
SUM |
3709 |
|
|
MEAN |
61.81666667 |
|
|
MODE |
69 |
|
|
MEDIAN |
61 |
|
|
MIDRANGE |
20.5 |
|
|
RANGE |
35 to 76 |
41 |
|
STANDARD DEVIATION |
8.924336687 |
|
|
VARIANCE |
79.64378531 |
|
Analysis
Any person suffering from the illness will have an average of 61 years. People suffering this illness must have their years lie between 35-76 years. This tells us that in that population only that age if affected. Standard deviation and variance tells us how this ages are varying from the mean of 61.
Conclusion
When examining a set of data, we use descriptive statistics to provide information about where the data is centered. The mode is a measure of the most frequently occurring number in a data set and is most useful for categorical data and data measured at the nominal level. The mean and median are two of the most commonly used measures of center. The mean, or average, is the sum of the data points divided by the total number of data points in the set. The median is the numeric middle of a data set. If there are an odd number of numbers, this middle value is easy to find. If there is an even number of data values, however, the median is the mean of the middle two values. The median is resistant, that is, it is not affected by the presence of outliers. The range is a measure of the difference between the smallest and largest numbers in a data set. The interquartile range is the difference between the upper and lower quartiles. The standard deviation is a measure of the “average” deviation for the entire data set. When we have the entire population, the sum of the squared deviations is divided by the population size. This quantity is called the variance. Taking the square root of the variance gives the standard deviation. For a population, the standard deviation is notated σ.
Infectious Disease Statistics Analysis
1. Discuss the importance of constructing confidence intervals for the population mean.
a. What are confidence intervals?
b. What is a point estimate?
c. What is the best point estimate for the population mean? Explain.
d. Why do we need confidence intervals?
Confidence intervals involve the range of values that are defined such that there is some probability that the value of a parameter lies within it. Confidence intervals provides some estimated range of values that are likely to involve an unknown population parameter, the estimated range that is calculated from some given set of sample data.
A point estimate of a population parameter involves the single value used to estimate the population parameter.
The sample mean x is a point estimate of the population mean μ. This is because for all populations, the sample mean x is an unbiased estimator of the population µ, which implies that the distribution of sample means tends to center about the value of the population mean µ. In addition, the distribution of sample means x for many populations tends to be quite consistent than the distributions of other sample statistics.
Confidence intervals are important to approximate the mean of the population. Confidence intervals address the problem of how well the sample statistics estimates the underlying population value. In essence, population intervals provide the range of values that are likely to contain the population parameter of interest.
2. Based on your selected topic, evaluate the following:
a. Find the best point estimate of the population mean.
b. Construct a 95% confidence interval for the population mean. Assume that your data is normally distributed and �ƒ is unknown.
i. Please show your work for the construction of this confidence interval and be sure to use the Equation Editor to format your equations.
c. Write a statement that correctly interprets the confidence interval in context of your selected topic.
From the information given, the disease data is here arranged in ascending order:
35, 45, 46, 48, 48, 49, 49, 50, 50, 52, 53, 54, 55, 55, 55, 55, 56, 58, 59, 59, 59, 59, 60, 60, 60, 60, 60, 61, 61, 61, 62, 63, 64, 64, 65, 68, 68, 68, 69, 69, 69, 69, 69, 69, 69, 70, 70, 70, 70, 70, 70, 71, 71, 71, 72, 72, 72, 73, 74, 76
Mean is calculated as the total sum divided by the count. Therefore, 3709 divided by 60. Mean is 61.82.
A confidence interval for the population mean with known standard deviation of size n, , where z* is the upper (1-C)/2 critical value for the standard normal distribution.
The standard deviation of the data is 8.92.
The standard deviation of the sample mean is equal to 8.92/square of 60 = 0.0025
The critical value for a 95 percent confidence interval is 1.96, where (1-0.95)/2 = 0.025.
A 95% confidence interval for the mean is [(61.82-(1.96*0.0025)], [(61.82+ (1.96*0.0025))] = (61.8151, 61.8249).
As the level of confidence decreases, the size of the corresponding interval has to decrease.
3. Based on your selected topic, evaluate the following:
a. Find the best point estimate of the population mean.
b. Construct a 99% confidence interval for the population mean. Assume that your data is normally distributed and �ƒ is unknown.
i. Please show your work for the construction of this confidence interval and be sure to use the Equation Editor to format your equations.
c. Write a statement that correctly interprets the confidence interval in context of your selected topic.
Mean is calculated as the total sum divided by the count. Therefore, 3709 divided by 60. Mean is 61.82.
A confidence interval for the population mean with known standard deviation of size n, , where z* is the upper (1-C)/2 critical value for the standard normal distribution.
The standard deviation of the data is 8.92.
The standard deviation of the sample mean is equal to 8.92/square of 60 = 0.0025
The critical value for 99 percent confidence interval is 2.576, where (1-0.99)/2 = 0.005
A 99% confidence interval for the mean is [(61.82-(2.576*0.0025)], [(61.82+ (2.576*0.0025))] = (61.8356, 61.83134).
In this case, an increase in sample size will decrease the length of the confidence interval with no reduction in the level of confidence.
4. Compare and contrast your findings for the 95% and 99% confidence interval.
a. Did you notice any changes in your interval estimate? Explain.
b. What conclusion(s) can be drawn about your interval estimates when the confidence level is increased? Explain.
From the 95% and 96%, 95 % confined intervals is quite almost correct while the other is far from correct. The 99 percent confidence interval is wider than the 95 percent confidence interval. This is because it allows one to be more confident that the unknown population parameter is in within the interval.
Reference:
Triola, M. F. (12/2012). Elementary Statistics, 12th Edition. [VitalSource Bookshelf Online]. Retrieved from http://bookshelf.vitalsource.com/#/books/9781323114896/
Smithson, M. (2003). Confidence intervals. Thousand Oaks, Calif: Sage Publications.
Albright, S. C., Winston, W. L., & Zappe, C. J. (2011). Data analysis and decision making. Mason, Ohio: South-Western/Cengage Learning.