WK 3 DIS, Data
NURS 8211 Research for an evidence-based practice Week3:Descriptive Statistics
Week 3: Descriptive Statistics
Objectives
Differentiate between measures of central tendency and measures of dispersion for continuous variables
Explain how frequencies and percentages can inform a research, QI or EBP project
Formulate a null, directional and nondirectional alternate hypothesis from a published research study
Explain the standardized infection ratio (SIR) metric and how it can be used in research, QI, and DNP projects
Explain the differences between type 1 and type 2 measurement errors and identify key tools to use to reduce both.
Levels of measurement
Categorical:
Nominal
Ordinal
Continuous:
Interval
Ratio
3
Descriptive statistics
Measures of Central Tendency
Mean
Mode
Median
Measures of Dispersion
Standard Deviation
Variance
Range
The mean is the arthmetric average. The mode is the most frequently occurring value in the dataset and the median is the point at which there is an equal number of observations above and below.
The standard deviation tells you how much variation there is, away from the mean. The variance is the SD squared, and the range is the difference between the minimun and maximum value.
Only the mode is useful for categorical data. Categorical data can only tell you “how many” and “what percentage” does that count represent.
You can “see” these data, often in published studies formatted in a table. For example, patient demographics help to describe the sample. Sometimes these variables are used in hypothesis testing with inferential statistics, sometimes they are not and just descriptive statistics are used. Sometimes you see both descriptive and inferential in the same table.
Depending on whether the data are categorical or continuous, you will see the M and the SD or the counts n and percentages (%).
4
Descriptive statistics
Means/Standard deviations
Counts/Percentages
Percentile rankings
Confidence intervals
5
Bar Chart vs. histogram
Sood et al. 2021
QI study performed at a hospital in NY with data from 2017 to 2019, focused on the incidence of SSI in C section patients (CB births) with a hospital-wide perioperative bundle implemented.
2,875 total CBs (N)
Of those, there were 1,086 in the period of time described as “prebundle” and you see from the table that there is a mean age and standard deviation are listed. There are also mean scores for gestational age (GA) with a standard deviation. Both age and GA are measured as continuous data, so an average makes good sense.
Now, sometimes the median is chosen as a better alternative to the mean. Often, that is because the mean is subject to outliers that can distort the distribution. A median score might be more useful. The range is used instead of the SD, b’c it is a measure of variability that is appropriate for the median. The SD shows deviation from the mean. So if the data are subject to broad variability with outliers, the median and range are better measures of central tendency and dispersion that the M and SD. The median is often used for ordinal level data that uses ranks.
This study is provided for your review and you can take a look and get a better idea of the variables ROM and Parity.
For now let’s take a look at ethnicity. This is a categorial variable that is essentially nominal level data. Note that for ethnicity you see counts and percentages. For example, you can see that there were 677 Asian patients in the prebundle that represents 62.3% of the total. You can easily do the math: 677+75+261+73= 1,086 CBs in the prebundle group. Of these, 677 were of Asian descent. So (677/1086) x 100 = 62.3%
7
Hypothesis testing: Part I
The null hypothesis states that there is no relationship or association between variables (the independent and the dependent)
The alternate hypothesis states that there is a relationship or association between variables (the independent and the dependent)
Hypotheses can be directional or nondirectional
The null hypothesis states that there is no relationship or association between variables (the independent and the dependent).
Here is an example from Beydoun et al. (2022)
There is no association between perioperative prophylaxis practices and surgical site infections (SSIs) in patients undergoing vascularized reconstruction of the upper aerodigestive tract (UADT)
The dependent variable is the outcome variable, the SSIs. The independent variable is the one that gets manipulated. In this case that would be the perioperative prophylaxis. That includes both topical prophylaxis and antibiotic use.
The alternate hypothesis states that there is a relationship or association between variables (the independent and the dependent)
There is an association between perioperative topical antisepsis and surgical site infections (SSIs) in patients undergoing vascularized reconstruction of the upper aerodigestive tract (UADT)
Now, alternate hypothesis can be directional or non directional. In this study, the researchers had participation from 12 academic medical centers over 11 months with 554 patients and tracked in the type of prophylaxis used (topical and/or antibiotics). They found a decreased risk of SSI with the use of topical and antibiotic prophylaxis.
Most research uses both descriptive and inferential statistics. Published research studies will not usually have an actual hypothesis. Sometimes a research question is posed, but you can always determine the hypothesis from the published work. Inferential statistics come in many, many types: both parametric (based on probability theory) and nonparametric, which is typically used in situations where the assumptions of the parametric test are not met. These include categorical data like dichotomous data which has only two choices. Ordinal data has a ranking within it; nonparametric inferential testing could also be useful.
Now you can usually determine the type of inferential used from a published work, whether parametric or nonparametric by the presence of a p value.
8
Standardized infection rate (SIR)
A summary measure used to track hospital acquired infections (HAIs)
The SIR has a risk adjustment component that allows for comparisons at the national, state, or local level.
The SIR compares to a prediction (the NHSN baseline) taking into account several risk factors
A SIR of > 1.0 indicates more HAI than predicted; a SIR of < 1.0 indicates fewer HAIs than predicted.
Let’s take a look at Hicks, T (2024) from the Walden repository.
Hicks implemented an educational process with a revised CLABSI prevention bundle and designed a pre and postintervention comparison to determine if the outcome (SIR) changed after the bundle was implemented and to determine whether or not the process indicators (like, documentation of dressing change, saline flush, etc.) were improved pre to post. He found a statistically significant improvement in compliance with the bundle and four consecutive quarters with a zero incidence of CLABSI.
9
Atlantic Health System Morristown Medical Center
Data are often publicly reported. The source of this data is the State of NJ DOH.
https://web.doh.state.nj.us/apps2/hpr/profile.aspx?num=11403
10
Hypothesis testing: Part II
We typically set a level of significance, that is essentially a cut-off point. The is the point at which we decide that we will reject the null hypothesis (there is NO difference) in favor of the alternate hypothesis (there IS a difference). Now, you can visualize a difference easily but looking at the means and standard deviations, but this won’t tell you whether the difference is due to chance or something else, some other factor, for example, the result of an intervention.
PP. 215-217 in Salkind and Frey 5th ed provides a keen overview of tests of significance. These tests are ways to determine whether or not there is a statistically significant outcome.
Typically, we set a level of significance at .05 and that is because it is based on the graphic above which is figure 10.2 on p217. The level of significance is set based on the statistic chosen. Identifying a critical value….the point at which decisions came be made about the null hypothesis.
11
Sood et al. (2021)
Sood et al. (2021) examined the impact of a hospital-wide perioperative bundle on the SSI rate in Caesarean Births
Total cases = counts
SSI (infection rate, %) per 10,000
# comparisons…..prebundle, transition, postbundle.
P=0.009……considerably less that .05.
These are counts, and percentages (how many) descriptive statistics.
The p value signifies the inferential comparison. (there were two inferential tests used in the study both of which we will take a closer look at later on.) The comparison above with the p value of .009 is a chi square test.
12
Measurement Error
Type 1:
The type of error that is made when statistical significance is found, but it is not really there.
The all-important p value is the protection against a type I error.
Type II:
A type II error is made when you don’t find statistical significance but it is really there.
Sample size is the biggest protection against a type II error (in general, bigger is better)
Discuss clinical vs. statistical significance.
13
Key Points
Descriptives vs. Inferential statistics
Levels of measurement: continuous vs. categorical
Hypothesis testing
The standardized infection rate (SIR) and how it might be used
Measurement error types