Statistic Assignment

profileDanyah
stat_lab_2-1.pptx

LAB 2:

Descriptive Statistics

1

Descriptive statistics are numerical estimates that organize and sum up or present the data.

For quantitative variables (scale)

Mean with Standard deviation are used to summarize non-skewed scale variables

Median with range or interquartile range are used to summarize skewed scale variables

The three steps to evaluate the normality assumption are:

Compare the statistics values ( mean versus median)

Obtain the histogram with normal curve

Obtain the Box-Whiskers plot

For this class,

If there is any extreme outliers, median with range should be used to summarize the variable of interest

If there is any outliers (regular outliers), you need to based your decision regarding the best measure (mean with SD or median with range) to summarize the variable of interest on the shape of the histogram

Introduction

2

For qualitative variable (nominal or ordinal)

Frequency distributions (number with percentages) are used to summarize qualitative variables

Descriptive statistics for multiple groups:

Use Split file option in SPSS to obtain the measures of central tendency and the measures of variation for quantitative variables.

After you split your file by the grouping variable, you should follow the previous steps to select the most appropriate measures to summarize your variable of interest.

Please note that you have to un-split the data before running further analysis

Use Crosstabs option in SPSS to obtain the frequency distributions (number with percentages) for the qualitative variables

Introduction

3

4

Types of variables

Continuous (Quantitative) Variables

Qualitative (Categorical) Variables

Nominal/ Ordinal

Interval/Ratio

Number and Percent

N (%)

Normal distribution

1- Statistics {Mean and Median}

2- Histogram with Normal Curve

3- Box-Whiskers Plot

No

Median with Range

Yes

Mean with Standard Deviation

Box-Whiskers Plot

5

6

7

Box-Whiskers Plot

Source: http://support.sas.com/documentation/cdl/en/statug/63347/HTML/default/viewer.htm#statug_boxplot_sect017.htm

8

Extreme outliers (values greater than 3 IQR from Q1/Q3)

Outliers (values between 1.5 and 3 IQRs from Q1/Q3)

Whisker extends to furthest observation within Q3 + 1.5*IQR

Whisker extends to furthest observation within Q1 - 1.5*IQR

9

Example

Types of Variables

Procedures

Un-split the data before running further analysis

Descriptive Statistics for Multiple Groups

10

Split File

Qualitative or Categorical Variable (Grouping variable)

Crosstabs

Qualitative or Categorical Variable

Qualitative or Categorical Variable

Continuous or Quantitative Variable

Men (Gender)

Baseline Pulse

Females (Gender)

Widowed (Marital Status)

Example:

Following is a dictionary for a data set. The data collected on a number of people from Cornwall, Ontario, Canada who attended a lifestyle intervention program (Coronary Health Improvement Project, or better known as CHIP) consisting of a series of lectures and personal counseling sessions every day for a five day period. This data set consists of a number of demographic and clinical variables. Create the lifestyle dataset using the table below and answer the following questions.

Variable View

Data View

Table 1
  Variable   Mean ± SD   Median (Range)   N (%)
  Age      
  Exercise            
  Smoke      
  Weight      
  Glucose      

Question1: Summarize using the most appropriate measure the following variables presented in the Table 1.1. Choose either the mean, median or n(%) as the most appropriate measure for each variable.

Answer: 1.1

For quantitative variables:

Step1: Compare statistics values (mean versus median) for all variables

Step2: Evaluate the normal curve

Answer: 1.1

For quantitative variables:

Step3: Evaluate the Box - Whiskers Plot

Variable Statistics (mean vs. median Histogram with normal curve Box- Whiskers Plot Decision
Age Close Normal One regular outlier (no extreme outliers) Mean ± SD
Weight Close Normal No outliers Mean ± SD
Glucose Close Skewed Extreme outlier Median (Range)

Normality Assumption Checklist

For qualitative variables

Question1: Summarize using the most appropriate measure the following variables presented in the Table 1.1.

Question2: Summarize using the most appropriate measure the following variables presented in the Table 1.2.

Step1: Split file by grouping variable (Gender)

Question2: Summarize using the most appropriate measure the following variables presented in the Table 1.2.

Step2: Compare statistics values (mean versus median) for all variables

Question2: Summarize using the most appropriate measure the following variables presented in the Table 1.2.

Step3: Evaluate the normal curve

Male

Female

Question2: Summarize using the most appropriate measure the following variables presented in the Table 1.2.

Step4: Evaluate the Box - Whiskers Plot

Male

Female

Question2: Summarize using the most appropriate measure the following variables presented in the Table 1.2.

Note: Un-split the data file before running further analysis

Question3: Summarize using the most appropriate measure the following variables presented in the Table 1.3.

  Qualitative Variables   Male : n (%)   Female : n (%)
  Frame Small        
  Medium        
  Large        
  Exercise: None        
  Mild        
  Moderate        
  Vigorous        

Question3: Summarize using the most appropriate measure the following variables presented in the Table 1.3.

Step1: Use Crosstabs option in SPSS to obtain the frequency distributions for the qualitative variable by the grouping variable

Question3: Summarize using the most appropriate measure the following variables presented in the Table 1.3.

  Qualitative Variables   Male : n (%)   Female : n (%)
  Frame Small    0 (0.0)    0 (0.0)
  Medium    4 (40.0)    6 (60.0)
  Large    6 (60.0)    4 (40.0)
  Exercise: None    4 (40.0)    6 (60.0)
  Mild    1 (10.0)    0 (0.0)
  Moderate   3 (30.0)    3 (30.0)
  Vigorous    2 (20.0)    1 (10.0)

Table 1.1

Variables

Mean ± SD

Median (Range)

N (%)

Age

51.60 ± 12.75

Baseline Weight 176.55 ± 38.46

Baseline Glucose

5.25 (5)

Exercise:

None

Mild

Moderate

Vigorous

10 (50.0)

1 (5.0)

6 (30.0)

3 (15.0)

Smoking Status:

Non-smoker

Smoker

20 (100.0)

0 (0.0)

Table 1.1

Variables

Mean ± SD

Median (Range)

N (%)

Age

51.60 ± 12.75

Baseline Weight

176.55 ± 38.46

Baseline Glucose

5.25 (5)

Exercise: None

Mild

Moderate Vigorous

10 (50.0)

1 (5.0)

6 (30.0)

3 (15.0)

Smoking Status: Non-smoker Smoker

20 (100.0)

0 (0.0)

Males n =

Females n =

Quantitative Variables

Mean ± SD Median (Range)

Mean ± SD Median (Range)

Age

Baseline Weight

Baseline Glucose

Males n =

Females n =

Quantitative Variables

Mean ± SD Median (Range)

Mean ± SD Median (Range)

Age

Baseline Weight

Baseline Glucose

Males n =

Females n =

Quantitative Variables

Mean ± SD Median (Range)

Mean ± SD Median (Range)

Age

Baseline Weight

Baseline Glucose

Males n =10

Females n =10

Quantitative Variables

Mean ± SD Median (Range)

Mean ± SD Median (Range)

Age

56.1 ± 14.3

47.1 ± 9.6

Baseline Weight

189.3 ± 35.8

150.5 (132)

Baseline Glucose

5.5 (5)

5.01 (1.7)

Males n =10

Females n =10

Quantitative Variables

Mean ± SD Median (Range)

Mean ± SD Median (Range)

Age

56.1 ± 14.3

47.1 ± 9.6

Baseline Weight

189.3 ± 35.8

150.5 (132)

Baseline Glucose

5.5 (5)

5.01 (1.7)