How do various drinks influence athlete performance [caffeinated, potassium-rich, protein-rich, vitamin-water, carbohydrate-rich, exercise drinks, sport drinks, energy drinks, fruit juice, soft drinks]?
PED 598 – Issues in Reliability and Validity
Dr. Terry Conkle
Spring 2020
1
Validity
The degree to which scores from a test, or instrument, measures what they are purported to measure
It is the soundness of the interpretation of scores from a test
It is the most vital consideration in measurement
In research it is also defined as the degree to which a researcher’s conclusions aftually match the study’s findings/results – rather from erroneous or chance sources
Validity is generally associated with quantitative studies
2
Trustworthiness
A term associated with qualitative studies.
The degree to which a researcher convinces the audience that the research was completed using appropriate techniques, that the findings are credible, and interpretations are appropriate and fully developed
3
Reliability
Relates to both validity and trustworthiness
The degree to which a study can be repeated with similar results. Is it consistent?
If a study is replicated, if similar outcomes are likely, then a reader can be confident that the results are meaningful – if very different results are found then one cannot be as confident that earlier results (or possibly the new results) are reliable
4
Types of Validity
Internal validity – the approximate validity with which we infer that a relationship between 2 variables is causal, or absence of relationship implies absence of cause (in an experimental cause-and-effect” study). In other words did A really cause B, or did some other variable?
External validity – the extent to which a study’s results can be generalized to other similar groups, individuals, or situations. One question of concern is, how well did the sample in a study truly represent the population?
Construct validity – the degree to which a researcher truly measures the construct of focus in a study.
5
Internal Reliability
The degree to which data collection, analysis, and explanations or conclusions are similar under comparable research conditions.
Two major areas to address relative to Internal Reliability are: Instrumentation Reliability and Observation Reliability. The reliability of a measure chosen for evaluating the DV in a study is crucial to reliability and validity of a study. Researchers can use a measure that has been previously published and standardized, or they can create their own measure. Previously published instrumentation should have a reported reliability coefficient; but if a new measure is created, the researcher should determine the reliability and report its coefficient
6
Reliability of Instruments or Instrumentation
Reliability of instrumentation is described by parallel forms reliability, test-re-test reliability, split-half reliability, or Cronbach’s alpha.
Parallel Forms Reliability – the degree to which one’s score is similar when given 2 different forms of the same test
Test-Retest Reliability – the degree to which one achieves a similar score on an assessment measure when the entire measure is administered once and then administered again at some later date
Split-half Reliability – the degree to which a person achieves a similar score on one half of the test items compared to the other half (e.g., even # items compared to odd # items
Cronbach’s Alpha – a statistical formula used to determine reliability based on at least 2 parts of a test
7
Reliability of Instrumentation
Instrument reliability is measured by a reliability coefficient (r = 0.xx)– these indicate the relationship between multiple administrations, multiple items, or other analyses of evaluation measures. Reliability coefficients range from 0.00 to 1.00, with zero = to no relationship and 1.00 = to perfect relationship. 0.80 or higher is generally considered an adequate relationship.
Various guides exist, but this is a commonly-used one:
Excellent = .90 – 1.00
High = .80 – .89
Average or Fair = .60 – .79
Unacceptable = .00 – .59
Score approaching ZERO is unreliable, from either direction, Score approaching (1.00) shows a reliable correlation
A scholarly article will typically report reliability according to the statistical test used: These are common - Cronbach's Alpha, Pearson's r, Spearman-Brown r, Kuder-Richardson r
8
External Reliability
The extent to which an independent researcher could replicate the study in other settings.
9