quantatative reasoning

profilebillyjates88
qr_background_info.docx

Part I

METHODS OF DATA COLLECTION & THEIR STRENGTHS AND WEAKNESSES

The ongoing struggle to avoid BIAS:

In every part of the enterprise of performing research in health science, a researcher needs to take great pains to avoid the dreaded possibility of BIAS.

BIAS, or error, can come about in any number of ways during the process of defining the question, collecting the data and analyzing it. It can also happen from random causes; what I like to refer to the "stuff happens" effect.  But this is by definition beyond the researcher's  control. 

In every way that can possibly be anticipated, there is a need to control for known sources of bias. If the data is BIASED towards a certain outcome that does not reflect reality, then a meaningful or useful answer to the original question has not been obtained.

Once the researcher has defined the question, the next step will be to find a way to obtain subjects that minimizes the potential for creating bias through the selection procedure.

Obtaining subjects for study - data collection methods:

Data is the word we use for the information that we collect in order to do our research (the singular for this word is datum but we rarely use it.)

( Click here for a Presentation on Types of Data )

Data collection is also known as sampling. It might not seem obvious, but HOW you go about obtaining your subjects can be as crucial to the validity of your outcome as the question you ask and the type of statistical procedure you decide to use to analyze your data.

There are two broad categories of data collection in research:

· Probability sampling

· Non-probability sampling

Probability sampling is also called random sampling and is considered to be the most powerful and desirable method because theoretically each member of the larger population from which the sample is drawn had an equal chance of being chosen.

Of course, it may occur to you that this can be very easy to imagine, but very hard to execute. Even if you have complete control over the sampling procedure (let's say you have 3,000+ experimental rats to test out your new cancer treatment)  you can see right away that any subjects you pull from this sample are NOT by definition random. They may be randomly chosen from your subject pool, but the fact that they were in your pool to begin with makes them by definition NOT randomly selected. How can we randomly sample human beings in similar studies?  If they have the cancer we are trying to treat, they are also by definition NOT randomly selected.

Systematic sampling might get us around some (but not all) of these problems. In a more benign example, let's say we are surveying hospital patients to determine what factors cause them to perceive their interactions with the nursing staff as positive and comfortable. If we surveyed all the patients in several hospitals, we would not be creating a random sample, however, if we chose every ith (let's say 10th) patient admitted to all 20 hospitals within 30 miles of our university, then we would come closer to obtaining some of the advantages of a probabilistic selection without being truly probabilistic in our procedures. Every patient in all 20 hospitals had a 10% chance of being chosen - that's still not random.

Stratified sampling is useful when we know that the larger population, to which we wish to generalize our conclusions, has two or more subpopulations. For example, let's say we are curious about whether or not nursing students feel adequately prepared for their quantitative analysis studies by their high school mathematics coursework. It might occur to you that our population of nursing students has a large female and smaller but still substantial male subpopulation. So we might want to stratify our sample relative to the proportion of females and males at the school - if your school has 400 female and 180 male students, you might want to take 10% from each group (40 females and 18 males.)  Or, in this case, because mathematics education techniques and trends changes from generation to generation, we might want to look at our 18 to 25-year-olds as contrasted with our 26-to-35 year olds as contrasted with our 36-to-45 year olds etc. and we would take 10% of each group.

Non-probability sampling means that there will be no way to even approximate a chance to be selected, or that you don't try to approximate it. 

Contrast the method of the quota sample with the stratified sampling described above. You decide to just find 5 nursing students - any five - in each age group and ask them about their perceptions of how well-prepared by their high school math courses they feel to take quantitative analysis, and not even bother with the relative proportions of age groups.

Or finally, the convenience sample is just what the name says: convenient. The subjects who just happen to be there and available. If I want to know how my Introductory Psychology students at Santa Monica College like the Virtual Office Hours system for posting student questions for faculty, I merely survey them at the end of the course. Perhaps you can tell me why this survey would not be very informative. Think about these aspects:

· Demand characteristics: The students are able to guess what my agenda in surveying them on this question is, and they either deliberately answer in a way that will help me or hurt me. In either case, the information I get will be distorted or biased.

· Experimenter bias: Any other impact that my behavior towards them might have on how they answer.

· Representativeness of the sample: Will the fact that they are one small segment of the much larger population of students at the college matter? How so?

Also please visit these links:

Statistics Glossary. Retrieved Jan 1, 2012 from  http://www.stats.gla.ac.uk/steps/glossary/sampling.html

Trochim, W.K. (2006). Research Methods Knowledge Base. Retrieved Jan 1, 2012 from http://www.socialresearchmethods.net/kb/contents.php

PART II

WAYS TO APPROACH YOUR LITERATURE REVIEW

The literature on health and medicine is extensive and always expanding and changing.  The first time you go to the Online Library or to your local "physical" university-level library, you may not feel like you know how to proceed and where best to direct your energies.

There are two broad, general directions in which to go:

· The "top-down" search

· The "bottom-up" search

The "top-down" search begins with actual references from academic and scientific journals, in other words. The strategy assumes that you already have a high level of familiarity with the research area and the issues and knowledge that relate directly and indirectly to the area. As such, "top-down" searches tend to be less systematic than "bottom-up" searches, and for a novice researcher, the omitted source material can translate into important missing information. 

The "bottom-up" method is strongly suggested for those who are new to the process of investigating a research question. It is the more effective strategy when one is still trying to build a general knowledge base in the field of interest, and it is the one that I will recommend that you choose as a novice health sciences researcher. It will allow you to become more familiar with broad concepts that you are just now mastering in other courses, and how these essential concepts related to current issues and ongoing areas of debate and uncertainty. 

STEPS INVOLVED IN A BASIC "BOTTOM-UP" LITERATURE REVIEW

1. Try to list all possible terms that might be useful "index terms" in checking broad references and databases regarding your area of research interest. Use the Glossary in the Trident Online Library to assist you in covering all possible relevant terms.

2. Look up your topic and terms related to it in a good general reference.

3. Use the index terms and information from the general reference to do either a Computerized Literature Search or a Manual Search of the Literature (or both, if the resources are available.)

4. Skim the abstracts, tables of contents and outlines of the articles and books you initially select in order to determine which will be most directly relevant, informative and helpful to you in understanding your topic and refining your research question.

5. Obtain actual electronic or print copies of the references that appear to fit the above-stated criteria, select the best from among them, and begin outlining and note-taking.

It is not at all unusual for the process of reviewing the literature to cause you to consider changing your research question. In my experience, novice researchers usually start with an overly broad question, and end up refining, focusing and working to a more specific and testable research problem.  Feel free to contact me by e-mail or post your questions to the course discussion area. The latter action will allow your peers to learn from your questions and comments.

< top >

 Part I 

Measures of Central Tendency

When we are representing a sample quantitatively, we rely on expressions that will summarize the data for our audience. These expressions or measures are called typical values.

If we are trying to summarize and analyze strictly quantitative (interval or ratio data) we may start by creating a histogram and draw a curve over the top, connecting the center of each of the bars.  When we do this, we create a curve that represents the distribution of the scores. Sometimes our measure of central tendency describes the center of this curve; sometimes it describes the most likely score to occur; sometimes our measure is the balance point, a mathematically symmetrical place (though it may not appear to you that way when you look at the graph.)

THE BALANCE POINT - THE "AVERAGE" OR ARITHMETICAL MEAN

The most commonly used measure of central tendency (I am sure you are familiar with it) and the one that people usually "mean" when they say average is the MEAN (as it were.) I too will refer to the mean as the "average" for the remainder of this course.

You can calculate a mean easily by hand for very small samples, or with the aid of a calculator or spreadsheet for large distributions:

Here is a small sample that can be done by hand:

2, 3, 4, 5, 6

Step 1: Count the number of observations (scores) in the sample. We call this the "N" 

N = 5

Step 2: Add the scores together 

2 + 3 + 4 + 5 + 6 = 20

Step 3: Divide by "N" 

20/5 = 4

The mean of this small distribution is 4.

Do this one by hand or with a calculator =

2, 7, 3, 9, 3, 4, 7

N = ?

Sum of the scores = ?

Sum divided by "N" = ?

I will trust that you can figure it out, but feel free to email me for confirmation of the right answer if you like.

Generally, when we are working with large collections of quantitative data, we use the mean to describe the typical score. However, there are circumstances when we might want to consider a different summary measure. The mean is very (mathematically) sensitive to extreme scores or "outliers." The degree of distortion that a uncharacteristically high or low score can cause in an otherwise smooth or somewhat "normal" distribution can lead to a misinterpretation of the significance of the data. So it is good to be familiar with the other two measures of central tendency, even though you may rarely (if ever) use them.

THE CENTER OF THE CURVE - THE MEDIAN

If we line all our scores up in rank order, we can pick a middle value. 

When we have an odd number of scores, the middle value is right there:

1, 3, 5, 7, 9 - it's obvious that the middle value is 5.

5 is the median.

When we have an even number of scores, the middle value is "hidden":

1, 3, 5, 7, 9, 11 - it just takes a little simple middle school arithmetic to find it.

· Locate the middle two scores (in this case, that's 5 and 7)

· Add them together; 5 + 7 = 12

· Divide the resulting sum by 2; 12/2 = 6

6 is the MEDIAN of this small distribution.

When is it most appropriate to use the median?

When our distribution of scores is very uneven, especially with many scores bunched at one end and just a few stragglers or outliers at the other end. The outliers would likely distort our ability to see what a typical score in this distribution is. 

Let's say that we did a study of the effects of a particular herbal extract on the longevity of a sample of mice who were given the extract every day of their lives from birth. We had 8 mice and this is how long they lived in days:

499 502 510 523 524 530 539 815

If we merely calculated the arithmetic mean, the effect of the age of our one rodent Methuselah might cause that summary measure to suggest to our clinical audience that the herbal extract helped our little rodent subjects to actually live longer. Instead, it is more likely that his extended lifespan was a fluke, and the reporting the median would give a more honest and accurate accounting of the research results.

Just compare - 

The median age in this sample is 523 + 524 divided by 2 or 523.5 days old.

The mean age is 555.25

If the average age of the mice we have been raising in our lab is 525 days, the second result would be more misleading to anyone interested in the effects of the extract.

THE MOST COMMON SCORE - THE MODE

I mentioned in the earlier modules that we sometimes work with qualitative data called "nominal" data. An example of such data in the health sciences would be diagnostic categories:

Diabetes = 22

Asthma = 40

Psoriasis = 13

In this sample, asthma is the most common diagnosis that we encounter. We might use a numbering system of our own devising, or an established one like the ICD-9. But an any rate, whether we report it by name, by our own numerical label, or someone else's, if we report that asthma was the most typical diagnosis in this sample, we are reporting the MODE.

We may use the mode with plain old quantitative data too. Let's say we weighed our mice at death (the ones who took the herbal extract described above.) These are the weights we obtained (in ounces):

.5, .5, .5, .5. .47, .5, .51, .49

Although we could do a simple mean calculation, it would also be appropriate to use the mode and say that the average weight of one of our experimental mice was 1/2 ounce.

Please spend some time at the following websites to become more familiar with the measures of central tendency, how to use them, and how to interpret them most accurately:

DIG Stats. Retrieved September 1, 2012 from  http://www.cvgs.k12.va.us:81/DIGSTATS/

Using Data & Statistics (2006). Retrieved September 1, 2012 from  http://www.mathsisfun.com/data/index.html

PART II

Measures of dispersion

We won't get the whole picture accurate picture of what a distribution of scores looks like if we only notice where they are "clumping." We have to look at how they are spread out and the typical span in which we find most of our observations. This information is crucial because it tells us what the expected difference from the average is, and lets us know when we can be confident that the difference between a particular observation (or set of observations) and the average is unusual, possibly "significant."

The least complicated to calculate (but also least informative) measure of spread or dispersion is the RANGE. You can find the range by subtracting the lowest score in your data set from the highest.

2, 3, 4, 5, 6

What's the highest score?  6

What's the lowest score? 2

6 - 2 = 4

4 is the range for this data set.

You try it with this one:

6, 8, 3, 10, 12

What's the highest score? What's the lowest score? Subtract them and you get....(I think you can handle this.)

Obviously, we don't learn much about a distribution from this summary measure. We can't really see the difference between this -

1, 2, 4, 4, 4, 4, 4, 12

and this

1, 3, 9, 9, 9, 9, 9, 12

very different distributions that have exactly the same range!

There are variations on the range that are used and are easy to calculate and somewhat more informative. Called the Interquartile Range and Semi-interquartile Range, they rely on some easy to calculate measures of position to narrow down the scope so we can begin to get a picture of where most of the scores are falling and thus begin to blend the central tendency with the typical pattern of dispersion for a more accurate picture of the distribution we are examining. If you'd like to know more about these visit -

Basic Statistical Concepts (2011). Retrieved Jan 1, 2012 from  http://www.statsoft.com/textbook/elementary-concepts-in-statistics/

MOST VITAL TO YOUR UNDERSTANDING WHAT YOU READ IN A RESEARCH ARTICLE IS THE CONCEPT OF THE STANDARD DEVIATION.

The standard deviation defines an area below and above the mean about which it is expected that a majority of the scores will fall. Once we get outside that area, in either direction (+/-), we are approaching the zone in which it is rare to find an observation in our distribution, and quite possible that the reason we would find a score there is of interest to us in accomplishing our research goals.

( Click here for a brief PowerPoint demonstration of how to calculate the: STANDARD DEVIATION )

Please visit the following links to learn more about means, standard deviations and the STANDARD NORMAL DISTRIBUTION.

Summarizing & Presenting Data. Retrieved Jan 1, 2012 from  http://surfstat.anu.edu.au/surfstat-home/chap1ex.html

Normal Distribution. Retrieved Jan 1, 2012 from  http://davidmlane.com/hyperstat/A6929.html

The Normal Distribution. Retrieved Jan 1, 2012 from  http://www-stat.stanford.edu/~naras/jsm/NormalDensity/NormalDensity.html

Trochim, W.K. (2006). Research Methods Knowledge Base. Retrieved Jan 1, 2012 from http://www.socialresearchmethods.net/kb/contents.php