Brilliant Answer

profilehottboy561
SkillBuilderVisualDisplaysforCategorical.docx

Distinguish Categorical Variables

Words in orange represent glossary terms. You can locate the Glossary in Appendix 1.

Introduction

A researcher asks students how they perceived their body weight. They might respond with overweight, underweight, or just about right, in which case each student is a unit of analysis, the answer options represent categories of responses, each answer option is a value, and all of the students’ responses comprise a data set. One of the first steps in analyzing a sample of data such as this one is to examine what is referred to as the distribution of values for the data set’s variables.

Visual displays of data help researchers communicate the distribution and other key information (the story they are telling with their data) both effectively and efficiently, including for their own exploration. Put another way, visual displays of data allow researchers to quickly identify interesting aspects of their data (for example, are the study’s participants predominately satisfied with their body weight?), and to do so more efficiently than merely using words. Researchers take different approaches to visually displaying categorical and continuous variables. This Skill Builder focuses on visual displays for the former.

Identifying Categorical Variables

Categorical variables are those that have a small number of possible values. Usually, categorical variables involve nominal or ordinal levels of measurement. For example, political party affiliation is an example of a nominal, categorical variable. This variable places individuals into one of just a few categories (e.g. Democrat, Republican, or Independent). 

An example of an ordinal categorical variable is the highest grade completed, with categories of less than high school, high school diploma, and more than high school. Again, this variable has just a small number of possible values. You will typically use categorical methods of displaying data, such as a bar chart or a pie chart, when the number of categories is less than 10 or 12. 

If there are too many categories, the displays become messy and difficult to read. Also, keep in mind that the pie charts and bar charts are not typically used for non-categorical variables. An example of a non-categorical variable would-be students’ percentile ranking on a standardized math test; this variable has a large range of values and students aren’t simply placed into one of a limited number of categories.

Scenario: Body Image

Returning to the researcher asking students how they perceived their body weight, presume that he or she actually have a random sample of 1,200 U.S. college students who were asked the question of how they perceive their body weight as part of a larger survey. The following table shows part of the responses collected.

Student

Body Image

Student 25

Overweight

Student 26

About right

Student 27

Underweight

Student 28

About right

Student 29

About right

· bullet

What percentage of the sampled students fall into each category?

· bullet

How are students divided across the three body image categories? Are they equally divided? If not, do the percentages follow some other kind of pattern?

There is no way to answer these questions by looking at the raw data, which are in the form of a long list of 1,200 responses, and thus not very manageable. However, both of these questions can be easily answered once the researcher summarizes how often each of the categories occurs and looks at the frequency distribution of the different values for the variable Body Image.

Creating a table that presents the different values (categories) for the variable Body Image is the first step to take to summarize the distribution of a categorical variable. For example, the table below shows how many times the value “About right” occurs (count), and, more importantly, how often this value occurs (relative frequency) as a percentage. To convert the counts to percentages, divide the frequency (855) by the total number of observations (1200) to obtain the relative frequency, and multiply by 100 to convert to a percentage.

Drag and Drop Activity

Graphical Displays: Pie Charts and Bar Charts

Graphs

Now the researcher is ready to use a graphical display so that others can visualize the numerical summaries that were obtained. There are two simple graphical displays for visualizing the distribution of categorical data: The pie chart and the bar graph. While both the pie chart and the bar graph help researchers and those who use their results visualize the distribution of categorical variables, the pie chart (a circle divided into sections like a pie—shown below) emphasizes how the different categories relate to the whole, and the bar chart (side-by-side bars) emphasizes how the different categories compare with each other.

Take a close look at each chart below and then answer the questions that follow. In reviewing the two bar graphs,  note that “count” is often referred to as frequency, and the percentage is often called “relative frequency.”

Graphical Displays: Pictograms

The Use of Pictograms

A variation on the pie chart and bar graph that is commonly used in the media is the pictogram. The pictogram is an ideogram conveying significance through the pictorial resemblance to a physical object.  Here are two examples:

Visualization of the number of survivors and victims on the RMS Titanic by class and age/gender according to the British Board of Trade report on the disaster. Image credit: Cmglee [ CC BY-SA ]

In the example above, the right-hand table uses icons to represent the numbers seen in the left-hand table. Adapted from Pictogram by Wikipedia.

The following pictograph shows the amount of money spent on advertising in three magazines.

Effective and Ineffective Features of Visual Displays

Consider the ratio of advertising pictograph above. This graphic display is aimed at advertisers deciding where to spend their budgets and clearly suggests that Time magazine attracts by far the largest amount of advertising spending. Are the differences as dramatic as the pictogram suggests? Look carefully at the numbers above the pens, and you’ll find that advertisers spending in Time is only 1.64 ($4,433,879 / $2,698,386 = 1.64) times more than in Newsweek, and only 2.88 ($4,433,879 / $1,537,617 = 2.88) times more than in U.S. News. Just glancing at the pictogram, however, gives the impression that Time is much further ahead.

Why? Because the areas covered by the pen illustrations are not in proportion to the representative values. In order to magnify the pictures of the pens without distorting them, a designer increased both the height and width of each pen. By increasing both height and width, the area of Time's pen is 1.64 * 1.64 = 2.7 times larger than the Newsweek pen, and 2.88 * 2.88 = 8.3 times larger than the U.S. News pen. Viewers’ eyes capture the area of the pens rather than only the height, and so are misled to think that Time is a bigger winner than it actually is.