Problem Solving,

profilebenisd
chp_8_9_10_11.docx

8

Problem Solving

Human ability to solve novel problems greatly surpasses that of any other species,

and this ability depends on the advanced evolution of the prefrontal cortex in

humans. We have already noted the role of the prefrontal cortex in a number of

higher-level cognitive functions: language, imagery, and memory. It is generally

thought that the prefrontal cortex performs more than these specific functions, however,

and plays a major role in the overall organization of behavior. The regions of the

prefrontal cortex that we have discussed so far tend to be ventral (toward the bottom)

and posterior (toward the back), and many of these regions are left lateralized.

In contrast, dorsal (toward the top), anterior (toward the front), and right-hemisphere

prefrontal structures tend to be more involved in the organization of behavior. These

are the prefrontal regions that have expanded the most in the human brain.

Goel and Grafman (2000) describe a patient, PF, who suffered damage to his

right anterior prefrontal cortex as the result of a stroke. Like many patients with damage

to the prefrontal cortex, PF appears normal and even intelligent, and he scored in

the superior range on an intelligence test. In fact, he performed well on most tests,

although he did have difficulty with the Tower of Hanoi problem described later in this

chapter. Nonetheless, for all these surface appearances of normality, there were profound

intellectual deficits. He had been a successful architect before his stroke but

was forced to retire due to loss of the ability to design. He was able to get some work

as a draftsman. Goel and Grafman gave PF a problem that involved redesigning their

laboratory space. Although he was able to speak coherently about the problem, he

was unable to make any real progress on the solution. A comparably trained architect

without brain damage achieved a good solution in a couple of hours. It seems that the

stroke affected only PF’s most highly developed intellectual abilities.

This chapter and Chapter 9 will look at what we know about human problem

solving. In this chapter, we will answer the following questions: • What does it mean to characterize human problem solving as a search of a

problem space? • How do humans learn methods, called operators, for searching the problem

space?

209

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 209

210 | Problem Solving

• How do humans select among different operators for searching a problem

space? • How can past experience affect the availability of different operators and the

success of problem-solving efforts?

The Nature of Problem Solving

A Comparative Perspective on Problem Solving

Figure 8.1 shows the relative sizes of the prefrontal cortex in various mammals

and illustrates the dramatic increase in humans. This increase supports the

advanced problem solving that only humans are capable of. Nonetheless, one

can find instances of interesting problem solving in other species, particularly

in the higher apes such as chimpanzees. The study of problem solving in other

species offers perspective on our own abilities. Köhler (1927) performed some

of the classic studies on chimpanzee problem solving. Köhler was a famous

German gestalt psychologist who came to America in the 1930s. During World

War I, he found himself trapped on Tenerife in the Canary Islands. On the

island, he found a colony of captive chimpanzees, which he studied, taking

particular interest in the problem-solving behavior of the animals. His best

participant was a chimpanzee named Sultan. One problem posed to Sultan was

FIGURE 8.1 The relative proportions of the frontal lobe given over to the prefrontal cortex in

six mammals. Note that these brains are not drawn to scale and that the human brain is really

much larger in absolute size. (After Fuster, 1989. Adapted by permission of the publisher. © 1989 by Raven Press.)

Squirrel monkey Cat Rhesus monkey

Dog Chimpanzee Human

Brain Structures

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 210

The Nature of Problem Solving | 211

to get some bananas that were outside his cage. Sultan had no difficulty when

he was given a stick that could reach the bananas; he simply used the stick to

pull the bananas into the cage. The problem became harder when Sultan was

provided with two poles, neither of which could reach the food. After unsuccessfully

trying to use the poles to get to the food, the frustrated ape sulked in

his cage. Suddenly, he went over to the poles and put one inside the other, creating

a pole long enough to reach the bananas (Figure 8.2). Clearly, Sultan had

creatively solved the problem.

What are the essential features that qualify this episode as an instance of

problem solving? There seem to be three:

1. Goal directedness. The behavior is clearly organized toward a goal—in

this case, getting the food.

2. Subgoal decomposition. If Sultan could have obtained the food simply

by reaching for it, the behavior would have been problem solving, but

only in the most trivial sense. The essence of the problem solution is that

the ape had to decompose the original goal into subtasks, or subgoals,

such as getting the poles and putting them together.

3. Operator application. Decomposing the overall goal into subgoals is

useful because the ape knows operators that can help him achieve these

subgoals. The term operator refers to an action that will transform the

problem state into another problem state. The solution of the overall

problem is a sequence of these known operators.

Problem solving is goal-directed behavior that often involves setting subgoals

to enable the application of operators.

FIGURE 8.2 Köhler’s ape,

Sultan, solved the two-stick

problem by joining two short

sticks to form a pole long

enough to reach the food

outside his cage. (From Köhler, 1956.

Reprinted by permission of the publisher.

© 1956 by Routledge & Kegan Paul.)

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 211

The Problem-Solving Process: Problem Space and Search

Often, problem solving is described in terms of searching a problem space,

which consists of various states of the problem. A state is a representation of the

problem in some degree of solution. The initial situation of the problem is

referred to as the start state; the situations on the way to the goal, as intermediate

states; and the goal, as the goal state. Beginning from the start state, there are

many ways the problem solver can choose to change the state. Sultan could reach

for a stick, stand on his head, sulk, or try other approaches. Suppose he reaches

for a stick. Now he has entered a new state. He can transform it into another

state—for example, by letting go of the stick (thereby returning to the earlier

state), reaching for the food with the stick, throwing the stick at the food,

or reaching for the other stick. Suppose he reaches for the other stick. Again, he

has created a new state. From this state, Sultan can choose to try, say, walking on

the sticks, putting them together, or eating them. Suppose he chooses to put the

sticks together. He can then choose to reach for the food, throw the sticks away,

or separate them. If he reaches for the food, he will achieve the goal state.

The various states that the problem solver can achieve define a problem

space, also called a state space. Problem-solving operators can be thought of as

ways to change one state in the problem space into another. The challenge is to

find some possible sequence of operators in the problem space that leads from

the start state to the goal state.We can think of the problem space as a maze of

states and of the operators as paths for moving among them. In this model, the

solution to a problem is achieved through search; that is, the problem solver

must find an appropriate path through a maze of states. This conception of

problem solving as a search through a state space was developed by Allen

Newell and Herbert Simon, who were dominant figures in cognitive science

throughout their careers, and it has become the major problem-solving approach,

in both cognitive psychology and AI.

A problem space characterization consists of a set of states and operators

for moving among the states. A good example of problem-space characterization

is the eight-tile puzzle, which consists of eight numbered, movable tiles set

in a 3 _ 3 frame. One cell of the frame is always empty, making it possible to

move an adjacent tile into the empty cell and thereby to “move” the empty cell

as well. The goal is to achieve a particular configuration of tiles, starting from

a different configuration. For instance, a problem might be to transform

212 | Problem Solving

The possible states of this problem are represented as configurations of tiles in

the eight-tile puzzle. So, the first configuration shown is the start state, and the second

is the goal state. The operators that change the states are movements of tiles

into empty spaces. Figure 8.3 reproduces an attempt of mine to solve this problem.

into

2 1 6

4 8

7 5 3

1 2 3

8 4

7 6 5

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 212

My solution involved 26 moves, each move being an operator that changed the

state of the problem. This sequence of operators is considerably longer than necessary.

Try to find a shorter sequence of moves. (The shortest sequence possible is

given in the appendix at the end of the chapter, in Figure A8.1.)

Often, discussions of problem solving involve the use of search graphs or

search trees. Figure 8.4 gives a partial search tree for the following, simpler

eight-tile problem:

The Nature of Problem Solving | 213

(a) (b) (c) (d) (e) (f) (g)

(o) (p) (q) (r) (s) (t) (u)

(n) (m) (l) (k) ( j) (i) (h)

2 1 6

4 8

7 5 3

2 1 6

4 8

7 5 3

(w) (v)

2 6 4

7 5

8 1 3

(x)

2 4

7 6 5

8 1 3

(y)

2 4

7 6 5

8

1 3

(z)

2 4

7 6 5

8

1 3

Goal state

2

7 6 5

8 4

1 3

8 4

6

2 7 5

1 3

8

6

4

2 7 5

1 3

8 4

2 7 5

1 3

6 8 4

2 7 5

6

1 3 8

4

2 7 5

6

1 3

2

7 5

6 4

8 1 3

2 4

8

1 6

7 5

3

2

8

1 6

4

7 5

3

2

1 6

8

4

7 5

3

2 8

7 5

1 4 6

3 2

7 8 5

1 4 6

3 2 4

7 8 5

1 6

3

1 6

2 4

7 8 5

3

2

1 6

4 8

7 5 3

1 6

2 4 8

7 5 3

1 6

2 8

4

7 5 3

2 8

1 4 6

7 5 3

8 4

7 5

2

1 6 3

2

1

4

7 5

6 3

8

4

5

3

7

8 1

2

6

into

2

1

8

4

7 5

3 1 2 3

8 4

7 6 5

6

FIGURE 8.3 The author’s sequence of moves for solving an eight-tile puzzle.

Figure 8.4 is like an upside-down tree with a single trunk and branches leading

out from it. This tree begins with the start state and represents all states reachable

from this state, then all states reachable from those states, and so on. Any

path through such a tree represents a possible sequence of moves that a problem

solver might make. By generating a complete tree, we can also find the shortest

sequence of operators between the start state and the goal state. Figure 8.4 illustrates

some of the problem space. In discussions of such examples, often only a

path through the problem space that leads to the solution is presented (for

instance, see Figure 8.3). Figure 8.4 gives a better idea of the size of the problem

space of possible moves for this kind of problem.

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 213

This search space terminology describes possible steps that the problem

solver might take. It leaves two important questions that we need to answer

before we can explain the behavior of a particular problem solver. First, what

determines the operators available to the problem solver? Second, how does the

problem solver select a particular operator when there are several available? An

answer to the first question determines the search space in which the problem

solver is working. An answer to the second question determines which path the

problem solver takes. We will discuss these questions in the next two sections,

focusing first on the origins of the problem-solving operators and then on the

issue of operator selection.

214 | Problem Solving

2

1 6

8

4

7 5

3

2

1

6

8

4

7 5

3

2

1

6

8

4

7 5

3

2

1

6

8

4

7 5

3

2

1

8 6

4

7 5

3

2

1

6 8 4

7 5

3 2

1

6

8

4

7 5

3 2

1

6

8

7 4

5

3

2

1

6

8

4

7 5

3

2

1

6

8

4

7 5

3

2

1

6 8 4

7 5

3 2

1

6 8 4

7 5

3 2

1

6

8

4

7 5

3

2

1

6

8

4

7

5

3 2

1

6

8

7 4

5

3 2

1

6

8

7 4

5

3

FIGURE 8.4 Part of the search tree, five moves deep, for an eight-tile problem. (After Nilsson, 1971.

Adapted by permission of the publisher. © 1971 by McGraw-Hill.)

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 214

Problem-solving operators generate a space of possible states through which

the problem solver must search to find a path to the goal.

Problem-Solving Operators

Acquisition of Operators

There are at least three ways to acquire new problem-solving operators.We can

acquire new operators by discovery, by being told about them, or by observing

someone else use them.

Problem-Solving Operators | 215

2

1 6

8

4

7 5

2 3

1

6

8

4

7 5

3

2 1

6

8

4

7 5

3 2

1

6

8

7 4

5

3 2

1

6

8 4

7 5

3 2

1

6

8 4

7 5

3

2 1

6

8

4

7 5

3 2

1

6

8

7 4

5

3 1 2

6

8 4

7 5

3

2

1

6

8

4

7 5

3 2

1

6

8 4

7 5

3 2

1

6

8

4

7 5

3

2

1 6

8

4

7 5

3

2 1

6

8

4

7 5

3

2

1

6

8

4

7 5

3 2

6 1

8

7 4

5

3 2

1

6

8

7 4

5

3 1

7 5

2

8 4

6

3 1 2

6

7 8 4

5

3

Goal state

Start state

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 215

Discovery. We might find that a new service station has opened nearby and

so learn by discovery a new operator for repairing our car. Children might

discover that their parents are particularly susceptible to temper tantrums and

so learn a new way to get what they want. We might discover how a new

microwave oven works by playing with it and so learn a new way to prepare

food. Or a scientist might discover a new drug that kills bacteria and so invent a

new way of combating infections. Each of these examples involves a variety of

reasoning processes. These processes will be the topic of Chapter 10.

Although discovery can involve complex reasoning in humans, it is interesting

that it is the only method that most other creatures have to learn new operators,

and they certainly do not engage in complex reasoning. In a famous study

reported in 1898, Thorndike placed cats in “puzzle boxes.” The boxes could be

opened by various nonobvious means. For instance, in one box, if the cat hit

a loop of wire, the door would fall open. The cats, who were hungry, were rewarded

with food when they got out. Initially, a cat would move about randomly,

clawing at the box and behaving ineffectively in other ways until it

happened to hit the unlatching device. After repeated trials in the same puzzle

box, the cats eventually arrived at a point where they would immediately hit

the unlatching device and get out. A controversy exists to this day over whether

the cats ever really “understood” the new operator they had acquired or just

gradually formed a mindless association between being in the box and hitting

the unlatching device. More recently it has been argued that it need not be an

either–or situation. Daw, N.D., Niv, Y., and Dayan, P. (2005) review evidence that

there are two bases for learning such operators from experience—one involves

the basal ganglia (see Figure 1.8), where simple associations are gradually reinforced,

whereas the other involves the prefrontal cortex and a mental model

of how these operators work. It is reasonable to suppose that the second system

becomes more important in mammals with larger prefrontal cortices.

Learning by Being Told or by Example. We can acquire new operators by

being told about them or by observing someone else use them. These are examples

of social learning. The first method is a uniquely human accomplishment because

it depends on language. The second is a capacity thought to be common in primates:

“Monkey see, monkey do.” As we will see, however, the capacity of nonhuman

primates for learning by imitation has often been overestimated.

It might seem that the most efficient way to learn new problem-solving

operators would be simply to be told about them, but seeing an example is often

at least as effective as being told what to do. Table 8.1 shows two forms of

instruction about an algebraic concept, called a pyramid expression, which is

novel to most undergraduates. Students either study part (a), which gives a

semiformal specification of what a pyramid expression is, or they study part (b),

which gives the single example of a pyramid expression. After reading one

instruction or the other, they are asked to evaluate pyramid expressions like

10$2

Which form of instruction do you think would be most useful? Carnegie

Mellon undergraduates show comparable levels of learning from the single

216 | Problem Solving

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 216

example in part (b) to what they learn from the rigorous specification in part (a).

Sometimes, examples can be the superior means of instruction. For instance,

Reed and Bolstad (1991) had participants learn to solve problems such as the

following:

An expert can complete a technical task in five hours, but a novice requires

seven hours to do the same task. When they work together, the novice works

two hours more than the expert. How long does the expert work? (p. 765)

Participants received instruction in how to use the following equation to solve

the problem:

rate1 _ time1 _ rate2 _ time2 _ tasks

The participants needed to acquire problem-solving operators for assigning

values to the terms in this equation. The participants either received abstract

instruction about how to make these assignments or saw a simple example of

how the assignments were made. There was also a condition in which participants

saw both the abstract instruction and the example. Participants given the

abstract instruction were able to solve only 13% of a set of later problems; participants

given an example solved 28% of the problems; and participants given

both instruction and an example were able to solve 40%.

Why would giving examples be better for learning problem-solving operators

than telling someone what to do directly? The problem with direct instruction

is that it can often be difficult to understand what such quantities as rate1

refer to. This information can be clearer in the context of an example. On the

other hand, it can be difficult to see how to extend an example solution from

one problem to another problem. Thus, experiments like Reed and Bolstad’s

indicate that the best learning occurs when participants have access to both

methods. Similar results have been obtained by Fong, Krantz, and Nisbett

(1986) in the domain of statistics and by Cheng, Holyoak, Nisbett, and Oliver

(1986) in the domain of logic.

Problem-Solving Operators | 217

TABLE 8.1

Instruction for Pyramid Problems

(a) Direct Specification

N$M is a pyramid expression for designating repeated addition where each

term in the sum is one less than the previous.

N, the base, is first term in the sum.

M, the height, is number of terms you add to the base.

(b) Just an Example

7$3 is an example of a pyramid expression.

7$3 _ 7 _ 6 _ 5 _ 4 _ 22

7 is the 3 is the

base height

M

a

a

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 217

Problem-solving operators can be acquired by discovery, by modeling example

problem solutions, or by direct instruction.

Analogy and Imitation

Analogy is the process by which a problem solver extracts the operators used

to solve one problem and maps them onto a solution for another problem.

Sometimes, the analogy process can be straightforward. For instance, a student

may take the structure of an example worked out in a section of a mathematics

text and map it into the solution for a problem in the exercises at the end

of the section. At other times, the transformations can be more complex.

Rutherford, for example, used the solar system as a model for the structure of

the atom, in which electrons revolve around the nucleus of the atom in the

same way as the planets revolve around the sun (Koestler, 1964; Gentner,

1983—see Table 8.2). Although this is a particularly famous example of an

analogy, scientists and engineers use such analogies, if often more mundane,

with great frequency. For instance, Christensen and Schunn (2007) found engineers

making 102 analogies in 9 hours of problem solving (see also Dunbar &

Blanchette, 2001).

An example of the power of analogy in problem solving is provided in an

experiment of Gick and Holyoak (1980). They presented their participants with

the following problem, which is adapted from Duncker (1945):

Suppose you are a doctor faced with a patient who has a malignant tumor in

his stomach. It is impossible to operate on the patient, but unless the tumor is

destroyed, the patient will die. There is a kind of ray that can be used to

destroy the tumor. If the rays reach the tumor all at once at a sufficiently high

intensity, the tumor will be destroyed. Unfortunately, at this intensity the

healthy tissue that the rays pass through on the way to the tumor will also be

destroyed. At lower intensities the rays are harmless to healthy tissue, but they

will not affect the tumor either. What type of procedure might be used to

destroy the tumor with the rays, and at the same time avoid destroying the

healthy tissue? (pp. 307–308)

218 | Problem Solving

TABLE 8.2

The Solar System–Atom Analogy

Base Domain: Solar System Target Domain: Atom

The sun attracts the planets. The nucleus attracts the electrons.

The sun is larger than the planets. The nucleus is larger than the electrons.

The planets revolve around the sun. The electrons revolve around the nucleus.

The planets revolve around the sun The electrons revolve around the nucleus

because of the attraction and because of the attraction and weight

weight difference. difference.

The planet Earth has life on it. No transfer.

After Gentner (1983). Adapted by permission of the publisher. © 1983 by LEA, Ltd.

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 218

This is a very difficult problem, and few people are able to solve it. However,

Gick and Holyoak presented their participants with the following story:

A small country was ruled from a strong fortress by a dictator. The fortress was

situated in the middle of the country, surrounded by farms and villages. Many

roads led to the fortress through the countryside. A rebel general vowed to capture

the fortress. The general knew that an attack by his entire army would

capture the fortress. He gathered his army at the head of one of the roads,

ready to launch a full-scale direct attack. However, the general then learned

that the dictator had planted mines on each of the roads. The mines were set so

that small bodies of men could pass over them safely, since the dictator needed

to move his troops and workers to and from the fortress. However, any large

force would detonate the mines. Not only would this blow up the road, but it

would also destroy many neighboring villages. It therefore seemed impossible

to capture the fortress. However, the general devised a simple plan. He divided

his army into small groups and dispatched each group to the head of a different

road.When all was ready he gave the signal and each group marched down

a different road. Each group continued down its road to the fortress so that the

entire army arrived together at the fortress at the same time. In this way, the

general captured the fortress and overthrew the dictator. (p. 351)

Told to use this story as the model for a solution, most participants were able

to develop an analogous operation to solve the tumor problem.

An interesting example of a solution by analogy that did not quite work is a

geometry problem encountered by one student. Figure 8.5a illustrates the steps

of a solution that the text gave as an example, and Figure 8.5b illustrates the

student’s attempts to use that example proof to guide his solution to a homework

problem. In Figure 8.5a, two segments of a line are given as equal length,

and the goal is to prove that two larger segments have equal length. In Figure 8.5b,

the student is given two line segments with AB longer than CD, and his task is to

prove the same inequality for two larger segments, AC and BD.

Our participant noted the obvious similarity between the two problems and

proceeded to develop the apparent analogy. He thought he could

simply substitute points on one line for points on another, and

inequality for equality. That is, he tried to substitute A for R, B

for O, C for N, D for Y, and _ for _.With these substitutions, he

got the first line correct: Analogous to RO _ NY, he wrote AB _

CD. Then he had to write something analogous to ON _ ON, so

he wrote BC _ BC! This example illustrates how analogy can be

used to create operators for problem solving and also shows that

it requires a little sophistication to use analogy correctly.

Another difficulty with analogy is finding the appropriate

examples from which to analogize operators. Often, participants

do not notice when an analogy is possible. Gick and

Holyoak (1980) did an experiment in which they read participants

the story about the general and the dictator and then gave

them Duncker’s (1945) ray problem (both shown earlier in this

section). Very few participants spontaneously noticed the relevance

of the first story to solving the second. To achieve success,

Problem-Solving Operators | 219

R

(a)

O

N

Y

Given: RO = NY, RONY

Prove: RN = OY

RO = NY

ON = ON

RO + ON = ON + NY

RONY

RO + NY = RN

ON + NY = OY

RN = OY

A

(b)

B

C

D

AB > CD

BC > BC

!!!

Given: AB > CD, ABCD

Prove: AC > BD

FIGURE 8.5 (a) A worked-out

proof problem given in a geometry

text. (b) One student’s attempt

to use the structure of this

problem’s solution to guide his

solution of a similar problem.

This example illustrates how

analogy can be used (and

misused) for problem solving.

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 219

participants had to be explicitly told to use the general and dictator story as an

analogy for solving the ray problem.

When participants do spontaneously use previous examples to solve a problem,

they are often guided by superficial similarities in their choice of examples.

For instance, B. H. Ross (1984, 1987) taught participants several methods for

solving probability problems. These methods were taught by reference to specific

examples, such as finding the probability that a pair of tossed dice will sum

to 7. Participants were then tested with new problems that were superficially

similar to prior examples. The similarity was superficial because both the

example and the problem involved the same content (e.g., dice) but not necessarily

the same principle of probability. Participants tried to solve the new

problem by using the operators illustrated in the superficially similar prior

example. When that example illustrated the same principle as required in the

current problem, participants were able to solve the problem.When it did not,

they were unable to solve the current problem. Reed (1987) has found similar

results with algebra story problems.

In solving school problems, students use proximity as a cue to determine

which examples to use in analogy. For instance, a student working on physics

problems at the end of a chapter expects that problems solved as examples in

the chapter will use the same methods and so tries to solve the problems by

analogy to these examples (Chi, Bassok, Lewis, Riemann, & Glaser, 1989).

Analogy involves noticing that a past problem solution is relevant and then

mapping the elements from that solution to produce an operator for the

current problem.

Analogy and Imitation from an Evolutionary

and Brain Perspective

It has been argued that analogical reasoning is a hallmark of human cognition

(Halford, 1992). The capacity to solve analogical problems is almost uniquely

found in humans. There is some evidence for it in chimpanzees (Oden,

Thompson, & Premack, 2001), although lower primates such as monkeys seem

totally incapable of such tasks. For instance, Premack (1976) found that Sarah,

a chimpanzee used in studies of language (see Chapter 11), was able to solve

analogies such as the following: Key is to a padlock as what is to a tin can? The

answer: can opener. In more careful study of Sarah’s abilities, however, Oden et

al. found that although Sarah could solve these problems more often than

chance, she was much more prone to error than human participants.

Recent brain-imaging studies have looked at the cortical regions that are

activated in analogical reasoning. Figure 8.6 shows examples of the stimuli used

in a study by Christoff et al. (2001), adapted from the Raven’s Progressive

Matrices test, which is a standard test of intelligence. Only the problems in

Figure 8.6c, which require that the solver coordinate two dimensions, could be

said to tap true analogical reasoning. There is evidence that children under age 5

(in whom the frontal cortex has not yet matured), nonhuman primates, and

220 | Problem Solving

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 220

patients with frontal damage all have special difficulty with

problems like those shown in Figure 8.6c and often just cannot

solve them. Christoff et al. were interested in discovering

which brain regions would be activated when participants

were solving these problems. Consistent with the trends we

noted in the introduction to this chapter, they found that the

right anterior prefrontal cortex was activated only when participants

were solving these 2-D problems.

Examples like those shown in Figure 8.6 are cases in which

analogical reasoning is used for purposes other than acquiring

new problem-solving operators. From the perspective of this

chapter, however, the real importance of analogy is that it can

be used to acquire new problem-solving operators. We noted

earlier that people often learn more from studying an example

than from reading abstract instructions. Humans have a special

ability to mimic the problem solutions of others. When

we ask someone how to use a new device, that person tends to

show us how, not to tell us how. Despite the proverb “Monkey

see, monkey do,” even the higher apes are quite poor at imitation

(Tomasello & Call, 1997). Thus, it seems that one of

the things that make humans such effective problem solvers

is that we have special abilities to acquire new problem-solving

operators by analogical reasoning.

Analogical problem solving appears to be a capability nearly unique to

humans and to depend on the advanced development of the prefrontal

cortex.

Operator Selection

As noted earlier, in any particular state, multiple problem-solving operators can

be applicable, and a critical task is to select the one to apply. In principle, a

problem solver may select operators in many ways, and the field of AI has

succeeded in enumerating various powerful techniques. However, it seems that

most methods are not particularly natural as human problem-solving approaches.

Here we will review three criteria that humans use to select operators.

Backup avoidance biases the problem solver against any operator that undoes

the effect of the previous operators. For instance, in the eight-tile puzzle,

people show great reluctance to take back a step even if this might be necessary

to solve the problem. However, backup avoidance by itself provides no basis for

choosing among the remaining operators.

Humans tend to select the nonrepeating operator that most reduces the difference

between the current state and the goal. Difference reduction is a very

general principle and describes the behavior of many creatures. For instance,

Köhler (1927) described how a chicken will move directly toward desired food

Operator Selection | 221

(a)

1 2

3 4

1 2

3 4

1 2

3 4

(b)

(c)

FIGURE 8.6 Examples of stimuli

used by Christoff et al. to study

which brain regions would be

activated when participants

attempted to solve three

different types of analogy

problem: (a) 0-dimensional;

(b) 1-dimensional; and

(c) 2-dimensional. The task

in each case was to infer the

missing figure and select it

from among the four alternative

choices. (After Christoff et al., 2001.

Adapted by permission of the publisher.

© 2001 by Neuroimage.)

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 221

and will not go around a fence that is blocking it. The poor creature is effectively

paralyzed, being unable to move forward and unwilling to back up and

undo its approach to the fence. It does not seem to have any principles for selection

of operators other than difference reduction and backup avoidance.

This leaves it without a solution to the problem.

On the other hand, the chimpanzee Sultan (see Figure 8.2) did not just claw

at his cage trying to get the bananas. He sought to create a new tool to enable

him to obtain the food. In effect, his new goal became the creation of a new

means for achieving the old goal. Means-ends analysis is the term used to

describe the creation of a new goal (end) to enable an operator (means) to

apply. By using means-ends analysis, humans and other higher primates can be

more resourceful in achieving a goal than they could be if they used only difference

reduction. In the next sections, we will discuss the roles of both difference

reduction and means-ends analysis in operator selection.

Humans use backup avoidance, difference reduction, and means-ends

analysis to guide their selection of operators.

The Difference-Reduction Method

A common method of problem solving, particularly in unfamiliar domains, is

to try to reduce the difference between the current state and the goal state. For

instance, consider my solution to the eight-tile puzzle in Figure 8.3. There were

four options possible for the first move. One possible operator was to move the

1 tile into the empty square, another was to move the 8, a third was to move the 5,

and the fourth was to move the 4. I chose the last operator. Why? Because it

seemed to get me closer to my end goal. I was moving the 4 tile closer to its

final destination. Human problem solvers are often strongly governed by difference

reduction or, conversely, by similarity increase. That is, they choose operators

that transform the current state into a new state that reduces differences

and resembles the goal state more closely than the current state. Difference

reduction is sometimes called hill climbing. If we imagine the goal as the highest

point of land, one approach to reaching it is always to take steps that go up.

By reducing the difference between the goal and the current state, the problem

solver is taking a step “higher” toward the goal. Hill climbing has a potential

flaw, however: By following it, we might reach the top of some hill that is lower

than the highest point of land that is the goal. Thus, difference reduction is not

guaranteed to work. It is myopic in that it considers only whether the next step

is an improvement and not whether the larger plan will work. Means-ends

analysis, which we will discuss later, is an attempt to introduce a more global

perspective into problem solving.

One way problem solvers improve operator selection is by using more

sophisticated measures of similarity.My first move was intended simply to get a

tile closer to its final destination. After working with many tile problems, we

begin to notice the importance of sequence—that is, whether noncentral tiles are

followed by their appropriate successors. For instance, in state (o) of Figure 8.3,

the 3 and 4 tiles are in sequence because they are followed by their successors 4

222 | Problem Solving

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 222

and 5, but the 5 is not in sequence because it is followed by 7 rather than 6.

Trying first to move tiles into sequence proves to be more important than trying

to move them to their final destinations right away. Thus, using sequence as

a measure of increasing similarity leads to more effective problem solving based

on difference reduction (see Nilsson, 1971, for further discussion).

The difference-reduction technique relies on evaluations of the similarity

between the current state and the goal state. Although difference reduction

works more often than not, it can also lead the problem solver astray. In some

problem-solving situations, a correct solution involves going against the grain

of similarity. A good example is called the hobbits and orcs problem:

On one side of a river are three hobbits and three orcs. They have a boat on their

side that is capable of carrying two creatures at a time across the river. The goal is

to transport all six creatures across to the other side of the river. At no point on

either side of the river can orcs outnumber hobbits (or the orcs would eat the

outnumbered hobbits). The problem, then, is to find a method of transporting

all six creatures across the river without the hobbits ever being outnumbered.

Stop reading and try to solve this problem. Figure 8.7 shows a correct sequence

of moves. Illustrated are the locations of hobbits (H), orcs (O), and the boat

(b). The boat, the three hobbits, and the three orcs all start on one side of the

river. This condition is represented in state 1 by the fact that all are above the

line. Then a hobbit, an orc, and the boat proceed to the other side of the river.

The outcome of this action is represented in state 2 by placement of the boat,

the hobbit, and the orc below the line. In state 3, one hobbit has taken the boat

back, and the diagram continues in the same way. Each state in the figure represents

another configuration of hobbits, orcs, and boat. Participants have a particular

problem with the transition from state 6 to state 7. In a study by Jeffries,

Polson, Razran, and Atwood (1977), about a third of all participants chose to

back up to a previous state 5 rather than moving on to state 7 (see also Greeno,

1974). One reason for this difficulty is that the action involves moving two

creatures back to the wrong side of the river. The move seems to be away from a

solution. At this point, participants will go back to state 5, even though this undoes

their last move. They would rather undo a move than take a step that

moves them to a state that appears further from the goal.

Atwood and Polson (1976) provide another experimental demonstration of

participants’ reliance on similarity and how that reliance can sometimes be

harmful and sometimes beneficial. Participants were given the following water

jug problem:

You have three jugs, which we will call A, B, and C. Jug A can hold exactly

8 cups of water, B can hold exactly 5 cups, and C can hold exactly 3 cups. Jug

A is filled to capacity with 8 cups of water. B and C are empty.We want you to

find a way of dividing the contents of A equally between A and B so that both

have exactly 4 cups. You are allowed to pour water from jug to jug.

Figure 8.8 shows two paths for solving this problem. At the top of the illustration,

all the water is in jug A—represented by A(8); there is no water in jugs B or C

represented by B(0) C(0). The two possible actions are to pour A into C, in which

case we get A(5) B(0) C(3), or to pour A into B, in which case we get A(3) B(5)

Operator Selection | 223

b H H H O O O

b

b

H

H

H O

O

O

H H H O

O

O

b

H H H

O O O

b H H H

O O

O

b

H

H H O

O

O

b H

H

H

O

O O

b H H

O O

O

b

H H

H

H

O O O

b H H H

O

O O

b

H H H

O O

O

b O O O H H H

(1)

(2)

(3)

(4)

(5)

(6)

(7)

(8)

(9)

(10)

(11)

(12)

FIGURE 8.7 A diagram of the

successive states in a solution

to the hobbits and orcs

problem. H _ hobbits,

O _ orcs, b _ boat.

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 223

C(0). From these two states, more moves can be

made. Numerous other sequences of moves are

possible besides the two paths illustrated, but

these are the two shortest sequences to the goal.

Atwood and Polson used the representation in

Figure 8.8 to analyze participants’ behavior. For

instance, they asked which move participants

would prefer to make at the start state 1. That is,

would they prefer to pour jug A into C and get

state 2, or jug A into B and get state 9? The answer

is that participants preferred the latter move.

More than twice as many participants moved to

state 9 as moved to state 2. Note that state 9 is

quite similar to the goal. The goal is to have 4

cups in both A and B, and state 9 has 3 cups in A

and 5 cups in B. In contrast, state 2 has no cups of

water in B. Throughout the experiment, Atwood

and Polson found a strong tendency for participants

to move to states that were similar to the

goal state. Usually, similarity was a good heuristic,

but there are critical cases where similarity is misleading.

For instance, the transitions from state 5

to state 6 and from state 11 to state 12 both lead

to significant decreases in similarity to the goal.

However, both transitions are critical to their solution

paths. Atwood and Polson found that more

than 50% of the time, participants deviated from

the correct sequence of moves at these critical

points. They instead chose some move that seemed closer to the goal but actually

took them away from the solution.1

It is worth noting that people do not get stuck in suboptimal states only

while solving puzzles. Hill climbing can get us stuck when making serious life

choices. A classic example is someone trapped in a suboptimal job because he

or she is unwilling to get the education needed for a better job. The person is

unwilling to endure the temporary deviation from the goal (of earning as much

as possible) to get the skills to earn an even higher salary.

People experience difficulty in solving a problem at points where the correct

solution involves increasing the differences between the current state and the

goal state.

Means-Ends Analysis

Means-ends analysis is a more sophisticated method of operator selection. This

method has been extensively studied by Newell and Simon, who used it in a

computer simulation program (called the General Problem Solver—GPS) that

224 | Problem Solving

A(5)

A(5)

A(2)

A(2)

A(7)

A(7)

A(4)

A(4) (4)

B(5)

B(2)

B(2)

B(0)

B(0) C(0)

C(0)

C(2)

C(2)

C(3)

C(0)

C (3) A(3)

A(6)

A(6)

B(5)

B(4)

A(1)

A(1)

B(3) A(3) C(3)

B(5)

B(0)

B

B(3) C(3)

C(1)

C (0)

C(3)

C(0)

C(1)

B(1)

B(1)

(1) A(8)

(2)

(3)

(4)

(5)

(6)

(7)

(8)

(15)

(13)

(14)

(12)

(11)

(10)

(9)

B(0) C(0)

A C

C B

A C

C B

B A

C B

A C

C B

C A

B C

A B

B C

C A

B C

A B

FIGURE 8.8 Two paths of

solution for the water jug

problem posed in Atwood and

Polson (1976). Each state is

represented in terms of the

contents of the three jugs; for

example, in state 1, A(8) B(0)

C (0). The transitions between

states (e.g., A →C) are labeled

in terms of which jug is poured

into which other jug.

1 For instance, moving back to state 9 from either state 5 or state 11.

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 224

models human problem solving. The following is their description of meansends

analysis.

Means-ends analysis is typified by the following kind of commonsense

argument:

I want to take my son to nursery school.What’s the difference between what I

have and what I want? One of distance. What changes distance? My automobile.

My automobile won’t work. What is needed to make it work? A new

battery.What has new batteries? An auto repair shop. I want the repair shop to

put in a new battery; but the shop doesn’t know I need one.What is the difficulty?

One of communication. What allows communication? A telephone . . .

and so on.

This kind of analysis—classifying things in terms of the functions they serve

and oscillating among ends, functions required, and means that perform

them—forms the basic system of GPS. (Newell & Simon, 1972, p. 416)

Means-ends analysis can be viewed as a more sophisticated version of difference

reduction. Like difference reduction, it tries to eliminate the differences

between the current state and the goal state. For instance, in this example, it

tried to reduce the distance between the home and the nursery school. Meansends

analysis will also identify the biggest difference first and try to eliminate it.

Thus, in this example, the focus is on difference in the general location of home

and nursery school. The difference between where the car will be parked and

the classroom has not been considered.

Means-ends analysis offers a major advance over difference reduction because

it will not abandon an operator if it cannot be applied immediately. If the car did

not work, for example, difference reduction would have one start walking to the

nursery school. The essential feature of means-ends analysis is that it focuses on

enabling blocked operators. The means temporarily becomes the end. In effect,

the problem solver deliberately ignores the real goal and focuses on the goal of

enabling the means. In the example we have been discussing, the problem solver

set a subgoal of repairing the automobile, which was the means of achieving the

original goal of getting the child to nursery school. New operators can be selected

to achieve this subgoal. For instance, installing a new battery was chosen. If this

operator is blocked, enabling it can become yet another subgoal.

Figure 8.9 shows two flowcharts of the procedures used in the means-ends

analysis employed by GPS. A general feature of this analysis is that it breaks a

larger goal into subgoals. GPS creates subgoals in two ways. First, in flowchart 1,

GPS breaks the current state into a set of differences and sets the reduction of

each difference as a separate subgoal. First it tries to eliminate what it perceives

as the most important difference. Second, in flowchart 2, GPS tries to find an

operator that will eliminate the difference. However, GPS may not be able to

apply this operator immediately because a difference exists between the operator’s

condition and the state of the environment. Thus, before the operator can

be applied, it may be necessary to eliminate another difference. To eliminate the

difference that is blocking the operator’s application, flowchart 2 will have to be

called again to find another operator relevant to eliminating that difference.

The term operator subgoal is used to refer to a subgoal whose purpose is to

eliminate a difference that is blocking application of an operator.

Operator Selection | 225

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 225

Means-ends analysis involves creating subgoals to eliminate the difference

blocking the application of a desired operator.

The Tower of Hanoi Problem

Means-ends analysis has proved to be a generally applicable and extremely powerful

method of problem solving. Ernst and Newell (1969) discussed its application

to the modeling of monkey and bananas problems (such as Sultan’s predicament

described at the beginning of the chapter), algebra problems, calculus problems,

and logic problems. Here, however, we will illustrate means-ends analysis by

applying it to the Tower of Hanoi problem. Figure 8.10 illustrates a simple version

226 | Problem Solving

Match current state

to goal state to find the

most important difference

Flowchart 1 Goal: Transform current state into goal state

Flowchart 2 Goal: Eliminate the difference

Difference

NO DIFFERENCES

NO DIFFERENCE

NONE FOUND

FAIL

FAIL

FAIL APPLY OPERATOR

FAIL

SUCCESS

SUCCESS

Operator

found

SUCCESS

detected

Difference

detected

Search for operator

relevant to reducing

the difference

Match condition of

operator to current

state to find most

important difference

Subgoal: Eliminate

the difference

Subgoal:

Eliminate

the difference

1 2 3 1 2 3

A

B

C

A

Start Goal

B

C

FIGURE 8.9 The application of means-ends analysis by Newell and Simon’s General Problem

Solving (GPS) program. Flowchart 1 breaks a problem down into a set of differences and tries to

eliminate each one. Flowchart 2 searches for an operator that is relevant to eliminating a difference.

FIGURE 8.10 The three-disk version of the Tower of Hanoi problem.

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 226

of this problem. There are three pegs and three

disks of differing sizes, A, B, and C. The disks

have holes in them so they can be stacked on the

pegs. The disks can be moved from any peg to

any other peg. Only the top disk on a peg can be

moved, and it can never be placed on a smaller

disk. The disks all start out on peg 1, but the

goal is to move them all to peg 3, one disk at a

time, by transferring disks among pegs.

Figure 8.11 traces the application of the

GPS techniques to this problem. The first line

gives the general goal of moving disks A, B, and

C to peg 3. This goal leads us to the first flowchart

of Figure 8.9. One difference between the

goal and the current state is that disk C is not

on peg 3. This difference is chosen because

GPS tries to remove the most important difference

first, and we are assuming that the largest

misplaced disk will be viewed as the most

important difference. A subgoal set up to eliminate

this difference takes us to the second

flowchart of Figure 8.9, which tries to find an

operator to reduce the difference. The operator

chosen is to move C to peg 3. The condition for

applying a move operator is that nothing be on

the disk. Because A and B are on C, there is a

difference between the condition of the operator

and the current state. Therefore, a new subgoal

is created to reduce one of the differences—

B on C. This subgoal gets us back to

the start of flowchart 2, but now with the goal

of removing B from C (line 6 in Figure 8.11).2

The operator chosen the second time in

flowchart 2 is to move disk B to peg 2. However,

we cannot immediately apply the operator

of moving B to 2, because B is covered by A.

Therefore, another subgoal—removing A—is

set up, and flowchart 2 is used to remove this

difference. The operator relevant to achieving

this subgoal is to move disk A to peg 3. There

are no differences between the conditions for this operator and the current state.

Finally, we have an operator we can apply (line 12 in Figure 8.11), and we

achieve the subgoal of moving A to 3. Now we return to the earlier intention of

Operator Selection | 227

Goal: Move A, B, and C to peg 3

: Difference is that C is not on 3

: Subgoal: Make C on 3

: Operator is to move C to 3

: Difference is that A and B are on C

: Subgoal: Remove B from C

: Operator is to move B to 2

: Difference is that A is on B

: Subgoal: Remove A from B

: Operator is to move A to 3

: No difference with operator's condition

: No difference with operator's condition

: No difference with operator's condition

: No difference with operator's condition

: No difference with operator's condition

: Difference is that A is on 3

: Subgoal: Remove A from 3

: Operator is to move A to 2

: Apply operator (move A to 3)

: Apply operator (move A to 2)

: Apply operator (move C to 3)

: Apply operator (move B to 2)

: Apply operator (move B to 3)

: Subgoal achieved

: Subgoal achieved

: Subgoal achieved

: Subgoal achieved

: Subgoal achieved

: Subgoal achieved

: Subgoal achieved

: No difference

Goal achieved

: Difference is that A is not on 3

: Subgoal: Make A on 3

: Operator is to move A to 3

: No difference with operator's condition

: Apply operator (move A to 3 )

: Difference is that B is not on 3

: Subgoal: Make B on 3

: Operator is to move B to 3

: Difference is that A is on B

: Subgoal: Remove A from B

: Operator is to move A to 1

: No difference with operator's condition

: Apply operator (move A to 1)

1.

2.

3.

4.

5.

6.

7.

8.

9.

10.

11.

12.

13.

14.

15.

16.

17.

18.

19.

20.

21.

22.

23.

24.

25.

26.

27.

28.

29.

30.

31.

32.

33.

34.

35.

36.

37.

38.

39.

40.

41.

42.

43.

44.

45.

FIGURE 8.11 A trace of the

application of the GPS program,

as shown in Figure 8.9, to the

Tower of Hanoi problem shown

in Figure 8.10.

2 Note that we have gone from the use of flowchart 1 to the use of flowchart 2, to a new use of flowchart 2.

To apply flowchart 2 to find a way to move disk C to peg 3, we need to apply flowchart 2 to find a way to

remove disk B from disk C. Thus, one procedure is using itself as a subprocedure; such an action is called

recursion.

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 227

moving B to 2. There are no more differences between the condition for this operator

and the current state, and so the action takes place. The subgoal of

removing B from C is then satisfied (line 16 in Figure 8.11).

We have now returned to the original intention of moving disk C to peg 3.

However, disk A is now on peg 3, which prevents the action. Thus, we have

another difference to be eliminated between the now-current state and the

operator’s condition. We move A onto peg 2 to remove this difference. Now

the original operator of moving C to 3 can be applied (line 24 in Figure 8.11).

The state now is that disk C is on peg 3 and disks A and B are on peg 2. At

this point, GPS returns to its original goal of moving the three disks to peg 3. It

notes another difference—that B is not on 3—and sets another subgoal of eliminating

this difference. It achieves this subgoal by first moving A to 1 and then B

to 3. This gets us to line 37 in Figure 8.11. The remaining difference is that A

is not on 3. This difference is eliminated in lines 38 through 42.With this step,

no more differences exist and the original goal is achieved.

Note that subgoals are created in service of other subgoals. For instance,

to achieve the subgoal of moving the largest disk, GPS creates a subgoal of

moving the second-largest disk, which is on top of it. We indicated this logical

dependency of one subgoal on another in Figure 8.11 by indenting the processing

of the dependent subgoal. Before the first move in line 12 of the illustration,

three subgoals had to be created. It appears that creating such goals and subgoals

can be quite costly. Both Anderson, Kushmerick, and Lebiere (1993) and

Ruiz (1987) found that the time required to make one of the moves is a function

of the number of subgoals that must be created. For instance, before disk A

is moved to peg 3 in Figure 8.11 (the first move), three subgoals have to be created,

whereas no subgoals have to be created before the next move is taken—

moving B to peg 2. Correspondingly, Anderson et al. found that it took 8.95 s to

make the first move and 2.46 s to make the second move.

There are two problem-solving methods that participants could bring to

bear in solving the Tower of Hanoi problem. They could use a means-ends

approach as illustrated in Figure 8.11, or they could use the simpler differencereduction

method—in which case they would never set a subgoal to move a disk

that currently cannot be moved. In the Tower of Hanoi problem, such a simple

difference-reduction method would not be effective, because one needs to look

beyond what is currently possible and have a more global plan of attack on the

problem. The only step that difference reduction could take in Figure 8.10 would

be to move the top disk (A) to the target peg (3), but then it would provide no

further guidance because no other move would reduce the difference between

the current state and the goal state. Participants would have to make a random

move. Kotovsky, Hayes, and Simon (1985) studied the way people actually

approach the Tower of Hanoi problem. They found that there was an initial

problem-solving period during which participants did adopt this fruitless

difference-reduction strategy. Then they switched to a means-ends strategy, after

which the solution to the problem came quickly.

The Tower of Hanoi problem is solved by adopting a means-ends strategy

in which subgoals are created.

228 | Problem Solving

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 228

Goal Structuresand Prefrontal Cortex

It is significant that complex goal structures, particularly those involving operator

subgoaling, have been observed with any frequency only in humans and higher

primates.We have already discussed one instance of Sultan’s solution to the twostick

problem (see Figure 8.2). Novel tool building, a clear instance of operator

subgoaling, is almost unique to the higher apes (Beck, 1980). I (Anderson, 1993)

have speculated that the process of handling complex subgoals is performed by the

prefrontal cortex—which, as Figure 8.1 illustrates, is much larger in the higher

primates than in most other mammals, and is larger in humans than in most apes.

Chapter 6 discussed the role of the prefrontal cortex in holding information in

working memory. One of the major prerequisites to developing complex goal

structures is the ability to maintain these goal structures in working memory.

Goel and Grafman (1995) looked at how patients with frontal damage performed

in solving the Tower of Hanoi problem. These patients had suffered

severe damage to the prefrontal cortex. Many were veterans of the Vietnam War

who had lost large amounts of brain tissue as a result of penetrating missile

wounds. Although these patients had normal IQs, they showed much worse performance

than normal participants on the Tower of Hanoi task. There were certain

moves that these patients found particularly difficult to solve. As we noted in

discussing how means-ends analysis applies to the Tower of Hanoi problem, it is

necessary to make moves that deviate from the prescriptions of hill climbing.

One might have a disk at the correct position but have to move it away to enable

another disk to be moved to that position. It was exactly at these points where the

patients had to move “backward” that they had their problems. Only by maintaining

a set of goals can one see that a backward move is necessary for a solution.

More generally, it has been noted that patients with frontal damage have

difficulty inhibiting a predominant response (e.g., Roberts, Hager, & Heron,

1994). For instance, in the Stroop task (see Chapter 3), these patients have

trouble not saying the word itself when they are supposed to say the color of the

word. Apparently, they find it hard to keep in mind that their goal is to say the

color and not the word.

There is increased blood flow in the

prefrontal cortex during many tasks that

involve organizing novel and complex

behavior (Gazzaniga, Ivry, & Mangun, 1998).

Fincham, Carter, van Veen, Stenger, and

Anderson (2002) did an fMRI study of

students while they were solving Tower of

Hanoi problems and looked at brain activation

as a function of the number of goals that

the student had to set. These students were

solving much more complicated problems

than the simple one shown in Figure 8.10.

A problem might involve placing five disks

and could require maintaining as many as

five goals to reach a solution. Figure 8.12

shows the fMRI BOLD response of a region

Operator Selection | 229

1

0.0

0.05

0.10

0.15

0.20

BOLD response

Number of goals

2 3

Step in problem

Number of goals on stack

Increase in BOLD response (%)

4 5 6 7 8

1

2

3

4

FIGURE 8.12 Results from a

study by Fincham et al. to

examine brain activation as a

function of steps while solving

a Tower of Hanoi problem. The

red line shows the magnitude

of fMRI BOLD response in a

region in the right, anterior,

dorsolateral prefrontal cortex

during a sequence of eight

problem-solving steps in which

the number of goals being held

varied from 1 to 4. The black

shows the number of goals

being held at each point.

(Data from Fincham et al., 2002.)

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 229

in the right, anterior, dorsolateral prefrontal cortex during a sequence of eight

problem-solving steps in which the number of goals being held varied from 1

to 4. It also shows the number of goals being held at each point. There seems

to be a striking match between the goal load and the magnitude of the fMRI

response.

The prefrontal cortex plays a critical role in maintaining goal structures.

Problem Representation

The Importance of the Correct Representation

We have analyzed a problem solution as consisting of problem states and operators

for changing states. So far, we have discussed problem solving as if the

only tasks involved were to acquire operators and select the appropriate ones.

However, there are also important effects of how one represents the problem. A

famous example illustrating the importance of representation is the mutilatedcheckerboard

problem (Kaplan & Simon, 1990). Suppose we have a checkerboard

from which two diagonally opposite corner squares have been cut out.

Figure 8.13 illustrates this mutilated checkerboard, on which 62 squares remain.

Now suppose that we have 31 dominoes, each of which covers exactly two

squares of the board. Can you find some way of arranging these 31 dominoes

on the board so that they cover all 62 squares? If it can be done, explain how. If

it cannot be done, prove that it cannot. Perhaps you would like to ponder this

problem before reading on. Relatively few people are able to solve it without

some hints, and very few see the answer quickly.

The answer is that the checkerboard cannot be covered by the dominoes.

The trick to seeing this is to include in your representation of the problem the

fact that each domino must cover one black and one white

square, not just any two squares. There is just no way to place a

domino on two squares of the checkerboard without having it

cover one black and one white square. So with 31 dominoes, we

can cover 31 black squares and 31 white squares. But the

mutilation has removed two white squares. Thus, there are 30

white squares and 32 black squares. It follows that the mutilated

checkerboard cannot be covered by 31 dominoes.

Why is the mutilated-checkerboard problem easier to solve

when we represent each domino as covering a white and a black

square? The answer is that in so representing the problem, we are

encouraged to compare the number of white and black squares

on the board. Thus, the effect of the problem representation is

that it allows the critical operators to apply (i.e., checking for

parity).

Another problem that depends on correct representation is

the 27-apples problem. Imagine 27 apples packed together in a

crate 3 apples high, 3 apples wide, and 3 apples deep. A worm is

230 | Problem Solving

FIGURE 8.13 The mutilated

checkerboard used in the

problem posed by Kaplan and

Simon (1990) to illustrate the

importance of representation.

(After Wickelgren, 1974. Adapted by

permission of the publisher. © 1974

by W. H. Freeman and Company.)

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 230

in the center apple. Its life’s ambition is to eat its way through all the apples in

the crate, but it does not want to waste time by visiting any apple twice. The

worm can move from apple to apple only by going from the side of one into the

side of another. This means it can move only into the apples directly above, below,

or beside it. It cannot move diagonally. Can you find some path by which

the worm, starting from the center apple, can reach all the apples without going

through any apple twice? If not, can you prove it is impossible? The solution is

left to you. (Hint: The solution is based on a partial 3-D analogy to the solution

for the mutilated-checkerboard problem; it is given in the appendix at the end

of the chapter.) Inappropriate problem representations often cause students to

fail to solve problems even though they have been taught the appropriate

knowledge. This fact often frustrates teachers. Bassok (1990) and Bassok and

Holyoak (1989) studied high-school students who had learned to solve such

physics problems as the following:

What is the acceleration (increase in speed each second) of a train, if its speed

increases uniformly from 15 m/s at the beginning of the 1st second, to 45 m/s

at the end of the 12th second?

Students were taught such physics problems and became very effective at solving

them. However, they had very little success in transferring that knowledge

to solving such algebra problems as this one:

Juanita went to work as a teller in a bank at a salary of $12,400 per year and

received constant yearly increases, coming up with a $16,000 salary during

her 13th year of work.What was her yearly salary increase?

The students failed to see that their experience with the physics problems was

relevant to solving such algebra problems, which actually have the same structure.

This happened because students did not appreciate that knowledge associated

with continuous quantities such as speed (m/s) was relevant to problems

posed in terms of discrete quantities such as dollars.

Successful problem solving depends on representing problems in such a way

that appropriate operators can be seen to apply.

Functional Fixedness

Sometimes solutions to problems depend on the solver’s ability to represent the

objects in his or her environment in novel ways. This fact has been demonstrated

in a series of studies by different experimenters. A typical experiment in

the series is the two-string problem of Maier (1931), illustrated in Figure 8.14.

Two strings hanging from the ceiling are to be tied together, but they are so far

apart that the participant cannot grasp both at once. Among the objects in the

room are a chair and a pair of pliers. Participants try various solutions involving

the chair, but these do not work. The only solution that works is to tie the

pliers to one string and set that string swinging like a pendulum; then get the

second string, bring it to the center of the room, and wait for the first string to

swing close enough to grasp. Only 39% of Maier’s participants were able to see

this solution within 10 minutes. The difficulty is that the participants did not

Problem Representation | 231

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 231

perceive the pliers as a weight that could be used as a pendulum. This phenomenon

is called functional fixedness. It is so named because people are fixed on

representing an object according to its conventional function and fail to represent

its novel function.

Another demonstration of functional fixedness is an experiment by Duncker

(1945). The task he posed to participants was to support a candle on a door,

ostensibly for an experiment on vision. The problem is illustrated in Figure 8.15.

On the table are a box of tacks, some matches, and the candle. The solution is to

tack the box to the door and use the box as a platform for the candle. This task is

difficult because participants see the box as a container, not as a platform. They

have greater difficulty with the task if the box is filled with tacks, reinforcing the

perception of the box as a container.

These demonstrations of functional fixedness are consistent with the interpretation

that representation has an effect on operator selection. For instance,

to solve Duncker’s candle problem, participants needed to represent the tack box

in such a way that it could be used by the problem-solving operators that were

232 | Problem Solving

FIGURE 8.14 The two-string problem used by Maier to demonstrate functional fixedness.

Only 39% of Maier’s participants were able to see the solution within 10 minutes. A large

majority of the participants did not perceive the pliers as a weight that could be used

as a pendulum. (After Maier, 1931. Adapted by permission of the publisher. © 1931 by the Journal of

Comparative Psychology.)

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 232

looking for a support for the candle. When the box was conceived of as a container

and not as a support, it was not available to the support-seeking operators.

Functional fixedness refers to people’s tendency to see objects as serving

conventional problem-solving functions and thus failing to see possible

novel functions.

Set Effects

People can become biased to prefer certain operators when solving a problem by

their experiences. Such biasing of the problem solution is referred to as a set

effect. A good illustration involves the water jug problem studied by Luchins

(1942) and Luchins and Luchins (1959). In these water jug problems—which are

different from the Atwood and Polson (1976) problem shown in Figure 8.8—

participants were given a set of jugs of various capacities and an unlimited water

supply. The task was to measure out a specified quantity of water. Two examples

are given below:

Set Effects | 233

FIGURE 8.15 The candle problem used by Duncker (1945) in another study of functional

fixedness. (After Glucksberg & Weisberg, 1966. Adapted by permission of the publisher. Copyright © 1966 by

the American Psychological Association.)

Capacity of Capacity of Capacity of Desired

Problem Jug A Jug B Jug C Quantity

1 5 cups 40 cups 18 cups 28 cups

2 21 cups 127 cups 3 cups 100 cups

Anderson7e_Chapter_08.qxd 8/21/09 7:53 PM Page 233

Assume that participants have a tap and a sink so that they can fill jugs and

empty them. The jugs start out empty. Participants are allowed only to fill the

jugs to capacity, empty them completely, and pour water from one jug to another.

In problem 1, participants are told that they have three jugs: jug A, with a

capacity of 5 cups; jug B, with a capacity of 40 cups; and jug C, with a capacity

of 18 cups. To solve this problem, participants would fill jug A and pour it into

B, fill A again and pour it into B, and fill C and pour it into B. The solution

to this problem is denoted by 2A _ C. The solution for the second problem is

to fill jug B with 127 cups; fill A from B so that 106 cups are left in B; fill C from

B so that 103 cups are left in B; empty C; and fill C again from B so that the goal

of 100 cups in jug B is achieved. The solution to this problem can be denoted

by B _ A _ 2C. The first solution is called an addition solution because it

involves adding the contents of the jugs together; the second is called a subtraction

solution because it involves subtracting the contents of one jug from

another. Luchins first gave participants a series of problems that all could be

solved by addition, thus creating an “addition set.” These participants then

solved new addition problems faster, and subtraction problems slower, than

control participants who had no practice.

The set effect that Luchins (1942) is most famous for demonstrating is the

Einstellung effect, or mechanization of thought, which is illustrated by the series

of problems shown in Table 8.3. Participants were given these problems in this

order and were required to find solutions for each. Take time out from reading

this text and try to solve each problem.

All problems except number 8 can be solved by using a B _ 2C A method

(i.e., filling B, twice pouring B into C, and once pouring B into A). For problems 1

through 5, this solution is the simplest; but for problems 7 and 9, the simpler

234 | Problem Solving

TABLE 8.3

Luchins’s Water Jug Problems Used to Illustrate the Set Effect

Capacity (cups)

Problem Jug A Jug B Jug C Desired Quantity

1 21 127 3 100

2 14 163 25 99

3 18 43 10 5

4 9 42 6 21

5 20 59 4 31

6 23 49 3 20

7 15 39 3 18

8 28 76 3 25

9 18 48 4 22

10 14 36 8 6

After Luchins (1942). Adapted by permission of the publisher. © 1942 by Psychological

Monographs.

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 234

solution of A _ C also applies. Problem 8 cannot be solved by the B _ 2C A

method but can be solved by the simpler solution of A _ C. Problems 6 and 10

are also solved more simply by A _ C than by B _ 2C _ A. Of Luchins’s participants

who received the whole setup of 10 problems, 83% used the B _ 2C _ A

method on problems 6 and 7, 64% failed to solve problem 8, and 79% used the

B _ 2C _ A method for problems 9 and 10. The performance of participants

who worked on all 10 problems was compared with that of control participants

who saw only the last 5 problems. These control participants did not see the

biasing B _ 2C _ A problems. Fewer than 1% of the control participants used

B _ 2C _ A solutions, and only 5% failed to solve problem 8. Thus, the first 5

problems created a powerful bias for a particular solution. This bias hurt the

solution of problems 6 through 10. Although these effects are quite dramatic, they

are relatively easy to reverse with the exercise of cognitive control. Luchins found

that simply warning participants by saying, “Don’t be blind” after problem 5 allowed

more than 50% of them to overcome the set for the B _ 2C _ A solution.

Another kind of set effect in problem solving has to do with the influence of

general semantic factors. This effect is well illustrated in the experiment of

Safren (1962) on anagram solutions. Safren presented participants with lists

such as the following, in which each set of letters was to be unscrambled and

made into a word:

kmli graus teews recma foefce ikrdn

This is an example of an organized list, in which the individual words are all

associated with drinking coffee. Safren compared solution times for organized

lists with times for unorganized lists. Median solution time was 12.2 s for

anagrams from unorganized lists and 7.4 s for anagrams from organized lists.

Presumably, the facilitation evident with the organized lists occurred because

the earlier items in the list associatively primed, and so made more available,

the later words. Note that this anagram experiment contrasts with the water jug

experiment in that no particular procedure was being strengthened. Rather,

what was being strengthened was part of the participant’s factual (declarative)

knowledge about spellings of associatively related words.

In general, set effects occur when some knowledge structures become more

available than others. These structures can be either procedures, as in the water

jug problem, or declarative information, as in the anagram problem. If the

available knowledge is what participants need to solve the problem, their problem

solving will be facilitated. If the available knowledge is not what is needed,

problem solving will be inhibited. It is good to realize that sometimes set effects

can be dissipated easily (as with Luchins’s “Don’t be blind” instruction). If you

find yourself stuck on a problem and you keep generating similar unsuccessful

approaches, it is often useful to force yourself to back off, change set, and try a

different kind of solution.

Set effects result when the knowledge relevant to a particular type of problem

solution is strengthened.

Set Effects | 235

Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 235

Incubation Effects

People often report that after trying to solve a problem and getting nowhere,

they can put it aside for hours, days, or weeks and then, upon returning to it,

can see the solution quickly. Many examples of this pattern were reported by

the famous French mathematician Poincaré (1929), including, for instance, the

following:

Then I turned my attention to the study of some arithmetical questions apparently

without much success and without a suspicion of any connection with

my preceding researches. Disgusted with my failure, I went to spend a few days

at the seaside, and thought of something else. One morning, walking on the

bluff, the idea came to me, with just the same characteristics of brevity, suddenness,

and immediate certainty, that the arithmetic transformations of indeterminate

ternary quadratic forms were identical with those of non-Euclidean

geometry. (p. 388)

Such phenomena are called incubation effects.

An incubation effect was nicely demonstrated in an experiment by Silveira

(1971). The problem she posed to participants, called the cheap-necklace

problem, is illustrated in Figure 8.16. Participants were given the following

instructions:

You are given four separate pieces of chain that are each three links in length.

It costs 2¢ to open a link and 3¢ to close a link. All links are closed at the

beginning of the problem. Your goal is to join all 12 links of chain into a

single circle at a cost of no more than 15¢.

Try to solve this problem yourself. (A solution is provided in the appendix at

the end of this chapter.) Silveira tested three groups. A control group worked

on the problem for half an hour; 55% of these participants solved the

problem. For one experimental group, the half hour spent on the problem

was interrupted by a half-hour break in which the participants did other activities;

64% of these participants solved the problem. A second experimental

group had a 4-hour break, and 85% of these participants solved the problem.

Silveira required her participants to speak aloud as they solved the cheapnecklace

problem. She found that they did not come back to the problem

after a break with solutions completely worked out. Rather, they began by

trying to solve the problem much as before. This result is evidence against

a common misbelief that people are subconsciously

solving the problem in the period that they are away

from it.

The best explanation for incubation effects relates them

to set effects. During initial attempts to solve a problem,

people set themselves to think about the problem in certain

ways and bring to bear certain knowledge structures. If this

initial set is appropriate, they will solve the problem. If the

initial set is not appropriate, however, they will be stuck

throughout the session with inappropriate procedures.

Going away from the problem allows activation of the

236 | Problem Solving

chain A

Given state Goal state

chain B

chain C

chain D

FIGURE 8.16 The cheapnecklace

problem used by

Silveira (1971) to investigate

the incubation effect. (After

Wickelgren, 1974. Adapted by permission of

the publisher. © 1974 by W. H. Freeman.)

Anderson7e_Chapter_08.qxd 8/20/09 9:49 AM Page 236

inappropriate knowledge structures to dissipate, and

people are able to take a fresh approach.

The basic argument is that incubation effects occur

because people “forget” inappropriate ways of solving

problems. Smith and Blakenship (1989, 1991) performed

a fairly direct test of this hypothesis. They had participants

solve problems like those shown in Figure 8.17. They

provided half of their participants, the fixation group,

with inappropriate ways to think about the problems. For

instance, with respect to the third problem, they told

participants to think about chemicals. Thus, in the fixation

condition, they deliberately induced incorrect sets.

Not surprisingly, the fixation participants solved fewer of

the problems than the control participants. The interesting

issue, however, was how much incubation effect these

two populations of participants showed. Half of both the

fixation and control participants worked on the problems

for a continuous period of time, whereas the other half

had an incubation period inserted in the middle of their

problem-solving efforts. The fixation participants showed

a greater benefit of the incubation period. Thus, Smith

and Blakenship were able to show a greater incubation

effect in participants who had started with an inappropriate

way of solving the problem. Also, when they asked the

fixation participants what the misleading clue had been,

they found that more of the participants who had an incubation

period had forgotten the inappropriate clue.

Incubation effects occur when people forget the inappropriate strategies they

were using to solve a problem.

Insight

A common misbelief about learning and problem solving is that there are magical

moments of insight when everything falls into place and we suddenly see a

solution. This is called the “aha” experience, and many of us can report uttering

that very exclamation after a long struggle with a problem that we suddenly solve.

The incubation effects just discussed have been used to argue that the subconscious

is deriving this insight during the incubation period. As we saw, however, what

really happens is that participants simply let go of poor ways of solving problems.

Metcalfe and Wiebe (1987) came up with an interesting way to define

insight problems. The insight problems they used included ones like the

cheap-necklace problem (see Figure 8.16). Their noninsight problems required

multistep solutions, as in the Tower of Hanoi problem (see Figure 8.10). They

asked participants to judge every 15 s how close they felt they were to the solution.

Fifteen seconds before they actually solved a noninsight problem, participants

were fairly confident they were close to a solution. In contrast, on the

Set Effects | 237

lines reading lines

oholene

or

or

search

and

FIGURE 8.17 Puzzles used by Smith and Blakenship to test

the hypothesis that incubation effects occur because people

“forget” inappropriate ways of solving problems. Participants

had to figure out what familiar phrase was represented by

each image. For example, the first picture represents the

phrase “reading between the lines”; the second, “search high

and low”; the third, “a hole in one”; the fourth “double or

nothing.” (After Smith & Blakenship, 1989, 1991. Adapted by permission

of the publishers. © 1989 by the Bulletin of the Psychonomic Society. © 1991

by the American Journal of Psychology.)

Anderson7e_Chapter_08.qxd 8/20/09 9:49 AM Page 237

insight problems, participants had little idea they were close to a solution, even

15 s before they actually solved the problem. Metcalfe and Wiebe suggested

that we use this difference as a definition of insight problems. That is, an

insight problem is one in which people are not aware that they are close to a

solution.

This definition would seem to support the notion that a solution comes in a

single moment. However, what Metcalfe and Wiebe actually showed was that

participants did not know when they were close to a solution to an insight

problem. They did not show that the solution came in a single moment. Kaplan

and Simon (1990) studied participants while they solved the mutilatedcheckerboard

problem (see Figure 8.13), which is another insight problem.

They found that some participants noticed key features of the solution to the

problem—such as that a domino covers one square of each color—early on.

Sometimes, though, these participants did not judge those features to be critical

and went off and tried other methods of solution; only later did they come

back to the key feature. So, it is not that solutions to insight problems cannot

come in pieces, but rather that participants do not recognize which pieces are

key until they see the final solution. It reminds me of the time I tried to find my

way through a maze, cut off from all cues as to where the exit was. I searched

for a very long time, was quite frustrated, and was wondering if I was ever going

to get out—and then I made a turn and there was the exit. I believe I even

exclaimed, “Aha!” It was not that I solved the maze in a single turn; it was that

I did not appreciate which turns were on the way to the solution until I made

that final turn.

Sometimes, insight problems require only a single step (or turn) to solve,

and it is just a matter of finding that step. What is so difficult about these

problems is just finding that one step, which can be a bit like trying to find

a needle in a haystack. As an example of such a problem, consider the

following:

What is greater than God

More evil than the Devil

The poor have it

The rich want it

And if you eat it, you’ll die.

Reportedly, schoolchildren find this problem easier than college undergraduates.

If so, it is because they consider fewer possibilities as an answer. (If you are

frustrated and cannot solve this problem, it turns out that, like many things,

one can find the answer by searching the Web—many people have posted this

problem on their Web pages.)

As a final example of insight problems consider the remote association

problems introduced by Mednick (1962). In one version of these problems

(Mednick, 1962), participants are asked to find some word that can be combined

with three words to make a compound word. So, for instance, given

fox, man, and peep, the solution is hole (foxhole, manhole, peephole). Here are a

238 | Problem Solving

Anderson7e_Chapter_08.qxd 8/20/09 9:49 AM Page 238

number of problems for you to try (the solutions

are given in the appendix):

print/berry/bird

dress/dial/flower

pine/crab/sauce

Studies of brain activity (Jung-Beeman et al.,

2004) have been conducted while people try to

solve these problems. Characteristic of insight

problems, people often get a sudden feeling

of insight when they solve them. Figure 8.18

shows the imaging results from our laboratory,

which say a lot about what is happening. The

region is plotting activity in the left prefrontal

region whose activity has been associated

with retrieval from declarative memory (e.g.,

Figures 1.16c, 7.6). The figure compares activity in cases where participants

are able to solve the problem with cases where they are not. Time 0 in the

figure marks the point where the solution was obtained in the successful

case. Both functions for the successful and unsuccessful cases are increasing,

reflecting increasing effort as the search progresses, but there is an abrupt drop

(time-lagged as we would expect with the BOLD response) after the insight. It

should be emphasized that other regions such as the motor region show a rise

at this point associated with the generation of the response. In dropping off,

the prefrontal is showing a strikingly different response compared to other

brain regions and is reflecting the end to the search of memory for the answer.

The participant had been trying retrieval after retrieval and finally retrieved the

right answer. The feeling of insight corresponds to the moment when retrieval

finally succeeds and activity drops in the retrieval area.

Insight problems are ones in which solvers cannot recognize when they are

getting close to the solution.

Conclusions

This chapter has been built around the Newell and Simon model of problem

solving as a search through a state space defined by operators. We have

looked at problem-solving success as determined by the operators available

and the methods used to guide the search for operators. This analysis is

particularly appropriate for first-time problems, whether a chimpanzee’s

quandary (see Figure 8.2) or a human’s predicament when shown a Tower of

Hanoi problem for the first time (see Figure 8.10). The next chapter will

focus on the other factors that come into play with repeated problem-solving

practice.

Conclusions | 239

−10

−0.1

0.3

0.2

0.0

0.1

0.4

0.5

0.7

0.6

0.8

Baseline

Solution

−5 0

Time (sec.) from response

LIPFC (retrieval)

5 10

FIGURE 8.18 A comparison of

brain activity for successful and

unsuccessful attempts to solve

a remote association problem.

The activity plotted is from a

prefrontal region that is sensitive

to retrieval. Activity increases

with increasing time on task but

drops off for successful problems

shortly after the solution

(at time 0).

Anderson7e_Chapter_08.qxd 8/20/09 9:49 AM Page 239

240 | Problem Solving

1. Recent research (e.g., Pizlo et al., 2006) has been

conducted on the so-called “traveling salesman

problem.” To make an example of such a problem, put

a number of dots (say, 10 to 20) randomly on a page

and pick one as your start dot. Now try to draw the

shortest path from this dot, visiting each dot just once

and arriving back at your start dot. If you were to

characterize this problem as a search space, what would

the states of the problem be and what would the operators

be? How do you select among the operators? Is this

particularly useful to characterize this problem in terms

of such a search space?

2. In the modern world, humans frequently want to learn

how to use devices like microwaves or software such as

a spreadsheet package.When do you try to learn these

things by discovery, by following an example, and by

following instructions? How often are your learning

experiences a mixture of these modes of learning?

3. A common goal for students is getting a good grade in a

course. There are many different things that you can do

to try to improve your grade. How do you select among

them? When are you engaged in hill climbing and when

are you engaged in means-ends analysis?

4. Figure 8.19 illustrates the nine-dots problem (Maier,

1931). The problem is to connect all 9 dots by drawing

4 lines, never lifting your pen from the page. Summarizing

a variety of studies, Kershaw and Ohlsson (2001)

report that given only a few minutes, only 5% of

undergraduates can solve this problem. Try to solve this

problem. If you get frustrated, you can find an answer

by Googling “nine-dots problem.”After you have tried

to solve the problem, use the terminology (see below)

of this chapter to describe the nature of the difficulties

posed by this problem and what people need to do to

successfully solve this problem.

Questions for Thought

Key Terms

analogy

backup avoidance

difference reductions

Einstellung effect

functional fixedness

General Problem Solver

(GPS)

goal state

hill climbing

incubation effect

insight problem

means-ends analysis

operator

problem space

search

search tree

set effect

state

subgoal

Tower of Hanoi problem

Appendix: Solutions

Figure A8.1 gives the minimum-path solution to the problem solved less

efficiently in Figure 8.3.

With regard to the problem of the 27 apples, the worm cannot succeed. To

see that this is the case, imagine that the apples alternate in color, green and

red, in a 3-D checkerboard pattern. If the center apple from which the worm

starts is red, there are 13 red apples and 14 green apples in all. Each time the

worm moves from one apple to another, it will be changing colors. Because

FIGURE 8.19 The nine-dots problem.

Anderson7e_Chapter_08.qxd 8/20/09 9:49 AM Page 240

the worm starts from a red apple, it cannot reach more green apples than red

apples. Thus, it cannot visit all 14 green apples if it also visits each of the 13 red

apples just once.

To solve the cheap-necklace problem shown in Figure 8.16, open all three

links in one chain (at a cost of 6¢) and then use the three open links to connect

the remaining three chains (at a cost of 9¢).

The solutions to the three remote association problems are blue, sun, and

apple.

Appendix: Solutions | 241

2 1

(a) (b) (c) (d) (e) (f) (g)

(o) (p) (q) (r) Goal state

(n) (m) (l) (k) ( j) (i) (h)

6 2 1 6 2 1 1 1

8 8 6 6 6

8 8

2 2

8 8

4 4 8

2 2

8

8

4 4 4

4 4 8

6 6

4 8 4 8 4

1 1

6 6

1

6

4 4 2

2 2

4 4

7

7

5

5

2 2

7 5

2

7 5

2

7 5

2

7 5

7 5 7 5

8

7 5

8 8 2 8

7 5 7 5 7 5

3 7 5 3

2 4

7 5

3

1 3 1 3 1 3

6 6

1 3

3 3

1

6

6

1

2 1

6

8

4

4

6

1

4

7 5

3

3 3 3 3

7 5 3 7 5 3 7 5 3

2 8 1

4 6

7 5 3

6

1 3

6

8 1

FIGURE A8.1 The minimum-path solution for the eight-tile problem that was solved less

efficiently in Figure 8.3.

Anderson7e_Chapter_08.qxd 8/20/09 9:49 AM Page 241

242

9Expertise

It has been speculated that the expansion of the human brain from Homo erectus

to modern Homo sapiens was driven by the need to acquire expertise in novel environments

(Skoyles, 1999). This ability allowed humans to spread throughout the

world and permitted the development of the technology that has created modern

civilization. Humans are the only species that display this kind of behavioral plasticity—

being able to become experts at driving a car in modern society, navigating the

oceans in Polynesian society, or designing search engines for the World Wide Web.

William G. Chase, late of Carnegie Mellon University, was one of our local experts on

human expertise. He emphasized two famous mottos that summarize much of the

nature of expertise and its development:

• No pain, no gain. • When the going gets tough, the tough get going.

The first motto refers to the fact that no one develops expertise without a great

deal of hard work. John R. Hayes (1985), another Carnegie Mellon faculty member,

has studied geniuses in fields varying from music to science to chess. He found that

no one reached genius levels of performance without at least 10 years of practice.

Chase’s second motto refers to the fact that the difference between relative novices

and relative experts increases as we look at more difficult problems. For instance,

there are many chess duffers who could play a credible, if losing, game against a

master when they are given unlimited time to choose moves. However, they would

lose embarrassingly if forced to play lightning chess, where they are permitted only

5 s per move.

Chapter 8 reviewed some of the general principles governing problem solving,

particularly in novel domains. This research has provided a framework for analyzing

the development of expertise in problem solving. Research on expertise has been a

major development in cognitive science in the past 30 years. This research is particularly

exciting because it has important contributions to make to the instruction of

technical or formal skills in areas such as mathematics, science, and engineering, as

will be reviewed at the end of this chapter.

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 242

Brain Changes with Skill Acquisition | 243

This chapter will address the following questions about the nature of human

expertise: • What are the stages in the development of expertise? • How does the organization of a skill change as one becomes expert? • What are the contributions of practice versus talent to the development of skill? • How much can skill in one domain transfer to a new domain? • What are the implications of our knowledge about expertise for teaching

new skills?

Brain Changes with Skill Acquisition

As people become more proficient at a task, they seem to use less of their brains

to perform that task. Figure 9.1 shows fMRI some data from Qin et al. (2003)

looking at areas of the brain activated as college students learned to perform

derivations in a new synthetic domain of mathematics. Figure 9.1a shows the regions

activated on their first day of doing the task and Figure 9.1b shows the

regions activated on the fifth day. As the students achieved greater efficiency in

the performance of the task, regions of activity dropped out or shrank. These

regions of activity correspond to metabolic expenditure, and it is quite apparent

that, with expertise, we spend less mental energy doing these tasks.

A general goal of research on expertise is to characterize both the qualitative

and the quantitative changes that take place with expertise. The result in

Figure 9.1 can be considered a quantitative result—more practice means more

efficient mental execution.We will look at a number of quantitative measures,

particularly latency, that indicate this increased efficiency. However, there are

also qualitative changes in how a skill is performed with practice. Figure 9.1

does not reveal such changes—in this study, it just seems that fewer areas,

FIGURE 9.1 Regions activated

in the symbol-manipulation task

of Qin et al. (2003): (a) day 1 of

practice; (b) day 5 of practice.

Note that these images depict

“transparent brains,” and the

activation that we see is not

just on the surface but also

below the surface.

Brain Structures

(a) (b)

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 243

rather than different areas, take part. However, this chapter will describe the

results of other brain imaging and behavioral studies that indicate that, indeed,

the way in which we perform a task can change as we become expert at it.

Through extensive practice, we can develop the high levels of expertise in

novel domains that have supported the evolution of human civilization.

General Characteristics of Skill Acquisition

Three Stages of Skill Acquisition

The development of a skill typically comprises three stages (Anderson, 1983;

Fitts & Posner, 1967). Fitts and Posner call the first stage the cognitive stage.

In this stage, participants develop a declarative encoding (see the distinction

between declarative and procedural representations at the end of Chapter 7) of

the skill; that is, they commit to memory a set of facts relevant to the skill.

Essentially these facts define the operators of the task (see Chapter 8). Learners

typically rehearse these facts as they first perform the skill. For instance, when I

was first learning to shift gears in a standard transmission car, I memorized the

location of the gears (e.g., “reverse is up, left”—for an old 3-speed transmission)

and the correct sequence of engaging the clutch and moving the stick

shift. I rehearsed this information as I performed the skill.

The information that I had learned about the location and function of the

gears amounted to a set of problem-solving operators for driving the car. For

instance, if I wanted to get the car into reverse, there was the operator of moving

the gear to the upper left. Despite the fact that the knowledge about what to do

next was unambiguous, one would hardly have judged my driving performance

as skilled. My use of the knowledge was very slow because that knowledge was

still in a declarative form. I had to retrieve specific facts and interpret them to

solve my driving problems. I did not have the knowledge in a procedural form.

The second stage of skill acquisition is called the associative stage. Two

main things happen in this second stage. First, errors in the initial understanding

are gradually detected and eliminated. So, I slowly learned to coordinate the

release of the clutch in first gear with the application of gas so as not to kill

the engine. Second, the connections among the various elements required for

successful performance are strengthened. Thus, I no longer had to sit for a few

seconds trying to remember how to get to second gear from first. Basically, the

outcome of the associative stage is a successful procedure for performing the

skill. However, it is not always the case that the procedural representation of

the knowledge replaces the declarative. Sometimes, the two forms of knowledge

can coexist side by side, as when we can speak a foreign language fluently and

still remember many rules of grammar. However, the procedural, not the

declarative, knowledge governs the skilled performance.

The third stage in the standard analysis of skill acquisition is the autonomous

stage. In this stage, the procedure becomes more and more automated

and rapid. The concept of automaticity was introduced in Chapter 3, where we

244 | Expertise

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 244

discussed how central cognition drops out of performance of a task as we become

more skilled at it. Complex skills such as driving a car or playing chess

gradually evolve in the direction of becoming more automated and requiring

fewer processing resources. For instance, driving a car can become so automatic

that people will engage in conversation with no memory for the traffic that they

have driven through.

The three stages of skill acquisition are the cognitive stage, the associative

stage, and the autonomous stage.

Power-Law Learning

Chapter 6 documented the way in which the retrieval of simple associations

improved as a function of practice according to a power law. It turns out that

the performance of complex skills, requiring the coordination of many such associations,

also improves according to a power law. Figure 9.2 illustrates a wellknown

instance of such skill acquisition. This study followed the development

of the cigar-making ability of a worker in a factory for 10 years. The figure plots

the time to make a cigar against number of years of practice. Both scales use

log-log coordinates to expose a power law (recall from Chapters 6 and 7 that a

linear function on log-log coordinates implies a power function in the original

scale). The data in this graph show an approximately linear function until

General Characteristics of Skill Acquisition | 245

10,000 100,000 1,000,000

Number of items produced (logarithmic scale)

100,000,000

10

(1 year) (7 years)

Minimum machine

cycle time

20

30

Cycle time (s, logarthmic scale)

5

FIGURE 9.2 Time required to produce a cigar as a function of amount of experience. (From

Crossman, 1959. Reprinted by permission from Taylor & Francis.)

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 245

about the fifth year, at which point the improvement appears to stop. It turns

out that the worker was approaching the cycle time of the machinery and could

improve no more. There is usually some limit to how much improvement can

be achieved, determined by the equipment, the capability of a person’s musculature,

age, and so on. However, except for these physical limits, there is no limit

on how much a skill can speed up. The time taken by the cognitive component

of a skill will go to zero, given enough practice.

Recall from Chapter 6 that a linear relation between log time T and log

practice P can be expressed as

ln T _ A _ b ln P

which can be transformed into

T _ aP_b

where A = ln a. In Chapter 6, we considered such power functions in memory

(see Figures 6.10 and 6.11). Basically, for these functions, the decrease in processing

time with further practice becomes small very rapidly.

Effects of practice have also been studied in domains of complex problem

solving, such as giving justifications for geometry-like proofs (Neves & Anderson,

1981). Figure 9.3 shows a power function for that domain, in both a normal

scale and a log-log scale. Such functions illustrate that the benefit of further

practice rapidly diminishes but that, no matter how much practice we have

had, further practice will help a little.

Kolers (1979) investigated the acquisition of reading skills, by using materials

such as those illustrated in Figure 9.4. The first type of text (N) is normal, but

the others have been transformed in various ways. In the R transformation, the

whole line has been turned upside down; in the I transformation, each letter

has been inverted; in the M transformation, the sentence has been set as a

mirror image of standard type. The rest are combinations of the several transformations.

In one study, Kolers looked at the effect of massive practice on

reading inverted (I) text. Participants took more than 16 min to read their first

page of inverted text compared with 1.5 min for normal text. After the initial

246 | Expertise

(a)

200

20

Number of problems Time to solution

40 60 80 100

400

600

800

1,000

1,200

1,400

(b)

2,000

1,000

400

200

100

2 4

Log (trials)

Log (s)

10 20 40 100

FIGURE 9.3 Time taken to

generate proofs in a geometrylike

proof system as a function

of the number of proofs

already done: (a) function on a

normal scale, RT _ 1,410P_55;

(b) function on a log-log scale.

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 246

reading-speed test, participants practiced on 200 pages of inverted text. Figure 9.5

provides a log-log plot of reading time against amount of practice. In this

figure, practice is measured as number of pages read. The change in speed with

practice is given by the curve labeled “Original training on inverted text.” Kolers

interspersed a few tests on normal text; data for these tests are given by the

curve labeled “Original tests on normal text.”We see the same kind of improvement

for inverted text as in Figures 9.2 and 9.3 (i.e., a straight-line function on

a log-log plot). After reading 200 pages, Kolers’s participants were reading at the

rate of 1.6 min per page—almost the same rate as that of participants reading

normal text.

A year later, Kolers had his participants read inverted text again. These data

are given by the curve in Figure 9.5 labeled “Retraining on inverted text.”

Participants now took about 3 min to read the first page of the inverted text.

Compared with their performance of 16 min on their first page a year earlier,

participants displayed an enormous savings, but it was now taking them almost

twice as long to read the text as it did after their 200 pages of training a year

earlier. They had clearly forgotten something. As the Figure 9.5 illustrates,

participants’ improvement on the retraining trials showed a log-log relation

between practice and performance, as had their original training. The same

level of performance that participants had initially reached after 200 pages of

General Characteristics of Skill Acquisition | 247

FIGURE 9.4 Examples of the

spatially transformed texts

used in Kolers’s studies of the

acquisition of reading skills. The

asterisks indicate the starting

point for reading. (From Kolers &

Perkins, 1975.)

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 247

training was now reached after 50 pages. Skills generally show very high levels

of retention. In many cases, such skills can be maintained for years with no

retention loss. Someone coming back to a skill—skiing, for example—after many

years of absence often requires just a short warm-up period before the skill is

reestablished (Schmidt, 1988).

Poldrack and Gabrieli (2001) investigated the brain correlates of the

changes taking place as participants learn to read transformed text such as that

in Figure 9.4. In an fMRI brain-imaging study, they found increased activity in

the basal ganglia and decreased activation in the hippocampus as learning progressed.

Recall from Chapters 6 and 7 that the basal ganglia are associated with

procedural knowledge, whereas the hippocampus is associated with declarative

knowledge. Similar changes in the activation of brain areas have been found by

Poldrack et al. (1999) in another skill-acquisition task that required the classification

of stimuli. As participants develop their skill, they appear to move to a

direct recognition of the text. Thus, the results of this brain-imaging research

reveal changes consistent with the switch between the cognitive and the associative

stages. Thus, qualitative changes appear to be contributing to the quantitative

changes captured by the power function.We will consider these qualitative

changes in more detail in the next section.

Performance of a cognitive skill improves as a power function of practice and

shows modest declines only over long retention intervals.

248 | Expertise

FIGURE 9.5 The results for readers in Kolers’s reading-skills experiment on two tests more

than a year apart. Participants were trained with 200 pages of inverted text in which pages of

normal text were occasionally interspersed. A year later, they were retrained with 100 pages

of inverted text, again with normal text occasionally interspersed. The results show the effect

of practice on the acquisition of the skill. Both reading time and number of pages practiced

are plotted on a logarithmic scale. (From Kolers, 1976. Copyright by the American Psychological Association.

Reprinted by permission.)

0

16

8

4

2

1

2 4 8

Page number (logarithmic scale)

Reading time (min, logarithmic scale)

16 32 64 128 256

Original training on inverted text

Retraining on inverted text

Original tests on normal text

Retraining tests on normal text

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 248

The Nature of Expertise

So far in this chapter, we have considered some of the phenomena associated

with skill acquisition. An understanding of the mechanisms behind these phenomena

has come from examining the nature of expertise in various fields of

endeavor. Since the mid-1970s, there has been a great deal of research looking

at expertise in such domains as mathematics, chess, computer programming,

and physics. This research compares people at various levels of development of

their expertise. Sometimes this research is truly longitudinal and follows students

from their introduction to a field to their development of some expertise.

More typically, such research samples people at different levels of expertise. For

instance, research on medical expertise might look at students just beginning

medical school, residents, and doctors with many years of medical practice.

This research has begun to identify some of the ways that problem solving

becomes more effective with experience. Let us consider some of these dimensions

of the development of expertise.

Proceduralization

The degree to which participants rely on declarative versus procedural knowledge

changes dramatically as expertise develops. It is illustrated in my own

work on the development of expertise in geometry (Anderson, 1982). One student

had just learned the side-side-side (SSS) and side-angle-side (SAS) postulates

for proving triangles congruent. The side-side-side postulate states that, if

three sides of one triangle are congruent to the corresponding sides of another

triangle, the triangles are congruent. The side-angle-side postulate states that, if

two sides and the included angle of one triangle are congruent to the corresponding

parts of another triangle, the triangles are congruent. Figure 9.6 illustrates

the first problem that the student had to solve. The first thing that he did

in trying to solve this problem was to decide which postulate to use. The following

is a part of his thinking-aloud protocol, during which he decided on the

appropriate postulate:

If you looked at the side-angle-side postulate (long pause) well RK and RJ

could almost be (long pause) what the missing (long pause) the missing side.

I think somehow the side-angle-side postulate works its way into here (long

pause). Let’s see what it says: “Two sides and the included angle.”What would

I have to have to have two sides JS and KS are one of them. Then you could go

back to RS = RS. So that would bring up the side-angle-side postulate (long

pause). But where would Angle 1 and Angle 2 are right angles fit in (long

pause) wait I see how they work (long pause). JS is congruent to KS (long

pause) and with Angle 1 and Angle 2 are right angles that’s a little problem

(long pause). OK, what does it say—check it one more time: “If two sides and

the included angle of one triangle are congruent to the corresponding parts.”

So I have got to find the two sides and the included angle.With the included

angle you get Angle 1 and Angle 2. I suppose (long pause) they are both right

angles, which means they are congruent to each other. My first side is JS is

to KS. And the next one is RS to RS. So these are the two sides. Yes, I think it is

the side-angle-side postulate. (Anderson, 1982, pp. 381–382)

The Nature of Expertise | 249

R

J

S

1

2

K

Given: ∠1 and ∠2 are right angles

Prove: RSJ ≅RSK

JS ≅KS

FIGURE 9.6 The first geometryproof

problem encountered by a

student after studying the sideside-

side and side-angle-side

postulates.

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 249

After reaching this point, the student still went through a long process of

actually writing out the proof, but the part of the protocol just given is germane

to assessing what goes into recognizing the relevance of the SAS postulate. After

a series of four more problems (two solved by SAS and two by SSS), the student

applied the SAS postulate in solving the problem illustrated in Figure 9.7. The

method-recognition part of the protocol was as follows:

Right off the top of my head I am going to take a guess at what I am

supposed to do: Angle DCK is congruent to Angle ABK. There is only one of

two and the side-angle-side postulate is what they are getting to. (Anderson,

1982, p. 382)

A number of things seem striking about the contrast between these two protocols.

One is that the application of the postulate has clearly sped up in the second

part of the protocol. A second is that there is no verbal rehearsal of the

statement of the postulate in the second case. The student is no longer calling a

declarative representation of the postulate into working memory. Note also

that, in the first protocol, working memory fails a number of times—points at

which the student had to recover information that he had forgotten. The third

feature of difference is that, in the first protocol, application of the postulate is

piecemeal; the student is separately identifying every element of the postulate.

Piecemeal application is absent in the second protocol. It appears that the postulate

is being matched in a single step.

These transitions are like the ones that Fitts and Posner characterized as

belonging to the associative stage of skill acquisition. The student is no longer

relying on verbal recall of the postulate but has advanced to the point where

he can simply recognize the application of the postulate as a pattern. Pattern

recognition is an important part of the procedural embodiment of a skill. We

no longer have to think about what to do next; we just recognize what is appropriate

for the situation. The process of converting the deliberate use of declarative

knowledge into pattern-driven application of procedural knowledge is

called proceduralization.

In Anderson (2007) I reported a meta-analysis of a number of studies in our

laboratory looking at the effects of practice on the performance of mathematical

problem-solving tasks like the ones we have been discussing in this section.We

were interested in the effects of this sort of practice on the three brain regions

illustrated in Figure 1.15:

Motor, which is involved in programming the actual motor movements in

writing out the solution;

Parietal, which is involved in representing the problem internally; and

Prefrontal, which is involved in retrieving things like the task instructions.

In addition we looked at a fourth region:

Anterior Cingulate Cortex (ACC), which is involved in the control of

cognition—see Figure 3.1 and later discussion in Chapter 3.

Figure 9.8 shows the mean level of activation in these regions initially and after

4 hours of practice. The motor or control demands of the tasks do not change

250 | Expertise

FIGURE 9.7 The sixth geometryproof

problem encountered by a

student after studying the sideside-

side and side-angle-side

postulates.

A

B

K

3

1

2

4

C

D Given: ∠1 ≅∠2

Prove: ABK ≅DCK

AB ≅DC

BK ≅CK

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 250

much and so there is comparable activation early versus late in the motor

cortex and the ACC. There is some reduction in the parietal suggesting that the

representational demands may be decreasing a bit. However, the dramatic

change is in the prefrontal, which is showing a major decrease because the task

instructions are no longer being retrieved. Rather, the knowledge is coming to

be directly applied.

Proceduralization refers to the process by which people switch from explicit

use of declarative knowledge to direct application of procedural knowledge,

which enables them do things such as riding a bike without thinking

about it.

Tactical Learning

As students practice problems, they come to learn the sequences of actions

required to solve a problem or parts of the problem. Learning to execute such

sequences of actions is called tactical learning. A tactic refers to a method that

accomplishes a particular goal. For instance, Greeno (1974) found that it took

only about four repetitions of the hobbits and orcs problem (see discussion

surrounding Figure 8.7) before participants could solve the problem perfectly.

In this experiment, participants were learning the sequence of moves to get the

creatures across the river. Once they had learned the sequence, they could simply

recall it and did not have to figure it out.

Logan (1988) argued that a general mechanism of skill acquisition involves

learning to recall solutions to problems that formerly had to be figured out. A

nice illustration of this mechanism is from a domain called alpha-arithmetic. It

entails solving problems such as F _ 3, in which the participant is supposed to

say the letter that is the number of letters forward in the alphabet—in this case,

The Nature of Expertise | 251

Motor Parietal

Brain region

Prefrontal ACC

0.10

0

0.05

0.15

0.20

0.25

0.30

Late

Early

FIGURE 9.8 Representation of the activity in four brain regions while performing tasks early on

versus after 5 days of practice.

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 251

F _3 _ I. Logan and Klapp (1991) performed an

experiment in which they gave participants problems

that included addends from 2 (e.g., C _ 2) through 5

(e.g., G _ 5). Figure 9.9 shows the time taken by participants

to answer these problems initially and then

after 12 sessions of practice. Initially, participants

took 1.5 s longer on the 5-addend problems than on

the 2-addend problems, because it takes longer to

count five letters forward in the alphabet than two

letters forward. However, the problems were repeated

again and again across the sessions. With repeated,

continued practice, participants became faster on all

problems, reaching the point where they could solve

the 5-addend problems as quickly as the 2-addend

problems. They had memorized the answers to these

problems and were not going through the procedure

of solving the problems by counting.1

There is evidence that, as people become more

practiced at a task and shift from computation to

retrieval, brain activation shifts from the prefrontal

cortex to more posterior areas of the cortex. For

instance, Jenkins, Brooks, Nixon, Frackowiak, and

Passingham (1994) looked at participants learning to key out various sequences

of finger presses such as “ring, index, middle, little, middle, index, ring, index.”

They compared participants initially learning these sequences with participants

practiced in these sequences. They used PET imaging studies and found that

there was more activation in frontal areas early in learning than late in learning.2

On the other hand, later in learning, there was more activation in the hippocampus,

which is a structure associated with memory. Such results indicate that, early

in a task, there is significant involvement of the anterior cingulate in organizing

the behavior but that, late in learning, participants are just recalling the answers

from memory. Thus, these neurophysiological data are consistent with Logan’s

proposal.

Tactical learning refers to a process by which people learn specific procedures

for solving specific problems.

Strategic Learning

The preceding subsection on tactical learning was concerned with how students

learn tactics by memorizing sequences of actions to solve problems. Many small

problems repeat so often that we can solve them this way. However, large and

252 | Expertise

Latency (s)

3.0

2.0

4.0

1.0

0.0

1 2

Addend

Session 1

Session 12

3 4 5

FIGURE 9.9 After 12 sessions,

participants solved alphaarithmetic

problems with

various-sized addends in

considerably less time.

(From Logan & Klapp, 1991).

1 Rabinowitz and Goldberg (1995) reported a study making a similar point.

2 This early-learning activation included the same anterior cingulate whose activity did not change in the

mathematical problem-solving tasks in Figure 9.8.However, in this simpler experiment the need for control

dramatically changes, and there is less activity later in the anterior cingulate.

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 252

complex problems do not repeat exactly, but they still have

similar structures, and one can learn how to organize one’s

solution to the overall problem. Learning how to organize

one’s problem solving to capitalize on the general structure of

a class of problems is referred to as strategic learning. The

contrast between strategic and tactical learning in skill acquisition

is analogous to the distinction between tactics and strategy

in the military. In the military, tactics refers to smaller-scale

battlefield maneuvers, whereas strategy refers to higher-level

organization of a military campaign. Similarly, tactical learning

involves learning new pieces of skill, whereas strategic learning

is concerned with putting them together.

One of the clearest demonstrations of such strategic changes is in the domain

of physics problem solving. Researchers have compared novice and expert solutions

to problems like the one depicted in Figure 9.10. A block is sliding down an

inclined plane of length l, and u is the angle between the plane and the horizontal.

The coefficient of friction is m. The participant’s task is to find the velocity of the

block when it reaches the bottom of the plane. The typical novices in these studies

are beginning college students and the typical experts are their teachers.

In one study comparing novices and experts, Larkin (1981) found a difference

in how they approached the problem. Table 9.1 shows a typical novice’s

The Nature of Expertise | 253

_

_ l

FIGURE 9.10 A sketch of

a sample physics problem.

(From Larkin, 1981.)

TABLE 9.1

Typical Novice Solution to a Physics Problem

To find the desired final speed v requires a principle with v in it—say

v _ v0 + 2 at

But both a and t are unknown; so that seems hopeless. Try instead

v2 _ v0

2 _ 2 ax

In that equation, v0 is zero and x is known; so it remains to find a. Therefore, try

F _ ma

In that equation, m is given and only F is unknown; therefore, use

F _ F s

which in this case means

F _ Fg _ f

where Fg and f can be found from

Fg _ mg sin

f _ N

N _ mg cos

With a variety of substitutions, a correct expression for speed,

can be found.

Adapted from Larkin (1981).

v = 12(g sin u - mg cos u)/

u

m

– u

© ¿

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 253

solution to the problem, whereas Table 9.2 shows a typical expert’s solution. The

novice’s solution typifies the reasoning backward method, which starts with

the unknown—in this case, the velocity v. Then the novice finds an equation for

calculating v.However, to calculate v by this equation, it is necessary to calculate a,

the acceleration. So the novice finds an equation for calculating a; and the novice

chains backward until a set of equations is found for solving the problem.

The expert, on the other hand, uses similar equations but in the completely

opposite order. The expert starts with quantities that can be directly computed,

such as gravitational force, and works toward the desired velocity. It is also apparent

that the expert is speaking a bit like the physics teacher that he is, leaving

the final substitutions for the student.

Another study by Priest and Lindsay (1992) failed to find a difference in

problem-solving direction between novices and experts. Their study included

British university students rather than American students, and they found that

both novices and experts predominantly reasoned forward. However, their

experts were much more successful in doing so. Priest and Lindsay suggest that

the experts have the necessary experience to know which forward inferences are

appropriate for a problem. It seems that novices have two choices—reason forward,

but fail (Priest & Lindsay’s students) or reason backward, which is hard

(Larkin’s students)

Reasoning backward is hard because it requires setting goals and subgoals

and keeping track of them. For instance, a student must remember that he

or she is calculating F so that a can be calculated and hence so that v can be

calculated. Thus, reasoning backward puts a severe strain on working memory

and this can lead to errors. Reasoning forward eliminates the need to keep

254 | Expertise

TABLE 9.2

Skilled Solution to a Physics Problem

The motion of the block is accounted for by the gravitational force,

Fg _ mg sin

directed downward along the plane, and the frictional force,

f _ mg cos

directed upward along the plane. The block’s acceleration a is then related to the

(signed) sum of these forces by

F _ ma

or

mg sin _ mg cos _ ma

Knowing the acceleration a, it is then possible to find the block’s final speed v from

the relations

and

v _ at

Adapted from Larkin (1981).

l =

1

2

at 2

u m u

m u

– u

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 254

track of subgoals. However, to successfully reason forward, one must know

which of the many possible forward inferences are relevant to the final solution,

which is what an expert learns with experience. He or she learns to associate

various inferences with various patterns of features in the problems. The

novices in Larkin’s study seemed to prefer to struggle with backward reasoning,

whereas the novices in Priest and Lindsay’s study tried forward reasoning

without success.

Not all domains show this advantage for forward problem solving. A good

counterexample is computer programming (Anderson, Farrell, & Sauers, 1984;

Jeffries, Turner, Polson, & Atwood, 1981; Rist, 1989). Both novice and expert programmers

develop programs in what is called a top-down manner; that is, they

work from the statement of the problem to subproblems to sub-subproblems,

and so on, until they solve the problem. This top-down development is basically

the same as what is called reasoning backward in the context of geometry

or physics. There are differences between expert programmers and novice

programmers, however. Experts tend to develop problem solutions breadth

first, whereas novices develop their solutions depth first. Physics and geometry

problems have a rich set of givens that are more predictive of solutions than

is the goal. In contrast, nothing in the typical statement of a programming

problem would guide a working forward or bottom-up solution. The typical

problem statement only describes the goal and often does so with information

that will guide a top-down solution. Thus, we see that expertise in different

domains requires the adoption of those approaches that will be successful for

those particular domains.

In summary, the transition from novices to experts does not entail the same

changes in strategy in all domains. Different problem domains have different

structures that make different strategies optimal. Physics experts learn to reason

forward; programming experts learn breadth-first expansion.

Strategic learning refers to a process by which people learn to organize their

problem solving.

Problem Perception

As they acquire expertise problem solvers learn to perceive problems in ways

that enable more effective problem-solving procedures to apply. This dimension

can be nicely demonstrated in the domain of physics. Physics, being an intellectually

deep subject, has principles that are only implicit in the surface features

of a physics problem. Experts learn to see these implicit principles and represent

problems in terms of them.

Chi, Feltovich, and Glaser (1981) asked participants to classify a large set of

problems into similar categories. Figure 9.11 shows sets of problems that

novices thought were similar and the novices’ explanations for the similarity

groupings. As can be seen, the novices chose surface features, such as rotations

or inclined planes, as their bases for classification. Being a physics novice myself,

I have to admit that these seem very intuitive bases for similarity. Contrast

The Nature of Expertise | 255

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 255

these classifications with the pairs of problems in Figure 9.12 that the expert

participants saw as similar. Problems that are completely different on the

surface were seen as similar because they both entailed conservation of energy

or they both used Newton’s second law. Thus, experts have the ability to map

surface features of a problem onto these deeper principles. This ability is very

useful because the deeper principles are more predictive of the method of

solution. This shift in classification from reliance on simple features to reliance

on more complex features has been found in a number of domains, including

mathematics (Silver, 1979; Schoenfeld & Herrmann, 1982), computer

programming (Weiser & Shertz, 1983), and medical diagnosis (Lesgold et al.,

1988).

256 | Expertise

Novice 2: "Angular velocity, momentum,

circular things."

Novice 3: "Rotation kinematics, angular

speeds, angular velocities."

Novice 6: "Problems that have something

rotating: angular speed."

_

T

R

10 M

M

V

m

Novice 1: "These deal with blocks on an incline plane."

Novice 5: "Inclined plane problems, coefficient of friction."

Novice 6: "Blocks on inclined planes with angles."

2 lb.

_ = 2

Length

2 ft

30

M 30

_

V4 ft/s

FIGURE 9.11 Diagrams depicting pairs of problems categorized by novices as similar and

samples of their explanations for the similarity. (Adapted from Chi et al., 1981.)

Expert 2: "These can be solved by Newton's

second law."

Expert 3: "F = ma; Newton's second law."

Expert 4: "Largely use F = ma; Newton's

second law."

.6 m

.15 m

Equilibrium

Expert 2: "Conservation of energy."

Expert 3: "Work energy theorem. They are all straightforward

problems."

Expert 4: "These can be done from energy considerations.

Either you should know the principle of conservation

of energy, or work is lost somewhere."

K = 200 nt/m

M 30

_

Length

T T

m

M

Mg

mg mg

Fp = Kv

FIGURE 9.12 Diagrams depicting pairs of problems categorized by experts as similar and

samples of their explanations for the similarity. (Adapted from Chi et al., 1981.)

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 256

A good example of this shift in processing of perceptual features is the interpretation

of X rays. Figure 9.13 is a schematic of one of the X rays diagnosed by

participants in the research by Lesgold et al. The sail-like area in the right lung is a

shadow (shown on the left side of the X ray) caused by a collapsed lobe of the

lung that created a denser shadow in the X ray than did other parts of the lung.

Medical students interpreted this shadow as an indication of a tumor because tumors

are the most common cause of shadows on the lung. Radiological experts,

on the other hand, were able to correctly interpret the shadow as an indication of

a collapsed lung. They saw counterindicative features such as the size of the saillike

region. Thus, experts no longer have a simple association between shadows

on the lungs and tumors, but rather can see a richer set of features in X rays.

An important dimension of growing expertise is the ability to learn to

perceive problems in ways that enable more effective problem-solving

procedures to apply.

Pattern Learning and Memory

A surprising discovery about expertise is that experts seem to display a special enhanced

memory for information about problems in their domains of expertise.

This enhanced memory was first discovered in the research of de Groot (1965,

1966), who was attempting to determine what separated master chess players from

weaker chess players. It turns out that chess masters are not particularly more

intelligent in domains other than chess. De Groot found hardly any differences between

expert players and weaker players—except, of course, that the expert players

chose much better moves. For instance, a chess master considers about the same

number of possible moves as does a weak chess player before selecting a move. In

fact, if anything, masters consider fewer moves than do chess duffers.

However, de Groot did find one intriguing difference between masters and

weaker players.He presented chess masters with chess positions (i.e., chessboards

The Nature of Expertise | 257

Novice: Tumor

Expert: Collapsed lung

What causes

this shadow ?

FIGURE 9.13 Schematic

representation of an X ray

showing a collapsed right middle

lung lobe. (From Lesgold et al., 1988.)

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 257

with pieces in a configuration that occurred in a game) for just 5 s and then removed

the chess pieces. The chess masters were able to reconstruct the positions

of more than 20 pieces after just 5 s of study. In contrast, the chess duffers could

reconstruct only 4 or 5 pieces—an amount much more in line with the traditional

capacity of working memory. Chess masters appear to have built up

patterns of 4 or 5 pieces that correspond to common board configurations as

a result of the massive amount of experience that they have had with chess.

Thus, they remember not individual pieces but these patterns. In line with this

analysis, if the players are presented with random chessboard positions rather

than ones that are actually encountered in games, no difference is demonstrated

between masters and duffers—both reconstruct only a few chess positions. The

masters also complain about being very uncomfortable and disturbed by such

chaotic board positions.

In a systematic analysis, Chase and Simon (1973) compared novices, Class A

players, and masters. They compared these different types of players with respect

to their ability to reproduce game positions such as those shown in Figure 9.14a

258 | Expertise

Middle game

White

End game

Random middle game

(a)

(b)

Random end game

Black

FIGURE 9.14 Examples of (a) middle and end games and (b) their randomized counterparts.

(From Chase & Simon, 1973.)

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 258

The Nature of Expertise | 259

and to reproduce random positions such as

those illustrated in Figure 9.14b. Figure 9.15

shows the results. Memory was poorer for

all groups for the random positions and, if

anything, masters were worse at reproducing

these positions. On the other hand, masters

showed a considerable advantage for the actual

board positions. This basic phenomenon

of superior expert memory for meaningful

problems has been demonstrated in a large

number of domains, including the game of Go

(Reitman, 1976), electronic circuit diagrams

(Egan & Schwartz, 1979), bridge hands (Engle

& Bukstel, 1978; Charness, 1979), and computer

programming (McKeithen, Reitman,

Rueter, & Hirtle, 1981; Schneiderman, 1976).

Chase and Simon (1973) also used a

chessboard-reproduction task to examine the

nature of the patterns, or chunks, used by

chess masters. The participants’ task was simply to reproduce the positions of

pieces of a target chessboard on a test chessboard. In this task, participants

glanced at the target board, placed some pieces on the test board, glanced back

to the target board, placed some more pieces on the test board, and so on.

Chase and Simon defined a chunk to be a group of pieces that participants

moved after one glance. They found that these chunks tended to define

meaningful game relations among the pieces. For instance, more than half of

the masters’ chunks were pawn chains (configurations of pawns that occur

frequently in chess).

Simon and Gilmartin (1973) estimated that chess masters have acquired

50,000 different chess patterns, that they can quickly recognize such patterns on

a chessboard, and that this ability is what underlies their superior memory performance

in chess. This 50,000 figure is not unreasonable when one considers

the years of dedicated study that becoming a chess master requires.What might

be the relation between memory for so many chess patterns and superior performance

in chess? Newell and Simon (1972) speculated that, in addition to

learning many patterns, masters have learned what to do in the presence of

such patterns. For instance, if the chunk pattern is symptomatic of a weak side,

the response might be to suggest an attack on the weak side. Thus, masters

effectively “see” possibilities for moves; they do not have to think them out,

which explains why chess masters do so well at lightning chess, in which they

have only a few seconds to move.

To summarize, chess experts have stored the solutions to many problems

that duffers must solve as novel problems. Duffers have to analyze different

configurations, try to figure out their consequences, and act accordingly.

Masters have all this information stored in memory, thereby claiming two

advantages. First, they do not risk making errors in solving these problems,

because they have stored the correct solution. Second, because they have stored

0

Beginner

Number of pieces correctly placed

Class A Master

2

4

6

8

10

12

14

16

18

Actual game positions

Random positions

FIGURE 9.15 Number of pieces

successfully recalled by chess

players after the first study

of a chessboard. (From Chase &

Simon, 1973).

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 259

correct analyses of so many positions, they can focus their problem-solving efforts

on more sophisticated aspects and strategies of chess. Thus, the experts’

pattern learning and better memory for board positions is a part of the tactical

learning discussed earlier. The way humans become expert at chess reflects the

fact that we are very good at pattern recognition but relatively poor at things

like mentally searching through sequences of possible moves. As the Implications

box describes, human strengths and weaknesses lead to a very different

way of achieving expertise at chess than we see in computer programs for playing

chess.

260 | Expertise

chess in the 1960s, was beaten by the program of an

MIT undergraduate, Richard Greenblatt, in 1966 (Boden,

2006, discusses the intrigue surrounding

these events). However, Dreyfus was a

chess duffer and the programs of the

1960s and 1970s performed poorly

against chess masters. As computers

became more powerful and could search

larger spaces, they became increasingly

competitive, and finally in May 1997,

IBM’s Deep Blue program defeated the

reigning world champion, Gary Kasparov.

Deep Blue evaluated 200 million imagined

chess positions per second. It also

had stored records of 4,000 opening

positions and 700,000 master games

(Hsu, 2002) and had many other optimizations

that took advantage of special computer hardware.

Today there are freely available chess programs

for your personal computer that can be downloaded

over the Web and will play highly competitive chess at

a master level. These developments have led to a profound

shift in the understanding of intelligence. It once

was thought that there was only one way to achieve

high levels of intelligent behavior, and that was the

human way. Nowadays it is increasingly being accepted

that intelligence can be achieved in different ways, and

the human way may not always be the best. Also, curiously,

as a consequence some researchers no longer

view the ability to play chess as a reflection of the

essence of human intelligence.

Implications

Computers achieve computer expertise differently than humans

In Chapter 8, we discussed how human problem solving

can be viewed as a search of a problem space, consisting

of various states. The initial situation

is the start state, the situations on the

way to the goal are the intermediate

states, and the solution is the goal state.

Chapter 8 also described how people

use certain methods, such as avoiding

backup, difference reduction, and meansends

analysis, to move through the

states. Often when humans search a

problem space, they are actually manipulating

the actual physical world, as in

the 8-puzzle (Figures 8.3 and 8.4).

However, sometimes they imagine states,

as when one plays chess and contemplates

how an opponent will react to

some move one is considering, how one might react to

the opponent’s move, and so on. Computers are very

effective at representing such hypothetical states and

searching through them for the optimal goal state.

Artificial intelligence algorithms have been developed

that are very successful at all sorts of problem-solving

applications, including playing chess. This has led to a

style of chess playing program that is very different from

human chess play, which relies much more on pattern

recognition. At first many people thought that, although

such computer programs could play competent and

modestly competitive chess games, they would be no

match for the best human players. The philosopher

Hubert Dreyfus, who was famously critical of computer

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 260

Experts can recognize patterns of elements that repeat in many problems,

and know what to do in the presence of such patterns without having to

think them through.

Long-Term Memory and Expertise

One might think that the memory advantage shown by experts is just a workingmemory

advantage, but research has shown that their advantage extends to

long-term memory. Charness (1976) compared experts’ memory for chess positions

immediately after they had viewed the positions or after a 30-s delay filled

with an interfering task. Class A chess players showed no loss in recall over the

30-s interval, unlike weaker participants, who showed a great deal of forgetting.

Thus, expert chess players, unlike duffers, have an increased capacity to store

information about the domain. Interestingly, these participants showed the

same poor memory for three-letter trigrams as do ordinary participants. Thus,

their increased long-term memory is only for the domain of expertise.

There is reason to believe that the memory advantage goes beyond experts’

ability to encode a problem in terms of familiar patterns. Experts appear to be

able to remember more patterns as well as larger patterns. For instance, Chase

and Simon (1973) in their study (see Figures 9.14 and 9.15) tried to identify the

patterns that their participants used to recall the chessboards. They found that

participants would tend to recall a pattern, pause, recall another pattern, pause,

and so on. They found that they could use a 2-s pause to identify boundaries

between patterns.With this objective definition of what a pattern is, they could

then explore how many patterns were recalled and how large these patterns

were. In comparing a master chess player with a beginner, they found large

differences in both measures. First, the pattern size of the master averaged

3.8 pieces, whereas it was only 2.4 for the beginner. Second, the master also

recalled an average of 7.7 patterns per board,

whereas the beginner recalled an average of only

5.3. Thus, it seems that the experts’ memory advantage

is based not only on larger patterns but

also on the ability to recall more of them.

The strongest evidence that expertise requires

the ability to remember more patterns as well as

larger patterns is from Chase and Ericsson (1982),

who studied the development of a simple but

remarkable skill. They watched a participant, S. F.,

increase his digit span, which is the number of

digits that he could repeat after one presentation.

As discussed in Chapter 6, the normal digit span is

about 7 or 8 items, just enough to accommodate a

telephone number. After about 200 hr of practice,

S. F. was able to recall 81 random digits presented

at the rate of 1 digit per second. Figure 9.16 illustrates

how his memory span grew with practice.

The Nature of Expertise | 261

10

20

40

60

80

20

Practice (5-day blocks)

Digit span

30 40 50

FIGURE 9.16 The growth in

S. F.’s memory span with

practice. Notice how the

number of digits that he can

recall increases gradually but

steadily with the number of

practice sessions. (From Chase &

Ericsson, 1982.)

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 261

What was behind this apparent superhuman feat of memory? In part, S. F.

was learning to chunk the digits into meaningful patterns. He was a longdistance

runner, and part of his technique was to convert digits into running

times. So, he would take 4 digits, such as 3492, and convert them into “Three

minutes, 49.2 seconds—near world-record mile time.” Using such a strategy, he

could convert a memory span for 7 digits into a memory span for 7-digit patterns

of length 3 or 4. This would get him to a digit span of more than 20, far

short of his eventual performance. In addition to this chunking, he developed

what Chase and Ericsson called a retrieval structure, which enabled him to recall

22 such patterns. This retrieval structure was very specific; it did not generalize

to retrieving letters rather than digits. Chase and Ericsson hypothesized

that part of what underlies the development of expertise in other domains such

as chess is the development of retrieval structures, which allows superior recall

for past patterns.

As people become more expert in a domain, they develop a better ability

to store problem information in long-term memory and to retrieve it.

The Role of Deliberate Practice

An implication of all the research that we have reviewed is that expertise comes

only with an investment of a great deal of time to learn the patterns, the problemsolving

rules, and the appropriate problem-solving organization for a domain.

As mentioned earlier, John Hayes found that geniuses in various fields produce

their best work only after 10 years of apprenticeship in a field. In another

research effort, Ericsson, Krampe, and Tesch-Römer (1993) compared the best

violinists at a music academy in Berlin with those who were only very good.

They looked at diaries and self-estimates to determine how much the two

populations had practiced and estimated that the best violinists had practiced

more than 7000 hr before coming to the academy, whereas the very good had

practiced only 5000 hr. Ericsson et al. reviewed a great many fields where, like

music, time spent practicing is critical. Not only is time on task important at

the highest levels of performance, but also it is essential to mastering school

subjects. For instance, Anderson, Reder, and Simon (1998) noted that a major

reason for the higher achievement in mathematics of students in Asian countries

is that those students spend twice as much time practicing mathematics.

Ericsson et al. (1993) make the strong claim that almost all of expertise is to

be accounted for by amount of practice, and there is virtually no role for natural

talent. They point to the research of Bloom (1985a, 1985b), who looked at the

histories of children who became great in fields such as music or tennis. Bloom

found that most of these children got started by playing around, but after a short

time they typically showed promise and were encouraged by their parents to

start serious training with a teacher. However, the early natural abilities of these

children were surprisingly modest and did not predict ultimate success in the

domain (Ericsson et al., 1993). Rather, what is critical seems to be that parents

come to believe that a child is talented and consequently pay for their child’s

instruction and equipment as well as support their time-consuming practice.

262 | Expertise

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 262

Ericsson et al. speculated that the resulting training is sufficient to account for

the development ofchildren’s success. There is almost certainly some role for

talent (considered in Chapter 13), but all the evidence indicates that genius is

90% perspiration and 10% inspiration.

Ericsson et al. are careful to note, however, that not all practice leads to the

development of expertise. They note that many people spend a lifetime playing

chess or some sport without ever getting any better.What is critical, according

to Ericsson et al., is what they call deliberate practice. In deliberate practice,

learners are motivated to learn, not just perform; they are given feedback on

their performance; and they carefully monitor how well their performance

corresponds to the correct performance and where the deviations exist. The

learners focus on eliminating these points of discrepancy. The importance of

deliberate practice is similar to the importance of deep and elaborative processing

of the to-be-learned material described in Chapters 6 and 7, in which

passive study was shown to yield few memory benefits.

An important function of deliberate practice in both children and adults

may be to drive the neural growth that is necessary to enable expertise. It had

once been thought that adults do not grow new neurons, but it now appears

that they do (Gross, 2000). An interesting recent discovery is that extensive

practice appears to drive neural growth in the adult brain. For instance, Elbert,

Pantev,Wienbruch, Rockstroh, and Taub (1995) found that violinists, who finger

strings with the left hand, show increased development of the right cortical

regions that correspond to their fingers. In another study already mentioned

in Chapter 4, Maguire et al. (2003) used imaging to examine the brains of

London taxi drivers. It takes at least 3 years for London taxi drivers to acquire

all of the knowledge necessary to navigate expertly through the streets of

London. The taxi drivers were found to have significantly more gray matter in

the hippocampal region than did matched controls. This finding corresponds to

the increased hippocampal volume reported in small mammals and birds that

engage in behavior requiring navigation (Lee, Miyasato, & Clayton, 1998). For

instance, food-storing birds show seasonal increases in hippocampal volume

corresponding to times of the year when they need to remember where they

store food.

A great deal of deliberate practice is necessary to develop expertise in any

field.

Transfer of Skill

Expertise can often be quite narrow. As noted, Chase and Ericsson’s participant

S. F. was unable to transfer memory span skill from digits to letters. This example

is an almost ridiculous extreme of a frequent pattern in the development

of cognitive skills—that these skills can be quite narrow and fail to transfer

to other activities. Chess grand masters do not appear to be better thinkers

for all their genius in chess. An amusing example of the narrowness of expertise

Transfer of Skill | 263

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 263

is a study by Carraher, Carraher, and Schliemann (1985). These researchers

investigated the mathematical strategies used by Brazilian schoolchildren who

also worked as street vendors. On the job, these children used quite sophisticated

strategies for calculating the total cost of orders consisting of different

numbers of different objects (e.g., the total cost of 4 coconuts and 12 lemons);

what’s more, they could perform such calculations reliably in their heads.

Carraher et al. actually went to the trouble of going to the streets and posing as

customers for these children, making certain kinds of purchases and recording

the percentage of correct calculations. The experimenters then asked the children

to come with them to the laboratory, where they were given written mathematics

tests that included the same numbers and mathematical operations that

they had manipulated successfully in the streets. For example, if a child had

correctly calculated the total cost of 5 lemons at 35 cruzeiros apiece on the

street, the child was given the following written problem:

5 _ 35 _ ?

Whereas children solved 98% of the problems presented in the real-world context,

they solved only 37% of the problems presented in the laboratory context.

It should be stressed that these problems included the exact same numbers and

mathematical operations. Interestingly, if the problems were stated in the form

of word problems in the laboratory, performance improved to 74%. This improvement

runs counter to the usual finding, which is that word problems are

more difficult than equivalent “number” problems (Carpenter & Moser, 1982).

Apparently, the additional context provided by the word problem allowed the

children to make contact with their pragmatic strategies.

The study of Carraher et al. showed a curious failure of expertise to transfer

from real life to the classroom, but the typical concern of educators is whether

what is taught in one class will transfer to other classes and the real world.

Early in the 20th century, educators were fairly optimistic on this matter. A

number of educational psychologists subscribed to what has been called the

doctrine of formal discipline (Angell, 1908; Pillsbury, 1908; Woodrow, 1927),

which held that studying such esoteric subjects as Latin and geometry was of

significant value because it served to discipline the mind. Formal discipline

subscribed to the faculty view of mind, which extends back to Aristotle and

was first formalized by Thomas Reid in the late 18th century (Boring, 1950).

The faculty position held that the mind is composed of a collection of general

faculties, such as observation, attention, discrimination, and reasoning, which

were exercised in much the same way as a set of muscles. The content of the

exercise made little difference; most important was the level of exertion (hence

the fondness for Latin and geometry). Transfer in such a view is broad and

takes place at a general level, sometimes spanning domains that have no content

in common.

Although it might be nice to believe that such general transfer is possible,

as envisioned by the doctrine of formal discipline, there has been effectively

no evidence for it, despite a century of research on the topic. Some of the

earliest research on this topic was performed by Thorndike (e.g., Thorndike &

Woodworth, 1901). In one study, no correlation was found between memory

264 | Expertise

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 264

for words and memory for numbers. In another, accuracy in spelling was not

correlated with accuracy in arithmetic. Thorndike interpreted these results as

evidence against the general faculties of memory and accuracy.

There is often failure to transfer skills to similar domains and virtually no

transfer to very different domains.

Theory of Identical Elements

In place of the doctrine of formal discipline, Thorndike proposed his theory

of identical elements. According to Thorndike, the mind is not composed of

general faculties, but rather of specific habits and associations, which provide a

person with a variety of narrow responses to very specific stimuli. In fact, the

mind was regarded as just a convenient name for countless special operations or

functions (Stratton, 1922). Thorndike’s theory stated that training in one kind of

activity would transfer to another only if the activities had situation-response

elements in common:

One mental function or activity improves others in so far as and because

they are in part identical with it, because it contains elements common to

them. Addition improves multiplication because multiplication is largely

addition; knowledge of Latin gives increased ability to learn French because

many of the facts learned in the one case are needed in the other. (Thorndike,

1906, p. 243)

Thus, Thorndike was happy to accept transfer between diverse skills as long as

the transfer could be shown to be mediated by identical elements. Generally,

however, he concluded that

The mind is so specialized into a multitude of independent capacities that we

alter human nature only in small spots, and any special school training has a

much narrower influence upon the mind as a whole than has commonly been

supposed. (p. 246)

Although the doctrine of formal discipline was too broad in its predictions

of transfer, Thorndike formulated his theory of identical elements in what

proved to be an overly narrow manner. For instance, he argued that if you

solved a geometry problem in which one set of letters is used to label the points

in a diagram, you would not be able to transfer to a geometry problem with a

different set of letters. The research on analogy examined in Chapter 8 indicated

that this is not true. Transfer is not tied to the identity of surface elements.

In some cases, there is very large positive transfer between two skills that

have the same logical structure even if they have different surface elements (see

Singley & Anderson, 1989, for a review). Thus, for instance, there is large positive

transfer between different word-processing systems, between different programming

languages, and between using calculus to solve economics problems

and using calculus to solve problems in solid geometry. However, all the available

evidence is that there are very definite bounds on how far skills will transfer

Theory of Identical Elements | 265

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 265

and that becoming an expert in one domain will have little positive benefit

on becoming an expert in a very different domain. There will be positive

transfer only to the extent that the two domains use the same facts, rules, and

patterns—that is, the same knowledge. Thus, Thorndike was right in saying

that there would be transfer between two skills to the extent that they have the

same elements in common. However, he was wrong in identifying these

“elements” with stimulus-response bonds. Modern cognitive psychology has

identified these elements as rather abstract knowledge structures that enjoy a

wider range of transfer.

There is a positive side to this specificity in transfer of skill: there seldom

seems to be negative transfer, in which learning one skill makes a person worse

at learning another skill. Interference, such as that which occurs in memory

for facts (see Chapter 7), is almost nonexistent in skill acquisition. Polson,

Muncher, and Kieras (1987) provided a good demonstration of lack of negative

transfer in the domain of text editing on a computer (using the commandbased

word processors that were common at the time). They asked participants

to learn one text editor and then learn a second, which was designed to be maximally

confusing. Whereas the command to go down a line of text might be n

and the command to delete a character might be k in one text editor, n would

mean to delete a character in another text editor and k would mean to go down

a line. However, participants experienced overwhelming positive transfer in going

from one text editor to the other because the two text editors worked in the

same way, even though the surface commands had been scrambled. There is

only one clearly documented kind of negative transfer in regard to cognitive

skills—the Einstellung effect discussed in Chapter 8. Students can learn ways of

solving problems in one domain that are no longer optimal for solving problems

in another domain. So, for instance, someone may learn tricks in algebra to

avoid having to perform difficult arithmetic computations. These tricks may no

longer be necessary when that person uses a calculator to perform these computations.

Still, students show a tendency to continue to perform these unnecessary

simplifications in their algebraic manipulations. This example is not a

case of failure to transfer; rather, it is a case of transferring knowledge that is no

longer useful.

There is transfer between skills only when these skills have the same abstract

knowledge elements.

Educational Implications

With this analysis of skill acquisition, we can ask the question: What are the

implications for the training of such skills? One implication is the importance of

problem decomposition. Traditional high-school algebra has been estimated

to require the acquisition of many thousands of rules (Anderson, 1992).

Instruction can be improved by an analysis of what these individual elements

266 | Expertise

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 266

are. Approaches to instruction that begin with an analysis of the elements

to be taught are called componential analyses. A description of the application

of componential approaches to the instruction of a number of topics

in reading and mathematics can be found in Anderson (2000). Generally,

higher achievement is obtained in programs that include such componential

analysis.

A particularly effective part of such componential programs is mastery

learning. The basic idea in mastery learning is to follow students’ performance

on each of the components underlying the cognitive skill and to ensure that all

components are mastered. Typical instruction, without mastery learning, leaves

some students not knowing some of the material. This failure to learn some of

the components can snowball in a course in which mastery of earlier material is

a prerequisite for mastery of later material. There is a good deal of evidence

that mastery learning leads to higher achievement (Guskey & Gates, 1986;

Kulik, Kulik, & Bangert-Downs, 1986).

Instruction is improved by approaches that identify the underlying knowledge

components and ensure that students master them all.

Intelligent Tutoring Systems

Probably the most extensive use of such componential analysis is for intelligent

tutoring systems (Sleeman & Brown, 1982). These computer systems

interact with students while they are learning and solving problems, much as a

human tutor would. An example of such a tutor is the LISP tutor (Anderson,

Conrad, & Corbett, 1989; Anderson & Reiser, 1985; Corbett & Anderson,

1990), which teaches LISP, the main programming language used in artificial

intelligence. The LISP tutor continuously taught LISP to students at Carnegie

Mellon University from 1984 to 2002 and served as a prototype for a generation

of intelligent tutors, many of which have focused on teaching middle-school

and high-school mathematics. The mathematics tutors are now distributed by a

company called Carnegie Learning, spun off by Carnegie Mellon University in

1998. The Carnegie Learning mathematics tutors have been deployed over

2,600 schools nationwide and interacts with approximately 500,000 students

each year (Koedinger & Corbett, 2006; Ritter, Anderson, Koedinger, & Corbett,

2007; you can visit the Web site www.carnegielearning.com for promotional

material that should be taken with a grain of salt). Color Plate 9.1 shows a

screen shot from its most widely used product, which is a tutor for high-school

algebra.

A motivation for research on intelligent tutoring is the evidence showing that

private human tutoring is very effective. The results of studies have shown that

giving students a private human tutor enables 98% of them to do better than the

average student in a standard classroom (Bloom, 1984). An ideal private tutor is

one who is with the student at all times while he or she is studying a particular

subject matter. To use the terms of Ericsson et al. (1993), a private tutor guarantees

the deliberate practice that is essential for learning. Having the tutor present

Educational Implications | 267

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 267

while solving problems in domains, such as LISP and mathematics, which

require complex problem-solving skills, is particularly important. In LISP,

problem solving takes the form of writing computer programs, or functions, as

they are often called in LISP. Therefore, in developing the LISP tutor, we chose to

focus on providing students with tutoring while they are writing computer programs.

Table 9.3 presents a short dialogue between a student and the LISP tutor

on an early problem in the curriculum. Note how carefully the tutor monitors

the student’s performance in solving the problem. It can do so because it knows

how to write LISP functions. As the student is writing the function, the tutor is

simultaneously trying to solve the same problem that the student is working on.

As soon as it sees the student making a mistake, the tutor can intervene with

remedial instruction.

Underlying the tutor’s ability to solve problems and monitor the student’s

problem solving is a set of rules that can solve the same LISP programming

268 | Expertise

TABLE 9.3

An Interaction Between the LISP Tutor and a Student

The student is working on a piece of code to satisfy the following problem

description. Before each point at which the tutor and student interact, the partial

code that the student has written is printed out.

The Problem Description

Define a function called ctof that takes a degree reading in Celsius degrees and

converts it to Fahrenheit degrees. First, you need to change the scale by multiplying

the argument by 1.8. Then you need to translate the scale by adding 32, because

0º Celsius _ 32º Fahrenheit. For example, (ctof 20) returns 68.

Cycles in Which the Tutor Communicates with the Student

Code: (defun ctof (cdeg) (*

Tutor : Yes, you will need to call “*” to get the degree on the enlarged scale, but you

are ultimately trying to get the Fahrenheit equivalent of the Celsius degree. So you

need to code a different function here that takes a call to “*” as an argument.

Code: (defun ctof (cdeg) (_

Student: Help.

Tutor: You need to multiply together the Celsius degree reading and the size of

Fahrenheit degrees using the function *. When you multiply these numbers, you will

get the degree on the enlarged scale, which is 36 in our example.

Code: (defun ctof (cdeg) (_ (* 1.8 cdge)

Tutor: Did you mean cdeg?

Student: Yes.

Code: (defun ctof (cdeg) (_ (* 1.8 cdeg) 32))

Tutor: You are done. Type “next” to go on after testing the function you have

defined.

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 268

problems that we expect students to be able to solve. In all, there are about

500 rules that encode the knowledge relating to LISP. A typical rule in the LISP

tutor is:

If the goal is to multiply one number by another,

Then use * and set subgoals to code the two numbers.

The basic goal of the LISP tutor is to communicate these 500 rules to the student,

monitor performance to see whether he or she possesses these rules in

correct form, and provide the student with practice on these rules. The success

of the LISP tutor is one piece of evidence that these 500 rules indeed underlie

coding skill in LISP.

Besides providing an instructional tool, the LISP tutor is a research tool for

studying the course of skill acquisition. The tutor can monitor how well a student

is doing on each of the 500 rules, recording statistics such as the number of

errors that a student is making and the time taken by a student to type the code

corresponding to each of these rules. These data have indicated that students

acquire the skill of LISP by independently acquiring each of the 500 rules.

Figure 9.17 displays the learning curves for these rules. The two dependent

measures are the number of errors made on a rule and the time taken to write

the code corresponding to a rule (when that rule is correctly coded). These

statistics are plotted as a function of learning opportunities, which present

themselves each time the student comes to a point in a problem where that rule

can be applied. As can be seen, performance on these rules dramatically improves

from first to second learning opportunity and improves more gradually

thereafter. These learning curves are similar to those identified in Chapter 6 for

the learning of simple associations.

Individual differences in the learning of these rules have been taken into

account. Students who have already learned a programming language are at a considerable

advantage compared with students for whom their first programming

Educational Implications | 269

1

(a) Opportunities

Number of errors

.20

.50

1.00

2 3−4 5−8

(b)

Coding time (s)

1

5

10

20

Opportunities

2 3−4 5−8

FIGURE 9.17 Data from the LISP tutor: (a) number of errors (maximum is three) per rule as

a function of the number of opportunities for practice; (b) time to correctly code rules as a

function of the amount of practice.

Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 269

language is that of the LISP tutor. The “identical elements model” of transfer,

in which rules for programming in one language transfer to programming in

another language, can account for this advantage.

We also analyzed the performance of individual students in the LISP tutor

and found evidence for two factors. Some students were able to learn new rules

in a lesson quite rapidly, whereas other students had more difficulty. More or

less independent of this acquisition factor, students could be classified according

to how well they retained rules from earlier lessons.3 Thus, students differ in

how rapidly they learn with the LISP tutor. However, the tutor employs a mastery

learning system in which slower students are given more practice and so

are brought to the same level of mastery of the material as that of the others.

Students emerge from their interactions with the LISP tutor having acquired

a complex and sophisticated skill. Their enhanced programming

abilities make them appear more intelligent among their peers. However, when

we examine what underlies that newfound intelligence, we find that it is the

methodical acquisition of some 500 rules of programming. Some students can

acquire these rules more easily than others because of past experience and specific

abilities. However, when they graduate from the LISP course, all students

have learned the 500 new rules.With the acquisition of these rules, few differences

remain among the students with respect to ability to program in LISP.

Thus, we see that, in the end, what is important with respect to individual

differences is how much information students have learned, not their native

ability.

By carefully monitoring individual components of a skill and providing feedback

on learning, intelligent tutors can help students rapidly master complex

skills.

Conclusions

This chapter began by noting the remarkable ability of humans to acquire the

complexities of culture and technology. In fact, in today’s world people can

expect to acquire a whole new set of skills over their lifetimes. For instance, I

now use my phone for instant messaging, GPS navigation, and surfing the

Web—none of which I imagined when I was a young man, let alone associated

with a phone. This chapter has emphasized the role of practice in acquiring

such skills, and certainly it has taken me some considerable practice to master

these new skills. However, human flexibility depends on more than time on

task—other creatures could never acquire such skills no matter how much they

practiced. Critical to human expertise are the higher-order problem-solving

skills that we reviewed in the previous chapter. Also critical is human ability to

reason, make decisions, and communicate by language. These are the topics of

the forthcoming chapters.

270 | Expertise

3 These acquisition and retention factors were strongly related to math SATs, but not to verbal SATs.

Anderson7e_Chapter_09.qxd 8/20/09 9:50 AM Page 270

Key Terms | 271

100

0.50

2.50

1.00

200

Number of books (log scale)

Months to complete a book (log scale)

300 500

FIGURE 9.18 Time to complete a book as a function of

practice, plotted with logarithmic coordinates on both axes.

(From Ohlsson, 1992).

1. An interesting case study of skill acquisition was reported

by Ohlsson (1992), who looked at the development of

Isaac Asimov’s writing skill. Asimov was one of the

most prolific authors of our time, writing approximately

500 books in a career that spanned 40 years.He sat

down at his keyboard every day at 7:30 A.M. and wrote

until 10:00 P.M. Figure 9.18 shows the average number

of months he took to write a book as a function of

practice on a log-log scale. It corresponds closely to a

power function. At what stage of skill acquisition do

you think Asimov was in terms of his writing skills?

2. The chapter discussed how chess experts have learned

to recognize appropriate moves just by looking at the

chessboard. It has been argued (Charness, 1981;

Holding, 1992; Roring, 2008) that experts also learn

to engage in more search and more effective search

for winning moves. Relate these two kinds of learning

(learning specific moves and learning how to search)

to the concepts of tactical and strategic learning.

3. In a 2006 New York Times article, Stephen J. Dubner

and Steven D. Levitt (of “Freakonomics” fame)

noted that elite soccer players are much more likely

to be born in the early months of the year than the

late months. Anders Ericsson argues they have an

advantage in youth soccer leagues, which organize

teams by birth year. Because they are older and tend

to be bigger than other children of the same birth year,

they are more likely to get selected for elite teams and

receive the benefit of deliberate practice. Can you

think of any other explanations for the fact that elite

soccer players tend to be born in the first months

of the year?

4. One reads frequent complaints about the performance

level of American students in studies of mathematics

achievement, where they are greatly outperformed

by children from other countries like Japan. Frequent

remedies point to changing the nature of the mathematics

curriculum or improving teacher quality.

Seldom mentioned is the fact that American children

actually spend much less time learning mathematics

(see Anderson, Reder, & Simon, 1998).What does this

chapter imply about the importance of instruction

versus amount of learning time? Can improvements

in one of these increase American achievement levels

without improvements in the other?

Questions for Thought

Key Terms

associative stage

autonomous stage

cognitive stage

componential analysis

deliberate practice

intelligent tutoring systems

mastery learning

negative transfer

proceduralization

strategic learning

tactical learning

theory of identical elements

Anderson7e_Chapter_09.qxd 8/20/09 9:50 AM Page 271

272

10Reasoning

As noted in Chapter 1, intelligence is thought to be the feature that distinguishes

humans as a species. In the last two chapters, we examined the enormous

capacity that we enjoy as a species to solve problems and acquire new intellectual

skills. In light of this particular capacity, we might expect that the research on human

reasoning (the topic of this chapter) and decision making (the topic of the next

chapter) would document how we achieve our superior intellectual performance.

Historically, however, most psychological research on reasoning and decision making

has started with prescriptions derived from logic and mathematics about how humans

should behave, compared these prescriptions to what humans actually do, and found

humans deficient compared to these standards.

The opposite conclusion seems to come from older research in artificial intelligence

(AI) where researchers tried to create artificial systems for reasoning and

decision making using the same prescriptions from logic and mathematics. For instance,

(Shortliffe, 1976) created an expert computer-based system for diagnosing

infectious diseases. Similar formal reasoning mechanisms were used in the first

generation of robots to help them reason about how to navigate through the

world. Researchers were very frustrated with such systems, noting that they

lacked common sense and would do the stupidest things that no human would do.

Faced with such frustrations, researchers are now creating systems based on less

logical computations, often emulating how neurons in the brain compute (e.g.,

Russell & Norvig, 2003).

Thus, we have a paradox: Human reasoning is judged as deficient when compared

against the standards of logic and mathematics, but AI systems built on these very

standards are judged as deficient when compared against the standards of humans.

This apparent contradiction might lead one to conclude either that logic and mathematics

are wrong or that humans have some mysterious intuition that guides their

thinking. However, the real problem seems to be with the way the principles of logic

and mathematics have been applied, not with the principles themselves. New research

is showing that the situations faced by people are more complex than often assumed.

We can better understand human behavior when we expand our analyses of human

reasoning to include the complexities. In this chapter and the next, we will review a

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 272

Reasoning and the Brain | 273

number of the models used to predict how people arrive at conclusions when presented

with certain evidence, research on how people deviated from these models,

followed by the newer and richer analyses of human reasoning.

This chapter will address the following questions about the way people reason: • How do they reason about situations described in conditional language

(e.g., “if–then”)? • How do they reason about situations described with quantifiers like all, some,

and none? • How do people reason from specific examples and pieces of evidence to

general conclusions?

Reasoning and the Brain

There has been some research investigating brain areas involved in reasoning,

and it suggests that people can bring different systems to bear on different

reasoning problems. Consider an fMRI experiment by Goel, Buchel, Frith, and

Dolan (2000). They had participants solve logical syllogisms, arguments consisting

of two premises and a conclusion, of the type that will be discussed in

the second section of this chapter. Participants were presented with congruent

problems such as

All poodles are pets.

All pets have names.

All poodles have names.

and had to judge whether the third statement followed from the first two. The

content of this example is more or less consistent with what people believe

about pets and poodles. Goel et al. contrasted this type of problem with incongruent

problems whose premises and conclusions violated standard beliefs

such as

All pets are poodles.

All poodles are vicious.

All pets are vicious.

They contrasted both of these types with reasoning about abstract concepts,

such as

All P are B.

All B are C.

All P are C.

Logicians would call all three kinds of syllogism valid.

Participants achieved 84% accuracy in their ability to judge the validity of

the first congruent syllogisms and only 74% in their judgments of the second

incongruent material. They achieved 77% accuracy in their ability to judge the

content-free material. The reader might wonder about the sensibility of judging

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 273

a participant as making a mistake in rejecting an incongruent conclusion such

as “All pets are vicious”; we will return to this matter in the second section of

the chapter. For now, of greater interest are the brain regions that were active

when participants were judging material with content and when they were

judging material without content; these areas are illustrated in Figure 10.1.

When participants were judging content-free material, parietal regions that

have been found to have roles in solving algebraic equations were active (see

Figure 1.16b).When they were judging meaningful content, left prefrontal and

temporal-parietal areas that are associated with language processing were active

(see Figure 4.1).We will find that areas such as the latter regions are frequently

active during reasoning about real problems and that they are associated both

with better performance on some problems and with what might be considered

worse performance on others. What this tells us is that people are capable of

approaching logical problems in two rather different ways.

Faced with logical problems, people can engage either brain regions associated

with the processing of meaningful content or regions associated with the processing

of more abstract information.

Reasoning about Conditionals

The first body of research we will cover concerns deductive reasoning. Deductive

reasoning is concerned with conclusions that follow with certainty from their

premises. It is distinguished from inductive reasoning,which is concerned with

274 | Reasoning

FIGURE 10.1 Comparison of

brain regions activated when

people reason about problems

with meaningful content versus

when they reason about

material without content.

Brain Structures

Posterior parietal:

Reasoning about

content-free material

Ventral prefrontal:

Reasoning about

meaningful content

Parietal-temporal:

Reasoning about

meaningful content

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 274

conclusions that probabilistically follow from their premises. To illustrate the

distinction, suppose someone is told, “Fred is the brother of Mary” and“Mary is

the mother of Lisa.” Then, one might conclude that “Fred is the uncle of Lisa”

and that “Fred is older than Lisa.” The first conclusion, “Fred is the uncle of

Lisa,” would be a correct deductive inference given the definition of familial

relationships. On the other hand, the second conclusion, “Fred is older than Lisa,”

is a good inductive inference, because it is probably true, but not a correct

deductive inference, because it is not necessarily true.

Our first topic will concern human deductive reasoning using the conditional

connective if. A conditional statement is an assertion, such as “If you

read this chapter, you will be wiser.” The if part (you read this chapter) is called

the antecedent and the then part (you will be wiser) is called the consequent.

Table 10.1 lays out the structure of conditional statements and various valid

and invalid rules of inference.We will discuss these rules of inference below.

A particularly central rule of inference in the logic of the conditional is

known as modus ponens (loosely translates from Latin as “method for affirming”) .

It allows us to infer the consequent of a conditional if we are given the

antecedent. Thus, given both the proposition If A, then B and the proposition

A, we can infer B. So, suppose we are told the following premises and

conclusion:

Modus Ponens

If Joan understood this book, then she would get a good grade.

Joan understood this book.

Therefore, Joan got a good grade.

This example is an instance of valid deduction. By valid, we mean that, if premises

1 and 2 are true, then conclusion 3 must be true. This example also illustrates

the artificiality of applying logic to real-world situations. How is one to

really know whether Joan understands the book? One can only assign a certain

probability to her understanding. Even if Joan does understand the book, at

Reasoning about Conditionals | 275

TABLE 10.1

Analysis of a Conditional Statement and Various Valid and Invalid Rules of Inference

A conditional statement:

The antecedent The consequent

(A) (B)

If you read this chapter, you will be wiser.

Name Rule of Inference

Valid deductions Modus ponens Given A is true, infer B is true.

Modus tollens Given B is false, infer A is false.

Invalid deductions Affirmation of the consequent Given B is true, infer A is true.

Denial of the antecedent Given A is false, infer B is false.

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 275

best it is only likely—not certain—that she will get a good grade. However,

participants are asked to suspend their knowledge about such matters and treat

these facts as if they are certainties. Or, more precisely, they are asked to reason

what would follow for certain if these facts were certain.1 Participants do not

find these instructions particularly strange, but, as we will see, they are not

always able to make logically correct inferences.

Another rule of inference is known in logic as modus tollens (loosely translates

as “method of denying”) . This rule states that, if we are given the proposition

A implies B and the fact that B is false, then we can infer that A is false. The

following inference exercise requires modus tollens:

Modus Tollens

If Joan understood this book, then she would get a good grade.

Joan did not get a good grade.

Therefore, Joan did not understand this book.

This conclusion might strike the reader as less than totally compelling because,

again, in the real world such statements are not typically treated as certain.

Modus ponens allows us to infer the consequent from the antecedent; modus

tollens allows us to infer the antecedent is false if the consequent is false.

Evaluation of Conditional Arguments

There are other inference patterns that people sometimes accept but which are

invalid. One is called affirmation of the consequent and is illustrated by the

following incorrect pattern of reasoning.

Affirmation of the Consequent

If Joan understood this book, then she would get a good grade.

Joan got a good grade.

Therefore, Joan did understand this book.

The other incorrect pattern is called denial of the antecedent and is illustrated

by the following pattern of reasoning.

Denial of the Antecedent

If Joan understood this book, then she would get a good grade.

Joan did not understand this book.

Therefore, Joan did not get a grade.

In both of these cases there could be other ways that Joan could get a good

grade, such as writing a great term essay. Evans (1993) reviewed a large

number of studies that compared the frequency with which people accept the

276 | Reasoning

1 Interestingly, the mathematical theory of probability includes conditional statements. In this case, the objects

of the conditional statement are statements about probabilities, which illustrates the fact that precise

mathematics requires the formal logic of the conditional; it should not be taken as an illustration of a way

of incorporating the formal logic of the conditional into a theory of everyday reasoning.

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 276

valid modus ponens and modus tollens inferences as well as the frequency with

which they accept the invalid inferences. The average percent acceptance over

these studies is plotted in Figure 10.2. As can be seen, people rarely fail to

accept a modus ponens inference but the frequency with which they accept the

valid modus tollens is only slightly greater than the frequencies with which

they accept the invalid affirmation of the consequent or the invalid denial of

the antecedent.

People are only able to show high levels of logical reasoning with modus

ponens.

Evaluating Conditional Arguments in a Larger Context

Byrne (1989) performed an interesting variation of the typical conditional reasoning

study that illustrates that human reasoning is sensitive to things that are

ignored in a simple classification like that shown in Table 10.1. In one condition,

she presented her participants with syllogisms like these:

If she has an essay to write, she will study late in the library.

(If she has textbooks to read, she will study late in the library.)

She will stay late in the library.

Therefore, she has an essay to write.

One group of participants did not see the premise in parentheses, whereas the

other group of participants did. Without the additional premise, her participants

accepted the conclusion 71% of the time, committing the fallacy of affirmation

of the consequent. On the other hand, given the parenthetical premise

in addition to the other premises, their acceptance of the conclusion went down

Reasoning about Conditionals | 277

Modus pones Modus tollens

Percent Acceptance

Affirmation of

the consequent

Denial of the

antecedent

20%

0%

40%

60%

80%

100%

Invalid inferences

Valid inferences

FIGURE 10.2 Frequency with which various conditional syllogisms are accepted—from Evans

(1993).

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 277

to 13%. So we see people can be much more accurate in their reasoning if the

material engages them to have a richer interpretation of the situation.

These results of Byrne are even more interesting when compared with another

situation in which she used examples like the following:

If she has an essay to write, she will study late in the library.

(If the library stays open, then she will study in the library.)

She has an essay to write.

Therefore, she will study late in the library.

Without the additional statement in parentheses, participants accepted the

modus ponens inference 96% of the time. However, with the additional statement,

their acceptance rate went down to 38%. In a narrow, logical sense, the

participants are making an error in not accepting the conclusion with the additional

premise. However, in the world outside of the laboratory, they would

be viewed as making the right judgment—how could she actually study in the

library if it were not open? AI researchers would be frustrated if their programs

still made the conclusion with this additional premise. People haveReasoning

As noted in Chapter 1, intelligence is thought to be the feature that distinguishes

humans as a species. In the last two chapters, we examined the enormous

capacity that we enjoy as a species to solve problems and acquire new intellectual

skills. In light of this particular capacity, we might expect that the research on human

reasoning (the topic of this chapter) and decision making (the topic of the next

chapter) would document how we achieve our superior intellectual performance.

Historically, however, most psychological research on reasoning and decision making

has started with prescriptions derived from logic and mathematics about how humans

should behave, compared these prescriptions to what humans actually do, and found

humans deficient compared to these standards.

The opposite conclusion seems to come from older research in artificial intelligence

(AI) where researchers tried to create artificial systems for reasoning and

decision making using the same prescriptions from logic and mathematics. For instance,

(Shortliffe, 1976) created an expert computer-based system for diagnosing

infectious diseases. Similar formal reasoning mechanisms were used in the first

generation of robots to help them reason about how to navigate through the

world. Researchers were very frustrated with such systems, noting that they

lacked common sense and would do the stupidest things that no human would do.

Faced with such frustrations, researchers are now creating systems based on less

logical computations, often emulating how neurons in the brain compute (e.g.,

Russell & Norvig, 2003).

Thus, we have a paradox: Human reasoning is judged as deficient when compared

against the standards of logic and mathematics, but AI systems built on these very

standards are judged as deficient when compared against the standards of humans.

This apparent contradiction might lead one to conclude either that logic and mathematics

are wrong or that humans have some mysterious intuition that guides their

thinking. However, the real problem seems to be with the way the principles of logic

and mathematics have been applied, not with the principles themselves. New research

is showing that the situations faced by people are more complex than often assumed.

We can better understand human behavior when we expand our analyses of human

reasoning to include the complexities. In this chapter and the next, we will review a

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 272

Reasoning and the Brain | 273

number of the models used to predict how people arrive at conclusions when presented

with certain evidence, research on how people deviated from these models,

followed by the newer and richer analyses of human reasoning.

This chapter will address the following questions about the way people reason: • How do they reason about situations described in conditional language

(e.g., “if–then”)? • How do they reason about situations described with quantifiers like all, some,

and none? • How do people reason from specific examples and pieces of evidence to

general conclusions?

Reasoning and the Brain

There has been some research investigating brain areas involved in reasoning,

and it suggests that people can bring different systems to bear on different

reasoning problems. Consider an fMRI experiment by Goel, Buchel, Frith, and

Dolan (2000). They had participants solve logical syllogisms, arguments consisting

of two premises and a conclusion, of the type that will be discussed in

the second section of this chapter. Participants were presented with congruent

problems such as

All poodles are pets.

All pets have names.

All poodles have names.

and had to judge whether the third statement followed from the first two. The

content of this example is more or less consistent with what people believe

about pets and poodles. Goel et al. contrasted this type of problem with incongruent

problems whose premises and conclusions violated standard beliefs

such as

All pets are poodles.

All poodles are vicious.

All pets are vicious.

They contrasted both of these types with reasoning about abstract concepts,

such as

All P are B.

All B are C.

All P are C.

Logicians would call all three kinds of syllogism valid.

Participants achieved 84% accuracy in their ability to judge the validity of

the first congruent syllogisms and only 74% in their judgments of the second

incongruent material. They achieved 77% accuracy in their ability to judge the

content-free material. The reader might wonder about the sensibility of judging

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 273

a participant as making a mistake in rejecting an incongruent conclusion such

as “All pets are vicious”; we will return to this matter in the second section of

the chapter. For now, of greater interest are the brain regions that were active

when participants were judging material with content and when they were

judging material without content; these areas are illustrated in Figure 10.1.

When participants were judging content-free material, parietal regions that

have been found to have roles in solving algebraic equations were active (see

Figure 1.16b).When they were judging meaningful content, left prefrontal and

temporal-parietal areas that are associated with language processing were active

(see Figure 4.1).We will find that areas such as the latter regions are frequently

active during reasoning about real problems and that they are associated both

with better performance on some problems and with what might be considered

worse performance on others. What this tells us is that people are capable of

approaching logical problems in two rather different ways.

Faced with logical problems, people can engage either brain regions associated

with the processing of meaningful content or regions associated with the processing

of more abstract information.

Reasoning about Conditionals

The first body of research we will cover concerns deductive reasoning. Deductive

reasoning is concerned with conclusions that follow with certainty from their

premises. It is distinguished from inductive reasoning,which is concerned with

274 | Reasoning

FIGURE 10.1 Comparison of

brain regions activated when

people reason about problems

with meaningful content versus

when they reason about

material without content.

Brain Structures

Posterior parietal:

Reasoning about

content-free material

Ventral prefrontal:

Reasoning about

meaningful content

Parietal-temporal:

Reasoning about

meaningful content

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 274

conclusions that probabilistically follow from their premises. To illustrate the

distinction, suppose someone is told, “Fred is the brother of Mary” and“Mary is

the mother of Lisa.” Then, one might conclude that “Fred is the uncle of Lisa”

and that “Fred is older than Lisa.” The first conclusion, “Fred is the uncle of

Lisa,” would be a correct deductive inference given the definition of familial

relationships. On the other hand, the second conclusion, “Fred is older than Lisa,”

is a good inductive inference, because it is probably true, but not a correct

deductive inference, because it is not necessarily true.

Our first topic will concern human deductive reasoning using the conditional

connective if. A conditional statement is an assertion, such as “If you

read this chapter, you will be wiser.” The if part (you read this chapter) is called

the antecedent and the then part (you will be wiser) is called the consequent.

Table 10.1 lays out the structure of conditional statements and various valid

and invalid rules of inference.We will discuss these rules of inference below.

A particularly central rule of inference in the logic of the conditional is

known as modus ponens (loosely translates from Latin as “method for affirming”) .

It allows us to infer the consequent of a conditional if we are given the

antecedent. Thus, given both the proposition If A, then B and the proposition

A, we can infer B. So, suppose we are told the following premises and

conclusion:

Modus Ponens

If Joan understood this book, then she would get a good grade.

Joan understood this book.

Therefore, Joan got a good grade.

This example is an instance of valid deduction. By valid, we mean that, if premises

1 and 2 are true, then conclusion 3 must be true. This example also illustrates

the artificiality of applying logic to real-world situations. How is one to

really know whether Joan understands the book? One can only assign a certain

probability to her understanding. Even if Joan does understand the book, at

Reasoning about Conditionals | 275

TABLE 10.1

Analysis of a Conditional Statement and Various Valid and Invalid Rules of Inference

A conditional statement:

The antecedent The consequent

(A) (B)

If you read this chapter, you will be wiser.

Name Rule of Inference

Valid deductions Modus ponens Given A is true, infer B is true.

Modus tollens Given B is false, infer A is false.

Invalid deductions Affirmation of the consequent Given B is true, infer A is true.

Denial of the antecedent Given A is false, infer B is false.

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 275

best it is only likely—not certain—that she will get a good grade. However,

participants are asked to suspend their knowledge about such matters and treat

these facts as if they are certainties. Or, more precisely, they are asked to reason

what would follow for certain if these facts were certain.1 Participants do not

find these instructions particularly strange, but, as we will see, they are not

always able to make logically correct inferences.

Another rule of inference is known in logic as modus tollens (loosely translates

as “method of denying”) . This rule states that, if we are given the proposition

A implies B and the fact that B is false, then we can infer that A is false. The

following inference exercise requires modus tollens:

Modus Tollens

If Joan understood this book, then she would get a good grade.

Joan did not get a good grade.

Therefore, Joan did not understand this book.

This conclusion might strike the reader as less than totally compelling because,

again, in the real world such statements are not typically treated as certain.

Modus ponens allows us to infer the consequent from the antecedent; modus

tollens allows us to infer the antecedent is false if the consequent is false.

Evaluation of Conditional Arguments

There are other inference patterns that people sometimes accept but which are

invalid. One is called affirmation of the consequent and is illustrated by the

following incorrect pattern of reasoning.

Affirmation of the Consequent

If Joan understood this book, then she would get a good grade.

Joan got a good grade.

Therefore, Joan did understand this book.

The other incorrect pattern is called denial of the antecedent and is illustrated

by the following pattern of reasoning.

Denial of the Antecedent

If Joan understood this book, then she would get a good grade.

Joan did not understand this book.

Therefore, Joan did not get a grade.

In both of these cases there could be other ways that Joan could get a good

grade, such as writing a great term essay. Evans (1993) reviewed a large

number of studies that compared the frequency with which people accept the

276 | Reasoning

1 Interestingly, the mathematical theory of probability includes conditional statements. In this case, the objects

of the conditional statement are statements about probabilities, which illustrates the fact that precise

mathematics requires the formal logic of the conditional; it should not be taken as an illustration of a way

of incorporating the formal logic of the conditional into a theory of everyday reasoning.

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 276

valid modus ponens and modus tollens inferences as well as the frequency with

which they accept the invalid inferences. The average percent acceptance over

these studies is plotted in Figure 10.2. As can be seen, people rarely fail to

accept a modus ponens inference but the frequency with which they accept the

valid modus tollens is only slightly greater than the frequencies with which

they accept the invalid affirmation of the consequent or the invalid denial of

the antecedent.

People are only able to show high levels of logical reasoning with modus

ponens.

Evaluating Conditional Arguments in a Larger Context

Byrne (1989) performed an interesting variation of the typical conditional reasoning

study that illustrates that human reasoning is sensitive to things that are

ignored in a simple classification like that shown in Table 10.1. In one condition,

she presented her participants with syllogisms like these:

If she has an essay to write, she will study late in the library.

(If she has textbooks to read, she will study late in the library.)

She will stay late in the library.

Therefore, she has an essay to write.

One group of participants did not see the premise in parentheses, whereas the

other group of participants did. Without the additional premise, her participants

accepted the conclusion 71% of the time, committing the fallacy of affirmation

of the consequent. On the other hand, given the parenthetical premise

in addition to the other premises, their acceptance of the conclusion went down

Reasoning about Conditionals | 277

Modus pones Modus tollens

Percent Acceptance

Affirmation of

the consequent

Denial of the

antecedent

20%

0%

40%

60%

80%

100%

Invalid inferences

Valid inferences

FIGURE 10.2 Frequency with which various conditional syllogisms are accepted—from Evans

(1993).

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 277

to 13%. So we see people can be much more accurate in their reasoning if the

material engages them to have a richer interpretation of the situation.

These results of Byrne are even more interesting when compared with another

situation in which she used examples like the following:

If she has an essay to write, she will study late in the library.

(If the library stays open, then she will study in the library.)

She has an essay to write.

Therefore, she will study late in the library.

Without the additional statement in parentheses, participants accepted the

modus ponens inference 96% of the time. However, with the additional statement,

their acceptance rate went down to 38%. In a narrow, logical sense, the

participants are making an error in not accepting the conclusion with the additional

premise. However, in the world outside of the laboratory, they would

be viewed as making the right judgment—how could she actually study in the

library if it were not open? AI researchers would be frustrated if their programs

still made the conclusion with this additional premise. People have a

rich ability to reason about the real world, and it can intrude and cause them

to make errors in these studies where they are told to reason by the strict rules

of logic. However, it can lead them to make the right decisions in the real

world.

When people’s ability to reason about real-world situations intrudes into

logical reasoning tasks, it can result in better or worse performance.

The Wason Selection Task

A series of experiments initially begun by Peter Wason (for a review of the early

research, see Wason & Johnson-Laird, 1972, Chapters 13 and 14) have been

taken as a striking demonstration of human inability to reason correctly. In a

typical experiment in this research, four cards showing the following symbols

were placed in front of participants:

278 | Reasoning

E K 4 7

Participants were told that a letter appeared on one side of each card and a

number on the other. Their task was to judge the validity of the following rule,

which referred only to these four cards:

If a card has a vowel on one side, then it has an even number on the other side.

The participants’ task was to turn over only those cards that had to be turned

over for the correctness of the rule to be judged. This task, typically referred to

as the selection task, has received a great deal of research.

Averaging over a large number of experiments (Oaksford & Chater, 1994),

about 90% of the participants have been found to select E, which is a logically

correct choice because an odd number on the other side would disconfirm the

rule. However, about 60% of the participants also choose to turn over the 4,

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 278

which is not logically informative because neither a vowel nor a consonant on

the other side would have falsified the rule. Only 25% elect to turn over the 7,

which is a logically informative choice because a vowel behind the 7 would have

falsified the rule. Only about 15% elect to turn over the K, which would not be

an informative choice.

Thus, participants display two types of logical errors in the task. First, they

often turn over the 4, an example of the fallacy of affirming the consequent.

Even more striking is the failure to take the modus tollens step of disconfirming

the consequent and determining whether the antecedent also is disconfirmed

(in other words, to turn over the 7).

The number of people that make the right combination of choices, turning

over only the E and 7, is often only 10%, which has been taken as a damning

indictment of human reasoning. Early in the history of research on the selection

task, Wason gave a talk at the IBM Research Center in which he presented this

same problem to an audience filled with Ph.D.s, many in mathematics and

physics.He got the same poor results from this audience and, reportedly, they were

so embarrassed that they harassed Wason with complaints about how the problem

was not accurately presented or the correct answer was not really correct. This

question of what the right answer is has been recently explored but, before considering

it, we will see what happens when one puts content into these problems.

When presented with neutral material in the Wason selection task, people

have particular difficulty in recognizing the importance of exploring if the

consequent is false.

Permission Interpretation of the Conditional

A person’s performance can sometimes be greatly enhanced when the material

to be judged has meaningful content. Griggs and Cox (1982) were among the

first to demonstrate this enhancement in a paradigm that is formally equivalent

to the Wason card-selection task. Participants were instructed to imagine that

they were police officers responsible for ensuring that the following regulation

was being followed: If a person is drinking beer, then the person must be over 19.

They were presented with four cards that represented people sitting around a

table. On one side of each card was the age of the person and on the other side

was the substance that the person was drinking. The cards were labeled “Drinking

beer,”“Drinking Coke,” “16 years of age,” and “22 years of age.” The task was

to select those people (cards to turn over) from whom further information was

needed to determine whether the drinking law was being violated. In this situation,

74% of the participants selected the logically correct cards (namely,

“Drinking beer” and “16 years of age”). Interestingly, patients with damage to

the ventromedial prefrontal cortex do not show this advantage with content

(Adolphs, Tranel, Bechara, Damasio, & Damasio, 1996). We will discuss this

patient population more thoroughly in the next chapter.

It has been argued that the better performance in this task depends on the

fact that the conditional statement is being interpreted as a rule about a social

norm called the permission schema. Society has many rules about how its

Reasoning about Conditionals | 279

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 279

members should conduct themselves, and the argument is that people are good

at applying such social rules (Cheng & Holyoak, 1985). An alternate possibility

is that better performance in this task depends not on the permission semantics

but on the greater familiarity of the participants with the rule. The participants

were Florida undergraduates, and this rule about drinking was in force in

Florida at the time. Would the participants have been able to reason about a

similar but unfamiliar law? To discriminate between these two possibilities,

Cheng and Holyoak (1985) performed the following experiment. One group of

participants was asked to evaluate the following apparently senseless rule

against a set of instances: “If the form says ‘entering’ on one side, then the other

side includes cholera among the list of diseases.” Another group was given the

same rule as well as the rationale that to satisfy immigration officials upon entering

a particular country, one must have been vaccinated for cholera. This rationale

should invoke people’s ability to reason about the permission schema.

The forms indicated on one side whether the passenger was entering the country

or in transit, whereas the other side listed the names of diseases for which he

or she was vaccinated. Participants were presented with a set of forms that said

“Transit,” or “Entering,” or “cholera, typhoid, hepatitis,” or “typhoid, hepatitis.”

The performance of the group given the rationale was much better than that of

the group given just the senseless rule without explanation; that is, the former

group knew to check the “Entering” form and the “typhoid, hepatitis” form.

Because the participants were not familiar with the rule, their good performance

apparently depended on evoking the concept of permission and not on practice

in applying the specific rule.

Cosmides (1989) and Gigerenzer and Hug (1992) argued that our good performance

with such rules (which they call social contract rules) depends on the

skill with which we have learned to detect cheaters. Gigerenzer and Hug had

participants evaluate the following rule:

If a student is assigned to Grover High School, then that student must live in

Grover City.

They saw cards that stated whether the students attended Grover High or not

on one side and whether they lived in Grover city or not on the other side. As in

the original Wason experiment, they had to decide which cards to turn over. In

the cheating condition, participants were asked to take the perspective of a

member of the Grover City School Board looking for students who were illegally

attending the high school. In the non-cheating condition, participants

were asked to take the perspective of a visiting official from the German government

who just wants to find out whether this rule is in effect at Grover High

School. Gigerenzer and Hug were interested in the frequency with which participants

would choose to turn over and check just the two cards: the student

marked as going to Grover High School and the student marked as a nonresident

of Grover City, which are the logically correct choices. In the cheating

condition, where they took the perspective of a school board member, 80% of

the participants chose just these two cards, replicating other results with permission

rules. In the no-cheating condition, where they took the perspective of

a disinterested visitor, only 45% of the participants chose just these two.

280 | Reasoning

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 280

When participants take the perspective of detecting whether a social rule

has been violated, they make a large proportion of logically correct choices

in the Wason card selection task.

Probabilistic Interpretation of the Conditional

The research just reviewed demonstrates that people can show good reasoning

when they adopt what is called the permission interpretation of the conditional.

However, how are we to understand their poor performance in the original

Wason demonstrations where participants are not taking this permission

interpretation? Oaksford and Chater (1994) argued that people tend to interpret

these statements not as strict logical statements but rather as probabilistic

statements about the world. Thus, when someone says, “If A, then B,” they

mean that B will probably occur when A occurs. Even more important to the

Oaksford and Chater argument is the idea that events A and B typically have

low probabilities of occurring in the world—which is what makes such statements

informative. To illustrate their argument, suppose you visited a city and a

friend told you that the following rule held about the cars driving in that city:

If a car has a broken headlight, it will have a broken taillight.

Events A and B (broken headlights and taillights) are rare, and consequently

asserting that one implies the other is informative. Suppose you go to a large

parking lot in which there are hundreds of cars; some are parked with their

fronts exposed and others with their rears exposed. Most do not have their

headlights broken or their taillights broken, but there are one or two with

broken headlights and taillights. On which cars would you check the end not

exposed to test your friend’s claim? Let us consider the following possibilities:

1. A car with a broken headlight: If you saw such a car, like participants in all

of these experiments, you would be inclined to check its taillight. Almost

everyone sees that it is the sensible thing to do.

2. A car without a broken headlight: You would not be inclined to check this

car, like most of the participants in these experiments, and, again,

everyone agrees that you are right.

3. A car with a broken taillight: You would be sorely tempted to see whether

that car did not have a broken headlight (despite the fact that it is supposedly

unnecessary or “illogical”), and Oaksford and Chater agree with

you. The reason is that a car with a broken taillight is so rare that, if it did

have a broken headlight, you would be inclined to believe your friend’s

claim. The coincidence would be too much to shrug off.

4. A car without a broken taillight: You would be reluctant to check every car

in the lot that met this condition (despite the fact that it is supposedly the

logical thing to do), and, again, Oaksford and Chater would agree with

you. The odds of finding a broken headlight on such a car are low because

a broken headlight is rare, and so many cars would have to be checked.

Checking those hundreds of normal cars just does not seem worthwhile.

Reasoning about Conditionals | 281

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 281

Oaksford and Chater developed a mathematical analysis of the optimal

behavior that explains why the typical errors in the original Wason task can

be sensible. Their analysis predicts the frequency of choices in the Wason task.

Their analysis depends on the assumption that properties such as “broken

headlight” and “broken taillight” are rare. For this reason, it is informative

to check the car with a broken taillight as in possibility 3 and is rather

uninformative to check a car without a broken taillight as in possibility 4.

Although the properties might not always be as rare as in this example,

Oaksford and Chater argued that they generally are rare. For instance, more

things are not dogs than are dogs and more things don’t bark than do, and so

the same analysis would apply to a rule such as “If an animal is a dog, then it

will bark” and many other such rules. There is a weakness in the Oaksford

and Chater argument, however, when applied to the original Wason experiment

where the participants were reasoning about even numbers: There are

not more odd numbers than even numbers. Nonetheless, Oaksford argued

that people carry their beliefs that properties are rare into the Wason

situation. There is evidence that manipulations of the probabilities of these

properties do change people’s behavior in the expected way (Oaksford &

Wakefield, 2003).

The behavior in the Wason card selection task can be explained if we assume

that participants select cards that will be informative under a probabilistic

model.

Final Thoughts on the Connective If

The logical connective if can evoke many different interpretations, which reflect

the richness of human cognition. We have considered evidence for its probabilistic

interpretation and its permission interpretation. People are capable of

adopting the logician’s interpretation of it as well, which, not surprisingly, is the

interpretation that logicians and students of logic take of it when doing logic.

Studies of their reasoning (Lewis, 1985; Scheines & Sieg, 1994) with the connective

find it to be similar to mathematical reasoning such as in the domain of

geometry discussed in Chapter 9. Basically, they take a problem-solving approach

to formal reasoning with the connective. Qin et al. (2003) looked at participants

solving abstract logic tasks and found activation in the same parietal

regions (see Figure 10.1) that Goel et al. (2000) found active with their contentfree

material.

An amusing result is that training in logic does not necessarily result in

better behavior on the original Wason selection task. In a study by Cheng,

Holyoak, Nisbett, and Oliver (1986), college students who had just taken a semester

course in logic did only 3% better on the card selection task than those

who had no formal training in logic. It was not that they did not know the rules

of logic; rather, they did choose to apply them in the experiment. When presented

with these problems outside of the logic classroom, the students chose to

adopt some other interpretation of the word if. However, this is not necessarily

a “flaw” in human reasoning. To repeat a point made before, many researchers

282 | Reasoning

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 282

in AI wish their programs were as adaptive in how they interpret the information

they are presented.

People use different problem-solving operators, depending on their interpretation

of the logical connective if.

Deductive Reasoning:

Reasoning about Quantifiers

Much of human knowledge is cast with logical quantifiers such as all or some.

Witness Lincoln’s famous statement: “You may fool all the people some of the

time; you can even fool some of the people all the time; but you can’t fool all of the

people all the time.” Scientific laws such as Newton’s third law, “For every action

there is always an opposite and equal reaction,” try to identify what is always the

case. It is important to understand how we reason with such quantifiers. This section

will report research on how people reason about such quantifiers when they

appear in simple sentences.As was the case for the logical connective if, we will see

that there are differences between the logician’s interpretation of quantifiers and

the way in which people frequently reason about them.

The Categorical Syllogism

Modern logic is greatly concerned with analyzing the meaning of quantifiers

such as all, no, and some, as in, for example:

All philosophers read some books.

Most of us might believe that this statement is true. The logician would then

say that we were committed to the belief that we could not find a philosopher

who did not read books, but most of us have no trouble accepting the idea that

there were philosophers in societies before there were books or that one still

might find somewhere in the world an illiterate person who professed sufficiently

profound ideas to deserve the title of “philosopher.” This example illustrates

the fact that frequently when we use all in real life, we mean “most” or

“with high probability.” Similarly, when we use no as in

No doctors are poor.

we often mean “hardly any” or “with small probability.” Logicians call both the all

and no statements universal statements because they interpret these statements

as blanket claims with no exceptions. Roger Schank, a famous AI researcher, was

once observed to make the assertion

No one uses universals.

which surely is a sign that people use these words in a richer or more complex

way than implied by the logical analysis.

By the beginning of the 20th century, the sophistication with which logicians

analyzed such quantified statements increased considerably (see Church, 1956,

Deductive Reasoning: Reasoning about Quantifiers | 283

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 283

for a historical discussion). This more advanced treatment of quantifiers is covered

in most modern logic courses. However, most of the research on quantifiers

in psychology has focused on a simpler and older kind of quantified deduction,

called the categorical syllogism. Much of Aristotle’s writing on reasoning concerned

the categorical syllogism. Extensive discussion of these types of syllogisms

can be found in old textbooks on logic such as Cohen and Nagel (1934).

Categorical syllogisms include statements containing the quantifiers some,

all, no, and some–not. Examples of such categorical statements are:

1. All doctors are rich.

2. Some lawyers are dishonest.

3. No politician is trustworthy.

4. Some actors are not handsome.

As a convenient shorthand, the categories (e.g., doctors, rich people, lawyers,

dishonest people) in such statements can be represented by letters—say, A, B, C,

and so on. Thus, the statements might be rendered in this way:

1. All A’s are B’s.

2. Some C’s are D’s.

3. No E’s are F’s.

4. Some G’s are not H’s.

Sometimes, as in the Goel et al. experiment described at the beginning of the

chapter, material is actually presented with such letters.

A categorical syllogism typically contains two premises and a conclusion.

A typical example that might be used in research follows:

1. No Pittsburgher is a Browns fan.

All Browns fans live in Cleveland.

No Pittsburgher lives in Cleveland.

Many people accept this syllogism as logically valid. To see that the conclusion

does not necessarily follow from the form of the premises, consider the following

equivalent syllogism:

2. No man is a woman.

Every woman is a human.

No man is a human.

The first example illustrates a frequent result in research on categorical syllogisms,

which is that people often accept invalid syllogisms. For instance, people accept

the invalid syllogism 1 almost asmuch as they do the following valid syllogism:

3. No Pittsburgher lives in Cleveland.

All Browns fans live in Cleveland.

‹ No Pittsburgher is a Browns fan.

284 | Reasoning

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 284

Research on reasoning with quantifiers has focused on trying to understand

why people accept many invalid categorical syllogisms.

The Atmosphere Hypothesis

The previous example 1 is a case where people are biased by the content of the

syllogism, but much of the research has focused on the tendency of people to

accept invalid syllogisms even when they have neutral content. People are generally

good at recognizing valid syllogisms when stated with neutral content.

For instance, almost everyone accepts

1. All A’s are B’s.

All B’s are C’s.

All A’s are C’s.

The problem is that people also accept many invalid syllogisms. For instance,

many people will accept

2. Some A’s are B’s.

Some B’s are C’s.

Some A’s are C’s.

(To see that this syllogism is invalid, consider replacing A with men, B with

humans, and C with women.) However, people are not completely indiscriminate

in what they accept as valid. For instance, even though they will accept example 2

in the preceding subsection, they will not accept example 3:

3. Some A’s are B’s.

Some B’s are C’s.

No A’s are C’s.

To account for the pattern of what participants accept and what they reject,

Woodworth and Sells (1935) proposed the atmosphere hypothesis. This

hypothesis states that the logical terms (some, all, no, and not) used in the

premises of a syllogism create an “atmosphere” that predisposes participants to

accept conclusions having the same terms. The atmosphere hypothesis consists

of two parts. One part asserts that participants tend to accept a positive conclusion

to positive premises and a negative conclusion to negative premises.When

the premises are mixed, participants tend to prefer a negative. Thus, they would

tend to accept the following invalid syllogism:

4. No A’s are B’s.

All B’s are C’s.

No A’s are C’s.

(This is an abstract form of the same invalid syllogism discussed on the previous

page.)

The other part of the atmosphere hypothesis concerns a participant’s response

to particular statements (some or some not) versus universal statements (all

or no). As example 4 illustrates, participants will accept a universal conclusion

if the premises are universal. They will tend to accept a particular conclusion if

the premises are particular, which accounts for their acceptance of syllogism 2

Deductive Reasoning: Reasoning about Quantifiers | 285

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 285

given earlier. When one premise is particular and the other universal, participants

prefer a particular conclusion. Thus they will accept the following invalid

syllogism:

5. All A’s are B’s.

Some B’s are C ’s.

Some A’s are C ’s.

(To see that this syllogism is invalid, consider replacing A with men, B with

humans, and C with women.)

The atmosphere hypothesis states that the logical terms (some, all, no,

and some not) used in the premises of a syllogism create an “atmosphere”

that predisposes participants to accept conclusions having the same terms.

Limitations of the Atmosphere Hypothesis

The atmosphere hypothesis provides a succinct characterization of participant

behavior with the various syllogisms, but it tells us little about what the participants

are actually thinking or why. It offers no explanation for why the content

of the syllogism (as in the Pittsburgh–Cleveland example) can have such a

strong effect on judgments. Its characterization of participant behavior is also

not always correct for content-free syllogisms. For example, according to the

atmosphere hypothesis, participants are just as likely to accept the atmospherefavored

conclusion when it is not valid as when it is valid. That is, it predicts

that participants would be just as likely to accept

6. All A’s are B’s.

Some B’s are C ’s.

Some A’s are C’s.

which is not valid, as they would be to accept

7. Some A’s are B’s.

All B’s are C’s.

Some A’s are C ’s.

which is valid. In fact, participants are more likely to accept the conclusion in

the valid case. Thus, participants do display some ability to evaluate a syllogism

accurately.

Another limitation of the atmosphere hypothesis is that it fails to predict

the effects that the form of a syllogism will have on participants’ validity judgments.

For instance, the hypothesis predicts that participants would be no more

likely to erroneously accept

8. Some A’s are B’s.

Some B’s are C ’s.

Some A’s are C ’s.

than they would be to erroneously accept

9. Some B’s are A’s.

Some C ’s are B’s.

‹ Some A’s are C ’s.

286 | Reasoning

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 286

In fact, participants are more willing to erroneously accept the conclusion in

the former case (Johnson-Laird & Steedman, 1978). In general, participants are

more willing to accept a conclusion from A to C if they can find a chain leading

from A to B in one premise and from B to C in the second premise.

Another problem with the atmosphere hypothesis is that it does not really

handle what participants do in the presence of two negatives. If participants are

given the following two premises,

No A’s are B’s.

No B’s are C ’s.

the atmosphere hypothesis would predict that participants should tend to accept

the invalid conclusion:

No A’s are C ’s.

Although a few participants do accept this conclusion, most refuse to accept

any conclusion when both premises are negative, which is the correct thing to

do (Dickstein, 1978).

Another problem with the atmosphere hypothesis is that it does not really

explain what people are thinking when they process such syllogisms. It just tries

to predict what conclusions they will accept. The next section will consider

some explanations of the thought processes that lead people to correct or

incorrect conclusions.

Participants only approximate the predictions of the atmosphere hypothesis

and are often more accurate than it would predict.

Process Explanations

One class of explanations is that participants choose not to do what the experimenters

think they are doing. For instance, it has been argued that it is not natural

for people to judge the logical validity of these arguments. Rather, people tend

to judge what conclusions are true. Consider the following pair of syllogisms:

All lawyers are human.

All Republicans are human.

Some lawyers are Republicans.

which has a true conclusion but is not valid (consider replacing lawyers by men

and Republicans by women) and

All bictoids are reptiles.

All bictoids are birds.

Some reptiles are birds.

which is a valid argument but has a false conclusion. People have a greater

tendency to accept the first, invalid argument having a true conclusion than

the second, valid argument having a false conclusion (Evans, Handley, &

Harper, 2001).

It is also argued that many people really do not understand what it means

for an argument to be valid and simply judge whether a conclusion is possible

Deductive Reasoning: Reasoning about Quantifiers | 287

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Pagegiven the premises. So, for example, although the preceding syllogism concerning

lawyers and Republicans is not valid, it is certainly possible given the

premises that the conclusion is true. Evans et al. showed that there is very

little difference in the judgments that participants make when they are asked

to judge when conclusions are necessarily true given the premises (the measure

of a valid argument) and when conclusions are possibly true given the

premises.

Johnson-Laird (Johnson-Laird, 1983; Johnson-Laird & Steedman, 1978)

proposed that participants judge whether a conclusion is possible by creating a

mental model of a world that satisfies the premises of the syllogism and inspecting

that model to see whether the conclusion is satisfied. This explanation

is called mental model theory. Consider these premises:

All the squares are striped.

Some of the striped objects have bold borders.

Figure 10.3a illustrates what a participant might imagine, according to Johnson-

Laird, as an instantiation of these premises. The participant has imagined a

group of objects, some of which are square, whereas others are round; some

of which are striped, whereas others are clear; and some of which have bold

borders, whereas others do not. This world represents one possible interpretation

of these premises. When the participant is asked to judge the following

conclusion,

Some of the squares have bold borders.

the participant inspects the diagram and sees that, indeed, the conclusion is

possible given the premises. The problem is that this response establishes only

that the conclusion is possible but not that it is necessary. For the conclusion

to be necessary, it must be true in all mental models that support

the premises. Figure 10.3b illustrates a case in which the

premises are true but the conclusion does not hold. Johnson-

Laird claimed that participants have considerable difficulty

developing alternative models. Thus, the participant is building

a specific model for the premises and is inspecting it to see what

is true in that model. Johnson-Laird (1983) developed a computer

simulation of this theory that reproduces many of the

errors that participants make. Johnson-Laird (1995) also argued

that there is neurological evidence in favor of the mental model

explanation. He noted that patients with right-hemisphere

damage are more impaired in reasoning tasks than are patients

with left-hemisphere damage. He noted that the right hemisphere

tends to take part in spatial processing of mental images.

In a brain-imaging study, Kroger, Cohen, Nystrom, and

Johnson-Laird (2008) found that the right frontal cortex was

more active than the left in processing such syllogisms but that

the opposite was true when people engaged in arithmetic calculation

(this left bias for arithmetic is also illustrated in the study

288 | Reasoning

(a)

(b)

FIGURE 10.3 Two possible

models that participants might

form for the premises of the

categorical syllogism dealing

with square and round objects.

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 288

described in Figure 1.16). Parsons and Osherson (2001) reported a similar finding,

with deductive reasoning being rightlocalized and probabilistic reasoning

being left localized.

Basically, Johnson-Laird’s argument is that people make errors in reasoning

because they overlook possible explanations of the premises. For example, a

participant imagines Figure 10.3a as an explanation and overlooks the possibility

of Figure 10.3b. Johnson-Laird (personal communication) argues that a great

many errors in human reasoning are produced by failures to consider possible

explanations of the data. For instance, a problem in the Chernobyl disaster was

that, for several hours, engineers failed to consider the possibility that the reactor

was no longer intact.

Evans et al. (2001) argued that people are inclined to look for only one

mental model for the premises and, if they can find such an interpretation, they

accept the argument if the conclusion is valid in that model. The effect of realworld

content is to bias that search for such a mental model. In this way, we can

explain the effect of real-world content. That is, people tend to accept the

invalid “Some lawyers are Republicans” syllogism because it naturally suggests a

model in which the conclusion holds. In contrast, they tend to reject the valid

“Some reptiles are birds” argument because, given its content, it is hard to

imagine a world where it is true.

To successfully adapt to our world, it is important to make correct inferences

about what is true in this world and not about what might be true in all

possible worlds (which is what judging logical validity is concerned with).

Given this perspective, the problem seems to lie as much with the questions

that participants are being asked to answer in these experiments as with the

participants themselves.

Errors in evaluating syllogisms can be explained by assuming that participants

fail to consider possible mental models of the syllogisms.

Inductive Reasoning and Hypothesis Testing

In contrast to deductive reasoning where logical rules allow one to infer certain

conclusions from premises, in inductive reasoning the conclusion does not necessarily

follow from the premises. Consider the following premises:

The first number in the series is 1.

The second number in the series is 2.

The third number in the series is 4.

What conclusion follows? The numbers are doubling and so possible conclusion

is that

The fourth number is 8.

However, a better conclusion might be to state the general rule:

Each number is twice the previous number.

Inductive Reasoning and Hypothesis Testing | 289

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 289

A characteristic of a good inductive inference like the second conclusion is that it

is a statement from which one can deduce all the premises. For example, because

we know each number is twice the previous number, we can now deduce what the

original three numbers must have been. Thus, in a certain sense it is deduction

turned around. The difficulty for inductive reasoning is that there is usually not a

single conclusion that would be consistent with the premises. For instance, in the

problem above one could have concluded that the difference between successive

numbers is increasing by one and that the fourth number would be 7.

Inductive reasoning is relevant to many aspects of everyday life: a detective

trying to solve a mystery given a set of clues, a doctor trying to diagnose the cause

of a set of symptoms, someone trying to determine what is wrong with a TV, or a

researcher trying to discover a new scientific law. In all these cases, one gets a set

of specific observations from which one is trying to infer some relevant conclusion.

Many of these cases involve the sort of probabilistic reasoning that will be

discussed in the next chapter (for instance, medical symptoms are typically only

associated probabilistically with disease). In this chapter, we will focus on cases,

like the above number example, where we are looking for a hypothesis that

implies the observations with certainty.Much of the interest in such cases is how

people seek evidence relevant to formulating such a hypothesis.

Hypothesis Formation

Bruner, Goodnow, and Austin (1956) performed a classic series of experiments

on hypothesis formation. Figure 10.4 illustrates the kind of material they used.

The stimuli were all rectangular boxes containing various objects. The stimuli

varied on four dimensions: number of objects (one, two, or three); number of

borders around the boxes (one, two, or three); shape (cross, circle, or square);

290 | Reasoning

FIGURE 10.4 Material used

by Bruner et al. in one of

their studies of concept

identification. The array

consists of instances formed by

combinations of four attributes,

each exhibiting three values.

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 290

and color (green, black, or red: represented here as white, black, or

red). Participants were told that they were to discover some concept

that described a particular subset of these instances. For instance,

the concept might have been black crosses. Participants were to discover

the correct concept on the basis of information they were

given about what were and what were not instances of the concept.

Figure 10.5 contains three illustrations (the three columns) of

the information participants might have been presented. Each column

consists of a sequence of instances identified either as members

of the concept (positive, _) or not (negative, _). Each column

represents a different concept. Participants would be presented with

the instances in a column one at a time. From these instances they

would determine what the concept was. Stop reading and try to

determine the concept for each column.

The concept in the first example is two crosses. This concept is referred

to as a conjunctive concept, since the conjunction of a number

of features (in this case the features are two and cross) must be present

for the instance to be positive. People typically find conjunctive

concepts easiest to discover. In some sense conjunctive hypotheses

seem to be the most natural kind of hypotheses. They are the kind

that have been researched most extensively. The solution to the second

example is two borders or circles. This kind of concept is referred to as

a disjunctive concept because an instance is a member of the concept

if either of the features is present. In the final example, the solution is that the

number of objects must equal the number of borders. This example is a relational

concept because it specifies a relationship between two dimensions.

The problems in this series are particularly difficult because to identify the

concept, you must both determine which features are relevant and discover

the kind of rule that connects the features (e.g., conjunctive, disjunctive, or

relational). The former problem is referred to as attribute identification and

the latter as rule learning (Haygood & Bourne, 1965). In many experiments,

either the form of the rule or the relevant attributes are identified for the participant.

For instance, in the Bruner et al. (1956) experiments, participants had

to identify only the correct attributes. They knew that they would be identifying

conjunctive concepts.

Forming a hypothesis involves identifying both what features are relevant

to the hypothesis and how these features are related.

Hypothesis Testing

In the experiment illustrated in Figure 10.5, participants receive a sequence of

instances and have to figure out what the concept is. Some problems in real life

are like this—we have no control over what evidence we see but must figure out

the rules that govern it. For instance, when there is an outbreak of food poisoning

in the United States, medical health researchers check on what the victims

ate, looking for some common pattern. They have no control over what the

Inductive Reasoning and Hypothesis Testing | 291

Concept 1 Concept 2 Concept 3

_

_

_

_

_

_

_

_

_

_

_

_

_

_

_

_

_

_

_

_

_

_

FIGURE 10.5 Examples of

sequences of instances from

which a participant is to identify

concepts. Each column gives

a sequence of instances and

non-instances for a different

concept. A plus (_) signals a

positive instance and minus (_)

sign a negative instance.

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 291

victims ate. On the other hand, in other situations one can do experiments and

test certain possibilities. For instance, when medical researchers want to determine

the most effective combination of drugs to treat a disease, they will perform

clinical trials where different groups of patients receive different drug

combinations. Scientific research can reach more certain conclusions more

quickly if the researchers can choose the cases to test rather than having to take

the cases that the situation presents to them.

In their classic research, Bruner, Goodnow, and Austin (1956) also studied

situations where participants could choose which instances they get information

about. For example, in one condition, Bruner et al. gave participants a single positive

instance of concept, and then the participants could select various instances

and ask whether they were also instances of the concept. For example, if you were

told that the middle stimulus in Figure 10.4 (fifth row, fifth column) was an

instance of a conjunctive concept that you had to discover, what stimuli would

you choose to select? The approach advocated in science would be to test each

dimension, one at a time, and determine whether it was critical to the hypothesis.

For instance, you could choose to test first the dimension of number of borders

and choose a stimulus that differed from the initial stimulus only on this

dimension. If the stimulus were not an instance, you would know that the value

of the original stimulus (in this case, two borders) was relevant and if not, you

would know that it was irrelevant. Then you could try another dimension. After

four stimuli, you would have identified the conjunctive concept with certainty.

Bruner et al. called this strategy “conservative focusing,” and some of their participants

(Harvard undergraduates of the 1950s) followed it. However, many participants

practiced less well-behaved strategies. For instance, given the same initial

stimulus, they might test an instance that changed both the color and the number

of borders. If the stimulus were an instance, they would know that neither dimension

was relevant.However, if wrong, they would have learned relatively little.

A well-known case where people seem to test their hypotheses less than

optimally is the 2-4-6 task introduced by Wason (1960—the same psychologist

who introduced the card selection task that we described earlier). In this experiment,

participants are told that “2 4 6” is an instance of a triad that is consistent

with a rule and are instructed to find out what the rule is by asking whether

other triples of numbers are instances of the rule. What triples would you try?

The protocol below comes from one of Wason’s participants. The protocol gives

each triad that the participant produced and the reason for the choice, along

with the experimenter’s feedback as to whether the triad conformed to the rule.

The sequence of triads was occasionally broken when the participant decided

to announce a hypothesis. The experimenter’s feedback for each hypothesis is

given in parenthesis:

Triad Reason Given for Triad Feedback

8 10 12 2 added each time. Yes

14 16 18 Even numbers in order of magnitude. Yes

20 22 24 Same reason. Yes

1 3 5 2 added to preceding number Yes

292 | Reasoning

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 292

Announcement: The rule is that by starting with any number, 2 is added each

time to form the next number. (Incorrect)

2 6 10 The middle number is the arithmetic mean of the

other two. Yes

1 50 99 Same reason. Yes

Announcement: The rule is that the middle number is the arithmetic mean of

the other two. (Incorrect)

3 10 17 Same number, 7, added each time. Yes

0 3 6 Three added each time. Yes

Announcement: The rule is that the difference between two numbers next to

each other is the same. (Incorrect)

12 8 4 The same number is subtracted each time to form

the next number. No

Announcement: The rule is adding a number, always the same one, to form

the next number. (Incorrect)

1 4 9 Any three numbers in order of magnitude. Yes

Announcement: The rule is any three numbers in order of magnitude.

(Correct)

The important feature to note about this protocol is that the participant tested

the hypothesis by almost exclusively generating sequences consistent with it. The

better procedure in this case would have been to also try sequences that were

inconsistent. That is, the participant should have looked sooner for negative

evidence as well as positive evidence. This would have exposed the fact that

the participant had started out with a hypothesis that was too narrow and was

missing the more general correct hypothesis. The only way to discover this

error is to try examples that disconfirm the hypothesis, but this is what people

have great difficulty doing,

In another experiment, Wason (1968) asked 16 participants, after they had

announced their hypotheses, what they would do to determine whether their

hypotheses were incorrect. Nine participants said they would generate only

instances consistent with their hypotheses and wait for one to be identified as

not an instance of the rule. Only four participants said that they would generate

instances inconsistent with the hypothesis to see whether they were identified

as members of the rule. The remaining three insisted that their hypotheses

could not be incorrect.

This strategy to select only positive instances has been called the confirmation

bias. It has been argued that confirmation bias is not necessarily a

mistaken strategy (Fischhoff & Neyth-Marom, 1983; Kayman & Ha, 1987). In

many situations, selecting instances consistent with a hypothesis is an effective

way to disconfirm the hypothesis. For instance, if one did well on an exam after

drinking a glass of orange juice and entertained the hypothesis that orange

juice led to good exam performance, drinking orange juice before a couple

more exams might quickly disabuse one of that hypothesis. What made this

Inductive Reasoning and Hypothesis Testing | 293

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 293

strategy so ineffective in Wason’s experiment is simply that the correct hypothesis

was very general. It would be like finding out that consuming any drink

would improve exam performance (particularly unlikely if we include alcoholic

drinks).

In choosing instances to test a hypothesis, people often focus on instances

consistent with their hypothesis, and this can cause difficulties if their

hypothesis is too narrow.

Scientific Discovery

Whether participants are trying to infer a concept by selecting instances from a

set of options like those in Figure 10.4 or trying to infer a rule that describes a

set of examples as in the protocol we just reviewed, participants are engaged in

problem-solving searches like those we discussed in Chapter 8 (such as in

Figure 8.4 or Figure 8.8). The difference, however, is that they are searching two

problem spaces. One problem space is the space of possible hypotheses and the

other is the space of possible instances. It has been argued (e.g., Simon & Lea,

1974; Klahr & Dunbar, 1998) that this is exactly the situation that scientists face

in discovering a new theory—they search through a space of possible theories

and a space of possible experiments to test these theories.

The term “confirmation bias” has been used to describe failures in the way

people test scientific theories. In the hypothesis-testing example we described, it

just referred to a tendency to test only instances that were an example of one’s

hypothesis. However, in the broader context of testing scientific theories, it refers

to a host of behaviors that serve to protect one’s favored theory from disconfirmation.

In one study, Dunbar (1993) had undergraduates try to discover how

genes were controlled by redoing, in a highly simplified form, the research that

won Jacques Monod and Francois Jacob the 1965 Nobel Prize for medicine.

They provided the participants with computer simulations that could mimic

some of the critical experiments. They were told that their task was to determine

how one set of genes controlled another set of genes that produced an

enzyme only when lactose was present. (This enzyme serves to break down the

lactose into glucose.) All the undergraduates initially thought that there must

be a mechanism by which the first set of genes responded to the presence of

lactose and activated the second set of genes. This is the hypothesis that Monod

and Jacob had initially as well, but in fact the mechanism is an inhibitory mechanism

by which the first set of genes inhibit the enzyme-producing genes when

lactose is absent but are blocked from inhibiting when lactose is present. Showing

the confirmation bias, these undergraduates tried to find experiments that

would confirm their activation hypothesis. The majority of the participants

continued to search the experimental space for some combination of genes that

would support their activation hypothesis, but a minority began to search for

alternative hypotheses about what was in control.

Science as an institution has a way of protecting us from scientists whose confirmation

bias leads them too strongly in the wrong direction. Other scientists are

294 | Reasoning

Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 294

often strongly motivated to find problems with the theories of a particular scientist

(Nickerson, 1998). There is also considerable variation in how individual

scientists practice. Michael Faraday, a famous 19th-century chemist, made his

discoveries by early focusing on collecting confirmatory evidence and then

switching to focusing on disconfirmatory evidence (Tweney, 1989). Dunbar

(1997) studied scientists in three immunology laboratories and one biology

laboratory at Stanford and noted that they were quite ready to attend to unexpected

results and modify their theory to accommodate these.

Fugelsang and Dunbar (2005) performed fMRI imaging studies looking at

participants as they tried to integrate data with specific hypothesis. For instance,

participants were told that they were seeing results from a clinical trial

looking at the effect of an antidepressant on mood. They either saw patient

records that indicated the drug had an effect on mood (consistent) or that it

did not have an effect (inconsistent). Participants started out believing the drug

had an effect and so found consistent evidence more plausible. When viewing

the inconsistent evidence, participants showed greater activation in their anterior

cingulate cortex (ACC) (see Figure 3.1). As we noted in Chapter 3, the ACC

is highly active when participants are engaged in a task that requires strong cognitive

control, such as dealing with an inconsistent trial in a Stroop task. These

same basic brain mechanisms seem to be invoked when participants must deal

with inconsistent data in a scientific context, and the results suggest that scientific

reasoning evokes basic cognitive processes.

Inductive Reasoning and Hypothesis Testing | 295

dropped a rock from a 100-meter tower and timed its

fall as 1 second, it would be wise not to accept the

theory that acceleration due to gravity

was 200 meters (using the formula

distance _ 1/2 _ acceleration _ time2)

rather than the standard accepted value

of approximately 10 meters on earth.

Almost certainly something was wrong in

the measurements and the experiment

needs to be repeated. On the other hand,

the Pasteur case might seem rather extreme,

ignoring 90% of the experiments

on a theory that was much in doubt at

the time. In this case, however, he turned

out to be right.

Implications

How convincing is a 90% result?

Scientists can be subject to a confirmation bias. For

instance, Louis Pasteur was involved in a major debate

with other scientists about whether new

organisms could spontaneously generate.

It was argued that the appearance of

bacteria in apparently sterilized organic

material was evidence for spontaneous

generation of life. Pasteur performed

many experiments trying to disprove this

and 90% of these failed, but he chose

only to publish the successful experiment,

claiming that the rest were due

to experimental errors (Geison, 1995).

Scientists frequently question their

experiment if results seem to contradict

establish theory. For instance, if one

Anderson7e_Chapter_10.qxd 8/20/09 9:51 AM Page 295

1. Johnson-Laird and Goldvarg (1997) presented

Princeton undergraduates with reasoning problems

like this one:

Only one of the following premises is true about a

particular hand of cards:

There is a king in the hand or there is an ace or both.

There is a queen in the hand or there is an ace or both.

There is a jack in the hand or there is a 10, or both.

Is it possible that there is an ace in the hand?

They report that the students only were correct on 1%

of such problems.What is the correct answer for the

problem above? Why is it so hard? Johnson-Laird and

Goldvarg attribute the difficulty that people have in

creating mental models of what is not the case.

2. Johnson-Laird and Steedman (1978) presented the

following syllogisms to participants drawn from

Columbia Teachers College:

All gourmets are shopkeepers.

All bowlers are shopkeepers.

And asked them what followed. The following is

the distribution of answers:

17 agreed that no conclusion followed.

2 thought that “Some gourmets are bowlers” followed.

4 thought that “All bowlers are gourmets” followed.

Questions for Thought

In studies of scientific discovery, participants tend to focus on experiments

consistent with their favorite hypothesis and show a reluctance to search for

alternative hypotheses.

Conclusions

Much of the research on human reasoning has found it wanting when compared

to the rules and implications of formal logic. As we just noted, this might even

be said of the process by which scientists engage in their research. However, this

dismal characterization of human reasoning fails to properly appreciate the

full context in which it occurs. In many actual reasoning situations, people do

quite well, in part because they take in the full complexity and implications of

the actual real-world content. Despite a tendency towards confirmation bias,

science as a whole has progressed with great success. To some extent, this is

because science is a social activity carried out by a community of researchers.

Competitive scientists are quick to find mistakes in each other’s approach, but

there is also a cooperative nature to science. Research takes place in teams of

researchers and they often rely on each other’s help. Okada and Simon (1997)

found that pairs of undergraduates were much more successful than individual

students at finding the inhibition mechanism in Dunbar’s (1993) genetic control

task. As Okada and Simon note, “In a collaborative situation, subjects must often

be more explicit than in an individual learning situation, to make partners

understand their ideas and to convince them. This can prompt subjects to entertain

requests for explanation and construct deeper explanations” (p. 130). The

bottom line of this chapter is that human reasoning normally takes place in

a world of complexities (both factual and social) and that what appears deficient

in the laboratory may be exquisitely tuned to that world.

296 | Reasoning

Anderson7e_Chapter_10.qxd 8/20/09 9:51 AM Page 296

Key Terms | 297

7 thought that “Some bowlers are gourmets” followed.

8 thought that “All gourmets are bowlers” followed.

Use the concepts of this chapter to help explain the

answers these participants gave and did not give.

3. Consider the third column in Figure 10.5, which was

described in the chapter as satisfying the rule that

“the number of borders is the same as the number of

objects.”An alternative rule that describes the instances

is “3 white objects or 2 black objects or 1 object with

one border.”Which is the better description of the

category and why? Is it possible to know for certain

which is the correct rule?

Key Terms

affirmation of the

consequent

antecedent

atmosphere hypothesis

attribute identification

categorical syllogism

conditional statement

confirmation bias

consequent

deductive reasoning

denial of the antecedent

inductive reasoning

logical quantifiers

mental model theory

modus ponens

modus tollens

particular statements

permission schema

rule learning

selection task

syllogisms

universal statements

Anderson7e_Chapter_10.qxd 8/20/09 9:51 AM Page 297

298

11

Judgment and

Decision Making

As we saw in Chapter 10, most of the research on human reasoning has compared

it to various prescriptive models from logic and mathematics. The prescriptive

models assume that people have access to information about which they can be

certain and that they can coolly reflect on the information. However, in the real world,

people have to make decisions in the face of incomplete and uncertain information.

Furthermore, in contrast to the relatively neutral character of the syllogisms of the

previous chapter, in real life our decisions can have important consequences for our

lives. Consider the simple task of deciding what to eat—we have all been frustrated

by the medical reports that pronounce formerly “healthy” food as “unhealthy” and vice

versa. In decision making, we must also deal with the unpleasant consequences of

what might be good decisions such as going on a diet or giving up a pleasurable

activity like smoking.

This chapter will focus on research on judgment and decision making that

comes closer to such real-life circumstances. As before, we will discuss research

showing how the performance of normal humans is wanting compared to models

that were developed for rational behavior. However, we will also see how much of

that research is incomplete, missing the complexity of everyday human decision

making. Recent research has developed a more nuanced characterization of the situations

that people face in their everyday life, and a better appreciation of the

nature of their judgments.

In this chapter, we will answer the questions: • How well do people judge the probability of uncertain events? • How do people use their past experiences to make judgments? • How do people decide among uncertain options that offer different rewards

and costs? • How does the brain support such decision making?

Anderson7e_Chapter_11.qxd 8/20/09 9:51 AM Page 298

The Brain and Decision Making | 299

The Brain and Decision Making

In 1848, Phineas Gage, a railroad worker in Vermont, suffered a bizarre accident:

He was using an iron bar to pack gunpowder down into a hole drilled into

a rock that had to be blasted to clear a roadbed for the

railroad. The powder unexpectedly exploded and sent the

iron bar flying through his head before landing 80 feet

away. Figure 11.1 shows a reconstruction of the trajectory

of the bar through his skull (Damasio, Grabowski, Frank,

Galabruda, & Damasio, 1994). (For a more detailed reconstruction,

see Color Plate 11.1.) The bar managed to miss

any vital areas and spared most of his brain but tore

through the middle of the very front of the brain—a

region called the ventromedial prefrontal cortex. Amazingly,

he not only survived, he was even was able to talk

and walk away from the accident after being unconscious

for a few minutes. His recovery was difficult, largely because

of infections, but he eventually was able to hold jobs

such as a coach driver. Henry Jacob Bigelow, a Professor of

Surgery at Harvard University, declared him “quite recovered

in faculties of body and mind” (MacMillan, 2000).

Based on such a report, one might have thought that this

part of the brain performed no function.

However, all was not well. His personality had undergone major changes.

Before his injury he had been polite, respectful, popular, and reliable, and generally

displayed ideal behavior for an American man of that time. Afterward he

became just the opposite—as his own physician, Harlow, later described him:

fitful, irreverent, indulging at times in the grossest profanity (which was not

previously his custom), manifesting but little deference for his fellows, impatient

of restraint or advice when it conflicts with his desires, at times pertinaciously

obstinate, yet capricious and vacillating, devising many plans of future

operations, which are no sooner arranged than they are abandoned in turn

for others appearing more feasible. A child in his intellectual capacity and

manifestations, he has the animal passions of a strong man. Previous to his

injury, although untrained in the schools, he possessed a well-balanced mind,

and was looked upon by those who knew him as a shrewd, smart businessman,

very energetic and persistent in executing all his plans of operation. In

this regard his mind was radically changed, so decidedly that his friends and

acquaintances said he was “no longer Gage.” (Harlow, 1868, p. 327)

Gage is the classic case demonstrating the importance of the ventromedial

prefrontal cortex to human personality. Subsequently, a number of other similar

patients have been described and they all show the same sorts of personality

disorders. Family members and friends will describe them with phrases like

“socially incompetent,”“decides against his best interest,” and “doesn’t learn from

his mistakes” (Sanfey, Hastie, Colvin, & Grafman, 2003). Earlier in Chapter 8, we

Brain Structures

FIGURE 11.1 A representation

of the passage of the bar

through Phineas Gage’s brain.

Note that only the middle

of the frontal-most portion

has been damaged. (From Damasio

et al., 1994).

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 299

discussed the case of the patient PF who also suffered damage to his anterior

prefrontal region, like Gage. However, in his case the damage also included

lateral portions of the anterior prefrontal region, and his difficulty was more

with organizing complex problem solving than decision making. In general, it is

thought that the more medial portion of the anterior prefrontal region, where

Gage’s injury was localized, is important to motivation, emotional regulation,

and social sensitivity (Gilbert, Spengler, Simons, Frith, & Burgess, 2006).

The ventromedial prefrontal cortex plays an important role in achieving the

motivational balance and social sensitivity that is key to making successful

judgments.

Probabilistic Judgment

There is a prescriptive model for how people should reason about probabilities

as they collect relevant evidence. This is called Bayes’s theorem, which is based

on a mathematical analysis of the nature of probability.Much of the research in

the field has been concerned with showing that human participants do not

match up with the prescriptions of Bayes’s theorem.

Bayes’s Theorem

As an example of the application of Bayes’s theorem, suppose I come home and

find the door to my house ajar. I am interested in the hypothesis that it might

be the work of a burglar. How do I evaluate this hypothesis? I might treat it as

a conditional syllogism of the following sort:

If a burglar is in the house, then the door will be ajar.

The door is ajar.

A burglar is in the house.

As a conditional syllogism, it would be judged as the erroneous affirmation of

the consequent. However, it does have a certain plausibility as an inductive

argument. Bayes’s theorem provides a way of assessing just how plausible it is

by combining what are called a prior probability and a conditional probability

to produce what is called a posterior probability, which is a measure of the

strength of the conclusion.

A prior probability is the probability that a hypothesis is true before consideration

of the evidence (e.g., the door is ajar). The less likely the hypothesis was

before the evidence, the less likely it should be after the evidence. Let us refer to

the hypothesis that my house has been burglarized as H. Suppose that I know

from police statistics that the probability of a house in my neighborhood being

burglarized on any particular day is 1 in 1,000.1 This probability is expressed as:

Prob(H) _ .001

300 | Judgment and Decision Making

1 Although this makes for easy calculation, the actual number for Pittsburgh is closer to 1 burglary per

100,000 households per day.

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 300

This equation expresses the prior probability of the hypothesis, or the probability

that the hypothesis is true before the evidence is considered. The other prior

probability needed for the application of Bayes’s theorem is the probability that

the house has not been burglarized. This alternate hypothesis is denoted ~H.

This value is 1 minus Prob(H) and is expressed:

Prob(~H) _ .999

A conditional probability is the probability that a particular type of evidence

is true if a particular hypothesis is true. Let us consider what the conditional

probabilities of the evidence (door ajar) would be under the two hypotheses.

First, suppose I believe that the probability of the door’s being ajar is quite high

if I have been burglarized, for example, 4 out of 5. Let E denote the evidence, or

the event of the door being ajar. Then, we will denote this conditional probability

of E given that H is true as

Prob(E_H) _ .8

Second, we determine the probability of E if H is not true. Suppose I know that

chances are only 1 out of 100 that the door would be ajar if no burglary had taken

place (e.g., by accident, neighbors with a key).We denote this probability by

Prob(E|~H) _ .01

the probability of E given that H is not true.

The posterior probability is the probability that a hypothesis is true after

consideration of the evidence. The notation Prob(H|E) is the posterior probability

of hypothesis H given evidence E. According to Bayes’s theorem, we can

calculate the posterior probability of H, that the house has been burglarized

given the evidence, thus:

Prob(E|H) • Prob(H)

Bayes equation: Prob(H|E) _ ______________________________________

Prob(E|H) • Prob(H) _ Prob(E|~H) • Prob(~H)

Given our assumed values, we can solve for Prob(H|E) by substituting into the

preceding equation:

(.8) (.001)

Prob(H|E) _____________________ _ .074

(.8) (.001) _ (.01) (.999)

Thus, the probability that my house has been burglarized is still less than 8 in

100. Note that the posterior probability is this low even though an open door

is good evidence for a burglary and not for a

normal state of affairs: Prob(E|H) _ .8 versus

Prob(E|~H) _ .01. The posterior probability

is still quite low because the prior probability

of H—Prob(H) _ .001—was very low to begin

with. Relative to that low start, the posterior

probability of .074 is a considerable increase.

Table 11.1 offers an illustration of Bayes’s

theorem as applied to the burglary example.

It offers an analysis of 100,000 households,

Probabilistic Judgment | 301

TABLE 11.1

An Analysis of Bayes’s Theorem—100,000 Households

Burglarized Not Burglarized Sums

Door open 80 999 1,079

Door not open 20 98,901 98,921

Sums 100 99,900 100,000

Adapted from J. R. Hayes (1984).

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 301

assuming these statistics. There are four possible states of affairs, determined by

whether the burglary hypothesis is true or not and by whether there is evidence

of an open door or not. The frequency of each state of affairs is set forth in the

four cells of the table. Let’s consider the frequency in the upper-left cell, which

is the case I was worried about—the door is open and my house has been burglarized.

Because 1 in a 1,000 households are burglarized (Prob(H) is .001),

there should be 100 burglaries in the 100,000 households. This is the frequency

of both events in the left column. Because 8 times out of 10 the front door is

left open in a burglary (Prob(E|H) is .8), 80 of these 100 burglaries should leave

the door open—the number in the upper left. Similarly, in the upper-right cell,

we can calculate that of the 99,900 homes without burglary, the front door will

be left open 1 in 100 times, for 999 cases. Thus, in total there are 80_999_1,079

cases of front doors left open, and the probability of the house being burglarized

is 80/1,079 _ .074. The calculations in Bayes’s theorem perform the

same calculation as afforded by Table 11.1, but in terms of probabilities rather

than frequencies. As we will see, people find it easier to reason in terms of

frequencies.

Because Bayes’s theorem rests on a mathematical analysis of the nature of

probability, the formula can be proved to evaluate hypotheses correctly. Thus, it

enables us to precisely determine the posterior probability of a hypothesis given

the prior and conditional probabilities. The theorem serves as a prescriptive

model, or normative model, specifying the means of evaluating the probability

of a hypothesis. Such a model contrasts with a descriptive model, which specifies

what people actually do. People normally do not perform the calculations

that we have just gone through any more than they follow the steps prescribed

by formal logic. Nonetheless, they do hold various strengths of belief in assertions

such as “My house has been burglarized.” Moreover, their strength of

belief does vary with evidence such as whether the door has been found ajar.

The interesting question is whether the strength of their belief changes in accord

with Bayes’s theorem.

Bayes’s theorem specifies how to combine the prior probability of a hypothesis

with the conditional probabilities of the evidence to determine the posterior

probability of a hypothesis.

Base-Rate Neglect

Many people are surprised that the open door in the preceding example does

not provide as much evidence for a burglary as might have been expected. The

reason for the surprise is that they do not grasp the importance of the prior

probabilities. People sometimes ignore prior probabilities. In one demonstration,

Kahneman and Tversky (1973) told one group of participants that a person

had been chosen at random from a set of 100 people consisting of 70 engineers

and 30 lawyers. This group of participants was termed the engineer-high

group. A second group, the engineer-low group, was told that the person came

from a set of 30 engineers and 70 lawyers. Both groups were asked to determine

302 | Judgment and Decision Making

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 302

the probability that the person chosen at random from the group would be

an engineer, given no information about the person. Participants were able to

respond with the right prior probabilities: The engineer-high group estimated

.70 and the engineer-low group estimated .30. Then participants were told that

another person was chosen from the population and they were given the following

description:

Jack is a 45-year-old man. He is married and has four children. He is generally

conservative, careful, and ambitious. He shows no interest in political and

social issues and spends most of his free time on his many hobbies, which

include home carpentry, sailing, and mathematical puzzles.

Participants in both groups gave a .90 probability estimate to the hypothesis

that this person is an engineer. No difference was displayed between the two

groups, which had been given different prior probabilities for an engineer

hypothesis. But Bayes’s theorem prescribes that prior probability should have a

strong effect, resulting in a higher posterior probability from the engineer-high

group than from the engineer-low group.

In a second case, Kahneman and Tversky presented participants with the

following description:

Dick is a 30-year-old man. He is married with no children. A man of high

ability and high motivation, he promises to be quite successful in his field.

He is well liked by his colleagues.

This example was designed to provide no diagnostic information either way

with respect to Dick’s profession. According to Bayes’s theorem, the posterior

probability of the engineer hypothesis should be the same as the prior probability

because this description is not informative. However, both the engineer-high

and the engineer-low groups estimated that the probability was .50 that the

man described is an engineer. Thus, they allowed a completely uninformative

piece of information to change their probabilities. Once again, the participants

were shown to be completely unable to use prior probabilities in assessing the

posterior probability of a hypothesis.

The failure to take prior probabilities into account can lead people to make

some totally unwarranted conclusions. For instance, suppose you take a diagnostic

test for a cancer. Suppose also that this type of cancer, when present, results in

a positive test 95% of the time. On the other hand, if a person does not have the

cancer, the probability of a positive test result is only 5%. Suppose you are informed

that your result is positive. If you are like most people, you will assume

that your chances of dying of cancer are about 95 out of 100 (Hammerton,

1973). You would be overreacting in assuming that the cancer will be fatal, but

you would also be making a fundamental error in probability estimation.What

is the error?

You would have failed to consider the base rate (prior probability) for the

particular type of cancer in question. Suppose only 1 in 10,000 people have this

cancer. This percentage would be your prior probability. Now, with this information,

you would be able to determine the posterior probability of your

Probabilistic Judgment | 303

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 303

having the cancer. Bringing out the Bayesian formula, you would express the

problem in the following way:

Prob(H) • Prob(E|H)

Prob(H|E) _ _______________________________________

Prob(H) • Prob(E|H) _ Prob(~H) • Prob(E|~H)

where the prior probability of the cancer hypothesis is Prob(H) _ .0001, and

Prob(~H) _ .9999, Prob(E|H) _ .95, and Prob(E|~H) _ .05. Thus,

(.0001) (.95)

Prob(H|E) _ _______________________ _ .0019

(.0001) (.95) _ (.9999) (.05)

That is, the posterior probability of your having the cancer would still be less

than 1 in 500.

People often fail to take base rates into account in making probability

judgments.

Conservatism

The preceding examples show that people weigh the evidence too much and

ignore base rates. However, there are also situations in which people do not

weigh evidence enough, particularly as the evidence pointing to a conclusion

accumulates.Ward Edwards (1968) extensively investigated how people use new

information to adjust their estimates of the probabilities of various hypotheses.

In one experiment, he presented participants with two bags, each containing

100 poker chips. One of the bags contained 70 red chips and 30 blue; the other

contained 70 blue chips and 30 red. The experimenter chose one of the bags

at random and the participants’ task was to decide which bag had been chosen.

In the absence of any prior information, the probability of either bag having

been chosen was 50%. Thus,

Prob(HR) _ .50 and Prob(HB) _ .50

where HR is the hypothesis of a predominantly red bag and HB is the hypothesis

of a predominantly blue bag. To obtain further information, participants sampled

chips at random from the bag. Suppose the first chip drawn was red. The

conditional probability of a red chip drawn from each bag is

Prob(R|HR) _ .70 and Prob(R|HR) _ .30

Now, we can calculate the posterior probability of the bag’s being predominantly

red, given the red chip is drawn, by applying the Bayes equation to this situation:

Prob(R|HR) • Prob(HR)

Prob(R|HR) _ _______________________________________

Prob(R|HR) • Prob(HR) _ Prob(R|HR) • Prob(HR)

(.70) • (.50)

_ _____________________ _ .70

(.70) • (.50) _ (.30) • (.50)

304 | Judgment and Decision Making

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 304

This result seems, to both naive and sophisticated observers, to be a rather

sharp increase in probabilities. Typically, participants do not increase the probability

of a red-majority bag to .70; rather, they make a more conservative revision

to a value such as .60.

After this first drawing, the experiment continues: The poker chip is put back

in the bag and a second chip is drawn at random. Suppose this chip too is red.

Again, by applying Bayes’s theorem, we can show that the posterior probability

of a red bag is now .84. Suppose our observations continued for 10 more trials

and, after all 12 trials, we have observed eight reds and four blues. By continuing

the Bayesian analysis, we could show that the new posterior probability of the

hypothesis of a red bag is .97. Participants who see this sequence of 12 trials only

estimate subjectively a posterior probability of .75 or less for the red bag.

Edwards has used the term conservative to refer to the tendency to underestimate

full effect of available evidence. He estimates that we use between a fifth and a

half of the evidence available to us in situations like this experiment.

People frequently underestimate the cumulative force of evidence in making

probability judgments.

Correspondence to Bayes’s Theorem with Experience

All the preceding examples showed that participants can be quite far off in their

judgments of probability. One possibility is that participants really do not

understand probabilities or how to reason with respect to them. Certainly, it is

a rare participant in these experiments who could reproduce Bayes’s theorem,

let alone who would report engaging in Bayesian calculation. However, there is

evidence that, although participants cannot articulate the correct probabilities,

many aspects of their behavior are in accordance with Bayesian principles. To return

to the explicit-implicit distinction discussed in Chapter 7, people often seem

to display implicit knowledge of Bayesian principles even if they do not display

any explicit knowledge and make errors when asked to make explicit judgments.

Gluck and Bower (1988) performed an experiment that illustrates implicit

Bayesian behavior. Participants were given records of fictitious patients who

could display from one to four symptoms (bloody nose, stomach cramps, puffy

eyes, and discolored gums) and made discriminative diagnoses about which of

two hypothetical diseases the patients had. One of these diseases had a base rate

three times that of the other. Additionally, the conditional probabilities of displaying

the various symptoms, given the diseases, were varied. Participants were

not told directly about these base rates or conditional probabilities. They

merely looked at a series of 256 patient records, chose the disease they thought

the patient had, and were given feedback on the correctness of their judgments.

There are 15 possible combinations of one to four symptom patterns that

a patient might have. Gluck and Bower calculated the probability of each disease

for each pattern by using Bayes’s theorem and arranged it so that each

disease occurred with that probability when the symptoms were present. Thus,

the participants experienced the base probabilities and conditional probabilities

implicitly in terms of the frequencies of symptom–disease combinations.

Probabilistic Judgment | 305

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 305

Of interest is the probability with which they assigned

the rarer disease to various symptom combinations.

Gluck and Bower compared the participant

probabilities with the true Bayesian probabilities.

This correspondence is displayed by the scatterplot

in Figure 11.2. There we have, for each symptom

combination, the Bayesian probability (labeled objective

probability) and the proportion of times that

participants assigned the rare disease to that symptom

combination. As can be seen, these points fall

very close to a straight diagonal line with a slope of

1, which indicates that the proportion of the participants’

choices were very close to the true probabilities.

Thus, implicitly, the participants had become

quite good Bayesians in this experiment. The behavior

of choosing among alternatives in proportion to

their success is called probability matching.

After the experiment, Gluck and Bower presented

the participants with the four symptoms individually

and asked them how frequently the rare disease had

appeared with each symptom. This result is presented

in Figure 11.3 in a format similar to that of

Figure 11.2. As can be seen, participants showed

some neglect of the base rate, consistently having

overestimated the frequency of the rare disease. Still,

their judgments show some influence of base rate in

that their average estimated probability of the rare

disease is less than 50%.

Gigerenzer and Hoffrage (1995) showed that

base-rate neglect also decreases if events are stated

in terms of frequencies rather than in terms of

probabilities. Some of their participants were given

a description in terms of probabilities, such as the

one that follows:

The probability of breast cancer is 1% for

women at age 40 who participate in routine

screening. If a woman has breast cancer,

the probability is 80% that she will get a positive

mammography. If a woman does not have

breast cancer, the probability is 9.6% that she

also will get a positive mammography. A woman

in this age group had a positive mammography

in a routine screening. What is the probability

that she actually has breast cancer?

Fewer than 20 out of 100 (20%) of the participants

given such statements calculated the correct

Bayesian answer (which is about 8%). In the other

306 | Judgment and Decision Making

.2

.2

.4

.6

.8

1.0

.4 .6 .8 1.0

Proportion of choices by subjects

Objective probability

FIGURE 11.2 Participants’ proportion of choices corresponds

closely to the objective probabilities as determined by Bayes’s

theorem.

.2

.2

0

.4

.6

.8

1.0

.4 .6 .8 1.0

Estimated probability

True probability

FIGURE 11.3 Participants’ estimated probabilities systematically

overestimated the frequency of the rare disease, showing base-rate

neglect.

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 306

condition, participants were given descriptions in terms of frequencies, such as

the one that follows:

Ten out of every 1,000 women at age 40 who participate in routine screening

have breast cancer. Eight of every 10 women with breast cancer will get a

positive mammography. Ninety-five out of every 990 women without breast

cancer also will get a positive mammography. Here is a new representative

sample of women at age 40 who got a positive mammography in routine

screening. How many of these women do you expect to actually have breast

cancer?

Almost 50% of the participants given such statements calculated the correct

Bayesian answer. Gigerenzer and Hoffrage argued that we can reason better

with frequencies than with probabilities because we experience frequencies of

events, but not probabilities, in our daily lives.

There is also evidence that experience makes people more statistically

tuned. In a study of medical diagnosis,Weber, Böckenholt, Hilton, and Wallace

(1993) found that doctors were quite sensitive both to base rates and to the evidence

provided by the symptoms. Moreover, the more clinical experience the

doctors had, the more tuned were their judgments.

Although participants’ processing of abstract probabilities often does not

correspond with Bayes’s theorem, their behavior based on experience

often does.

Judgments of Probability

What are participants actually doing when they report probabilities of an event

such as the probability that someone who has bloody gums has a particular disease?

The evidence is that rather than thinking about probabilities,

they are thinking about relative frequencies. Thus they are

trying to judge the proportion of the patients that they saw with

bloody gums who had that particular disease. People are reasonably

accurate at making such proportionate judgments when they

do not have to rely on memory (Robinson, 1964; Shuford, 1961).

Consider an experiment by Shuford (1961). He presented arrays

such as that shown in Figure 11.4 to participants for 1 s. He then

asked participants to judge the proportion of vertical bars relative

to horizontal bars. The number of vertical bars varied from 10%

to 90% in different matrices. Shuford’s results are shown in Figure

11.5. As can be seen, participants’ estimates are quite close to

the true proportions.

The situation just described is one where the participants

can see the relevant information and make a judgment about

proportions.When participants cannot see the events and must

recall them from memory, they can give distorted judgments if

they recall too many of one kind from memory. A fair amount

of research has been done on the ways in which participants can

be biased in their estimation of the relative frequency of various

Probabilistic Judgment | 307

FIGURE 11.4 A random matrix presented to

participants to determine their accuracy in judging

proportions. The matrix is 90% vertical bars and

10% horizontal bars. (From Shuford, 1961. Copyright © 1961

by the American Psychological Association. Reprinted by permission.)

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 307

events in the population. Consider the following

experiment reported by Tversky and Kahneman

(1974), which demonstrates that judgments of

proportion can be biased by differential availability

of examples. These investigators asked participants

to judge the proportion of words in the language

that fit certain characteristics. For instance, they

asked participants to estimate the proportion of

English words that begin with the letter k versus

words with the letter k in the third position. How

might participants perform this task? One obvious

method is to briefly try to think of words that

satisfy the specification and words that do not and

to estimate the relative proportion of target words.

How many words can you think of that begin with

the letter k? How many words can you think of that

do not? What is your estimate of their proportion?

Now, how many words can you think of that have

the letter k in the third position? How many words

can you think of that do not? What is their relative

proportion? Participants estimated that more words begin with letter k than

have the letter k in the third position. In actual fact, three times as many words

have the letter k in the third position as begin with the letter k. Generally,

participants overestimate the frequency with which words begin with various

letters.

As in this experiment, many real-life circumstances require that we estimate

probabilities without having direct access to the population that these probabilities

describe. In such cases, we must rely on memory as the source for our estimates.

The memory factors that we studied in Chapters 6 and 7 serve to explain

how such estimates can be biased. Under the reasonable assumption that words

are more strongly associated with their first letter than with their third letter,

the bias exhibited in the experimental results can be explained by the spreadingactivation

theory (Chapter 6). With the focus of attention on the letter k, for

example, activation will spread from that letter to words beginning with it. This

process will tend to make words beginning with the letter k more available than

other words. Thus, these words will be overrepresented in the sample that participants

take from memory to estimate the true proportion in the population.

The same overestimation is not made for words with the letter k in the third

position because words are unlikely to be directly associated with the letters in

the third position. Therefore, these words cannot be associatively primed and

made more available.

Other factors besides memory lead to biases in probability estimates.

Consider another example from Tversky and Kahneman (1974).Which of the

following sequences of six tosses of a coin (where H denotes heads and T tails)

is more likely: H T H T T H or H H H H H H? Many people think the first

sequence is more probable, but both sequences are actually equally probable.

308 | Judgment and Decision Making

0

0

20

40

60

80

100

Horizontal

20

Porportion in display

Judged proportion

40 60 80 100

Vertical

FIGURE 11.5 Mean estimated

proportion as a function of the

true proportion. Participants

exhibited a fairly accurate ability

to estimate the proportions of

vertical and horizontal bars in

Figure 10.5. (From Shuford, 1961.

Copyright © 1961 by the American

Psychological Association. Reprinted

by permission.)

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 308

The probability of the first sequence is the probability of H on the first toss

(which is .50) times the probability of T on the second toss (which is .50),

times the probability of H on the third toss (which is .50), and so on. The

probability of the whole sequence is .50 * .50 * .50 * .50 * .50 * .50 _ .016.

Similarly, the probability of the second sequence is the product of the probabilities

of each coin toss, and the probability of a head on each coin toss is .50.

Thus, again, the final probability also is .50 * .50 * .50 * .50 * .50 * .50 _ .016.

Why do some people have the illusion that the first sequence is more probable?

It is because the first event seems similar to a lot of other events—for example,

H T H T H T or H T T H T H. These similar events serve to bias upward a person’s

probability estimate of the target event. On the other hand, H H H H H H,

straight heads, seems unlike any other event, and its probability will therefore

not be biased upward by other similar sequences. In conclusion, a person’s

estimate of the probability of an event will be biased by other events that are

similar to it.

A related phenomenon is what is called the gambler’s fallacy. The fallacy

is the belief that if an event has not occurred for a while, then it is more likely,

by the “law of averages,” to occur in the near future. This phenomenon can be

demonstrated in an experimental setting—for instance, one in which participants

see a sequence of coin tosses and must guess whether each toss will be a

head or a tail. If they see a string of heads, they become more and more likely to

guess that tails will come up on the next trial. Casino operators count on this

fallacy to help them make money. Players who have had a string of losses at a

table will keep playing, assuming that by the “law of averages” they will experience

a compensating string of wins. However, the game is set in favor of the

house. The dice do not know or care whether a gambler has had a string of

losses. The consequence is that players tend to lose more as they try to recoup

their losses. The “law of averages” is a fallacy.

The gambler’s fallacy can be used to advantage in certain situations—for

instance, at the racetrack. Most racetracks operate by a pari-mutuel system in

which the odds on a horse are determined by the number of people betting on

the horse. By the end of the day, if favorites have won all the races, people tend

to doubt that another favorite can win, and they switch their bets to the long

shots. As a consequence, the betting odds on the favorite deviate from what

they should be, and a person can sometimes make money by betting on the

favorite.

People can be biased in their estimates of probabilities when they must rely

on factors such as memory and similarity judgments.

The Adaptive Nature of the Recognition Heuristic

The examples in the previous section focused on cases where people came to

bad judgments relying on, for example, availability of events in memory.

Gigerenzer, Todd, and ABC Research Group (1999), in their book Simple

Heuristics That Make Us Smart, argue that such cases are the exception and not

Probabilistic Judgment | 309

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 309

the rule. They argue that people tend to identify the most valid cues for making

judgments and use these. For instance, through evolution people have

acquired a tendency to pay attention to availability of events in memory, which

is more often helpful than not.

Goldstein and Gigerenzer (1999, 2002) report studies of what they call the

recognition heuristic. This heuristic applies in cases where people can recognize

one thing and not another. This heuristic leads people to believe that the

recognized item is bigger and more important than the unrecognized item. In

one study, they looked at the ability of students at the University of Chicago

to judge the relative size of various German cities. For instance, which city is

larger—Bamberg or Heidelberg? Most of the students knew that Heidelberg is a

German city, but most do not recognize Bamberg—that is, one city is available

in memory and the other is not. Goldstein and Gigerenzer showed that when

faced with pairs like this, students almost always pick the city they can recognize.

One might think this shows another fallacy based on availability in memory.

However, Goldstein and Gigerenzer show that the students are actually

more accurate when they make their judgment for pairs of cities like this

(where they recognize one and not the other) than when they are given two

cities they can recognize and must use other bases for judging the size of the

cities (such as Munich versus Hamburg). This is because most American students

have little knowledge about the population of German cities. Thus, far

from a fallacy, this proves to be an effective basis for making judgments. Also,

American students do better at judging the relative size of German cities using

this heuristic than either American students do judging American cities or

German students do judging German cities, where this heuristic cannot be used

because almost all the cities are recognized.2 German students also do better

than American students in judging the population of American cities because

they can use the recognition heuristic and Americans cannot.

Figure 11.6 illustrates Goldstein and Gigerenzer’s explanation for why these

students were more accurate in judging the size of cities when they did not

know one. They looked at the frequency with which German cities were mentioned

in the Chicago Tribune and the frequency with which American cities

were mentioned in the German newspaper Die Zeit. It turns out that there is a

strong correlation between the actual size of the city and the frequency of mention

in these newspapers. Not surprisingly, people read about the larger cities in

other countries more frequently. Gigerenzer and Goldstein also show that there

is a strong correlation between the frequency of mention in the newspapers

(and the media more generally) and the probability that these students will

recognize the name. This is just the basic effect of frequency on memory. As a

consequence of these two strong correlations, there will be a strong correlation

between availability in memory and the actual size of the city.

310 | Judgment and Decision Making

2My German informant (Angela Brunstein) tells me that almost all Germans would recognize Bamberg and

Heidelberg, but many would be puzzled by which is larger. Interestingly, Google search on English texts

reports 37 million hits on Heidelberg and 3.5 million on Bamberg. Google search on German texts reports

30 million hits on Heidelberg and 12 million on Bamberg—a much closer ratio and many more hits

on Bamberg.

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 310

Goldstein and Gigerenzer argue that the recognition heuristic is useful in

many but not all domains. In some domains, researchers have shown that

people intelligently combine it with other information. For instance, Richter

and Späth (2006) had participants judge which of two animals has the larger

population size. For example, consider the following questions:

Are there more hainan partridges or arctic hares?

Are there more giant pandas or mottled umbers?

In the first case, most people have heard of the arctic hares and not hainan

partridges and would correctly choose arctic hares using the recognition

heuristic. In the second case, most people would recognize giant pandas and

not mottled umbers (a moth). Nonetheless, they also know giant pandas are

an endangered species and therefore correctly choose mottled umbers. This is

an example of how people can adaptively choose what aspects of information

to pay attention to.

People can use their ability to recognize an item, and combine this with other

information, to make good judgments.

Probabilistic Judgment | 311

Ecological correlation

Mediator

Recognition correlation

.66/.60

.72/.70

.86/.79

Criterion Recognition

Surrogate correlation

FIGURE 11.6 Ecological correlation (correlation between frequency of mention in newspapers

and population size), surrogate correlation (correlation between frequency of mention in

newspapers and probability of recognition), and recognition correlation (correlation between

probability of recognition and population size). The first value is for American cities and

the German newspaper Die Zeit as mediator, and the second value is for German cities and

the Chicago Tribune as mediator. (From Goldstein and Gigerenzer, 2002.)

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 311

312 | Judgment and Decision Making

Decision Making

An extension of the research on probabilistic reasoning is research on decision

making, which is concerned with the way in which people make choices. Sometimes,

the choices that we have to make are easy. If we are offered the choice between

$400 and $1000, most of us would not have much difficulty in figuring

out which to accept. However, if we were faced with the choice of a certainty of

$400 but only a 50% chance of $1000, which would we select then? Something

like it might happen if we inherited a risky stock that we could cash in for $400

or that we could hold on to and see whether the company would take off or

fold. A great deal of research on decision making under uncertainty requires

participants to make choices among gambles. For instance, a participant might

be asked to choose between the following two gambles:

A. $8 with a probability of 1/3

B. $3 with a probability of 5/6

In some cases, participants are just asked for their opinions; in other cases, they

actually play the gamble that they choose. As an example of the latter possibility,

a participant might roll a die and win in case A if he gets a 5 or 6 and win in

case B if he gets a number other than 1.Which gamble would you choose?

As in the other domains of reasoning, decision making has its own standard

prescriptive theory for the way that people should behave in such situations

(von Neumann & Morgenstern, 1944). This theory says that they should choose

the alternative with highest expected value. The expected value of an alternative

is to be calculated by multiplying the probability by the value. Thus, the expected

value of alternative A is $8 _ 1_3 _ $2.67, whereas the expected value of

alternative B is $3 _ 5_6 _ $2.50. Thus, the normative theory says that participants

should select gamble A. However, most participants will select gamble B.

As a perhaps more extreme example of the same result, suppose you are

given a choice between

A. 1 million dollars with a probability of 1

B. 2.5 million dollars with a probability of 1/2

Maybe, in this case, you are on a game show and are offered a choice between

this great wealth with certainty or the opportunity to toss a coin and get even

more. I (and I assume you) would take the money (1 million)

and run, but in fact, if we do the utility calculations, we should

prefer the second choice because its utility is .5 _ 2.5 million _

1.25 million. Are we really behaving irrationally?

Most people, when asked to justify their behavior in such

situations, will argue that there comes a point when one has

enough money (if we could only convince CEOs of this notion!)

and that there really isn’t much difference for them

between 1 million dollars and 2.5 million dollars. This idea has

been formalized in the terms of what is referred to as subjective

utility—the value that we place on money is not linear with the

face value of the money. Figure 11.7 shows a typical function

proposed for the relation of subjective utility to money. It has

Value

Gains Losses

FIGURE 11.7 A function that

relates subjective value to

magnitude of gain and loss.

(From Kahneman & Tversky, 1984.

Reprinted by permission from the

American Psychological Association.)

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 312

a couple of properties. The first property is that it is curvilinear

such that it takes more than a doubling in the amount of money

to double its utility. Thus, in the preceding example, we may value

2.5 million only 20% more than 1 million. Let us say that the utility

of 1 million is U. The utility of 2.5 million can then be expressed

as 1.2U. The expected value of gamble A is 1 _ U _ U,

and the expected value of gamble B is 1_2 _ 1.2U _ .6U. Thus, in

terms of subjective utility, gamble A is more valuable and is to be

preferred.

The second property of this utility function is that it is steeper

in the loss region than in the gain region. Thus, participants given

the following choice of gambles

A: Gain $10 with 1/2 probability and lose $10 with 1/2 probability

B: Nothing with certainty

prefer B because they weight the loss of $10 more strongly than the gain of $10.

Kahneman and Tversky (1984) also argued that, as with subjective utility,

people associate a subjective probability with an event that is not identical

with the objective probability. They proposed the function in Figure 11.8 to

relate subjective probability to objective probability. According to this function,

very low probabilities are overweighted relative to high probabilities, producing

a bowing in the function. Thus, a participant might prefer a 1% chance of $400

to a 2% chance of $200 because 1% is not represented as half of 2%. Kahneman

and Tversky (1979) showed that a great deal of human decision making can be

explained by assuming that participants are responding in terms of these subjective

utilities and subjective probabilities.

An interesting question is whether the subjective functions in Figures 11.7

and 11.8 represent irrational tendencies. Generally, the utility function in Figure

11.7 is thought to be reasonable. As we get more money, getting even more

seems less and less important. Certainly, the amount of happiness that a billion

dollars can buy is not 1,000 times the amount of happiness that a million

dollars can buy. It should be noted that not all the utility functions of different

people are like that shown in Figure 11.7, which represents a sort of average.

One can imagine someone needing $10,000 for an important medical procedure.

Then, all sums less than $10,000 would be rather useless, and all sums

greater than $10,000 would be relatively equally good. Thus, such a person

might have a step in the utility function at $10,000.

There is less agreement about how we should assess the subjective probability

function in Figure 11.8. I (Anderson, 1990) have argued that it might actually

make sense to discount the extremity of low probabilities in the way that function

does. The argument is that, sometimes when we are told that probabilities

are extreme, we are being misinformed (see the third Question for Thought at

the end of the chapter). However, there is little consensus in the field about how

to evaluate the subjective probability function.

People make decisions under uncertainty in terms of subjective utilities and

subjective probabilities.

Decision Making | 313

Subjective probability

Objective probability

0

.5

.5

1.0

1.0

FIGURE 11.8 A function that

relates subjective probability

to objective probability.

(From Kahneman & Tversky, 1984.

Reprinted by permission from the

American Psychological Association.)

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 313

Framing Effects

Although one might view the functions in Figures 11.7 and 11.8 as reasonable,

there is evidence that they can lead people to do rather strange things. These

demonstrations deal with framing effects. These effects refer to the fact that

people’s decisions vary, depending on where they perceive themselves to be on

the utility curve in Figure 11.7. In an example from Kahneman and Tversky

(1984), someone must purchase a $15 item versus a $125 item. If another store

offers a $5 discount off the $15, the person is likely to make an effort to go to

the other store, whereas he is not likely to do so if the same $5 discount is

offered on the $125 item. However, in both cases, it is the same $5 savings, and

the question is simply whether one’s time is worth the $5. However, the two

contexts place the person on different points of the utility curve, which is negatively

accelerated. According to that curve, the difference between $15 and $10

is larger than the difference between $125 and $120. Thus, in the first case, the

saving seems worth it, but in the second case, it does not.

Another example has to do with betting behavior. Consider someone who

has lost $140 at the racetrack and has an opportunity to bet $10 on a horse that

will pay 15 to 1. The bettor can view this choice in one of two ways. In one way,

it becomes this choice:

A. Refuse the bet and accept a certainty of losing $140.

B. Make the bet and face a good chance of losing $150 and a poor chance of breaking

even.

Because the subjective difference between losing $140 and $150 is small, the

person will likely choose B and make the bet. On the other hand, the bettor

could view it as the following choice:

C. Refuse the bet and face the certainty of having nothing change.

D. Make the bet and face a good chance of losing an additional $10 and a poor chance

of gaining $140.

In this case, because of the greater weight on losses than on gains and because

of the negatively accelerated utility function, the bettor is likely to avoid the bet.

The only difference is whether one places oneself at the –$140 point or the 0

point on the curve in Figure 11.7. However, one gets a different evaluation of

the two outcomes, depending on where one places oneself.

As an example that appears to be more consequential, consider this situation

described by Kahneman and Tversky (1984):

Problem 1: Imagine that the U.S. is preparing for the outbreak of an unusual

Asian disease, which is expected to kill 600 people. Two alternative programs

to combat the disease have been proposed. Assume that the exact scientific

estimates of the consequences of the programs are as follows:

If Program A is adopted, 200 people will be saved.

If Program B is adopted, there is a one-third probability that 600 people will be saved

and a two-thirds probability that no people will be saved.

Which of the two programs would you favor?

Seventy-two percent of the participants preferred program A, which guarantees

lives, to dealing with the risk of program B. However, consider what

314 | Judgment and Decision Making

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 314

happens when, rather than describing the two programs in regard to saving

lives, the two programs are described as follows:

If Program C is adopted, 400 people will die.

If Program D is adopted, there is a one-third probability nobody will die and a twothirds

probability that 600 people will die.

With this description, only 22% preferred program C, which the reader will

recognize as equivalent to A (and D is equivalent to B). Both of these choices

can be understood in terms of a negatively accelerated utility function for

lives. In the first case, the subjective value of 600 lives saved is less than three

times the subjective value of 200 lives saved, whereas in the second case, the

subjective value of 400 deaths is more than two-thirds the subjective value of

600 deaths. McNeil, Pauker, Cox, and Tversky (1982) found that this tendency

extended to actual medical treatment. What treatment a doctor will choose

depends on whether the treatment is described in terms of odds of living or

odds of dying.

Situations in which framing effects are most prevalent tend to have one

thing in common—no clear basis for choice. This commonality is true of the

three examples that we have reviewed. In the case in which the shopper has an

opportunity for a saving, whether $5 is worth going to another store is unclear.

In the gambling example, there is no clear basis for making a decision.3 The

stakes are very high in the third case, but it is, unfortunately, one of those social

policy decisions that defy a clear analysis. Thus, these cases are hard to decide

on their merits alone.

Shafir (1993) suggested that, in such situations, we may make a decision not

on the basis of which decision is actually the best one but on the basis of which

will be easiest to justify (to ourselves or to others). Different framings make it

easier or harder to justify an action. In the disease example, the first framing

focuses one on saving lives and the second framing focuses one on avoiding

deaths. In the first case, one would justify the action by pointing to the people

whose lives have been saved (therefore it is critical that there be some people to

point to). In the second case, a justification would have to explain why people

died (and it would be better if there were no such people).

This need to justify one’s action can lead one to pick the same alternative

whether asked to pick something to accept or something to reject. Consider the

example in Table 11.2 in which two parents are described in a divorce case and

participants are asked to play the role of a judge who must decide to which parent

to award custody of the child. In the award condition, participants are asked

to decide who is to be awarded custody; in the deny condition, they are asked to

decide who is to be denied custody. The parents are overall rather equivalent, but

parent B has rather more extreme positive and negative factors. Asked to make

an award decision, more participants choose to award custody to parent B; asked

to make a deny decision, they tend to deny custody, again, to parent B. The

reason, Shafir argued, is that parent B offers reasons, such as a close relation with

Decision Making | 315

3 That is, there is no basis for making the gambling decision that would not have rejected gambling as

irrational in the first place.

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 315

the child, that can be used to justify the awarding of custody, but parent B also

has reasons, such as time away from home, to justify denying custody of the

child to that parent.

An interesting study in framing was performed by Greene, Sommerville,

Nystrom, Darley, and Cohen (2001). They compared ethical dilemmas such as

the following pair. In the first dilemma, a runaway trolley is headed for five

people who will be killed if it proceeds on its current course. The only way to

save them is to hit a switch that will turn the trolley onto an alternate set of

tracks where it will kill one person instead of five. The second dilemma is like the

first, except that you are standing next to a large stranger on a footbridge

that spans the tracks in between the oncoming trolley and the five people. In

this scenario, the only way to save the five people is to push the stranger off the

bridge onto the tracks below. He will die, but his body will stop the trolley from

reaching the others. In the first case, most people are willing to sacrifice one

person to save five, but in the second case, they are not.

In an fMRI study, Greene et al. compared the brain areas activated when

people considered an impersonal dilemma such as the first case, with the brain

areas activated when people considered a personal dilemma such as the second.

In the impersonal case, the regions of the parietal cortex that are associated

with cold calculation were active. On the other hand, when they judged the personal

case, regions of the brain associated with emotion (such as the ventromedial

prefrontal cortex that we discussed in the beginning of the chapter) were

active. Thus, part of what can be involved in the different framing of problems

seems which brain regions are engaged.

316 | Judgment and Decision Making

TABLE 11.2

Imagine that you serve on the jury of an only-child sole-custody case following a

relatively messy divorce. The facts of the case are complicated by ambiguous

economic, social, and emotional considerations, and you decide to base your

decision entirely on the following few observations.

(To which parent would you award sole custody of the child?/To which parent would

you deny sole custody of the child?)

Decisions

Award Deny

Parent A Average income 36% 45%

Average health

Average working hours

Reasonable rapport with the child

Relatively stable social life

Parent B Above-average income 64% 55%

Very close relation with the child

Extremely active social life

Lots of work-related travel

Minor health problems

Adapted from Shafir, 1993. Reprinted by permission from the Psychonomic Society, Inc.

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 316

When there is no clear basis for making a decision, people are influenced by

the way in which the problem is framed.

Neural Representation of Subjective Utility and Probability

The subjective utility of an outcome appears to be related to the activity of

dopamine neurons in the basal ganglia. The importance of this region to motivation

has been known since the 1950s, when Olds and Milner (1954) discovered

that rats would press a lever to the point of exhaustion to receive electrical

Decision Making | 317

risk. Reyna and Farley argue that adults don’t think

through the potential costs and benefits of a risky

behavior, but rather they simply

recognize the risk and avoid the situation—

just as the chess masters

discussed in Chapter 9 could recognize

the risk of a potential chess

position. In contrast, adolescents

often have to try to reason through

the consequences of a situation,

much as a chess duffer does, and

can make errors in reasoning.

2. Different values and situations. Risky behavior has

benefits such as immediate pleasure, and adolescents

value these benefits more. Adolescents are

particularly likely to weight the benefits of risky

behavior heavily in the context of their peers,

where social acceptance is at stake. Thus their

utilities in computing expected value are different.

Reyna and Farley speculate that this is related

to the fact that brain regions like the

ventromedial prefrontal cortex continue to mature

into the early 20s. Fischhoff also notes that risky

behavior often arises when adolescents attempt

to establish independence and personal competence,

which are important to achieve. However,

this can put adolescents in situations where older

adults seldom find themselves. If adults found

themselves in similar situations, they might find

themselves also acting in a more risky manner.

Implications

Why are adolescents more likely to make bad decisions?

One of society’s great concerns is risk taking in adolescents.

Compared to older adults, adolescents are more

likely to engage in risky sexual behavior,

abuse drugs and alcohol, and

drive recklessly. Such poor adolescent

choices are the leading cause

of death in adolescence and can

lead to a lifetime of suffering due to

such things as failed education, destroyed

personal relationships, and

addiction to cigarettes, alcohol, and

other drugs. This has been a subject

of a great deal of research (e.g., Fischhoff, 2008; Reyna &

Farley, 2006) and the results are a bit surprising. Contrary

to common belief, adolescents do not perceive themselves

to be any more invulnerable than older adults do

and often perceive greater danger from risky behavior than

do older adults. Also in many laboratory studies, late adolescents

often show as good or better performance as

older adults on abstract tasks of reasoning and decision

making (this will be discussed further in Chapter 14). Thus,

it does not appear that adolescents are poorer thinkers

about risk than older adults. Rather, it appears that the

explanation is involves two classes of factors:

1. Knowledge and experience. Adolescents lack some

of the information that adults have. For instance,

adolescents may know it is important to “practice

safe sex” but not know all that they should about

how to practice safe sex. Also, through experience

adults have become experts on reasoning about

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 317

stimulation from electrodes near this region. This stimulation caused release of

dopamine in a region of the basal ganglia called the nucleus accumbens. Drugs

like heroin and cocaine have their effect by producing increased levels of

dopamine from this region. These dopamine neurons show increased activity

for all sorts of positive rewards including basic rewards like food and sex, but

also social rewards like money or sports cars (Camerer, Loewenstein, & Prelec,

2005). Thus they appear to be the neural equivalent of subjective utility.

Although one can record activity in dopamine neurons while animals do

things like touch morsels of food, much of the knowledge of the human system

comes from neural imaging studies. In one fMRI study, Knutson, Taylor,

Kaufman, Peterson, and Glover (2005) presented participants with various uncertain

outcomes. For instance, on one trial participants might be told that they

had a 50% chance of winning $5; on another trial that they had a 50% chance

of winning $1. Knutson et al. imaged the activity while participants contemplated

the gamble and before they actually saw the results of the gamble. The

magnitude of the fMRI response in the nucleus accumbens reflected the differential

magnitude of these rewards. However, this region does not respond differently

to probability of reward. For instance, it did not respond differentially

when participants were told on one trial that they had an 80% probability of a

reward versus a 20% probability on another trial. In this case, the ventromedial

prefrontal cortex responded to probability of the reward although it had not responded

to the magnitude of the reward. Figure 11.9 illustrates the contrasting

response of these regions to reward magnitude and reward probability.

Although the Knutson et al. study found the ventromedial prefrontal region

only responding to probabilities, other research suggests it is involved in integrating

probabilities and utilities. The ventromedial region is that portion that

was destroyed in Phineas Gage (see Figure 11.1) and his problems went beyond

318 | Judgment and Decision Making

(a)

_0.20

0.20

NAcc

cue ant rsp

MPFC

**

*

0.15

0.10

0.05

0

_0.05

_0.10

_0.15

0 2 4 6 8 10 12 14

Seconds

% Signal change (SEM)

(b)

_0.20

0.20

0.15

0.10

0.05

0

_0.05

_0.10

_0.15

0 2 4 6 8 10 12 14

Seconds

% Signal change (SEM)

*

_$5.00/50%

_$1.00/50%

_$5.00/80%

_$5.00/20%

FIGURE 11.9 (A) The magnitude of a reward is represented in the activity of the nucleus

accumbens (A); (B) the probability of a reward is represented in the activity of the

ventromedial prefrontal cortex. (From Knutson et al., 2005.)

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 318

judging probabilities. Subsequent research has confirmed that people who have

damage to this region do have difficulty in responding adaptively in situations

where they experience good and bad outcomes with different probabilities. For

instance, this has been studied extensively in a task known as the Iowa gambling

task (Bechara, Damasio, Damasio, & Anderson, 1994; Bechara, Damasio,

Tranel, & Damasio, 2005), illustrated in Figure 11.10. The participants choose

cards from four decks. In this version of the problem, decks A and B are equivalent

and decks C and D are equivalent. Every time one selects from deck A

or B, the participant will gain $100 dollars but 1 time out of 10 will also lose

$1250 dollars. So, applying our formula for expected value, the expected value

of selecting a card from one of these decks is

$100 _ 0.1 _ $1250 __$25

or equivalently if participants play these decks 10 trials they can expect to lose

$250. Every time they select a card from decks C and D, they get only $50, but

they also only lose $250 on that 1 out of every 10 draws. The expected value of

selecting from one of these desks is

$50 _ 0.1 _ $250 __$25

and so choosing from these decks, participants can expect to make $250 every

10 trials. Players are initially attracted to decks A and B because of their higher

payoff, but normal participants eventually learn to avoid them. In contrast,

patients with ventromedial damage keep coming back to the high-paying decks.

Also, unlike normal participants, they do not show measures of emotional

engagement (such as increased galvanic skin response) when they choose from

these desks.

Decision Making | 319

“Bad” decks

A B C D

The Iowa gambling task

Gain per card $100

$1250

_$250

$100

$1250

_$250

$50

$250

_$250

$50

$250

_$250

Loss per 10 cards

Net per 10 cards

“Good” decks

FIGURE 11.10 A schematic diagram of the Iowa Gambling Task. The participants are given four

decks of cards, a loan of $2000 facsimile U.S. bills, and asked to play so as to win the most

money. Turning each card carries an immediate reward ($100 in decks A and B and $50 in

decks C and D). Unpredictably, however, the turning of some cards also carries a penalty

(which is large in decks A and B and small in decks C and D). Playing mostly from decks A

and B leads to an overall loss. Playing mostly from decks C and D leads to an overall gain.

(From Bechara et al., 2005.)

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 319

Dopamine activityin the nucleus accumbens reflects the magnitude of

reward, whereas the human ventromedial cortex is involved in integrating

probabilities with reward.

Conclusions

Decision making deals with choosing actions that can have real consequences

in the presence of real uncertainty. All mammals have the dopamine system

that we just described, which gives them a basic ability to seek things that are

rewarding and avoid things that are harmful. However, humans, by virtue of

their greatly expanded prefrontal cortex, have the capacity to reflect on their

circumstances and select actions other than what their more primitive systems

might urge. Research suggests that the ventromedial portion of the human prefrontal

cortex, which is greatly expanded in size even in comparison to the

genetically similar apes, might play a particularly important role in such regulation.

Humans attempt acts of self-regulation—for example, diet plans—that are

far beyond the reach of any other species. However, we live in an uncertain

world, as witnessed by all the contradictory claims made for various diet plans.

Perhaps if we understood better how people responded to such uncertainty and

contradiction, we would also be in a better position to understand why there

are so many failures of our good resolutions.

320 | Judgment and Decision Making

1. Consider the Monte Hall problem:

Suppose you’re on a game show, and you’re given the choice of

three doors: Behind one door is a car; behind the others, goats.

You pick a door—for example, door 1—and the host, who

knows what’s behind the doors, opens another door—for example,

door 3—that has a goat. He then says to you, “Do you

want to pick door 2?” Is it to your advantage to switch your

choice? (Whitaker, 1990, p. 16)

This can be analyzed using the following form of

Bayes’s theorem:

P(H2)P(E3|H2)

P(H2|E3) _ ___________________________________________

P(H1)P(E3|H1) _ P(H2)P(E3|H2)+P(H3)P(E3|H3)

where P(H2|E3) is the probability that the car is behind

door 2 given that the host has opened door 3. P(H1),

P(H2), and P(H3) are the prior probabilities that the

car is behind each door and all three are 1_3. P(E3|H1),

P(E3|H2), and P(E3|H3) are the conditional probabilities

that the host opens each door given each hypothesis.

In calculating these probabilities, keep in mind that

the host cannot open the door you chose and must

open a door that has a goat.

2. Conservatism and Bayes rate neglect seem to be in conflict

(Fischhoff & Beyth-Marom, 1983; Gigerenzer et al.,

1989). Conservatism says that people pay too little

attention to data, whereas Bayes rate neglect says they

only pay attention to evidence and ignore base rates.

Could the contradiction be explained by differences

between studies like Edwards’s that show conservatism

and those like Kahneman and Tversky’s that demonstrate

base-rate neglect?

3. Consult the Web site

http://www.rense.com/general81/dw.htm for a list of

things that people said would never happen.What does

this imply about what our subjective probability should

be when someone informs us that the objective

probability is 0?

4. In the 1980s, it used to be recommended that a

pregnant woman 35 years or older be tested to find

Questions for Thought

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 320

Key Terms | 321

out whether the fetus had Down syndrome. The logic

behind this recommendation was that the probability

of having a Down syndrome baby increases with age

and is about 1_250 for when the expectant mother is

age 35, whereas the probability of the procedure

resulting in a miscarriage was also 1_250. Analyze the

assumptions behind this decision-making criterion

used in the 1980s in terms of the expected-value

calculations described in this chapter. Do you agree

with the recommendation?

Key Terms

Bayes’s theorem

conditional probability

descriptive model

framing effects

gambler’s fallacy

posterior probability

prescriptive model

prior probability

probability matching

recognition heuristic

subjective probability

subjective utility

ventromedial prefrontal

cortex

Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 321

322

12