Problem Solving,
8
Problem Solving
Human ability to solve novel problems greatly surpasses that of any other species,
and this ability depends on the advanced evolution of the prefrontal cortex in
humans. We have already noted the role of the prefrontal cortex in a number of
higher-level cognitive functions: language, imagery, and memory. It is generally
thought that the prefrontal cortex performs more than these specific functions, however,
and plays a major role in the overall organization of behavior. The regions of the
prefrontal cortex that we have discussed so far tend to be ventral (toward the bottom)
and posterior (toward the back), and many of these regions are left lateralized.
In contrast, dorsal (toward the top), anterior (toward the front), and right-hemisphere
prefrontal structures tend to be more involved in the organization of behavior. These
are the prefrontal regions that have expanded the most in the human brain.
Goel and Grafman (2000) describe a patient, PF, who suffered damage to his
right anterior prefrontal cortex as the result of a stroke. Like many patients with damage
to the prefrontal cortex, PF appears normal and even intelligent, and he scored in
the superior range on an intelligence test. In fact, he performed well on most tests,
although he did have difficulty with the Tower of Hanoi problem described later in this
chapter. Nonetheless, for all these surface appearances of normality, there were profound
intellectual deficits. He had been a successful architect before his stroke but
was forced to retire due to loss of the ability to design. He was able to get some work
as a draftsman. Goel and Grafman gave PF a problem that involved redesigning their
laboratory space. Although he was able to speak coherently about the problem, he
was unable to make any real progress on the solution. A comparably trained architect
without brain damage achieved a good solution in a couple of hours. It seems that the
stroke affected only PF’s most highly developed intellectual abilities.
This chapter and Chapter 9 will look at what we know about human problem
solving. In this chapter, we will answer the following questions: • What does it mean to characterize human problem solving as a search of a
problem space? • How do humans learn methods, called operators, for searching the problem
space?
209
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 209
210 | Problem Solving
• How do humans select among different operators for searching a problem
space? • How can past experience affect the availability of different operators and the
success of problem-solving efforts?
•The Nature of Problem Solving
A Comparative Perspective on Problem Solving
Figure 8.1 shows the relative sizes of the prefrontal cortex in various mammals
and illustrates the dramatic increase in humans. This increase supports the
advanced problem solving that only humans are capable of. Nonetheless, one
can find instances of interesting problem solving in other species, particularly
in the higher apes such as chimpanzees. The study of problem solving in other
species offers perspective on our own abilities. Köhler (1927) performed some
of the classic studies on chimpanzee problem solving. Köhler was a famous
German gestalt psychologist who came to America in the 1930s. During World
War I, he found himself trapped on Tenerife in the Canary Islands. On the
island, he found a colony of captive chimpanzees, which he studied, taking
particular interest in the problem-solving behavior of the animals. His best
participant was a chimpanzee named Sultan. One problem posed to Sultan was
FIGURE 8.1 The relative proportions of the frontal lobe given over to the prefrontal cortex in
six mammals. Note that these brains are not drawn to scale and that the human brain is really
much larger in absolute size. (After Fuster, 1989. Adapted by permission of the publisher. © 1989 by Raven Press.)
Squirrel monkey Cat Rhesus monkey
Dog Chimpanzee Human
Brain Structures
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 210
The Nature of Problem Solving | 211
to get some bananas that were outside his cage. Sultan had no difficulty when
he was given a stick that could reach the bananas; he simply used the stick to
pull the bananas into the cage. The problem became harder when Sultan was
provided with two poles, neither of which could reach the food. After unsuccessfully
trying to use the poles to get to the food, the frustrated ape sulked in
his cage. Suddenly, he went over to the poles and put one inside the other, creating
a pole long enough to reach the bananas (Figure 8.2). Clearly, Sultan had
creatively solved the problem.
What are the essential features that qualify this episode as an instance of
problem solving? There seem to be three:
1. Goal directedness. The behavior is clearly organized toward a goal—in
this case, getting the food.
2. Subgoal decomposition. If Sultan could have obtained the food simply
by reaching for it, the behavior would have been problem solving, but
only in the most trivial sense. The essence of the problem solution is that
the ape had to decompose the original goal into subtasks, or subgoals,
such as getting the poles and putting them together.
3. Operator application. Decomposing the overall goal into subgoals is
useful because the ape knows operators that can help him achieve these
subgoals. The term operator refers to an action that will transform the
problem state into another problem state. The solution of the overall
problem is a sequence of these known operators.
Problem solving is goal-directed behavior that often involves setting subgoals
to enable the application of operators.
FIGURE 8.2 Köhler’s ape,
Sultan, solved the two-stick
problem by joining two short
sticks to form a pole long
enough to reach the food
outside his cage. (From Köhler, 1956.
Reprinted by permission of the publisher.
© 1956 by Routledge & Kegan Paul.)
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 211
The Problem-Solving Process: Problem Space and Search
Often, problem solving is described in terms of searching a problem space,
which consists of various states of the problem. A state is a representation of the
problem in some degree of solution. The initial situation of the problem is
referred to as the start state; the situations on the way to the goal, as intermediate
states; and the goal, as the goal state. Beginning from the start state, there are
many ways the problem solver can choose to change the state. Sultan could reach
for a stick, stand on his head, sulk, or try other approaches. Suppose he reaches
for a stick. Now he has entered a new state. He can transform it into another
state—for example, by letting go of the stick (thereby returning to the earlier
state), reaching for the food with the stick, throwing the stick at the food,
or reaching for the other stick. Suppose he reaches for the other stick. Again, he
has created a new state. From this state, Sultan can choose to try, say, walking on
the sticks, putting them together, or eating them. Suppose he chooses to put the
sticks together. He can then choose to reach for the food, throw the sticks away,
or separate them. If he reaches for the food, he will achieve the goal state.
The various states that the problem solver can achieve define a problem
space, also called a state space. Problem-solving operators can be thought of as
ways to change one state in the problem space into another. The challenge is to
find some possible sequence of operators in the problem space that leads from
the start state to the goal state.We can think of the problem space as a maze of
states and of the operators as paths for moving among them. In this model, the
solution to a problem is achieved through search; that is, the problem solver
must find an appropriate path through a maze of states. This conception of
problem solving as a search through a state space was developed by Allen
Newell and Herbert Simon, who were dominant figures in cognitive science
throughout their careers, and it has become the major problem-solving approach,
in both cognitive psychology and AI.
A problem space characterization consists of a set of states and operators
for moving among the states. A good example of problem-space characterization
is the eight-tile puzzle, which consists of eight numbered, movable tiles set
in a 3 _ 3 frame. One cell of the frame is always empty, making it possible to
move an adjacent tile into the empty cell and thereby to “move” the empty cell
as well. The goal is to achieve a particular configuration of tiles, starting from
a different configuration. For instance, a problem might be to transform
212 | Problem Solving
The possible states of this problem are represented as configurations of tiles in
the eight-tile puzzle. So, the first configuration shown is the start state, and the second
is the goal state. The operators that change the states are movements of tiles
into empty spaces. Figure 8.3 reproduces an attempt of mine to solve this problem.
into
2 1 6
4 8
7 5 3
1 2 3
8 4
7 6 5
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 212
My solution involved 26 moves, each move being an operator that changed the
state of the problem. This sequence of operators is considerably longer than necessary.
Try to find a shorter sequence of moves. (The shortest sequence possible is
given in the appendix at the end of the chapter, in Figure A8.1.)
Often, discussions of problem solving involve the use of search graphs or
search trees. Figure 8.4 gives a partial search tree for the following, simpler
eight-tile problem:
The Nature of Problem Solving | 213
(a) (b) (c) (d) (e) (f) (g)
(o) (p) (q) (r) (s) (t) (u)
(n) (m) (l) (k) ( j) (i) (h)
2 1 6
4 8
7 5 3
2 1 6
4 8
7 5 3
(w) (v)
2 6 4
7 5
8 1 3
(x)
2 4
7 6 5
8 1 3
(y)
2 4
7 6 5
8
1 3
(z)
2 4
7 6 5
8
1 3
Goal state
2
7 6 5
8 4
1 3
8 4
6
2 7 5
1 3
8
6
4
2 7 5
1 3
8 4
2 7 5
1 3
6 8 4
2 7 5
6
1 3 8
4
2 7 5
6
1 3
2
7 5
6 4
8 1 3
2 4
8
1 6
7 5
3
2
8
1 6
4
7 5
3
2
1 6
8
4
7 5
3
2 8
7 5
1 4 6
3 2
7 8 5
1 4 6
3 2 4
7 8 5
1 6
3
1 6
2 4
7 8 5
3
2
1 6
4 8
7 5 3
1 6
2 4 8
7 5 3
1 6
2 8
4
7 5 3
2 8
1 4 6
7 5 3
8 4
7 5
2
1 6 3
2
1
4
7 5
6 3
8
4
5
3
7
8 1
2
6
into
2
1
8
4
7 5
3 1 2 3
8 4
7 6 5
6
FIGURE 8.3 The author’s sequence of moves for solving an eight-tile puzzle.
Figure 8.4 is like an upside-down tree with a single trunk and branches leading
out from it. This tree begins with the start state and represents all states reachable
from this state, then all states reachable from those states, and so on. Any
path through such a tree represents a possible sequence of moves that a problem
solver might make. By generating a complete tree, we can also find the shortest
sequence of operators between the start state and the goal state. Figure 8.4 illustrates
some of the problem space. In discussions of such examples, often only a
path through the problem space that leads to the solution is presented (for
instance, see Figure 8.3). Figure 8.4 gives a better idea of the size of the problem
space of possible moves for this kind of problem.
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 213
This search space terminology describes possible steps that the problem
solver might take. It leaves two important questions that we need to answer
before we can explain the behavior of a particular problem solver. First, what
determines the operators available to the problem solver? Second, how does the
problem solver select a particular operator when there are several available? An
answer to the first question determines the search space in which the problem
solver is working. An answer to the second question determines which path the
problem solver takes. We will discuss these questions in the next two sections,
focusing first on the origins of the problem-solving operators and then on the
issue of operator selection.
214 | Problem Solving
2
1 6
8
4
7 5
3
2
1
6
8
4
7 5
3
2
1
6
8
4
7 5
3
2
1
6
8
4
7 5
3
2
1
8 6
4
7 5
3
2
1
6 8 4
7 5
3 2
1
6
8
4
7 5
3 2
1
6
8
7 4
5
3
2
1
6
8
4
7 5
3
2
1
6
8
4
7 5
3
2
1
6 8 4
7 5
3 2
1
6 8 4
7 5
3 2
1
6
8
4
7 5
3
2
1
6
8
4
7
5
3 2
1
6
8
7 4
5
3 2
1
6
8
7 4
5
3
FIGURE 8.4 Part of the search tree, five moves deep, for an eight-tile problem. (After Nilsson, 1971.
Adapted by permission of the publisher. © 1971 by McGraw-Hill.)
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 214
Problem-solving operators generate a space of possible states through which
the problem solver must search to find a path to the goal.
•Problem-Solving Operators
Acquisition of Operators
There are at least three ways to acquire new problem-solving operators.We can
acquire new operators by discovery, by being told about them, or by observing
someone else use them.
Problem-Solving Operators | 215
2
1 6
8
4
7 5
2 3
1
6
8
4
7 5
3
2 1
6
8
4
7 5
3 2
1
6
8
7 4
5
3 2
1
6
8 4
7 5
3 2
1
6
8 4
7 5
3
2 1
6
8
4
7 5
3 2
1
6
8
7 4
5
3 1 2
6
8 4
7 5
3
2
1
6
8
4
7 5
3 2
1
6
8 4
7 5
3 2
1
6
8
4
7 5
3
2
1 6
8
4
7 5
3
2 1
6
8
4
7 5
3
2
1
6
8
4
7 5
3 2
6 1
8
7 4
5
3 2
1
6
8
7 4
5
3 1
7 5
2
8 4
6
3 1 2
6
7 8 4
5
3
Goal state
Start state
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 215
Discovery. We might find that a new service station has opened nearby and
so learn by discovery a new operator for repairing our car. Children might
discover that their parents are particularly susceptible to temper tantrums and
so learn a new way to get what they want. We might discover how a new
microwave oven works by playing with it and so learn a new way to prepare
food. Or a scientist might discover a new drug that kills bacteria and so invent a
new way of combating infections. Each of these examples involves a variety of
reasoning processes. These processes will be the topic of Chapter 10.
Although discovery can involve complex reasoning in humans, it is interesting
that it is the only method that most other creatures have to learn new operators,
and they certainly do not engage in complex reasoning. In a famous study
reported in 1898, Thorndike placed cats in “puzzle boxes.” The boxes could be
opened by various nonobvious means. For instance, in one box, if the cat hit
a loop of wire, the door would fall open. The cats, who were hungry, were rewarded
with food when they got out. Initially, a cat would move about randomly,
clawing at the box and behaving ineffectively in other ways until it
happened to hit the unlatching device. After repeated trials in the same puzzle
box, the cats eventually arrived at a point where they would immediately hit
the unlatching device and get out. A controversy exists to this day over whether
the cats ever really “understood” the new operator they had acquired or just
gradually formed a mindless association between being in the box and hitting
the unlatching device. More recently it has been argued that it need not be an
either–or situation. Daw, N.D., Niv, Y., and Dayan, P. (2005) review evidence that
there are two bases for learning such operators from experience—one involves
the basal ganglia (see Figure 1.8), where simple associations are gradually reinforced,
whereas the other involves the prefrontal cortex and a mental model
of how these operators work. It is reasonable to suppose that the second system
becomes more important in mammals with larger prefrontal cortices.
Learning by Being Told or by Example. We can acquire new operators by
being told about them or by observing someone else use them. These are examples
of social learning. The first method is a uniquely human accomplishment because
it depends on language. The second is a capacity thought to be common in primates:
“Monkey see, monkey do.” As we will see, however, the capacity of nonhuman
primates for learning by imitation has often been overestimated.
It might seem that the most efficient way to learn new problem-solving
operators would be simply to be told about them, but seeing an example is often
at least as effective as being told what to do. Table 8.1 shows two forms of
instruction about an algebraic concept, called a pyramid expression, which is
novel to most undergraduates. Students either study part (a), which gives a
semiformal specification of what a pyramid expression is, or they study part (b),
which gives the single example of a pyramid expression. After reading one
instruction or the other, they are asked to evaluate pyramid expressions like
10$2
Which form of instruction do you think would be most useful? Carnegie
Mellon undergraduates show comparable levels of learning from the single
216 | Problem Solving
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 216
example in part (b) to what they learn from the rigorous specification in part (a).
Sometimes, examples can be the superior means of instruction. For instance,
Reed and Bolstad (1991) had participants learn to solve problems such as the
following:
An expert can complete a technical task in five hours, but a novice requires
seven hours to do the same task. When they work together, the novice works
two hours more than the expert. How long does the expert work? (p. 765)
Participants received instruction in how to use the following equation to solve
the problem:
rate1 _ time1 _ rate2 _ time2 _ tasks
The participants needed to acquire problem-solving operators for assigning
values to the terms in this equation. The participants either received abstract
instruction about how to make these assignments or saw a simple example of
how the assignments were made. There was also a condition in which participants
saw both the abstract instruction and the example. Participants given the
abstract instruction were able to solve only 13% of a set of later problems; participants
given an example solved 28% of the problems; and participants given
both instruction and an example were able to solve 40%.
Why would giving examples be better for learning problem-solving operators
than telling someone what to do directly? The problem with direct instruction
is that it can often be difficult to understand what such quantities as rate1
refer to. This information can be clearer in the context of an example. On the
other hand, it can be difficult to see how to extend an example solution from
one problem to another problem. Thus, experiments like Reed and Bolstad’s
indicate that the best learning occurs when participants have access to both
methods. Similar results have been obtained by Fong, Krantz, and Nisbett
(1986) in the domain of statistics and by Cheng, Holyoak, Nisbett, and Oliver
(1986) in the domain of logic.
Problem-Solving Operators | 217
TABLE 8.1
Instruction for Pyramid Problems
(a) Direct Specification
N$M is a pyramid expression for designating repeated addition where each
term in the sum is one less than the previous.
N, the base, is first term in the sum.
M, the height, is number of terms you add to the base.
(b) Just an Example
7$3 is an example of a pyramid expression.
7$3 _ 7 _ 6 _ 5 _ 4 _ 22
7 is the 3 is the
base height
M
a
a
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 217
Problem-solving operators can be acquired by discovery, by modeling example
problem solutions, or by direct instruction.
Analogy and Imitation
Analogy is the process by which a problem solver extracts the operators used
to solve one problem and maps them onto a solution for another problem.
Sometimes, the analogy process can be straightforward. For instance, a student
may take the structure of an example worked out in a section of a mathematics
text and map it into the solution for a problem in the exercises at the end
of the section. At other times, the transformations can be more complex.
Rutherford, for example, used the solar system as a model for the structure of
the atom, in which electrons revolve around the nucleus of the atom in the
same way as the planets revolve around the sun (Koestler, 1964; Gentner,
1983—see Table 8.2). Although this is a particularly famous example of an
analogy, scientists and engineers use such analogies, if often more mundane,
with great frequency. For instance, Christensen and Schunn (2007) found engineers
making 102 analogies in 9 hours of problem solving (see also Dunbar &
Blanchette, 2001).
An example of the power of analogy in problem solving is provided in an
experiment of Gick and Holyoak (1980). They presented their participants with
the following problem, which is adapted from Duncker (1945):
Suppose you are a doctor faced with a patient who has a malignant tumor in
his stomach. It is impossible to operate on the patient, but unless the tumor is
destroyed, the patient will die. There is a kind of ray that can be used to
destroy the tumor. If the rays reach the tumor all at once at a sufficiently high
intensity, the tumor will be destroyed. Unfortunately, at this intensity the
healthy tissue that the rays pass through on the way to the tumor will also be
destroyed. At lower intensities the rays are harmless to healthy tissue, but they
will not affect the tumor either. What type of procedure might be used to
destroy the tumor with the rays, and at the same time avoid destroying the
healthy tissue? (pp. 307–308)
218 | Problem Solving
TABLE 8.2
The Solar System–Atom Analogy
Base Domain: Solar System Target Domain: Atom
The sun attracts the planets. The nucleus attracts the electrons.
The sun is larger than the planets. The nucleus is larger than the electrons.
The planets revolve around the sun. The electrons revolve around the nucleus.
The planets revolve around the sun The electrons revolve around the nucleus
because of the attraction and because of the attraction and weight
weight difference. difference.
The planet Earth has life on it. No transfer.
After Gentner (1983). Adapted by permission of the publisher. © 1983 by LEA, Ltd.
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 218
This is a very difficult problem, and few people are able to solve it. However,
Gick and Holyoak presented their participants with the following story:
A small country was ruled from a strong fortress by a dictator. The fortress was
situated in the middle of the country, surrounded by farms and villages. Many
roads led to the fortress through the countryside. A rebel general vowed to capture
the fortress. The general knew that an attack by his entire army would
capture the fortress. He gathered his army at the head of one of the roads,
ready to launch a full-scale direct attack. However, the general then learned
that the dictator had planted mines on each of the roads. The mines were set so
that small bodies of men could pass over them safely, since the dictator needed
to move his troops and workers to and from the fortress. However, any large
force would detonate the mines. Not only would this blow up the road, but it
would also destroy many neighboring villages. It therefore seemed impossible
to capture the fortress. However, the general devised a simple plan. He divided
his army into small groups and dispatched each group to the head of a different
road.When all was ready he gave the signal and each group marched down
a different road. Each group continued down its road to the fortress so that the
entire army arrived together at the fortress at the same time. In this way, the
general captured the fortress and overthrew the dictator. (p. 351)
Told to use this story as the model for a solution, most participants were able
to develop an analogous operation to solve the tumor problem.
An interesting example of a solution by analogy that did not quite work is a
geometry problem encountered by one student. Figure 8.5a illustrates the steps
of a solution that the text gave as an example, and Figure 8.5b illustrates the
student’s attempts to use that example proof to guide his solution to a homework
problem. In Figure 8.5a, two segments of a line are given as equal length,
and the goal is to prove that two larger segments have equal length. In Figure 8.5b,
the student is given two line segments with AB longer than CD, and his task is to
prove the same inequality for two larger segments, AC and BD.
Our participant noted the obvious similarity between the two problems and
proceeded to develop the apparent analogy. He thought he could
simply substitute points on one line for points on another, and
inequality for equality. That is, he tried to substitute A for R, B
for O, C for N, D for Y, and _ for _.With these substitutions, he
got the first line correct: Analogous to RO _ NY, he wrote AB _
CD. Then he had to write something analogous to ON _ ON, so
he wrote BC _ BC! This example illustrates how analogy can be
used to create operators for problem solving and also shows that
it requires a little sophistication to use analogy correctly.
Another difficulty with analogy is finding the appropriate
examples from which to analogize operators. Often, participants
do not notice when an analogy is possible. Gick and
Holyoak (1980) did an experiment in which they read participants
the story about the general and the dictator and then gave
them Duncker’s (1945) ray problem (both shown earlier in this
section). Very few participants spontaneously noticed the relevance
of the first story to solving the second. To achieve success,
Problem-Solving Operators | 219
R
(a)
O
N
Y
Given: RO = NY, RONY
Prove: RN = OY
RO = NY
ON = ON
RO + ON = ON + NY
RONY
RO + NY = RN
ON + NY = OY
RN = OY
A
(b)
B
C
D
AB > CD
BC > BC
!!!
Given: AB > CD, ABCD
Prove: AC > BD
FIGURE 8.5 (a) A worked-out
proof problem given in a geometry
text. (b) One student’s attempt
to use the structure of this
problem’s solution to guide his
solution of a similar problem.
This example illustrates how
analogy can be used (and
misused) for problem solving.
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 219
participants had to be explicitly told to use the general and dictator story as an
analogy for solving the ray problem.
When participants do spontaneously use previous examples to solve a problem,
they are often guided by superficial similarities in their choice of examples.
For instance, B. H. Ross (1984, 1987) taught participants several methods for
solving probability problems. These methods were taught by reference to specific
examples, such as finding the probability that a pair of tossed dice will sum
to 7. Participants were then tested with new problems that were superficially
similar to prior examples. The similarity was superficial because both the
example and the problem involved the same content (e.g., dice) but not necessarily
the same principle of probability. Participants tried to solve the new
problem by using the operators illustrated in the superficially similar prior
example. When that example illustrated the same principle as required in the
current problem, participants were able to solve the problem.When it did not,
they were unable to solve the current problem. Reed (1987) has found similar
results with algebra story problems.
In solving school problems, students use proximity as a cue to determine
which examples to use in analogy. For instance, a student working on physics
problems at the end of a chapter expects that problems solved as examples in
the chapter will use the same methods and so tries to solve the problems by
analogy to these examples (Chi, Bassok, Lewis, Riemann, & Glaser, 1989).
Analogy involves noticing that a past problem solution is relevant and then
mapping the elements from that solution to produce an operator for the
current problem.
Analogy and Imitation from an Evolutionary
and Brain Perspective
It has been argued that analogical reasoning is a hallmark of human cognition
(Halford, 1992). The capacity to solve analogical problems is almost uniquely
found in humans. There is some evidence for it in chimpanzees (Oden,
Thompson, & Premack, 2001), although lower primates such as monkeys seem
totally incapable of such tasks. For instance, Premack (1976) found that Sarah,
a chimpanzee used in studies of language (see Chapter 11), was able to solve
analogies such as the following: Key is to a padlock as what is to a tin can? The
answer: can opener. In more careful study of Sarah’s abilities, however, Oden et
al. found that although Sarah could solve these problems more often than
chance, she was much more prone to error than human participants.
Recent brain-imaging studies have looked at the cortical regions that are
activated in analogical reasoning. Figure 8.6 shows examples of the stimuli used
in a study by Christoff et al. (2001), adapted from the Raven’s Progressive
Matrices test, which is a standard test of intelligence. Only the problems in
Figure 8.6c, which require that the solver coordinate two dimensions, could be
said to tap true analogical reasoning. There is evidence that children under age 5
(in whom the frontal cortex has not yet matured), nonhuman primates, and
220 | Problem Solving
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 220
patients with frontal damage all have special difficulty with
problems like those shown in Figure 8.6c and often just cannot
solve them. Christoff et al. were interested in discovering
which brain regions would be activated when participants
were solving these problems. Consistent with the trends we
noted in the introduction to this chapter, they found that the
right anterior prefrontal cortex was activated only when participants
were solving these 2-D problems.
Examples like those shown in Figure 8.6 are cases in which
analogical reasoning is used for purposes other than acquiring
new problem-solving operators. From the perspective of this
chapter, however, the real importance of analogy is that it can
be used to acquire new problem-solving operators. We noted
earlier that people often learn more from studying an example
than from reading abstract instructions. Humans have a special
ability to mimic the problem solutions of others. When
we ask someone how to use a new device, that person tends to
show us how, not to tell us how. Despite the proverb “Monkey
see, monkey do,” even the higher apes are quite poor at imitation
(Tomasello & Call, 1997). Thus, it seems that one of
the things that make humans such effective problem solvers
is that we have special abilities to acquire new problem-solving
operators by analogical reasoning.
Analogical problem solving appears to be a capability nearly unique to
humans and to depend on the advanced development of the prefrontal
cortex.
•Operator Selection
As noted earlier, in any particular state, multiple problem-solving operators can
be applicable, and a critical task is to select the one to apply. In principle, a
problem solver may select operators in many ways, and the field of AI has
succeeded in enumerating various powerful techniques. However, it seems that
most methods are not particularly natural as human problem-solving approaches.
Here we will review three criteria that humans use to select operators.
Backup avoidance biases the problem solver against any operator that undoes
the effect of the previous operators. For instance, in the eight-tile puzzle,
people show great reluctance to take back a step even if this might be necessary
to solve the problem. However, backup avoidance by itself provides no basis for
choosing among the remaining operators.
Humans tend to select the nonrepeating operator that most reduces the difference
between the current state and the goal. Difference reduction is a very
general principle and describes the behavior of many creatures. For instance,
Köhler (1927) described how a chicken will move directly toward desired food
Operator Selection | 221
(a)
1 2
3 4
1 2
3 4
1 2
3 4
(b)
(c)
FIGURE 8.6 Examples of stimuli
used by Christoff et al. to study
which brain regions would be
activated when participants
attempted to solve three
different types of analogy
problem: (a) 0-dimensional;
(b) 1-dimensional; and
(c) 2-dimensional. The task
in each case was to infer the
missing figure and select it
from among the four alternative
choices. (After Christoff et al., 2001.
Adapted by permission of the publisher.
© 2001 by Neuroimage.)
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 221
and will not go around a fence that is blocking it. The poor creature is effectively
paralyzed, being unable to move forward and unwilling to back up and
undo its approach to the fence. It does not seem to have any principles for selection
of operators other than difference reduction and backup avoidance.
This leaves it without a solution to the problem.
On the other hand, the chimpanzee Sultan (see Figure 8.2) did not just claw
at his cage trying to get the bananas. He sought to create a new tool to enable
him to obtain the food. In effect, his new goal became the creation of a new
means for achieving the old goal. Means-ends analysis is the term used to
describe the creation of a new goal (end) to enable an operator (means) to
apply. By using means-ends analysis, humans and other higher primates can be
more resourceful in achieving a goal than they could be if they used only difference
reduction. In the next sections, we will discuss the roles of both difference
reduction and means-ends analysis in operator selection.
Humans use backup avoidance, difference reduction, and means-ends
analysis to guide their selection of operators.
The Difference-Reduction Method
A common method of problem solving, particularly in unfamiliar domains, is
to try to reduce the difference between the current state and the goal state. For
instance, consider my solution to the eight-tile puzzle in Figure 8.3. There were
four options possible for the first move. One possible operator was to move the
1 tile into the empty square, another was to move the 8, a third was to move the 5,
and the fourth was to move the 4. I chose the last operator. Why? Because it
seemed to get me closer to my end goal. I was moving the 4 tile closer to its
final destination. Human problem solvers are often strongly governed by difference
reduction or, conversely, by similarity increase. That is, they choose operators
that transform the current state into a new state that reduces differences
and resembles the goal state more closely than the current state. Difference
reduction is sometimes called hill climbing. If we imagine the goal as the highest
point of land, one approach to reaching it is always to take steps that go up.
By reducing the difference between the goal and the current state, the problem
solver is taking a step “higher” toward the goal. Hill climbing has a potential
flaw, however: By following it, we might reach the top of some hill that is lower
than the highest point of land that is the goal. Thus, difference reduction is not
guaranteed to work. It is myopic in that it considers only whether the next step
is an improvement and not whether the larger plan will work. Means-ends
analysis, which we will discuss later, is an attempt to introduce a more global
perspective into problem solving.
One way problem solvers improve operator selection is by using more
sophisticated measures of similarity.My first move was intended simply to get a
tile closer to its final destination. After working with many tile problems, we
begin to notice the importance of sequence—that is, whether noncentral tiles are
followed by their appropriate successors. For instance, in state (o) of Figure 8.3,
the 3 and 4 tiles are in sequence because they are followed by their successors 4
222 | Problem Solving
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 222
and 5, but the 5 is not in sequence because it is followed by 7 rather than 6.
Trying first to move tiles into sequence proves to be more important than trying
to move them to their final destinations right away. Thus, using sequence as
a measure of increasing similarity leads to more effective problem solving based
on difference reduction (see Nilsson, 1971, for further discussion).
The difference-reduction technique relies on evaluations of the similarity
between the current state and the goal state. Although difference reduction
works more often than not, it can also lead the problem solver astray. In some
problem-solving situations, a correct solution involves going against the grain
of similarity. A good example is called the hobbits and orcs problem:
On one side of a river are three hobbits and three orcs. They have a boat on their
side that is capable of carrying two creatures at a time across the river. The goal is
to transport all six creatures across to the other side of the river. At no point on
either side of the river can orcs outnumber hobbits (or the orcs would eat the
outnumbered hobbits). The problem, then, is to find a method of transporting
all six creatures across the river without the hobbits ever being outnumbered.
Stop reading and try to solve this problem. Figure 8.7 shows a correct sequence
of moves. Illustrated are the locations of hobbits (H), orcs (O), and the boat
(b). The boat, the three hobbits, and the three orcs all start on one side of the
river. This condition is represented in state 1 by the fact that all are above the
line. Then a hobbit, an orc, and the boat proceed to the other side of the river.
The outcome of this action is represented in state 2 by placement of the boat,
the hobbit, and the orc below the line. In state 3, one hobbit has taken the boat
back, and the diagram continues in the same way. Each state in the figure represents
another configuration of hobbits, orcs, and boat. Participants have a particular
problem with the transition from state 6 to state 7. In a study by Jeffries,
Polson, Razran, and Atwood (1977), about a third of all participants chose to
back up to a previous state 5 rather than moving on to state 7 (see also Greeno,
1974). One reason for this difficulty is that the action involves moving two
creatures back to the wrong side of the river. The move seems to be away from a
solution. At this point, participants will go back to state 5, even though this undoes
their last move. They would rather undo a move than take a step that
moves them to a state that appears further from the goal.
Atwood and Polson (1976) provide another experimental demonstration of
participants’ reliance on similarity and how that reliance can sometimes be
harmful and sometimes beneficial. Participants were given the following water
jug problem:
You have three jugs, which we will call A, B, and C. Jug A can hold exactly
8 cups of water, B can hold exactly 5 cups, and C can hold exactly 3 cups. Jug
A is filled to capacity with 8 cups of water. B and C are empty.We want you to
find a way of dividing the contents of A equally between A and B so that both
have exactly 4 cups. You are allowed to pour water from jug to jug.
Figure 8.8 shows two paths for solving this problem. At the top of the illustration,
all the water is in jug A—represented by A(8); there is no water in jugs B or C—
represented by B(0) C(0). The two possible actions are to pour A into C, in which
case we get A(5) B(0) C(3), or to pour A into B, in which case we get A(3) B(5)
Operator Selection | 223
b H H H O O O
b
b
H
H
H O
O
O
H H H O
O
O
b
H H H
O O O
b H H H
O O
O
b
H
H H O
O
O
b H
H
H
O
O O
b H H
O O
O
b
H H
H
H
O O O
b H H H
O
O O
b
H H H
O O
O
b O O O H H H
(1)
(2)
(3)
(4)
(5)
(6)
(7)
(8)
(9)
(10)
(11)
(12)
FIGURE 8.7 A diagram of the
successive states in a solution
to the hobbits and orcs
problem. H _ hobbits,
O _ orcs, b _ boat.
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 223
C(0). From these two states, more moves can be
made. Numerous other sequences of moves are
possible besides the two paths illustrated, but
these are the two shortest sequences to the goal.
Atwood and Polson used the representation in
Figure 8.8 to analyze participants’ behavior. For
instance, they asked which move participants
would prefer to make at the start state 1. That is,
would they prefer to pour jug A into C and get
state 2, or jug A into B and get state 9? The answer
is that participants preferred the latter move.
More than twice as many participants moved to
state 9 as moved to state 2. Note that state 9 is
quite similar to the goal. The goal is to have 4
cups in both A and B, and state 9 has 3 cups in A
and 5 cups in B. In contrast, state 2 has no cups of
water in B. Throughout the experiment, Atwood
and Polson found a strong tendency for participants
to move to states that were similar to the
goal state. Usually, similarity was a good heuristic,
but there are critical cases where similarity is misleading.
For instance, the transitions from state 5
to state 6 and from state 11 to state 12 both lead
to significant decreases in similarity to the goal.
However, both transitions are critical to their solution
paths. Atwood and Polson found that more
than 50% of the time, participants deviated from
the correct sequence of moves at these critical
points. They instead chose some move that seemed closer to the goal but actually
took them away from the solution.1
It is worth noting that people do not get stuck in suboptimal states only
while solving puzzles. Hill climbing can get us stuck when making serious life
choices. A classic example is someone trapped in a suboptimal job because he
or she is unwilling to get the education needed for a better job. The person is
unwilling to endure the temporary deviation from the goal (of earning as much
as possible) to get the skills to earn an even higher salary.
People experience difficulty in solving a problem at points where the correct
solution involves increasing the differences between the current state and the
goal state.
Means-Ends Analysis
Means-ends analysis is a more sophisticated method of operator selection. This
method has been extensively studied by Newell and Simon, who used it in a
computer simulation program (called the General Problem Solver—GPS) that
224 | Problem Solving
A(5)
A(5)
A(2)
A(2)
A(7)
A(7)
A(4)
A(4) (4)
B(5)
B(2)
B(2)
B(0)
B(0) C(0)
C(0)
C(2)
C(2)
C(3)
C(0)
C (3) A(3)
A(6)
A(6)
B(5)
B(4)
A(1)
A(1)
B(3) A(3) C(3)
B(5)
B(0)
B
B(3) C(3)
C(1)
C (0)
C(3)
C(0)
C(1)
B(1)
B(1)
(1) A(8)
(2)
(3)
(4)
(5)
(6)
(7)
(8)
(15)
(13)
(14)
(12)
(11)
(10)
(9)
B(0) C(0)
A C
C B
A C
C B
B A
C B
A C
C B
C A
B C
A B
B C
C A
B C
A B
FIGURE 8.8 Two paths of
solution for the water jug
problem posed in Atwood and
Polson (1976). Each state is
represented in terms of the
contents of the three jugs; for
example, in state 1, A(8) B(0)
C (0). The transitions between
states (e.g., A →C) are labeled
in terms of which jug is poured
into which other jug.
1 For instance, moving back to state 9 from either state 5 or state 11.
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 224
models human problem solving. The following is their description of meansends
analysis.
Means-ends analysis is typified by the following kind of commonsense
argument:
I want to take my son to nursery school.What’s the difference between what I
have and what I want? One of distance. What changes distance? My automobile.
My automobile won’t work. What is needed to make it work? A new
battery.What has new batteries? An auto repair shop. I want the repair shop to
put in a new battery; but the shop doesn’t know I need one.What is the difficulty?
One of communication. What allows communication? A telephone . . .
and so on.
This kind of analysis—classifying things in terms of the functions they serve
and oscillating among ends, functions required, and means that perform
them—forms the basic system of GPS. (Newell & Simon, 1972, p. 416)
Means-ends analysis can be viewed as a more sophisticated version of difference
reduction. Like difference reduction, it tries to eliminate the differences
between the current state and the goal state. For instance, in this example, it
tried to reduce the distance between the home and the nursery school. Meansends
analysis will also identify the biggest difference first and try to eliminate it.
Thus, in this example, the focus is on difference in the general location of home
and nursery school. The difference between where the car will be parked and
the classroom has not been considered.
Means-ends analysis offers a major advance over difference reduction because
it will not abandon an operator if it cannot be applied immediately. If the car did
not work, for example, difference reduction would have one start walking to the
nursery school. The essential feature of means-ends analysis is that it focuses on
enabling blocked operators. The means temporarily becomes the end. In effect,
the problem solver deliberately ignores the real goal and focuses on the goal of
enabling the means. In the example we have been discussing, the problem solver
set a subgoal of repairing the automobile, which was the means of achieving the
original goal of getting the child to nursery school. New operators can be selected
to achieve this subgoal. For instance, installing a new battery was chosen. If this
operator is blocked, enabling it can become yet another subgoal.
Figure 8.9 shows two flowcharts of the procedures used in the means-ends
analysis employed by GPS. A general feature of this analysis is that it breaks a
larger goal into subgoals. GPS creates subgoals in two ways. First, in flowchart 1,
GPS breaks the current state into a set of differences and sets the reduction of
each difference as a separate subgoal. First it tries to eliminate what it perceives
as the most important difference. Second, in flowchart 2, GPS tries to find an
operator that will eliminate the difference. However, GPS may not be able to
apply this operator immediately because a difference exists between the operator’s
condition and the state of the environment. Thus, before the operator can
be applied, it may be necessary to eliminate another difference. To eliminate the
difference that is blocking the operator’s application, flowchart 2 will have to be
called again to find another operator relevant to eliminating that difference.
The term operator subgoal is used to refer to a subgoal whose purpose is to
eliminate a difference that is blocking application of an operator.
Operator Selection | 225
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 225
Means-ends analysis involves creating subgoals to eliminate the difference
blocking the application of a desired operator.
The Tower of Hanoi Problem
Means-ends analysis has proved to be a generally applicable and extremely powerful
method of problem solving. Ernst and Newell (1969) discussed its application
to the modeling of monkey and bananas problems (such as Sultan’s predicament
described at the beginning of the chapter), algebra problems, calculus problems,
and logic problems. Here, however, we will illustrate means-ends analysis by
applying it to the Tower of Hanoi problem. Figure 8.10 illustrates a simple version
226 | Problem Solving
Match current state
to goal state to find the
most important difference
Flowchart 1 Goal: Transform current state into goal state
Flowchart 2 Goal: Eliminate the difference
Difference
NO DIFFERENCES
NO DIFFERENCE
NONE FOUND
FAIL
FAIL
FAIL APPLY OPERATOR
FAIL
SUCCESS
SUCCESS
Operator
found
SUCCESS
detected
Difference
detected
Search for operator
relevant to reducing
the difference
Match condition of
operator to current
state to find most
important difference
Subgoal: Eliminate
the difference
Subgoal:
Eliminate
the difference
1 2 3 1 2 3
A
B
C
A
Start Goal
B
C
FIGURE 8.9 The application of means-ends analysis by Newell and Simon’s General Problem
Solving (GPS) program. Flowchart 1 breaks a problem down into a set of differences and tries to
eliminate each one. Flowchart 2 searches for an operator that is relevant to eliminating a difference.
FIGURE 8.10 The three-disk version of the Tower of Hanoi problem.
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 226
of this problem. There are three pegs and three
disks of differing sizes, A, B, and C. The disks
have holes in them so they can be stacked on the
pegs. The disks can be moved from any peg to
any other peg. Only the top disk on a peg can be
moved, and it can never be placed on a smaller
disk. The disks all start out on peg 1, but the
goal is to move them all to peg 3, one disk at a
time, by transferring disks among pegs.
Figure 8.11 traces the application of the
GPS techniques to this problem. The first line
gives the general goal of moving disks A, B, and
C to peg 3. This goal leads us to the first flowchart
of Figure 8.9. One difference between the
goal and the current state is that disk C is not
on peg 3. This difference is chosen because
GPS tries to remove the most important difference
first, and we are assuming that the largest
misplaced disk will be viewed as the most
important difference. A subgoal set up to eliminate
this difference takes us to the second
flowchart of Figure 8.9, which tries to find an
operator to reduce the difference. The operator
chosen is to move C to peg 3. The condition for
applying a move operator is that nothing be on
the disk. Because A and B are on C, there is a
difference between the condition of the operator
and the current state. Therefore, a new subgoal
is created to reduce one of the differences—
B on C. This subgoal gets us back to
the start of flowchart 2, but now with the goal
of removing B from C (line 6 in Figure 8.11).2
The operator chosen the second time in
flowchart 2 is to move disk B to peg 2. However,
we cannot immediately apply the operator
of moving B to 2, because B is covered by A.
Therefore, another subgoal—removing A—is
set up, and flowchart 2 is used to remove this
difference. The operator relevant to achieving
this subgoal is to move disk A to peg 3. There
are no differences between the conditions for this operator and the current state.
Finally, we have an operator we can apply (line 12 in Figure 8.11), and we
achieve the subgoal of moving A to 3. Now we return to the earlier intention of
Operator Selection | 227
Goal: Move A, B, and C to peg 3
: Difference is that C is not on 3
: Subgoal: Make C on 3
: Operator is to move C to 3
: Difference is that A and B are on C
: Subgoal: Remove B from C
: Operator is to move B to 2
: Difference is that A is on B
: Subgoal: Remove A from B
: Operator is to move A to 3
: No difference with operator's condition
: No difference with operator's condition
: No difference with operator's condition
: No difference with operator's condition
: No difference with operator's condition
: Difference is that A is on 3
: Subgoal: Remove A from 3
: Operator is to move A to 2
: Apply operator (move A to 3)
: Apply operator (move A to 2)
: Apply operator (move C to 3)
: Apply operator (move B to 2)
: Apply operator (move B to 3)
: Subgoal achieved
: Subgoal achieved
: Subgoal achieved
: Subgoal achieved
: Subgoal achieved
: Subgoal achieved
: Subgoal achieved
: No difference
Goal achieved
: Difference is that A is not on 3
: Subgoal: Make A on 3
: Operator is to move A to 3
: No difference with operator's condition
: Apply operator (move A to 3 )
: Difference is that B is not on 3
: Subgoal: Make B on 3
: Operator is to move B to 3
: Difference is that A is on B
: Subgoal: Remove A from B
: Operator is to move A to 1
: No difference with operator's condition
: Apply operator (move A to 1)
1.
2.
3.
4.
5.
6.
7.
8.
9.
10.
11.
12.
13.
14.
15.
16.
17.
18.
19.
20.
21.
22.
23.
24.
25.
26.
27.
28.
29.
30.
31.
32.
33.
34.
35.
36.
37.
38.
39.
40.
41.
42.
43.
44.
45.
FIGURE 8.11 A trace of the
application of the GPS program,
as shown in Figure 8.9, to the
Tower of Hanoi problem shown
in Figure 8.10.
2 Note that we have gone from the use of flowchart 1 to the use of flowchart 2, to a new use of flowchart 2.
To apply flowchart 2 to find a way to move disk C to peg 3, we need to apply flowchart 2 to find a way to
remove disk B from disk C. Thus, one procedure is using itself as a subprocedure; such an action is called
recursion.
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 227
moving B to 2. There are no more differences between the condition for this operator
and the current state, and so the action takes place. The subgoal of
removing B from C is then satisfied (line 16 in Figure 8.11).
We have now returned to the original intention of moving disk C to peg 3.
However, disk A is now on peg 3, which prevents the action. Thus, we have
another difference to be eliminated between the now-current state and the
operator’s condition. We move A onto peg 2 to remove this difference. Now
the original operator of moving C to 3 can be applied (line 24 in Figure 8.11).
The state now is that disk C is on peg 3 and disks A and B are on peg 2. At
this point, GPS returns to its original goal of moving the three disks to peg 3. It
notes another difference—that B is not on 3—and sets another subgoal of eliminating
this difference. It achieves this subgoal by first moving A to 1 and then B
to 3. This gets us to line 37 in Figure 8.11. The remaining difference is that A
is not on 3. This difference is eliminated in lines 38 through 42.With this step,
no more differences exist and the original goal is achieved.
Note that subgoals are created in service of other subgoals. For instance,
to achieve the subgoal of moving the largest disk, GPS creates a subgoal of
moving the second-largest disk, which is on top of it. We indicated this logical
dependency of one subgoal on another in Figure 8.11 by indenting the processing
of the dependent subgoal. Before the first move in line 12 of the illustration,
three subgoals had to be created. It appears that creating such goals and subgoals
can be quite costly. Both Anderson, Kushmerick, and Lebiere (1993) and
Ruiz (1987) found that the time required to make one of the moves is a function
of the number of subgoals that must be created. For instance, before disk A
is moved to peg 3 in Figure 8.11 (the first move), three subgoals have to be created,
whereas no subgoals have to be created before the next move is taken—
moving B to peg 2. Correspondingly, Anderson et al. found that it took 8.95 s to
make the first move and 2.46 s to make the second move.
There are two problem-solving methods that participants could bring to
bear in solving the Tower of Hanoi problem. They could use a means-ends
approach as illustrated in Figure 8.11, or they could use the simpler differencereduction
method—in which case they would never set a subgoal to move a disk
that currently cannot be moved. In the Tower of Hanoi problem, such a simple
difference-reduction method would not be effective, because one needs to look
beyond what is currently possible and have a more global plan of attack on the
problem. The only step that difference reduction could take in Figure 8.10 would
be to move the top disk (A) to the target peg (3), but then it would provide no
further guidance because no other move would reduce the difference between
the current state and the goal state. Participants would have to make a random
move. Kotovsky, Hayes, and Simon (1985) studied the way people actually
approach the Tower of Hanoi problem. They found that there was an initial
problem-solving period during which participants did adopt this fruitless
difference-reduction strategy. Then they switched to a means-ends strategy, after
which the solution to the problem came quickly.
The Tower of Hanoi problem is solved by adopting a means-ends strategy
in which subgoals are created.
228 | Problem Solving
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 228
Goal Structuresand Prefrontal Cortex
It is significant that complex goal structures, particularly those involving operator
subgoaling, have been observed with any frequency only in humans and higher
primates.We have already discussed one instance of Sultan’s solution to the twostick
problem (see Figure 8.2). Novel tool building, a clear instance of operator
subgoaling, is almost unique to the higher apes (Beck, 1980). I (Anderson, 1993)
have speculated that the process of handling complex subgoals is performed by the
prefrontal cortex—which, as Figure 8.1 illustrates, is much larger in the higher
primates than in most other mammals, and is larger in humans than in most apes.
Chapter 6 discussed the role of the prefrontal cortex in holding information in
working memory. One of the major prerequisites to developing complex goal
structures is the ability to maintain these goal structures in working memory.
Goel and Grafman (1995) looked at how patients with frontal damage performed
in solving the Tower of Hanoi problem. These patients had suffered
severe damage to the prefrontal cortex. Many were veterans of the Vietnam War
who had lost large amounts of brain tissue as a result of penetrating missile
wounds. Although these patients had normal IQs, they showed much worse performance
than normal participants on the Tower of Hanoi task. There were certain
moves that these patients found particularly difficult to solve. As we noted in
discussing how means-ends analysis applies to the Tower of Hanoi problem, it is
necessary to make moves that deviate from the prescriptions of hill climbing.
One might have a disk at the correct position but have to move it away to enable
another disk to be moved to that position. It was exactly at these points where the
patients had to move “backward” that they had their problems. Only by maintaining
a set of goals can one see that a backward move is necessary for a solution.
More generally, it has been noted that patients with frontal damage have
difficulty inhibiting a predominant response (e.g., Roberts, Hager, & Heron,
1994). For instance, in the Stroop task (see Chapter 3), these patients have
trouble not saying the word itself when they are supposed to say the color of the
word. Apparently, they find it hard to keep in mind that their goal is to say the
color and not the word.
There is increased blood flow in the
prefrontal cortex during many tasks that
involve organizing novel and complex
behavior (Gazzaniga, Ivry, & Mangun, 1998).
Fincham, Carter, van Veen, Stenger, and
Anderson (2002) did an fMRI study of
students while they were solving Tower of
Hanoi problems and looked at brain activation
as a function of the number of goals that
the student had to set. These students were
solving much more complicated problems
than the simple one shown in Figure 8.10.
A problem might involve placing five disks
and could require maintaining as many as
five goals to reach a solution. Figure 8.12
shows the fMRI BOLD response of a region
Operator Selection | 229
1
0.0
0.05
0.10
0.15
0.20
BOLD response
Number of goals
2 3
Step in problem
Number of goals on stack
Increase in BOLD response (%)
4 5 6 7 8
1
2
3
4
FIGURE 8.12 Results from a
study by Fincham et al. to
examine brain activation as a
function of steps while solving
a Tower of Hanoi problem. The
red line shows the magnitude
of fMRI BOLD response in a
region in the right, anterior,
dorsolateral prefrontal cortex
during a sequence of eight
problem-solving steps in which
the number of goals being held
varied from 1 to 4. The black
shows the number of goals
being held at each point.
(Data from Fincham et al., 2002.)
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 229
in the right, anterior, dorsolateral prefrontal cortex during a sequence of eight
problem-solving steps in which the number of goals being held varied from 1
to 4. It also shows the number of goals being held at each point. There seems
to be a striking match between the goal load and the magnitude of the fMRI
response.
The prefrontal cortex plays a critical role in maintaining goal structures.
•Problem Representation
The Importance of the Correct Representation
We have analyzed a problem solution as consisting of problem states and operators
for changing states. So far, we have discussed problem solving as if the
only tasks involved were to acquire operators and select the appropriate ones.
However, there are also important effects of how one represents the problem. A
famous example illustrating the importance of representation is the mutilatedcheckerboard
problem (Kaplan & Simon, 1990). Suppose we have a checkerboard
from which two diagonally opposite corner squares have been cut out.
Figure 8.13 illustrates this mutilated checkerboard, on which 62 squares remain.
Now suppose that we have 31 dominoes, each of which covers exactly two
squares of the board. Can you find some way of arranging these 31 dominoes
on the board so that they cover all 62 squares? If it can be done, explain how. If
it cannot be done, prove that it cannot. Perhaps you would like to ponder this
problem before reading on. Relatively few people are able to solve it without
some hints, and very few see the answer quickly.
The answer is that the checkerboard cannot be covered by the dominoes.
The trick to seeing this is to include in your representation of the problem the
fact that each domino must cover one black and one white
square, not just any two squares. There is just no way to place a
domino on two squares of the checkerboard without having it
cover one black and one white square. So with 31 dominoes, we
can cover 31 black squares and 31 white squares. But the
mutilation has removed two white squares. Thus, there are 30
white squares and 32 black squares. It follows that the mutilated
checkerboard cannot be covered by 31 dominoes.
Why is the mutilated-checkerboard problem easier to solve
when we represent each domino as covering a white and a black
square? The answer is that in so representing the problem, we are
encouraged to compare the number of white and black squares
on the board. Thus, the effect of the problem representation is
that it allows the critical operators to apply (i.e., checking for
parity).
Another problem that depends on correct representation is
the 27-apples problem. Imagine 27 apples packed together in a
crate 3 apples high, 3 apples wide, and 3 apples deep. A worm is
230 | Problem Solving
FIGURE 8.13 The mutilated
checkerboard used in the
problem posed by Kaplan and
Simon (1990) to illustrate the
importance of representation.
(After Wickelgren, 1974. Adapted by
permission of the publisher. © 1974
by W. H. Freeman and Company.)
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 230
in the center apple. Its life’s ambition is to eat its way through all the apples in
the crate, but it does not want to waste time by visiting any apple twice. The
worm can move from apple to apple only by going from the side of one into the
side of another. This means it can move only into the apples directly above, below,
or beside it. It cannot move diagonally. Can you find some path by which
the worm, starting from the center apple, can reach all the apples without going
through any apple twice? If not, can you prove it is impossible? The solution is
left to you. (Hint: The solution is based on a partial 3-D analogy to the solution
for the mutilated-checkerboard problem; it is given in the appendix at the end
of the chapter.) Inappropriate problem representations often cause students to
fail to solve problems even though they have been taught the appropriate
knowledge. This fact often frustrates teachers. Bassok (1990) and Bassok and
Holyoak (1989) studied high-school students who had learned to solve such
physics problems as the following:
What is the acceleration (increase in speed each second) of a train, if its speed
increases uniformly from 15 m/s at the beginning of the 1st second, to 45 m/s
at the end of the 12th second?
Students were taught such physics problems and became very effective at solving
them. However, they had very little success in transferring that knowledge
to solving such algebra problems as this one:
Juanita went to work as a teller in a bank at a salary of $12,400 per year and
received constant yearly increases, coming up with a $16,000 salary during
her 13th year of work.What was her yearly salary increase?
The students failed to see that their experience with the physics problems was
relevant to solving such algebra problems, which actually have the same structure.
This happened because students did not appreciate that knowledge associated
with continuous quantities such as speed (m/s) was relevant to problems
posed in terms of discrete quantities such as dollars.
Successful problem solving depends on representing problems in such a way
that appropriate operators can be seen to apply.
Functional Fixedness
Sometimes solutions to problems depend on the solver’s ability to represent the
objects in his or her environment in novel ways. This fact has been demonstrated
in a series of studies by different experimenters. A typical experiment in
the series is the two-string problem of Maier (1931), illustrated in Figure 8.14.
Two strings hanging from the ceiling are to be tied together, but they are so far
apart that the participant cannot grasp both at once. Among the objects in the
room are a chair and a pair of pliers. Participants try various solutions involving
the chair, but these do not work. The only solution that works is to tie the
pliers to one string and set that string swinging like a pendulum; then get the
second string, bring it to the center of the room, and wait for the first string to
swing close enough to grasp. Only 39% of Maier’s participants were able to see
this solution within 10 minutes. The difficulty is that the participants did not
Problem Representation | 231
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 231
perceive the pliers as a weight that could be used as a pendulum. This phenomenon
is called functional fixedness. It is so named because people are fixed on
representing an object according to its conventional function and fail to represent
its novel function.
Another demonstration of functional fixedness is an experiment by Duncker
(1945). The task he posed to participants was to support a candle on a door,
ostensibly for an experiment on vision. The problem is illustrated in Figure 8.15.
On the table are a box of tacks, some matches, and the candle. The solution is to
tack the box to the door and use the box as a platform for the candle. This task is
difficult because participants see the box as a container, not as a platform. They
have greater difficulty with the task if the box is filled with tacks, reinforcing the
perception of the box as a container.
These demonstrations of functional fixedness are consistent with the interpretation
that representation has an effect on operator selection. For instance,
to solve Duncker’s candle problem, participants needed to represent the tack box
in such a way that it could be used by the problem-solving operators that were
232 | Problem Solving
FIGURE 8.14 The two-string problem used by Maier to demonstrate functional fixedness.
Only 39% of Maier’s participants were able to see the solution within 10 minutes. A large
majority of the participants did not perceive the pliers as a weight that could be used
as a pendulum. (After Maier, 1931. Adapted by permission of the publisher. © 1931 by the Journal of
Comparative Psychology.)
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 232
looking for a support for the candle. When the box was conceived of as a container
and not as a support, it was not available to the support-seeking operators.
Functional fixedness refers to people’s tendency to see objects as serving
conventional problem-solving functions and thus failing to see possible
novel functions.
•Set Effects
People can become biased to prefer certain operators when solving a problem by
their experiences. Such biasing of the problem solution is referred to as a set
effect. A good illustration involves the water jug problem studied by Luchins
(1942) and Luchins and Luchins (1959). In these water jug problems—which are
different from the Atwood and Polson (1976) problem shown in Figure 8.8—
participants were given a set of jugs of various capacities and an unlimited water
supply. The task was to measure out a specified quantity of water. Two examples
are given below:
Set Effects | 233
FIGURE 8.15 The candle problem used by Duncker (1945) in another study of functional
fixedness. (After Glucksberg & Weisberg, 1966. Adapted by permission of the publisher. Copyright © 1966 by
the American Psychological Association.)
Capacity of Capacity of Capacity of Desired
Problem Jug A Jug B Jug C Quantity
1 5 cups 40 cups 18 cups 28 cups
2 21 cups 127 cups 3 cups 100 cups
Anderson7e_Chapter_08.qxd 8/21/09 7:53 PM Page 233
Assume that participants have a tap and a sink so that they can fill jugs and
empty them. The jugs start out empty. Participants are allowed only to fill the
jugs to capacity, empty them completely, and pour water from one jug to another.
In problem 1, participants are told that they have three jugs: jug A, with a
capacity of 5 cups; jug B, with a capacity of 40 cups; and jug C, with a capacity
of 18 cups. To solve this problem, participants would fill jug A and pour it into
B, fill A again and pour it into B, and fill C and pour it into B. The solution
to this problem is denoted by 2A _ C. The solution for the second problem is
to fill jug B with 127 cups; fill A from B so that 106 cups are left in B; fill C from
B so that 103 cups are left in B; empty C; and fill C again from B so that the goal
of 100 cups in jug B is achieved. The solution to this problem can be denoted
by B _ A _ 2C. The first solution is called an addition solution because it
involves adding the contents of the jugs together; the second is called a subtraction
solution because it involves subtracting the contents of one jug from
another. Luchins first gave participants a series of problems that all could be
solved by addition, thus creating an “addition set.” These participants then
solved new addition problems faster, and subtraction problems slower, than
control participants who had no practice.
The set effect that Luchins (1942) is most famous for demonstrating is the
Einstellung effect, or mechanization of thought, which is illustrated by the series
of problems shown in Table 8.3. Participants were given these problems in this
order and were required to find solutions for each. Take time out from reading
this text and try to solve each problem.
All problems except number 8 can be solved by using a B _ 2C – A method
(i.e., filling B, twice pouring B into C, and once pouring B into A). For problems 1
through 5, this solution is the simplest; but for problems 7 and 9, the simpler
234 | Problem Solving
TABLE 8.3
Luchins’s Water Jug Problems Used to Illustrate the Set Effect
Capacity (cups)
Problem Jug A Jug B Jug C Desired Quantity
1 21 127 3 100
2 14 163 25 99
3 18 43 10 5
4 9 42 6 21
5 20 59 4 31
6 23 49 3 20
7 15 39 3 18
8 28 76 3 25
9 18 48 4 22
10 14 36 8 6
After Luchins (1942). Adapted by permission of the publisher. © 1942 by Psychological
Monographs.
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 234
solution of A _ C also applies. Problem 8 cannot be solved by the B _ 2C – A
method but can be solved by the simpler solution of A _ C. Problems 6 and 10
are also solved more simply by A _ C than by B _ 2C _ A. Of Luchins’s participants
who received the whole setup of 10 problems, 83% used the B _ 2C _ A
method on problems 6 and 7, 64% failed to solve problem 8, and 79% used the
B _ 2C _ A method for problems 9 and 10. The performance of participants
who worked on all 10 problems was compared with that of control participants
who saw only the last 5 problems. These control participants did not see the
biasing B _ 2C _ A problems. Fewer than 1% of the control participants used
B _ 2C _ A solutions, and only 5% failed to solve problem 8. Thus, the first 5
problems created a powerful bias for a particular solution. This bias hurt the
solution of problems 6 through 10. Although these effects are quite dramatic, they
are relatively easy to reverse with the exercise of cognitive control. Luchins found
that simply warning participants by saying, “Don’t be blind” after problem 5 allowed
more than 50% of them to overcome the set for the B _ 2C _ A solution.
Another kind of set effect in problem solving has to do with the influence of
general semantic factors. This effect is well illustrated in the experiment of
Safren (1962) on anagram solutions. Safren presented participants with lists
such as the following, in which each set of letters was to be unscrambled and
made into a word:
kmli graus teews recma foefce ikrdn
This is an example of an organized list, in which the individual words are all
associated with drinking coffee. Safren compared solution times for organized
lists with times for unorganized lists. Median solution time was 12.2 s for
anagrams from unorganized lists and 7.4 s for anagrams from organized lists.
Presumably, the facilitation evident with the organized lists occurred because
the earlier items in the list associatively primed, and so made more available,
the later words. Note that this anagram experiment contrasts with the water jug
experiment in that no particular procedure was being strengthened. Rather,
what was being strengthened was part of the participant’s factual (declarative)
knowledge about spellings of associatively related words.
In general, set effects occur when some knowledge structures become more
available than others. These structures can be either procedures, as in the water
jug problem, or declarative information, as in the anagram problem. If the
available knowledge is what participants need to solve the problem, their problem
solving will be facilitated. If the available knowledge is not what is needed,
problem solving will be inhibited. It is good to realize that sometimes set effects
can be dissipated easily (as with Luchins’s “Don’t be blind” instruction). If you
find yourself stuck on a problem and you keep generating similar unsuccessful
approaches, it is often useful to force yourself to back off, change set, and try a
different kind of solution.
Set effects result when the knowledge relevant to a particular type of problem
solution is strengthened.
Set Effects | 235
Anderson7e_Chapter_08.qxd 8/20/09 9:48 AM Page 235
Incubation Effects
People often report that after trying to solve a problem and getting nowhere,
they can put it aside for hours, days, or weeks and then, upon returning to it,
can see the solution quickly. Many examples of this pattern were reported by
the famous French mathematician Poincaré (1929), including, for instance, the
following:
Then I turned my attention to the study of some arithmetical questions apparently
without much success and without a suspicion of any connection with
my preceding researches. Disgusted with my failure, I went to spend a few days
at the seaside, and thought of something else. One morning, walking on the
bluff, the idea came to me, with just the same characteristics of brevity, suddenness,
and immediate certainty, that the arithmetic transformations of indeterminate
ternary quadratic forms were identical with those of non-Euclidean
geometry. (p. 388)
Such phenomena are called incubation effects.
An incubation effect was nicely demonstrated in an experiment by Silveira
(1971). The problem she posed to participants, called the cheap-necklace
problem, is illustrated in Figure 8.16. Participants were given the following
instructions:
You are given four separate pieces of chain that are each three links in length.
It costs 2¢ to open a link and 3¢ to close a link. All links are closed at the
beginning of the problem. Your goal is to join all 12 links of chain into a
single circle at a cost of no more than 15¢.
Try to solve this problem yourself. (A solution is provided in the appendix at
the end of this chapter.) Silveira tested three groups. A control group worked
on the problem for half an hour; 55% of these participants solved the
problem. For one experimental group, the half hour spent on the problem
was interrupted by a half-hour break in which the participants did other activities;
64% of these participants solved the problem. A second experimental
group had a 4-hour break, and 85% of these participants solved the problem.
Silveira required her participants to speak aloud as they solved the cheapnecklace
problem. She found that they did not come back to the problem
after a break with solutions completely worked out. Rather, they began by
trying to solve the problem much as before. This result is evidence against
a common misbelief that people are subconsciously
solving the problem in the period that they are away
from it.
The best explanation for incubation effects relates them
to set effects. During initial attempts to solve a problem,
people set themselves to think about the problem in certain
ways and bring to bear certain knowledge structures. If this
initial set is appropriate, they will solve the problem. If the
initial set is not appropriate, however, they will be stuck
throughout the session with inappropriate procedures.
Going away from the problem allows activation of the
236 | Problem Solving
chain A
Given state Goal state
chain B
chain C
chain D
FIGURE 8.16 The cheapnecklace
problem used by
Silveira (1971) to investigate
the incubation effect. (After
Wickelgren, 1974. Adapted by permission of
the publisher. © 1974 by W. H. Freeman.)
Anderson7e_Chapter_08.qxd 8/20/09 9:49 AM Page 236
inappropriate knowledge structures to dissipate, and
people are able to take a fresh approach.
The basic argument is that incubation effects occur
because people “forget” inappropriate ways of solving
problems. Smith and Blakenship (1989, 1991) performed
a fairly direct test of this hypothesis. They had participants
solve problems like those shown in Figure 8.17. They
provided half of their participants, the fixation group,
with inappropriate ways to think about the problems. For
instance, with respect to the third problem, they told
participants to think about chemicals. Thus, in the fixation
condition, they deliberately induced incorrect sets.
Not surprisingly, the fixation participants solved fewer of
the problems than the control participants. The interesting
issue, however, was how much incubation effect these
two populations of participants showed. Half of both the
fixation and control participants worked on the problems
for a continuous period of time, whereas the other half
had an incubation period inserted in the middle of their
problem-solving efforts. The fixation participants showed
a greater benefit of the incubation period. Thus, Smith
and Blakenship were able to show a greater incubation
effect in participants who had started with an inappropriate
way of solving the problem. Also, when they asked the
fixation participants what the misleading clue had been,
they found that more of the participants who had an incubation
period had forgotten the inappropriate clue.
Incubation effects occur when people forget the inappropriate strategies they
were using to solve a problem.
Insight
A common misbelief about learning and problem solving is that there are magical
moments of insight when everything falls into place and we suddenly see a
solution. This is called the “aha” experience, and many of us can report uttering
that very exclamation after a long struggle with a problem that we suddenly solve.
The incubation effects just discussed have been used to argue that the subconscious
is deriving this insight during the incubation period. As we saw, however, what
really happens is that participants simply let go of poor ways of solving problems.
Metcalfe and Wiebe (1987) came up with an interesting way to define
insight problems. The insight problems they used included ones like the
cheap-necklace problem (see Figure 8.16). Their noninsight problems required
multistep solutions, as in the Tower of Hanoi problem (see Figure 8.10). They
asked participants to judge every 15 s how close they felt they were to the solution.
Fifteen seconds before they actually solved a noninsight problem, participants
were fairly confident they were close to a solution. In contrast, on the
Set Effects | 237
lines reading lines
oholene
or
or
search
and
FIGURE 8.17 Puzzles used by Smith and Blakenship to test
the hypothesis that incubation effects occur because people
“forget” inappropriate ways of solving problems. Participants
had to figure out what familiar phrase was represented by
each image. For example, the first picture represents the
phrase “reading between the lines”; the second, “search high
and low”; the third, “a hole in one”; the fourth “double or
nothing.” (After Smith & Blakenship, 1989, 1991. Adapted by permission
of the publishers. © 1989 by the Bulletin of the Psychonomic Society. © 1991
by the American Journal of Psychology.)
Anderson7e_Chapter_08.qxd 8/20/09 9:49 AM Page 237
insight problems, participants had little idea they were close to a solution, even
15 s before they actually solved the problem. Metcalfe and Wiebe suggested
that we use this difference as a definition of insight problems. That is, an
insight problem is one in which people are not aware that they are close to a
solution.
This definition would seem to support the notion that a solution comes in a
single moment. However, what Metcalfe and Wiebe actually showed was that
participants did not know when they were close to a solution to an insight
problem. They did not show that the solution came in a single moment. Kaplan
and Simon (1990) studied participants while they solved the mutilatedcheckerboard
problem (see Figure 8.13), which is another insight problem.
They found that some participants noticed key features of the solution to the
problem—such as that a domino covers one square of each color—early on.
Sometimes, though, these participants did not judge those features to be critical
and went off and tried other methods of solution; only later did they come
back to the key feature. So, it is not that solutions to insight problems cannot
come in pieces, but rather that participants do not recognize which pieces are
key until they see the final solution. It reminds me of the time I tried to find my
way through a maze, cut off from all cues as to where the exit was. I searched
for a very long time, was quite frustrated, and was wondering if I was ever going
to get out—and then I made a turn and there was the exit. I believe I even
exclaimed, “Aha!” It was not that I solved the maze in a single turn; it was that
I did not appreciate which turns were on the way to the solution until I made
that final turn.
Sometimes, insight problems require only a single step (or turn) to solve,
and it is just a matter of finding that step. What is so difficult about these
problems is just finding that one step, which can be a bit like trying to find
a needle in a haystack. As an example of such a problem, consider the
following:
What is greater than God
More evil than the Devil
The poor have it
The rich want it
And if you eat it, you’ll die.
Reportedly, schoolchildren find this problem easier than college undergraduates.
If so, it is because they consider fewer possibilities as an answer. (If you are
frustrated and cannot solve this problem, it turns out that, like many things,
one can find the answer by searching the Web—many people have posted this
problem on their Web pages.)
As a final example of insight problems consider the remote association
problems introduced by Mednick (1962). In one version of these problems
(Mednick, 1962), participants are asked to find some word that can be combined
with three words to make a compound word. So, for instance, given
fox, man, and peep, the solution is hole (foxhole, manhole, peephole). Here are a
238 | Problem Solving
Anderson7e_Chapter_08.qxd 8/20/09 9:49 AM Page 238
number of problems for you to try (the solutions
are given in the appendix):
print/berry/bird
dress/dial/flower
pine/crab/sauce
Studies of brain activity (Jung-Beeman et al.,
2004) have been conducted while people try to
solve these problems. Characteristic of insight
problems, people often get a sudden feeling
of insight when they solve them. Figure 8.18
shows the imaging results from our laboratory,
which say a lot about what is happening. The
region is plotting activity in the left prefrontal
region whose activity has been associated
with retrieval from declarative memory (e.g.,
Figures 1.16c, 7.6). The figure compares activity in cases where participants
are able to solve the problem with cases where they are not. Time 0 in the
figure marks the point where the solution was obtained in the successful
case. Both functions for the successful and unsuccessful cases are increasing,
reflecting increasing effort as the search progresses, but there is an abrupt drop
(time-lagged as we would expect with the BOLD response) after the insight. It
should be emphasized that other regions such as the motor region show a rise
at this point associated with the generation of the response. In dropping off,
the prefrontal is showing a strikingly different response compared to other
brain regions and is reflecting the end to the search of memory for the answer.
The participant had been trying retrieval after retrieval and finally retrieved the
right answer. The feeling of insight corresponds to the moment when retrieval
finally succeeds and activity drops in the retrieval area.
Insight problems are ones in which solvers cannot recognize when they are
getting close to the solution.
•Conclusions
This chapter has been built around the Newell and Simon model of problem
solving as a search through a state space defined by operators. We have
looked at problem-solving success as determined by the operators available
and the methods used to guide the search for operators. This analysis is
particularly appropriate for first-time problems, whether a chimpanzee’s
quandary (see Figure 8.2) or a human’s predicament when shown a Tower of
Hanoi problem for the first time (see Figure 8.10). The next chapter will
focus on the other factors that come into play with repeated problem-solving
practice.
Conclusions | 239
−10
−0.1
0.3
0.2
0.0
0.1
0.4
0.5
0.7
0.6
0.8
Baseline
Solution
−5 0
Time (sec.) from response
LIPFC (retrieval)
5 10
FIGURE 8.18 A comparison of
brain activity for successful and
unsuccessful attempts to solve
a remote association problem.
The activity plotted is from a
prefrontal region that is sensitive
to retrieval. Activity increases
with increasing time on task but
drops off for successful problems
shortly after the solution
(at time 0).
Anderson7e_Chapter_08.qxd 8/20/09 9:49 AM Page 239
240 | Problem Solving
1. Recent research (e.g., Pizlo et al., 2006) has been
conducted on the so-called “traveling salesman
problem.” To make an example of such a problem, put
a number of dots (say, 10 to 20) randomly on a page
and pick one as your start dot. Now try to draw the
shortest path from this dot, visiting each dot just once
and arriving back at your start dot. If you were to
characterize this problem as a search space, what would
the states of the problem be and what would the operators
be? How do you select among the operators? Is this
particularly useful to characterize this problem in terms
of such a search space?
2. In the modern world, humans frequently want to learn
how to use devices like microwaves or software such as
a spreadsheet package.When do you try to learn these
things by discovery, by following an example, and by
following instructions? How often are your learning
experiences a mixture of these modes of learning?
3. A common goal for students is getting a good grade in a
course. There are many different things that you can do
to try to improve your grade. How do you select among
them? When are you engaged in hill climbing and when
are you engaged in means-ends analysis?
4. Figure 8.19 illustrates the nine-dots problem (Maier,
1931). The problem is to connect all 9 dots by drawing
4 lines, never lifting your pen from the page. Summarizing
a variety of studies, Kershaw and Ohlsson (2001)
report that given only a few minutes, only 5% of
undergraduates can solve this problem. Try to solve this
problem. If you get frustrated, you can find an answer
by Googling “nine-dots problem.”After you have tried
to solve the problem, use the terminology (see below)
of this chapter to describe the nature of the difficulties
posed by this problem and what people need to do to
successfully solve this problem.
Questions for Thought
Key Terms
analogy
backup avoidance
difference reductions
Einstellung effect
functional fixedness
General Problem Solver
(GPS)
goal state
hill climbing
incubation effect
insight problem
means-ends analysis
operator
problem space
search
search tree
set effect
state
subgoal
Tower of Hanoi problem
•Appendix: Solutions
Figure A8.1 gives the minimum-path solution to the problem solved less
efficiently in Figure 8.3.
With regard to the problem of the 27 apples, the worm cannot succeed. To
see that this is the case, imagine that the apples alternate in color, green and
red, in a 3-D checkerboard pattern. If the center apple from which the worm
starts is red, there are 13 red apples and 14 green apples in all. Each time the
worm moves from one apple to another, it will be changing colors. Because
FIGURE 8.19 The nine-dots problem.
Anderson7e_Chapter_08.qxd 8/20/09 9:49 AM Page 240
the worm starts from a red apple, it cannot reach more green apples than red
apples. Thus, it cannot visit all 14 green apples if it also visits each of the 13 red
apples just once.
To solve the cheap-necklace problem shown in Figure 8.16, open all three
links in one chain (at a cost of 6¢) and then use the three open links to connect
the remaining three chains (at a cost of 9¢).
The solutions to the three remote association problems are blue, sun, and
apple.
Appendix: Solutions | 241
2 1
(a) (b) (c) (d) (e) (f) (g)
(o) (p) (q) (r) Goal state
(n) (m) (l) (k) ( j) (i) (h)
6 2 1 6 2 1 1 1
8 8 6 6 6
8 8
2 2
8 8
4 4 8
2 2
8
8
4 4 4
4 4 8
6 6
4 8 4 8 4
1 1
6 6
1
6
4 4 2
2 2
4 4
7
7
5
5
2 2
7 5
2
7 5
2
7 5
2
7 5
7 5 7 5
8
7 5
8 8 2 8
7 5 7 5 7 5
3 7 5 3
2 4
7 5
3
1 3 1 3 1 3
6 6
1 3
3 3
1
6
6
1
2 1
6
8
4
4
6
1
4
7 5
3
3 3 3 3
7 5 3 7 5 3 7 5 3
2 8 1
4 6
7 5 3
6
1 3
6
8 1
FIGURE A8.1 The minimum-path solution for the eight-tile problem that was solved less
efficiently in Figure 8.3.
Anderson7e_Chapter_08.qxd 8/20/09 9:49 AM Page 241
242
9Expertise
It has been speculated that the expansion of the human brain from Homo erectus
to modern Homo sapiens was driven by the need to acquire expertise in novel environments
(Skoyles, 1999). This ability allowed humans to spread throughout the
world and permitted the development of the technology that has created modern
civilization. Humans are the only species that display this kind of behavioral plasticity—
being able to become experts at driving a car in modern society, navigating the
oceans in Polynesian society, or designing search engines for the World Wide Web.
William G. Chase, late of Carnegie Mellon University, was one of our local experts on
human expertise. He emphasized two famous mottos that summarize much of the
nature of expertise and its development:
• No pain, no gain. • When the going gets tough, the tough get going.
The first motto refers to the fact that no one develops expertise without a great
deal of hard work. John R. Hayes (1985), another Carnegie Mellon faculty member,
has studied geniuses in fields varying from music to science to chess. He found that
no one reached genius levels of performance without at least 10 years of practice.
Chase’s second motto refers to the fact that the difference between relative novices
and relative experts increases as we look at more difficult problems. For instance,
there are many chess duffers who could play a credible, if losing, game against a
master when they are given unlimited time to choose moves. However, they would
lose embarrassingly if forced to play lightning chess, where they are permitted only
5 s per move.
Chapter 8 reviewed some of the general principles governing problem solving,
particularly in novel domains. This research has provided a framework for analyzing
the development of expertise in problem solving. Research on expertise has been a
major development in cognitive science in the past 30 years. This research is particularly
exciting because it has important contributions to make to the instruction of
technical or formal skills in areas such as mathematics, science, and engineering, as
will be reviewed at the end of this chapter.
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 242
Brain Changes with Skill Acquisition | 243
This chapter will address the following questions about the nature of human
expertise: • What are the stages in the development of expertise? • How does the organization of a skill change as one becomes expert? • What are the contributions of practice versus talent to the development of skill? • How much can skill in one domain transfer to a new domain? • What are the implications of our knowledge about expertise for teaching
new skills?
•Brain Changes with Skill Acquisition
As people become more proficient at a task, they seem to use less of their brains
to perform that task. Figure 9.1 shows fMRI some data from Qin et al. (2003)
looking at areas of the brain activated as college students learned to perform
derivations in a new synthetic domain of mathematics. Figure 9.1a shows the regions
activated on their first day of doing the task and Figure 9.1b shows the
regions activated on the fifth day. As the students achieved greater efficiency in
the performance of the task, regions of activity dropped out or shrank. These
regions of activity correspond to metabolic expenditure, and it is quite apparent
that, with expertise, we spend less mental energy doing these tasks.
A general goal of research on expertise is to characterize both the qualitative
and the quantitative changes that take place with expertise. The result in
Figure 9.1 can be considered a quantitative result—more practice means more
efficient mental execution.We will look at a number of quantitative measures,
particularly latency, that indicate this increased efficiency. However, there are
also qualitative changes in how a skill is performed with practice. Figure 9.1
does not reveal such changes—in this study, it just seems that fewer areas,
FIGURE 9.1 Regions activated
in the symbol-manipulation task
of Qin et al. (2003): (a) day 1 of
practice; (b) day 5 of practice.
Note that these images depict
“transparent brains,” and the
activation that we see is not
just on the surface but also
below the surface.
Brain Structures
(a) (b)
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 243
rather than different areas, take part. However, this chapter will describe the
results of other brain imaging and behavioral studies that indicate that, indeed,
the way in which we perform a task can change as we become expert at it.
Through extensive practice, we can develop the high levels of expertise in
novel domains that have supported the evolution of human civilization.
•General Characteristics of Skill Acquisition
Three Stages of Skill Acquisition
The development of a skill typically comprises three stages (Anderson, 1983;
Fitts & Posner, 1967). Fitts and Posner call the first stage the cognitive stage.
In this stage, participants develop a declarative encoding (see the distinction
between declarative and procedural representations at the end of Chapter 7) of
the skill; that is, they commit to memory a set of facts relevant to the skill.
Essentially these facts define the operators of the task (see Chapter 8). Learners
typically rehearse these facts as they first perform the skill. For instance, when I
was first learning to shift gears in a standard transmission car, I memorized the
location of the gears (e.g., “reverse is up, left”—for an old 3-speed transmission)
and the correct sequence of engaging the clutch and moving the stick
shift. I rehearsed this information as I performed the skill.
The information that I had learned about the location and function of the
gears amounted to a set of problem-solving operators for driving the car. For
instance, if I wanted to get the car into reverse, there was the operator of moving
the gear to the upper left. Despite the fact that the knowledge about what to do
next was unambiguous, one would hardly have judged my driving performance
as skilled. My use of the knowledge was very slow because that knowledge was
still in a declarative form. I had to retrieve specific facts and interpret them to
solve my driving problems. I did not have the knowledge in a procedural form.
The second stage of skill acquisition is called the associative stage. Two
main things happen in this second stage. First, errors in the initial understanding
are gradually detected and eliminated. So, I slowly learned to coordinate the
release of the clutch in first gear with the application of gas so as not to kill
the engine. Second, the connections among the various elements required for
successful performance are strengthened. Thus, I no longer had to sit for a few
seconds trying to remember how to get to second gear from first. Basically, the
outcome of the associative stage is a successful procedure for performing the
skill. However, it is not always the case that the procedural representation of
the knowledge replaces the declarative. Sometimes, the two forms of knowledge
can coexist side by side, as when we can speak a foreign language fluently and
still remember many rules of grammar. However, the procedural, not the
declarative, knowledge governs the skilled performance.
The third stage in the standard analysis of skill acquisition is the autonomous
stage. In this stage, the procedure becomes more and more automated
and rapid. The concept of automaticity was introduced in Chapter 3, where we
244 | Expertise
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 244
discussed how central cognition drops out of performance of a task as we become
more skilled at it. Complex skills such as driving a car or playing chess
gradually evolve in the direction of becoming more automated and requiring
fewer processing resources. For instance, driving a car can become so automatic
that people will engage in conversation with no memory for the traffic that they
have driven through.
The three stages of skill acquisition are the cognitive stage, the associative
stage, and the autonomous stage.
Power-Law Learning
Chapter 6 documented the way in which the retrieval of simple associations
improved as a function of practice according to a power law. It turns out that
the performance of complex skills, requiring the coordination of many such associations,
also improves according to a power law. Figure 9.2 illustrates a wellknown
instance of such skill acquisition. This study followed the development
of the cigar-making ability of a worker in a factory for 10 years. The figure plots
the time to make a cigar against number of years of practice. Both scales use
log-log coordinates to expose a power law (recall from Chapters 6 and 7 that a
linear function on log-log coordinates implies a power function in the original
scale). The data in this graph show an approximately linear function until
General Characteristics of Skill Acquisition | 245
10,000 100,000 1,000,000
Number of items produced (logarithmic scale)
100,000,000
10
(1 year) (7 years)
Minimum machine
cycle time
20
30
Cycle time (s, logarthmic scale)
5
FIGURE 9.2 Time required to produce a cigar as a function of amount of experience. (From
Crossman, 1959. Reprinted by permission from Taylor & Francis.)
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 245
about the fifth year, at which point the improvement appears to stop. It turns
out that the worker was approaching the cycle time of the machinery and could
improve no more. There is usually some limit to how much improvement can
be achieved, determined by the equipment, the capability of a person’s musculature,
age, and so on. However, except for these physical limits, there is no limit
on how much a skill can speed up. The time taken by the cognitive component
of a skill will go to zero, given enough practice.
Recall from Chapter 6 that a linear relation between log time T and log
practice P can be expressed as
ln T _ A _ b ln P
which can be transformed into
T _ aP_b
where A = ln a. In Chapter 6, we considered such power functions in memory
(see Figures 6.10 and 6.11). Basically, for these functions, the decrease in processing
time with further practice becomes small very rapidly.
Effects of practice have also been studied in domains of complex problem
solving, such as giving justifications for geometry-like proofs (Neves & Anderson,
1981). Figure 9.3 shows a power function for that domain, in both a normal
scale and a log-log scale. Such functions illustrate that the benefit of further
practice rapidly diminishes but that, no matter how much practice we have
had, further practice will help a little.
Kolers (1979) investigated the acquisition of reading skills, by using materials
such as those illustrated in Figure 9.4. The first type of text (N) is normal, but
the others have been transformed in various ways. In the R transformation, the
whole line has been turned upside down; in the I transformation, each letter
has been inverted; in the M transformation, the sentence has been set as a
mirror image of standard type. The rest are combinations of the several transformations.
In one study, Kolers looked at the effect of massive practice on
reading inverted (I) text. Participants took more than 16 min to read their first
page of inverted text compared with 1.5 min for normal text. After the initial
246 | Expertise
(a)
200
20
Number of problems Time to solution
40 60 80 100
400
600
800
1,000
1,200
1,400
(b)
2,000
1,000
400
200
100
2 4
Log (trials)
Log (s)
10 20 40 100
FIGURE 9.3 Time taken to
generate proofs in a geometrylike
proof system as a function
of the number of proofs
already done: (a) function on a
normal scale, RT _ 1,410P_55;
(b) function on a log-log scale.
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 246
reading-speed test, participants practiced on 200 pages of inverted text. Figure 9.5
provides a log-log plot of reading time against amount of practice. In this
figure, practice is measured as number of pages read. The change in speed with
practice is given by the curve labeled “Original training on inverted text.” Kolers
interspersed a few tests on normal text; data for these tests are given by the
curve labeled “Original tests on normal text.”We see the same kind of improvement
for inverted text as in Figures 9.2 and 9.3 (i.e., a straight-line function on
a log-log plot). After reading 200 pages, Kolers’s participants were reading at the
rate of 1.6 min per page—almost the same rate as that of participants reading
normal text.
A year later, Kolers had his participants read inverted text again. These data
are given by the curve in Figure 9.5 labeled “Retraining on inverted text.”
Participants now took about 3 min to read the first page of the inverted text.
Compared with their performance of 16 min on their first page a year earlier,
participants displayed an enormous savings, but it was now taking them almost
twice as long to read the text as it did after their 200 pages of training a year
earlier. They had clearly forgotten something. As the Figure 9.5 illustrates,
participants’ improvement on the retraining trials showed a log-log relation
between practice and performance, as had their original training. The same
level of performance that participants had initially reached after 200 pages of
General Characteristics of Skill Acquisition | 247
FIGURE 9.4 Examples of the
spatially transformed texts
used in Kolers’s studies of the
acquisition of reading skills. The
asterisks indicate the starting
point for reading. (From Kolers &
Perkins, 1975.)
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 247
training was now reached after 50 pages. Skills generally show very high levels
of retention. In many cases, such skills can be maintained for years with no
retention loss. Someone coming back to a skill—skiing, for example—after many
years of absence often requires just a short warm-up period before the skill is
reestablished (Schmidt, 1988).
Poldrack and Gabrieli (2001) investigated the brain correlates of the
changes taking place as participants learn to read transformed text such as that
in Figure 9.4. In an fMRI brain-imaging study, they found increased activity in
the basal ganglia and decreased activation in the hippocampus as learning progressed.
Recall from Chapters 6 and 7 that the basal ganglia are associated with
procedural knowledge, whereas the hippocampus is associated with declarative
knowledge. Similar changes in the activation of brain areas have been found by
Poldrack et al. (1999) in another skill-acquisition task that required the classification
of stimuli. As participants develop their skill, they appear to move to a
direct recognition of the text. Thus, the results of this brain-imaging research
reveal changes consistent with the switch between the cognitive and the associative
stages. Thus, qualitative changes appear to be contributing to the quantitative
changes captured by the power function.We will consider these qualitative
changes in more detail in the next section.
Performance of a cognitive skill improves as a power function of practice and
shows modest declines only over long retention intervals.
248 | Expertise
FIGURE 9.5 The results for readers in Kolers’s reading-skills experiment on two tests more
than a year apart. Participants were trained with 200 pages of inverted text in which pages of
normal text were occasionally interspersed. A year later, they were retrained with 100 pages
of inverted text, again with normal text occasionally interspersed. The results show the effect
of practice on the acquisition of the skill. Both reading time and number of pages practiced
are plotted on a logarithmic scale. (From Kolers, 1976. Copyright by the American Psychological Association.
Reprinted by permission.)
0
16
8
4
2
1
2 4 8
Page number (logarithmic scale)
Reading time (min, logarithmic scale)
16 32 64 128 256
Original training on inverted text
Retraining on inverted text
Original tests on normal text
Retraining tests on normal text
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 248
•The Nature of Expertise
So far in this chapter, we have considered some of the phenomena associated
with skill acquisition. An understanding of the mechanisms behind these phenomena
has come from examining the nature of expertise in various fields of
endeavor. Since the mid-1970s, there has been a great deal of research looking
at expertise in such domains as mathematics, chess, computer programming,
and physics. This research compares people at various levels of development of
their expertise. Sometimes this research is truly longitudinal and follows students
from their introduction to a field to their development of some expertise.
More typically, such research samples people at different levels of expertise. For
instance, research on medical expertise might look at students just beginning
medical school, residents, and doctors with many years of medical practice.
This research has begun to identify some of the ways that problem solving
becomes more effective with experience. Let us consider some of these dimensions
of the development of expertise.
Proceduralization
The degree to which participants rely on declarative versus procedural knowledge
changes dramatically as expertise develops. It is illustrated in my own
work on the development of expertise in geometry (Anderson, 1982). One student
had just learned the side-side-side (SSS) and side-angle-side (SAS) postulates
for proving triangles congruent. The side-side-side postulate states that, if
three sides of one triangle are congruent to the corresponding sides of another
triangle, the triangles are congruent. The side-angle-side postulate states that, if
two sides and the included angle of one triangle are congruent to the corresponding
parts of another triangle, the triangles are congruent. Figure 9.6 illustrates
the first problem that the student had to solve. The first thing that he did
in trying to solve this problem was to decide which postulate to use. The following
is a part of his thinking-aloud protocol, during which he decided on the
appropriate postulate:
If you looked at the side-angle-side postulate (long pause) well RK and RJ
could almost be (long pause) what the missing (long pause) the missing side.
I think somehow the side-angle-side postulate works its way into here (long
pause). Let’s see what it says: “Two sides and the included angle.”What would
I have to have to have two sides JS and KS are one of them. Then you could go
back to RS = RS. So that would bring up the side-angle-side postulate (long
pause). But where would Angle 1 and Angle 2 are right angles fit in (long
pause) wait I see how they work (long pause). JS is congruent to KS (long
pause) and with Angle 1 and Angle 2 are right angles that’s a little problem
(long pause). OK, what does it say—check it one more time: “If two sides and
the included angle of one triangle are congruent to the corresponding parts.”
So I have got to find the two sides and the included angle.With the included
angle you get Angle 1 and Angle 2. I suppose (long pause) they are both right
angles, which means they are congruent to each other. My first side is JS is
to KS. And the next one is RS to RS. So these are the two sides. Yes, I think it is
the side-angle-side postulate. (Anderson, 1982, pp. 381–382)
The Nature of Expertise | 249
R
J
S
1
2
K
Given: ∠1 and ∠2 are right angles
Prove: RSJ ≅RSK
JS ≅KS
FIGURE 9.6 The first geometryproof
problem encountered by a
student after studying the sideside-
side and side-angle-side
postulates.
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 249
After reaching this point, the student still went through a long process of
actually writing out the proof, but the part of the protocol just given is germane
to assessing what goes into recognizing the relevance of the SAS postulate. After
a series of four more problems (two solved by SAS and two by SSS), the student
applied the SAS postulate in solving the problem illustrated in Figure 9.7. The
method-recognition part of the protocol was as follows:
Right off the top of my head I am going to take a guess at what I am
supposed to do: Angle DCK is congruent to Angle ABK. There is only one of
two and the side-angle-side postulate is what they are getting to. (Anderson,
1982, p. 382)
A number of things seem striking about the contrast between these two protocols.
One is that the application of the postulate has clearly sped up in the second
part of the protocol. A second is that there is no verbal rehearsal of the
statement of the postulate in the second case. The student is no longer calling a
declarative representation of the postulate into working memory. Note also
that, in the first protocol, working memory fails a number of times—points at
which the student had to recover information that he had forgotten. The third
feature of difference is that, in the first protocol, application of the postulate is
piecemeal; the student is separately identifying every element of the postulate.
Piecemeal application is absent in the second protocol. It appears that the postulate
is being matched in a single step.
These transitions are like the ones that Fitts and Posner characterized as
belonging to the associative stage of skill acquisition. The student is no longer
relying on verbal recall of the postulate but has advanced to the point where
he can simply recognize the application of the postulate as a pattern. Pattern
recognition is an important part of the procedural embodiment of a skill. We
no longer have to think about what to do next; we just recognize what is appropriate
for the situation. The process of converting the deliberate use of declarative
knowledge into pattern-driven application of procedural knowledge is
called proceduralization.
In Anderson (2007) I reported a meta-analysis of a number of studies in our
laboratory looking at the effects of practice on the performance of mathematical
problem-solving tasks like the ones we have been discussing in this section.We
were interested in the effects of this sort of practice on the three brain regions
illustrated in Figure 1.15:
Motor, which is involved in programming the actual motor movements in
writing out the solution;
Parietal, which is involved in representing the problem internally; and
Prefrontal, which is involved in retrieving things like the task instructions.
In addition we looked at a fourth region:
Anterior Cingulate Cortex (ACC), which is involved in the control of
cognition—see Figure 3.1 and later discussion in Chapter 3.
Figure 9.8 shows the mean level of activation in these regions initially and after
4 hours of practice. The motor or control demands of the tasks do not change
250 | Expertise
FIGURE 9.7 The sixth geometryproof
problem encountered by a
student after studying the sideside-
side and side-angle-side
postulates.
A
B
K
3
1
2
4
C
D Given: ∠1 ≅∠2
Prove: ABK ≅DCK
AB ≅DC
BK ≅CK
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 250
much and so there is comparable activation early versus late in the motor
cortex and the ACC. There is some reduction in the parietal suggesting that the
representational demands may be decreasing a bit. However, the dramatic
change is in the prefrontal, which is showing a major decrease because the task
instructions are no longer being retrieved. Rather, the knowledge is coming to
be directly applied.
Proceduralization refers to the process by which people switch from explicit
use of declarative knowledge to direct application of procedural knowledge,
which enables them do things such as riding a bike without thinking
about it.
Tactical Learning
As students practice problems, they come to learn the sequences of actions
required to solve a problem or parts of the problem. Learning to execute such
sequences of actions is called tactical learning. A tactic refers to a method that
accomplishes a particular goal. For instance, Greeno (1974) found that it took
only about four repetitions of the hobbits and orcs problem (see discussion
surrounding Figure 8.7) before participants could solve the problem perfectly.
In this experiment, participants were learning the sequence of moves to get the
creatures across the river. Once they had learned the sequence, they could simply
recall it and did not have to figure it out.
Logan (1988) argued that a general mechanism of skill acquisition involves
learning to recall solutions to problems that formerly had to be figured out. A
nice illustration of this mechanism is from a domain called alpha-arithmetic. It
entails solving problems such as F _ 3, in which the participant is supposed to
say the letter that is the number of letters forward in the alphabet—in this case,
The Nature of Expertise | 251
Motor Parietal
Brain region
Prefrontal ACC
0.10
0
0.05
0.15
0.20
0.25
0.30
Late
Early
FIGURE 9.8 Representation of the activity in four brain regions while performing tasks early on
versus after 5 days of practice.
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 251
F _3 _ I. Logan and Klapp (1991) performed an
experiment in which they gave participants problems
that included addends from 2 (e.g., C _ 2) through 5
(e.g., G _ 5). Figure 9.9 shows the time taken by participants
to answer these problems initially and then
after 12 sessions of practice. Initially, participants
took 1.5 s longer on the 5-addend problems than on
the 2-addend problems, because it takes longer to
count five letters forward in the alphabet than two
letters forward. However, the problems were repeated
again and again across the sessions. With repeated,
continued practice, participants became faster on all
problems, reaching the point where they could solve
the 5-addend problems as quickly as the 2-addend
problems. They had memorized the answers to these
problems and were not going through the procedure
of solving the problems by counting.1
There is evidence that, as people become more
practiced at a task and shift from computation to
retrieval, brain activation shifts from the prefrontal
cortex to more posterior areas of the cortex. For
instance, Jenkins, Brooks, Nixon, Frackowiak, and
Passingham (1994) looked at participants learning to key out various sequences
of finger presses such as “ring, index, middle, little, middle, index, ring, index.”
They compared participants initially learning these sequences with participants
practiced in these sequences. They used PET imaging studies and found that
there was more activation in frontal areas early in learning than late in learning.2
On the other hand, later in learning, there was more activation in the hippocampus,
which is a structure associated with memory. Such results indicate that, early
in a task, there is significant involvement of the anterior cingulate in organizing
the behavior but that, late in learning, participants are just recalling the answers
from memory. Thus, these neurophysiological data are consistent with Logan’s
proposal.
Tactical learning refers to a process by which people learn specific procedures
for solving specific problems.
Strategic Learning
The preceding subsection on tactical learning was concerned with how students
learn tactics by memorizing sequences of actions to solve problems. Many small
problems repeat so often that we can solve them this way. However, large and
252 | Expertise
Latency (s)
3.0
2.0
4.0
1.0
0.0
1 2
Addend
Session 1
Session 12
3 4 5
FIGURE 9.9 After 12 sessions,
participants solved alphaarithmetic
problems with
various-sized addends in
considerably less time.
(From Logan & Klapp, 1991).
1 Rabinowitz and Goldberg (1995) reported a study making a similar point.
2 This early-learning activation included the same anterior cingulate whose activity did not change in the
mathematical problem-solving tasks in Figure 9.8.However, in this simpler experiment the need for control
dramatically changes, and there is less activity later in the anterior cingulate.
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 252
complex problems do not repeat exactly, but they still have
similar structures, and one can learn how to organize one’s
solution to the overall problem. Learning how to organize
one’s problem solving to capitalize on the general structure of
a class of problems is referred to as strategic learning. The
contrast between strategic and tactical learning in skill acquisition
is analogous to the distinction between tactics and strategy
in the military. In the military, tactics refers to smaller-scale
battlefield maneuvers, whereas strategy refers to higher-level
organization of a military campaign. Similarly, tactical learning
involves learning new pieces of skill, whereas strategic learning
is concerned with putting them together.
One of the clearest demonstrations of such strategic changes is in the domain
of physics problem solving. Researchers have compared novice and expert solutions
to problems like the one depicted in Figure 9.10. A block is sliding down an
inclined plane of length l, and u is the angle between the plane and the horizontal.
The coefficient of friction is m. The participant’s task is to find the velocity of the
block when it reaches the bottom of the plane. The typical novices in these studies
are beginning college students and the typical experts are their teachers.
In one study comparing novices and experts, Larkin (1981) found a difference
in how they approached the problem. Table 9.1 shows a typical novice’s
The Nature of Expertise | 253
_
_ l
FIGURE 9.10 A sketch of
a sample physics problem.
(From Larkin, 1981.)
TABLE 9.1
Typical Novice Solution to a Physics Problem
To find the desired final speed v requires a principle with v in it—say
v _ v0 + 2 at
But both a and t are unknown; so that seems hopeless. Try instead
v2 _ v0
2 _ 2 ax
In that equation, v0 is zero and x is known; so it remains to find a. Therefore, try
F _ ma
In that equation, m is given and only F is unknown; therefore, use
F _ F s
which in this case means
F _ Fg _ f
where Fg and f can be found from
Fg _ mg sin
f _ N
N _ mg cos
With a variety of substitutions, a correct expression for speed,
can be found.
Adapted from Larkin (1981).
v = 12(g sin u - mg cos u)/
u
m
– u
–
–
© ¿
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 253
solution to the problem, whereas Table 9.2 shows a typical expert’s solution. The
novice’s solution typifies the reasoning backward method, which starts with
the unknown—in this case, the velocity v. Then the novice finds an equation for
calculating v.However, to calculate v by this equation, it is necessary to calculate a,
the acceleration. So the novice finds an equation for calculating a; and the novice
chains backward until a set of equations is found for solving the problem.
The expert, on the other hand, uses similar equations but in the completely
opposite order. The expert starts with quantities that can be directly computed,
such as gravitational force, and works toward the desired velocity. It is also apparent
that the expert is speaking a bit like the physics teacher that he is, leaving
the final substitutions for the student.
Another study by Priest and Lindsay (1992) failed to find a difference in
problem-solving direction between novices and experts. Their study included
British university students rather than American students, and they found that
both novices and experts predominantly reasoned forward. However, their
experts were much more successful in doing so. Priest and Lindsay suggest that
the experts have the necessary experience to know which forward inferences are
appropriate for a problem. It seems that novices have two choices—reason forward,
but fail (Priest & Lindsay’s students) or reason backward, which is hard
(Larkin’s students)
Reasoning backward is hard because it requires setting goals and subgoals
and keeping track of them. For instance, a student must remember that he
or she is calculating F so that a can be calculated and hence so that v can be
calculated. Thus, reasoning backward puts a severe strain on working memory
and this can lead to errors. Reasoning forward eliminates the need to keep
254 | Expertise
TABLE 9.2
Skilled Solution to a Physics Problem
The motion of the block is accounted for by the gravitational force,
Fg _ mg sin
directed downward along the plane, and the frictional force,
f _ mg cos
directed upward along the plane. The block’s acceleration a is then related to the
(signed) sum of these forces by
F _ ma
or
mg sin _ mg cos _ ma
Knowing the acceleration a, it is then possible to find the block’s final speed v from
the relations
and
v _ at
Adapted from Larkin (1981).
l =
1
2
at 2
u m u
m u
– u
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 254
track of subgoals. However, to successfully reason forward, one must know
which of the many possible forward inferences are relevant to the final solution,
which is what an expert learns with experience. He or she learns to associate
various inferences with various patterns of features in the problems. The
novices in Larkin’s study seemed to prefer to struggle with backward reasoning,
whereas the novices in Priest and Lindsay’s study tried forward reasoning
without success.
Not all domains show this advantage for forward problem solving. A good
counterexample is computer programming (Anderson, Farrell, & Sauers, 1984;
Jeffries, Turner, Polson, & Atwood, 1981; Rist, 1989). Both novice and expert programmers
develop programs in what is called a top-down manner; that is, they
work from the statement of the problem to subproblems to sub-subproblems,
and so on, until they solve the problem. This top-down development is basically
the same as what is called reasoning backward in the context of geometry
or physics. There are differences between expert programmers and novice
programmers, however. Experts tend to develop problem solutions breadth
first, whereas novices develop their solutions depth first. Physics and geometry
problems have a rich set of givens that are more predictive of solutions than
is the goal. In contrast, nothing in the typical statement of a programming
problem would guide a working forward or bottom-up solution. The typical
problem statement only describes the goal and often does so with information
that will guide a top-down solution. Thus, we see that expertise in different
domains requires the adoption of those approaches that will be successful for
those particular domains.
In summary, the transition from novices to experts does not entail the same
changes in strategy in all domains. Different problem domains have different
structures that make different strategies optimal. Physics experts learn to reason
forward; programming experts learn breadth-first expansion.
Strategic learning refers to a process by which people learn to organize their
problem solving.
Problem Perception
As they acquire expertise problem solvers learn to perceive problems in ways
that enable more effective problem-solving procedures to apply. This dimension
can be nicely demonstrated in the domain of physics. Physics, being an intellectually
deep subject, has principles that are only implicit in the surface features
of a physics problem. Experts learn to see these implicit principles and represent
problems in terms of them.
Chi, Feltovich, and Glaser (1981) asked participants to classify a large set of
problems into similar categories. Figure 9.11 shows sets of problems that
novices thought were similar and the novices’ explanations for the similarity
groupings. As can be seen, the novices chose surface features, such as rotations
or inclined planes, as their bases for classification. Being a physics novice myself,
I have to admit that these seem very intuitive bases for similarity. Contrast
The Nature of Expertise | 255
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 255
these classifications with the pairs of problems in Figure 9.12 that the expert
participants saw as similar. Problems that are completely different on the
surface were seen as similar because they both entailed conservation of energy
or they both used Newton’s second law. Thus, experts have the ability to map
surface features of a problem onto these deeper principles. This ability is very
useful because the deeper principles are more predictive of the method of
solution. This shift in classification from reliance on simple features to reliance
on more complex features has been found in a number of domains, including
mathematics (Silver, 1979; Schoenfeld & Herrmann, 1982), computer
programming (Weiser & Shertz, 1983), and medical diagnosis (Lesgold et al.,
1988).
256 | Expertise
Novice 2: "Angular velocity, momentum,
circular things."
Novice 3: "Rotation kinematics, angular
speeds, angular velocities."
Novice 6: "Problems that have something
rotating: angular speed."
_
T
R
10 M
M
V
m
Novice 1: "These deal with blocks on an incline plane."
Novice 5: "Inclined plane problems, coefficient of friction."
Novice 6: "Blocks on inclined planes with angles."
2 lb.
_ = 2
Length
2 ft
30
M 30
_
V4 ft/s
FIGURE 9.11 Diagrams depicting pairs of problems categorized by novices as similar and
samples of their explanations for the similarity. (Adapted from Chi et al., 1981.)
Expert 2: "These can be solved by Newton's
second law."
Expert 3: "F = ma; Newton's second law."
Expert 4: "Largely use F = ma; Newton's
second law."
.6 m
.15 m
Equilibrium
Expert 2: "Conservation of energy."
Expert 3: "Work energy theorem. They are all straightforward
problems."
Expert 4: "These can be done from energy considerations.
Either you should know the principle of conservation
of energy, or work is lost somewhere."
K = 200 nt/m
M 30
_
Length
T T
m
M
Mg
mg mg
Fp = Kv
FIGURE 9.12 Diagrams depicting pairs of problems categorized by experts as similar and
samples of their explanations for the similarity. (Adapted from Chi et al., 1981.)
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 256
A good example of this shift in processing of perceptual features is the interpretation
of X rays. Figure 9.13 is a schematic of one of the X rays diagnosed by
participants in the research by Lesgold et al. The sail-like area in the right lung is a
shadow (shown on the left side of the X ray) caused by a collapsed lobe of the
lung that created a denser shadow in the X ray than did other parts of the lung.
Medical students interpreted this shadow as an indication of a tumor because tumors
are the most common cause of shadows on the lung. Radiological experts,
on the other hand, were able to correctly interpret the shadow as an indication of
a collapsed lung. They saw counterindicative features such as the size of the saillike
region. Thus, experts no longer have a simple association between shadows
on the lungs and tumors, but rather can see a richer set of features in X rays.
An important dimension of growing expertise is the ability to learn to
perceive problems in ways that enable more effective problem-solving
procedures to apply.
Pattern Learning and Memory
A surprising discovery about expertise is that experts seem to display a special enhanced
memory for information about problems in their domains of expertise.
This enhanced memory was first discovered in the research of de Groot (1965,
1966), who was attempting to determine what separated master chess players from
weaker chess players. It turns out that chess masters are not particularly more
intelligent in domains other than chess. De Groot found hardly any differences between
expert players and weaker players—except, of course, that the expert players
chose much better moves. For instance, a chess master considers about the same
number of possible moves as does a weak chess player before selecting a move. In
fact, if anything, masters consider fewer moves than do chess duffers.
However, de Groot did find one intriguing difference between masters and
weaker players.He presented chess masters with chess positions (i.e., chessboards
The Nature of Expertise | 257
Novice: Tumor
Expert: Collapsed lung
What causes
this shadow ?
FIGURE 9.13 Schematic
representation of an X ray
showing a collapsed right middle
lung lobe. (From Lesgold et al., 1988.)
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 257
with pieces in a configuration that occurred in a game) for just 5 s and then removed
the chess pieces. The chess masters were able to reconstruct the positions
of more than 20 pieces after just 5 s of study. In contrast, the chess duffers could
reconstruct only 4 or 5 pieces—an amount much more in line with the traditional
capacity of working memory. Chess masters appear to have built up
patterns of 4 or 5 pieces that correspond to common board configurations as
a result of the massive amount of experience that they have had with chess.
Thus, they remember not individual pieces but these patterns. In line with this
analysis, if the players are presented with random chessboard positions rather
than ones that are actually encountered in games, no difference is demonstrated
between masters and duffers—both reconstruct only a few chess positions. The
masters also complain about being very uncomfortable and disturbed by such
chaotic board positions.
In a systematic analysis, Chase and Simon (1973) compared novices, Class A
players, and masters. They compared these different types of players with respect
to their ability to reproduce game positions such as those shown in Figure 9.14a
258 | Expertise
Middle game
White
End game
Random middle game
(a)
(b)
Random end game
Black
FIGURE 9.14 Examples of (a) middle and end games and (b) their randomized counterparts.
(From Chase & Simon, 1973.)
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 258
The Nature of Expertise | 259
and to reproduce random positions such as
those illustrated in Figure 9.14b. Figure 9.15
shows the results. Memory was poorer for
all groups for the random positions and, if
anything, masters were worse at reproducing
these positions. On the other hand, masters
showed a considerable advantage for the actual
board positions. This basic phenomenon
of superior expert memory for meaningful
problems has been demonstrated in a large
number of domains, including the game of Go
(Reitman, 1976), electronic circuit diagrams
(Egan & Schwartz, 1979), bridge hands (Engle
& Bukstel, 1978; Charness, 1979), and computer
programming (McKeithen, Reitman,
Rueter, & Hirtle, 1981; Schneiderman, 1976).
Chase and Simon (1973) also used a
chessboard-reproduction task to examine the
nature of the patterns, or chunks, used by
chess masters. The participants’ task was simply to reproduce the positions of
pieces of a target chessboard on a test chessboard. In this task, participants
glanced at the target board, placed some pieces on the test board, glanced back
to the target board, placed some more pieces on the test board, and so on.
Chase and Simon defined a chunk to be a group of pieces that participants
moved after one glance. They found that these chunks tended to define
meaningful game relations among the pieces. For instance, more than half of
the masters’ chunks were pawn chains (configurations of pawns that occur
frequently in chess).
Simon and Gilmartin (1973) estimated that chess masters have acquired
50,000 different chess patterns, that they can quickly recognize such patterns on
a chessboard, and that this ability is what underlies their superior memory performance
in chess. This 50,000 figure is not unreasonable when one considers
the years of dedicated study that becoming a chess master requires.What might
be the relation between memory for so many chess patterns and superior performance
in chess? Newell and Simon (1972) speculated that, in addition to
learning many patterns, masters have learned what to do in the presence of
such patterns. For instance, if the chunk pattern is symptomatic of a weak side,
the response might be to suggest an attack on the weak side. Thus, masters
effectively “see” possibilities for moves; they do not have to think them out,
which explains why chess masters do so well at lightning chess, in which they
have only a few seconds to move.
To summarize, chess experts have stored the solutions to many problems
that duffers must solve as novel problems. Duffers have to analyze different
configurations, try to figure out their consequences, and act accordingly.
Masters have all this information stored in memory, thereby claiming two
advantages. First, they do not risk making errors in solving these problems,
because they have stored the correct solution. Second, because they have stored
0
Beginner
Number of pieces correctly placed
Class A Master
2
4
6
8
10
12
14
16
18
Actual game positions
Random positions
FIGURE 9.15 Number of pieces
successfully recalled by chess
players after the first study
of a chessboard. (From Chase &
Simon, 1973).
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 259
correct analyses of so many positions, they can focus their problem-solving efforts
on more sophisticated aspects and strategies of chess. Thus, the experts’
pattern learning and better memory for board positions is a part of the tactical
learning discussed earlier. The way humans become expert at chess reflects the
fact that we are very good at pattern recognition but relatively poor at things
like mentally searching through sequences of possible moves. As the Implications
box describes, human strengths and weaknesses lead to a very different
way of achieving expertise at chess than we see in computer programs for playing
chess.
260 | Expertise
chess in the 1960s, was beaten by the program of an
MIT undergraduate, Richard Greenblatt, in 1966 (Boden,
2006, discusses the intrigue surrounding
these events). However, Dreyfus was a
chess duffer and the programs of the
1960s and 1970s performed poorly
against chess masters. As computers
became more powerful and could search
larger spaces, they became increasingly
competitive, and finally in May 1997,
IBM’s Deep Blue program defeated the
reigning world champion, Gary Kasparov.
Deep Blue evaluated 200 million imagined
chess positions per second. It also
had stored records of 4,000 opening
positions and 700,000 master games
(Hsu, 2002) and had many other optimizations
that took advantage of special computer hardware.
Today there are freely available chess programs
for your personal computer that can be downloaded
over the Web and will play highly competitive chess at
a master level. These developments have led to a profound
shift in the understanding of intelligence. It once
was thought that there was only one way to achieve
high levels of intelligent behavior, and that was the
human way. Nowadays it is increasingly being accepted
that intelligence can be achieved in different ways, and
the human way may not always be the best. Also, curiously,
as a consequence some researchers no longer
view the ability to play chess as a reflection of the
essence of human intelligence.
Implications
Computers achieve computer expertise differently than humans
In Chapter 8, we discussed how human problem solving
can be viewed as a search of a problem space, consisting
of various states. The initial situation
is the start state, the situations on the
way to the goal are the intermediate
states, and the solution is the goal state.
Chapter 8 also described how people
use certain methods, such as avoiding
backup, difference reduction, and meansends
analysis, to move through the
states. Often when humans search a
problem space, they are actually manipulating
the actual physical world, as in
the 8-puzzle (Figures 8.3 and 8.4).
However, sometimes they imagine states,
as when one plays chess and contemplates
how an opponent will react to
some move one is considering, how one might react to
the opponent’s move, and so on. Computers are very
effective at representing such hypothetical states and
searching through them for the optimal goal state.
Artificial intelligence algorithms have been developed
that are very successful at all sorts of problem-solving
applications, including playing chess. This has led to a
style of chess playing program that is very different from
human chess play, which relies much more on pattern
recognition. At first many people thought that, although
such computer programs could play competent and
modestly competitive chess games, they would be no
match for the best human players. The philosopher
Hubert Dreyfus, who was famously critical of computer
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 260
Experts can recognize patterns of elements that repeat in many problems,
and know what to do in the presence of such patterns without having to
think them through.
Long-Term Memory and Expertise
One might think that the memory advantage shown by experts is just a workingmemory
advantage, but research has shown that their advantage extends to
long-term memory. Charness (1976) compared experts’ memory for chess positions
immediately after they had viewed the positions or after a 30-s delay filled
with an interfering task. Class A chess players showed no loss in recall over the
30-s interval, unlike weaker participants, who showed a great deal of forgetting.
Thus, expert chess players, unlike duffers, have an increased capacity to store
information about the domain. Interestingly, these participants showed the
same poor memory for three-letter trigrams as do ordinary participants. Thus,
their increased long-term memory is only for the domain of expertise.
There is reason to believe that the memory advantage goes beyond experts’
ability to encode a problem in terms of familiar patterns. Experts appear to be
able to remember more patterns as well as larger patterns. For instance, Chase
and Simon (1973) in their study (see Figures 9.14 and 9.15) tried to identify the
patterns that their participants used to recall the chessboards. They found that
participants would tend to recall a pattern, pause, recall another pattern, pause,
and so on. They found that they could use a 2-s pause to identify boundaries
between patterns.With this objective definition of what a pattern is, they could
then explore how many patterns were recalled and how large these patterns
were. In comparing a master chess player with a beginner, they found large
differences in both measures. First, the pattern size of the master averaged
3.8 pieces, whereas it was only 2.4 for the beginner. Second, the master also
recalled an average of 7.7 patterns per board,
whereas the beginner recalled an average of only
5.3. Thus, it seems that the experts’ memory advantage
is based not only on larger patterns but
also on the ability to recall more of them.
The strongest evidence that expertise requires
the ability to remember more patterns as well as
larger patterns is from Chase and Ericsson (1982),
who studied the development of a simple but
remarkable skill. They watched a participant, S. F.,
increase his digit span, which is the number of
digits that he could repeat after one presentation.
As discussed in Chapter 6, the normal digit span is
about 7 or 8 items, just enough to accommodate a
telephone number. After about 200 hr of practice,
S. F. was able to recall 81 random digits presented
at the rate of 1 digit per second. Figure 9.16 illustrates
how his memory span grew with practice.
The Nature of Expertise | 261
10
20
40
60
80
20
Practice (5-day blocks)
Digit span
30 40 50
FIGURE 9.16 The growth in
S. F.’s memory span with
practice. Notice how the
number of digits that he can
recall increases gradually but
steadily with the number of
practice sessions. (From Chase &
Ericsson, 1982.)
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 261
What was behind this apparent superhuman feat of memory? In part, S. F.
was learning to chunk the digits into meaningful patterns. He was a longdistance
runner, and part of his technique was to convert digits into running
times. So, he would take 4 digits, such as 3492, and convert them into “Three
minutes, 49.2 seconds—near world-record mile time.” Using such a strategy, he
could convert a memory span for 7 digits into a memory span for 7-digit patterns
of length 3 or 4. This would get him to a digit span of more than 20, far
short of his eventual performance. In addition to this chunking, he developed
what Chase and Ericsson called a retrieval structure, which enabled him to recall
22 such patterns. This retrieval structure was very specific; it did not generalize
to retrieving letters rather than digits. Chase and Ericsson hypothesized
that part of what underlies the development of expertise in other domains such
as chess is the development of retrieval structures, which allows superior recall
for past patterns.
As people become more expert in a domain, they develop a better ability
to store problem information in long-term memory and to retrieve it.
The Role of Deliberate Practice
An implication of all the research that we have reviewed is that expertise comes
only with an investment of a great deal of time to learn the patterns, the problemsolving
rules, and the appropriate problem-solving organization for a domain.
As mentioned earlier, John Hayes found that geniuses in various fields produce
their best work only after 10 years of apprenticeship in a field. In another
research effort, Ericsson, Krampe, and Tesch-Römer (1993) compared the best
violinists at a music academy in Berlin with those who were only very good.
They looked at diaries and self-estimates to determine how much the two
populations had practiced and estimated that the best violinists had practiced
more than 7000 hr before coming to the academy, whereas the very good had
practiced only 5000 hr. Ericsson et al. reviewed a great many fields where, like
music, time spent practicing is critical. Not only is time on task important at
the highest levels of performance, but also it is essential to mastering school
subjects. For instance, Anderson, Reder, and Simon (1998) noted that a major
reason for the higher achievement in mathematics of students in Asian countries
is that those students spend twice as much time practicing mathematics.
Ericsson et al. (1993) make the strong claim that almost all of expertise is to
be accounted for by amount of practice, and there is virtually no role for natural
talent. They point to the research of Bloom (1985a, 1985b), who looked at the
histories of children who became great in fields such as music or tennis. Bloom
found that most of these children got started by playing around, but after a short
time they typically showed promise and were encouraged by their parents to
start serious training with a teacher. However, the early natural abilities of these
children were surprisingly modest and did not predict ultimate success in the
domain (Ericsson et al., 1993). Rather, what is critical seems to be that parents
come to believe that a child is talented and consequently pay for their child’s
instruction and equipment as well as support their time-consuming practice.
262 | Expertise
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 262
Ericsson et al. speculated that the resulting training is sufficient to account for
the development ofchildren’s success. There is almost certainly some role for
talent (considered in Chapter 13), but all the evidence indicates that genius is
90% perspiration and 10% inspiration.
Ericsson et al. are careful to note, however, that not all practice leads to the
development of expertise. They note that many people spend a lifetime playing
chess or some sport without ever getting any better.What is critical, according
to Ericsson et al., is what they call deliberate practice. In deliberate practice,
learners are motivated to learn, not just perform; they are given feedback on
their performance; and they carefully monitor how well their performance
corresponds to the correct performance and where the deviations exist. The
learners focus on eliminating these points of discrepancy. The importance of
deliberate practice is similar to the importance of deep and elaborative processing
of the to-be-learned material described in Chapters 6 and 7, in which
passive study was shown to yield few memory benefits.
An important function of deliberate practice in both children and adults
may be to drive the neural growth that is necessary to enable expertise. It had
once been thought that adults do not grow new neurons, but it now appears
that they do (Gross, 2000). An interesting recent discovery is that extensive
practice appears to drive neural growth in the adult brain. For instance, Elbert,
Pantev,Wienbruch, Rockstroh, and Taub (1995) found that violinists, who finger
strings with the left hand, show increased development of the right cortical
regions that correspond to their fingers. In another study already mentioned
in Chapter 4, Maguire et al. (2003) used imaging to examine the brains of
London taxi drivers. It takes at least 3 years for London taxi drivers to acquire
all of the knowledge necessary to navigate expertly through the streets of
London. The taxi drivers were found to have significantly more gray matter in
the hippocampal region than did matched controls. This finding corresponds to
the increased hippocampal volume reported in small mammals and birds that
engage in behavior requiring navigation (Lee, Miyasato, & Clayton, 1998). For
instance, food-storing birds show seasonal increases in hippocampal volume
corresponding to times of the year when they need to remember where they
store food.
A great deal of deliberate practice is necessary to develop expertise in any
field.
•Transfer of Skill
Expertise can often be quite narrow. As noted, Chase and Ericsson’s participant
S. F. was unable to transfer memory span skill from digits to letters. This example
is an almost ridiculous extreme of a frequent pattern in the development
of cognitive skills—that these skills can be quite narrow and fail to transfer
to other activities. Chess grand masters do not appear to be better thinkers
for all their genius in chess. An amusing example of the narrowness of expertise
Transfer of Skill | 263
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 263
is a study by Carraher, Carraher, and Schliemann (1985). These researchers
investigated the mathematical strategies used by Brazilian schoolchildren who
also worked as street vendors. On the job, these children used quite sophisticated
strategies for calculating the total cost of orders consisting of different
numbers of different objects (e.g., the total cost of 4 coconuts and 12 lemons);
what’s more, they could perform such calculations reliably in their heads.
Carraher et al. actually went to the trouble of going to the streets and posing as
customers for these children, making certain kinds of purchases and recording
the percentage of correct calculations. The experimenters then asked the children
to come with them to the laboratory, where they were given written mathematics
tests that included the same numbers and mathematical operations that
they had manipulated successfully in the streets. For example, if a child had
correctly calculated the total cost of 5 lemons at 35 cruzeiros apiece on the
street, the child was given the following written problem:
5 _ 35 _ ?
Whereas children solved 98% of the problems presented in the real-world context,
they solved only 37% of the problems presented in the laboratory context.
It should be stressed that these problems included the exact same numbers and
mathematical operations. Interestingly, if the problems were stated in the form
of word problems in the laboratory, performance improved to 74%. This improvement
runs counter to the usual finding, which is that word problems are
more difficult than equivalent “number” problems (Carpenter & Moser, 1982).
Apparently, the additional context provided by the word problem allowed the
children to make contact with their pragmatic strategies.
The study of Carraher et al. showed a curious failure of expertise to transfer
from real life to the classroom, but the typical concern of educators is whether
what is taught in one class will transfer to other classes and the real world.
Early in the 20th century, educators were fairly optimistic on this matter. A
number of educational psychologists subscribed to what has been called the
doctrine of formal discipline (Angell, 1908; Pillsbury, 1908; Woodrow, 1927),
which held that studying such esoteric subjects as Latin and geometry was of
significant value because it served to discipline the mind. Formal discipline
subscribed to the faculty view of mind, which extends back to Aristotle and
was first formalized by Thomas Reid in the late 18th century (Boring, 1950).
The faculty position held that the mind is composed of a collection of general
faculties, such as observation, attention, discrimination, and reasoning, which
were exercised in much the same way as a set of muscles. The content of the
exercise made little difference; most important was the level of exertion (hence
the fondness for Latin and geometry). Transfer in such a view is broad and
takes place at a general level, sometimes spanning domains that have no content
in common.
Although it might be nice to believe that such general transfer is possible,
as envisioned by the doctrine of formal discipline, there has been effectively
no evidence for it, despite a century of research on the topic. Some of the
earliest research on this topic was performed by Thorndike (e.g., Thorndike &
Woodworth, 1901). In one study, no correlation was found between memory
264 | Expertise
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 264
for words and memory for numbers. In another, accuracy in spelling was not
correlated with accuracy in arithmetic. Thorndike interpreted these results as
evidence against the general faculties of memory and accuracy.
There is often failure to transfer skills to similar domains and virtually no
transfer to very different domains.
•Theory of Identical Elements
In place of the doctrine of formal discipline, Thorndike proposed his theory
of identical elements. According to Thorndike, the mind is not composed of
general faculties, but rather of specific habits and associations, which provide a
person with a variety of narrow responses to very specific stimuli. In fact, the
mind was regarded as just a convenient name for countless special operations or
functions (Stratton, 1922). Thorndike’s theory stated that training in one kind of
activity would transfer to another only if the activities had situation-response
elements in common:
One mental function or activity improves others in so far as and because
they are in part identical with it, because it contains elements common to
them. Addition improves multiplication because multiplication is largely
addition; knowledge of Latin gives increased ability to learn French because
many of the facts learned in the one case are needed in the other. (Thorndike,
1906, p. 243)
Thus, Thorndike was happy to accept transfer between diverse skills as long as
the transfer could be shown to be mediated by identical elements. Generally,
however, he concluded that
The mind is so specialized into a multitude of independent capacities that we
alter human nature only in small spots, and any special school training has a
much narrower influence upon the mind as a whole than has commonly been
supposed. (p. 246)
Although the doctrine of formal discipline was too broad in its predictions
of transfer, Thorndike formulated his theory of identical elements in what
proved to be an overly narrow manner. For instance, he argued that if you
solved a geometry problem in which one set of letters is used to label the points
in a diagram, you would not be able to transfer to a geometry problem with a
different set of letters. The research on analogy examined in Chapter 8 indicated
that this is not true. Transfer is not tied to the identity of surface elements.
In some cases, there is very large positive transfer between two skills that
have the same logical structure even if they have different surface elements (see
Singley & Anderson, 1989, for a review). Thus, for instance, there is large positive
transfer between different word-processing systems, between different programming
languages, and between using calculus to solve economics problems
and using calculus to solve problems in solid geometry. However, all the available
evidence is that there are very definite bounds on how far skills will transfer
Theory of Identical Elements | 265
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 265
and that becoming an expert in one domain will have little positive benefit
on becoming an expert in a very different domain. There will be positive
transfer only to the extent that the two domains use the same facts, rules, and
patterns—that is, the same knowledge. Thus, Thorndike was right in saying
that there would be transfer between two skills to the extent that they have the
same elements in common. However, he was wrong in identifying these
“elements” with stimulus-response bonds. Modern cognitive psychology has
identified these elements as rather abstract knowledge structures that enjoy a
wider range of transfer.
There is a positive side to this specificity in transfer of skill: there seldom
seems to be negative transfer, in which learning one skill makes a person worse
at learning another skill. Interference, such as that which occurs in memory
for facts (see Chapter 7), is almost nonexistent in skill acquisition. Polson,
Muncher, and Kieras (1987) provided a good demonstration of lack of negative
transfer in the domain of text editing on a computer (using the commandbased
word processors that were common at the time). They asked participants
to learn one text editor and then learn a second, which was designed to be maximally
confusing. Whereas the command to go down a line of text might be n
and the command to delete a character might be k in one text editor, n would
mean to delete a character in another text editor and k would mean to go down
a line. However, participants experienced overwhelming positive transfer in going
from one text editor to the other because the two text editors worked in the
same way, even though the surface commands had been scrambled. There is
only one clearly documented kind of negative transfer in regard to cognitive
skills—the Einstellung effect discussed in Chapter 8. Students can learn ways of
solving problems in one domain that are no longer optimal for solving problems
in another domain. So, for instance, someone may learn tricks in algebra to
avoid having to perform difficult arithmetic computations. These tricks may no
longer be necessary when that person uses a calculator to perform these computations.
Still, students show a tendency to continue to perform these unnecessary
simplifications in their algebraic manipulations. This example is not a
case of failure to transfer; rather, it is a case of transferring knowledge that is no
longer useful.
There is transfer between skills only when these skills have the same abstract
knowledge elements.
•Educational Implications
With this analysis of skill acquisition, we can ask the question: What are the
implications for the training of such skills? One implication is the importance of
problem decomposition. Traditional high-school algebra has been estimated
to require the acquisition of many thousands of rules (Anderson, 1992).
Instruction can be improved by an analysis of what these individual elements
266 | Expertise
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 266
are. Approaches to instruction that begin with an analysis of the elements
to be taught are called componential analyses. A description of the application
of componential approaches to the instruction of a number of topics
in reading and mathematics can be found in Anderson (2000). Generally,
higher achievement is obtained in programs that include such componential
analysis.
A particularly effective part of such componential programs is mastery
learning. The basic idea in mastery learning is to follow students’ performance
on each of the components underlying the cognitive skill and to ensure that all
components are mastered. Typical instruction, without mastery learning, leaves
some students not knowing some of the material. This failure to learn some of
the components can snowball in a course in which mastery of earlier material is
a prerequisite for mastery of later material. There is a good deal of evidence
that mastery learning leads to higher achievement (Guskey & Gates, 1986;
Kulik, Kulik, & Bangert-Downs, 1986).
Instruction is improved by approaches that identify the underlying knowledge
components and ensure that students master them all.
Intelligent Tutoring Systems
Probably the most extensive use of such componential analysis is for intelligent
tutoring systems (Sleeman & Brown, 1982). These computer systems
interact with students while they are learning and solving problems, much as a
human tutor would. An example of such a tutor is the LISP tutor (Anderson,
Conrad, & Corbett, 1989; Anderson & Reiser, 1985; Corbett & Anderson,
1990), which teaches LISP, the main programming language used in artificial
intelligence. The LISP tutor continuously taught LISP to students at Carnegie
Mellon University from 1984 to 2002 and served as a prototype for a generation
of intelligent tutors, many of which have focused on teaching middle-school
and high-school mathematics. The mathematics tutors are now distributed by a
company called Carnegie Learning, spun off by Carnegie Mellon University in
1998. The Carnegie Learning mathematics tutors have been deployed over
2,600 schools nationwide and interacts with approximately 500,000 students
each year (Koedinger & Corbett, 2006; Ritter, Anderson, Koedinger, & Corbett,
2007; you can visit the Web site www.carnegielearning.com for promotional
material that should be taken with a grain of salt). Color Plate 9.1 shows a
screen shot from its most widely used product, which is a tutor for high-school
algebra.
A motivation for research on intelligent tutoring is the evidence showing that
private human tutoring is very effective. The results of studies have shown that
giving students a private human tutor enables 98% of them to do better than the
average student in a standard classroom (Bloom, 1984). An ideal private tutor is
one who is with the student at all times while he or she is studying a particular
subject matter. To use the terms of Ericsson et al. (1993), a private tutor guarantees
the deliberate practice that is essential for learning. Having the tutor present
Educational Implications | 267
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 267
while solving problems in domains, such as LISP and mathematics, which
require complex problem-solving skills, is particularly important. In LISP,
problem solving takes the form of writing computer programs, or functions, as
they are often called in LISP. Therefore, in developing the LISP tutor, we chose to
focus on providing students with tutoring while they are writing computer programs.
Table 9.3 presents a short dialogue between a student and the LISP tutor
on an early problem in the curriculum. Note how carefully the tutor monitors
the student’s performance in solving the problem. It can do so because it knows
how to write LISP functions. As the student is writing the function, the tutor is
simultaneously trying to solve the same problem that the student is working on.
As soon as it sees the student making a mistake, the tutor can intervene with
remedial instruction.
Underlying the tutor’s ability to solve problems and monitor the student’s
problem solving is a set of rules that can solve the same LISP programming
268 | Expertise
TABLE 9.3
An Interaction Between the LISP Tutor and a Student
The student is working on a piece of code to satisfy the following problem
description. Before each point at which the tutor and student interact, the partial
code that the student has written is printed out.
The Problem Description
Define a function called ctof that takes a degree reading in Celsius degrees and
converts it to Fahrenheit degrees. First, you need to change the scale by multiplying
the argument by 1.8. Then you need to translate the scale by adding 32, because
0º Celsius _ 32º Fahrenheit. For example, (ctof 20) returns 68.
Cycles in Which the Tutor Communicates with the Student
Code: (defun ctof (cdeg) (*
Tutor : Yes, you will need to call “*” to get the degree on the enlarged scale, but you
are ultimately trying to get the Fahrenheit equivalent of the Celsius degree. So you
need to code a different function here that takes a call to “*” as an argument.
Code: (defun ctof (cdeg) (_
Student: Help.
Tutor: You need to multiply together the Celsius degree reading and the size of
Fahrenheit degrees using the function *. When you multiply these numbers, you will
get the degree on the enlarged scale, which is 36 in our example.
Code: (defun ctof (cdeg) (_ (* 1.8 cdge)
Tutor: Did you mean cdeg?
Student: Yes.
Code: (defun ctof (cdeg) (_ (* 1.8 cdeg) 32))
Tutor: You are done. Type “next” to go on after testing the function you have
defined.
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 268
problems that we expect students to be able to solve. In all, there are about
500 rules that encode the knowledge relating to LISP. A typical rule in the LISP
tutor is:
If the goal is to multiply one number by another,
Then use * and set subgoals to code the two numbers.
The basic goal of the LISP tutor is to communicate these 500 rules to the student,
monitor performance to see whether he or she possesses these rules in
correct form, and provide the student with practice on these rules. The success
of the LISP tutor is one piece of evidence that these 500 rules indeed underlie
coding skill in LISP.
Besides providing an instructional tool, the LISP tutor is a research tool for
studying the course of skill acquisition. The tutor can monitor how well a student
is doing on each of the 500 rules, recording statistics such as the number of
errors that a student is making and the time taken by a student to type the code
corresponding to each of these rules. These data have indicated that students
acquire the skill of LISP by independently acquiring each of the 500 rules.
Figure 9.17 displays the learning curves for these rules. The two dependent
measures are the number of errors made on a rule and the time taken to write
the code corresponding to a rule (when that rule is correctly coded). These
statistics are plotted as a function of learning opportunities, which present
themselves each time the student comes to a point in a problem where that rule
can be applied. As can be seen, performance on these rules dramatically improves
from first to second learning opportunity and improves more gradually
thereafter. These learning curves are similar to those identified in Chapter 6 for
the learning of simple associations.
Individual differences in the learning of these rules have been taken into
account. Students who have already learned a programming language are at a considerable
advantage compared with students for whom their first programming
Educational Implications | 269
1
(a) Opportunities
Number of errors
.20
.50
1.00
2 3−4 5−8
(b)
Coding time (s)
1
5
10
20
Opportunities
2 3−4 5−8
FIGURE 9.17 Data from the LISP tutor: (a) number of errors (maximum is three) per rule as
a function of the number of opportunities for practice; (b) time to correctly code rules as a
function of the amount of practice.
Anderson7e_Chapter_09.qxd 8/20/09 9:49 AM Page 269
language is that of the LISP tutor. The “identical elements model” of transfer,
in which rules for programming in one language transfer to programming in
another language, can account for this advantage.
We also analyzed the performance of individual students in the LISP tutor
and found evidence for two factors. Some students were able to learn new rules
in a lesson quite rapidly, whereas other students had more difficulty. More or
less independent of this acquisition factor, students could be classified according
to how well they retained rules from earlier lessons.3 Thus, students differ in
how rapidly they learn with the LISP tutor. However, the tutor employs a mastery
learning system in which slower students are given more practice and so
are brought to the same level of mastery of the material as that of the others.
Students emerge from their interactions with the LISP tutor having acquired
a complex and sophisticated skill. Their enhanced programming
abilities make them appear more intelligent among their peers. However, when
we examine what underlies that newfound intelligence, we find that it is the
methodical acquisition of some 500 rules of programming. Some students can
acquire these rules more easily than others because of past experience and specific
abilities. However, when they graduate from the LISP course, all students
have learned the 500 new rules.With the acquisition of these rules, few differences
remain among the students with respect to ability to program in LISP.
Thus, we see that, in the end, what is important with respect to individual
differences is how much information students have learned, not their native
ability.
By carefully monitoring individual components of a skill and providing feedback
on learning, intelligent tutors can help students rapidly master complex
skills.
•Conclusions
This chapter began by noting the remarkable ability of humans to acquire the
complexities of culture and technology. In fact, in today’s world people can
expect to acquire a whole new set of skills over their lifetimes. For instance, I
now use my phone for instant messaging, GPS navigation, and surfing the
Web—none of which I imagined when I was a young man, let alone associated
with a phone. This chapter has emphasized the role of practice in acquiring
such skills, and certainly it has taken me some considerable practice to master
these new skills. However, human flexibility depends on more than time on
task—other creatures could never acquire such skills no matter how much they
practiced. Critical to human expertise are the higher-order problem-solving
skills that we reviewed in the previous chapter. Also critical is human ability to
reason, make decisions, and communicate by language. These are the topics of
the forthcoming chapters.
270 | Expertise
3 These acquisition and retention factors were strongly related to math SATs, but not to verbal SATs.
Anderson7e_Chapter_09.qxd 8/20/09 9:50 AM Page 270
Key Terms | 271
100
0.50
2.50
1.00
200
Number of books (log scale)
Months to complete a book (log scale)
300 500
FIGURE 9.18 Time to complete a book as a function of
practice, plotted with logarithmic coordinates on both axes.
(From Ohlsson, 1992).
1. An interesting case study of skill acquisition was reported
by Ohlsson (1992), who looked at the development of
Isaac Asimov’s writing skill. Asimov was one of the
most prolific authors of our time, writing approximately
500 books in a career that spanned 40 years.He sat
down at his keyboard every day at 7:30 A.M. and wrote
until 10:00 P.M. Figure 9.18 shows the average number
of months he took to write a book as a function of
practice on a log-log scale. It corresponds closely to a
power function. At what stage of skill acquisition do
you think Asimov was in terms of his writing skills?
2. The chapter discussed how chess experts have learned
to recognize appropriate moves just by looking at the
chessboard. It has been argued (Charness, 1981;
Holding, 1992; Roring, 2008) that experts also learn
to engage in more search and more effective search
for winning moves. Relate these two kinds of learning
(learning specific moves and learning how to search)
to the concepts of tactical and strategic learning.
3. In a 2006 New York Times article, Stephen J. Dubner
and Steven D. Levitt (of “Freakonomics” fame)
noted that elite soccer players are much more likely
to be born in the early months of the year than the
late months. Anders Ericsson argues they have an
advantage in youth soccer leagues, which organize
teams by birth year. Because they are older and tend
to be bigger than other children of the same birth year,
they are more likely to get selected for elite teams and
receive the benefit of deliberate practice. Can you
think of any other explanations for the fact that elite
soccer players tend to be born in the first months
of the year?
4. One reads frequent complaints about the performance
level of American students in studies of mathematics
achievement, where they are greatly outperformed
by children from other countries like Japan. Frequent
remedies point to changing the nature of the mathematics
curriculum or improving teacher quality.
Seldom mentioned is the fact that American children
actually spend much less time learning mathematics
(see Anderson, Reder, & Simon, 1998).What does this
chapter imply about the importance of instruction
versus amount of learning time? Can improvements
in one of these increase American achievement levels
without improvements in the other?
Questions for Thought
Key Terms
associative stage
autonomous stage
cognitive stage
componential analysis
deliberate practice
intelligent tutoring systems
mastery learning
negative transfer
proceduralization
strategic learning
tactical learning
theory of identical elements
Anderson7e_Chapter_09.qxd 8/20/09 9:50 AM Page 271
272
10Reasoning
As noted in Chapter 1, intelligence is thought to be the feature that distinguishes
humans as a species. In the last two chapters, we examined the enormous
capacity that we enjoy as a species to solve problems and acquire new intellectual
skills. In light of this particular capacity, we might expect that the research on human
reasoning (the topic of this chapter) and decision making (the topic of the next
chapter) would document how we achieve our superior intellectual performance.
Historically, however, most psychological research on reasoning and decision making
has started with prescriptions derived from logic and mathematics about how humans
should behave, compared these prescriptions to what humans actually do, and found
humans deficient compared to these standards.
The opposite conclusion seems to come from older research in artificial intelligence
(AI) where researchers tried to create artificial systems for reasoning and
decision making using the same prescriptions from logic and mathematics. For instance,
(Shortliffe, 1976) created an expert computer-based system for diagnosing
infectious diseases. Similar formal reasoning mechanisms were used in the first
generation of robots to help them reason about how to navigate through the
world. Researchers were very frustrated with such systems, noting that they
lacked common sense and would do the stupidest things that no human would do.
Faced with such frustrations, researchers are now creating systems based on less
logical computations, often emulating how neurons in the brain compute (e.g.,
Russell & Norvig, 2003).
Thus, we have a paradox: Human reasoning is judged as deficient when compared
against the standards of logic and mathematics, but AI systems built on these very
standards are judged as deficient when compared against the standards of humans.
This apparent contradiction might lead one to conclude either that logic and mathematics
are wrong or that humans have some mysterious intuition that guides their
thinking. However, the real problem seems to be with the way the principles of logic
and mathematics have been applied, not with the principles themselves. New research
is showing that the situations faced by people are more complex than often assumed.
We can better understand human behavior when we expand our analyses of human
reasoning to include the complexities. In this chapter and the next, we will review a
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 272
Reasoning and the Brain | 273
number of the models used to predict how people arrive at conclusions when presented
with certain evidence, research on how people deviated from these models,
followed by the newer and richer analyses of human reasoning.
This chapter will address the following questions about the way people reason: • How do they reason about situations described in conditional language
(e.g., “if–then”)? • How do they reason about situations described with quantifiers like all, some,
and none? • How do people reason from specific examples and pieces of evidence to
general conclusions?
•Reasoning and the Brain
There has been some research investigating brain areas involved in reasoning,
and it suggests that people can bring different systems to bear on different
reasoning problems. Consider an fMRI experiment by Goel, Buchel, Frith, and
Dolan (2000). They had participants solve logical syllogisms, arguments consisting
of two premises and a conclusion, of the type that will be discussed in
the second section of this chapter. Participants were presented with congruent
problems such as
All poodles are pets.
All pets have names.
All poodles have names.
and had to judge whether the third statement followed from the first two. The
content of this example is more or less consistent with what people believe
about pets and poodles. Goel et al. contrasted this type of problem with incongruent
problems whose premises and conclusions violated standard beliefs
such as
All pets are poodles.
All poodles are vicious.
All pets are vicious.
They contrasted both of these types with reasoning about abstract concepts,
such as
All P are B.
All B are C.
All P are C.
Logicians would call all three kinds of syllogism valid.
Participants achieved 84% accuracy in their ability to judge the validity of
the first congruent syllogisms and only 74% in their judgments of the second
incongruent material. They achieved 77% accuracy in their ability to judge the
content-free material. The reader might wonder about the sensibility of judging
‹
‹
‹
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 273
a participant as making a mistake in rejecting an incongruent conclusion such
as “All pets are vicious”; we will return to this matter in the second section of
the chapter. For now, of greater interest are the brain regions that were active
when participants were judging material with content and when they were
judging material without content; these areas are illustrated in Figure 10.1.
When participants were judging content-free material, parietal regions that
have been found to have roles in solving algebraic equations were active (see
Figure 1.16b).When they were judging meaningful content, left prefrontal and
temporal-parietal areas that are associated with language processing were active
(see Figure 4.1).We will find that areas such as the latter regions are frequently
active during reasoning about real problems and that they are associated both
with better performance on some problems and with what might be considered
worse performance on others. What this tells us is that people are capable of
approaching logical problems in two rather different ways.
Faced with logical problems, people can engage either brain regions associated
with the processing of meaningful content or regions associated with the processing
of more abstract information.
•Reasoning about Conditionals
The first body of research we will cover concerns deductive reasoning. Deductive
reasoning is concerned with conclusions that follow with certainty from their
premises. It is distinguished from inductive reasoning,which is concerned with
274 | Reasoning
FIGURE 10.1 Comparison of
brain regions activated when
people reason about problems
with meaningful content versus
when they reason about
material without content.
Brain Structures
Posterior parietal:
Reasoning about
content-free material
Ventral prefrontal:
Reasoning about
meaningful content
Parietal-temporal:
Reasoning about
meaningful content
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 274
conclusions that probabilistically follow from their premises. To illustrate the
distinction, suppose someone is told, “Fred is the brother of Mary” and“Mary is
the mother of Lisa.” Then, one might conclude that “Fred is the uncle of Lisa”
and that “Fred is older than Lisa.” The first conclusion, “Fred is the uncle of
Lisa,” would be a correct deductive inference given the definition of familial
relationships. On the other hand, the second conclusion, “Fred is older than Lisa,”
is a good inductive inference, because it is probably true, but not a correct
deductive inference, because it is not necessarily true.
Our first topic will concern human deductive reasoning using the conditional
connective if. A conditional statement is an assertion, such as “If you
read this chapter, you will be wiser.” The if part (you read this chapter) is called
the antecedent and the then part (you will be wiser) is called the consequent.
Table 10.1 lays out the structure of conditional statements and various valid
and invalid rules of inference.We will discuss these rules of inference below.
A particularly central rule of inference in the logic of the conditional is
known as modus ponens (loosely translates from Latin as “method for affirming”) .
It allows us to infer the consequent of a conditional if we are given the
antecedent. Thus, given both the proposition If A, then B and the proposition
A, we can infer B. So, suppose we are told the following premises and
conclusion:
Modus Ponens
If Joan understood this book, then she would get a good grade.
Joan understood this book.
Therefore, Joan got a good grade.
This example is an instance of valid deduction. By valid, we mean that, if premises
1 and 2 are true, then conclusion 3 must be true. This example also illustrates
the artificiality of applying logic to real-world situations. How is one to
really know whether Joan understands the book? One can only assign a certain
probability to her understanding. Even if Joan does understand the book, at
Reasoning about Conditionals | 275
TABLE 10.1
Analysis of a Conditional Statement and Various Valid and Invalid Rules of Inference
A conditional statement:
The antecedent The consequent
(A) (B)
If you read this chapter, you will be wiser.
Name Rule of Inference
Valid deductions Modus ponens Given A is true, infer B is true.
Modus tollens Given B is false, infer A is false.
Invalid deductions Affirmation of the consequent Given B is true, infer A is true.
Denial of the antecedent Given A is false, infer B is false.
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 275
best it is only likely—not certain—that she will get a good grade. However,
participants are asked to suspend their knowledge about such matters and treat
these facts as if they are certainties. Or, more precisely, they are asked to reason
what would follow for certain if these facts were certain.1 Participants do not
find these instructions particularly strange, but, as we will see, they are not
always able to make logically correct inferences.
Another rule of inference is known in logic as modus tollens (loosely translates
as “method of denying”) . This rule states that, if we are given the proposition
A implies B and the fact that B is false, then we can infer that A is false. The
following inference exercise requires modus tollens:
Modus Tollens
If Joan understood this book, then she would get a good grade.
Joan did not get a good grade.
Therefore, Joan did not understand this book.
This conclusion might strike the reader as less than totally compelling because,
again, in the real world such statements are not typically treated as certain.
Modus ponens allows us to infer the consequent from the antecedent; modus
tollens allows us to infer the antecedent is false if the consequent is false.
Evaluation of Conditional Arguments
There are other inference patterns that people sometimes accept but which are
invalid. One is called affirmation of the consequent and is illustrated by the
following incorrect pattern of reasoning.
Affirmation of the Consequent
If Joan understood this book, then she would get a good grade.
Joan got a good grade.
Therefore, Joan did understand this book.
The other incorrect pattern is called denial of the antecedent and is illustrated
by the following pattern of reasoning.
Denial of the Antecedent
If Joan understood this book, then she would get a good grade.
Joan did not understand this book.
Therefore, Joan did not get a grade.
In both of these cases there could be other ways that Joan could get a good
grade, such as writing a great term essay. Evans (1993) reviewed a large
number of studies that compared the frequency with which people accept the
276 | Reasoning
1 Interestingly, the mathematical theory of probability includes conditional statements. In this case, the objects
of the conditional statement are statements about probabilities, which illustrates the fact that precise
mathematics requires the formal logic of the conditional; it should not be taken as an illustration of a way
of incorporating the formal logic of the conditional into a theory of everyday reasoning.
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 276
valid modus ponens and modus tollens inferences as well as the frequency with
which they accept the invalid inferences. The average percent acceptance over
these studies is plotted in Figure 10.2. As can be seen, people rarely fail to
accept a modus ponens inference but the frequency with which they accept the
valid modus tollens is only slightly greater than the frequencies with which
they accept the invalid affirmation of the consequent or the invalid denial of
the antecedent.
People are only able to show high levels of logical reasoning with modus
ponens.
Evaluating Conditional Arguments in a Larger Context
Byrne (1989) performed an interesting variation of the typical conditional reasoning
study that illustrates that human reasoning is sensitive to things that are
ignored in a simple classification like that shown in Table 10.1. In one condition,
she presented her participants with syllogisms like these:
If she has an essay to write, she will study late in the library.
(If she has textbooks to read, she will study late in the library.)
She will stay late in the library.
Therefore, she has an essay to write.
One group of participants did not see the premise in parentheses, whereas the
other group of participants did. Without the additional premise, her participants
accepted the conclusion 71% of the time, committing the fallacy of affirmation
of the consequent. On the other hand, given the parenthetical premise
in addition to the other premises, their acceptance of the conclusion went down
Reasoning about Conditionals | 277
Modus pones Modus tollens
Percent Acceptance
Affirmation of
the consequent
Denial of the
antecedent
20%
0%
40%
60%
80%
100%
Invalid inferences
Valid inferences
FIGURE 10.2 Frequency with which various conditional syllogisms are accepted—from Evans
(1993).
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 277
to 13%. So we see people can be much more accurate in their reasoning if the
material engages them to have a richer interpretation of the situation.
These results of Byrne are even more interesting when compared with another
situation in which she used examples like the following:
If she has an essay to write, she will study late in the library.
(If the library stays open, then she will study in the library.)
She has an essay to write.
Therefore, she will study late in the library.
Without the additional statement in parentheses, participants accepted the
modus ponens inference 96% of the time. However, with the additional statement,
their acceptance rate went down to 38%. In a narrow, logical sense, the
participants are making an error in not accepting the conclusion with the additional
premise. However, in the world outside of the laboratory, they would
be viewed as making the right judgment—how could she actually study in the
library if it were not open? AI researchers would be frustrated if their programs
still made the conclusion with this additional premise. People haveReasoning
As noted in Chapter 1, intelligence is thought to be the feature that distinguishes
humans as a species. In the last two chapters, we examined the enormous
capacity that we enjoy as a species to solve problems and acquire new intellectual
skills. In light of this particular capacity, we might expect that the research on human
reasoning (the topic of this chapter) and decision making (the topic of the next
chapter) would document how we achieve our superior intellectual performance.
Historically, however, most psychological research on reasoning and decision making
has started with prescriptions derived from logic and mathematics about how humans
should behave, compared these prescriptions to what humans actually do, and found
humans deficient compared to these standards.
The opposite conclusion seems to come from older research in artificial intelligence
(AI) where researchers tried to create artificial systems for reasoning and
decision making using the same prescriptions from logic and mathematics. For instance,
(Shortliffe, 1976) created an expert computer-based system for diagnosing
infectious diseases. Similar formal reasoning mechanisms were used in the first
generation of robots to help them reason about how to navigate through the
world. Researchers were very frustrated with such systems, noting that they
lacked common sense and would do the stupidest things that no human would do.
Faced with such frustrations, researchers are now creating systems based on less
logical computations, often emulating how neurons in the brain compute (e.g.,
Russell & Norvig, 2003).
Thus, we have a paradox: Human reasoning is judged as deficient when compared
against the standards of logic and mathematics, but AI systems built on these very
standards are judged as deficient when compared against the standards of humans.
This apparent contradiction might lead one to conclude either that logic and mathematics
are wrong or that humans have some mysterious intuition that guides their
thinking. However, the real problem seems to be with the way the principles of logic
and mathematics have been applied, not with the principles themselves. New research
is showing that the situations faced by people are more complex than often assumed.
We can better understand human behavior when we expand our analyses of human
reasoning to include the complexities. In this chapter and the next, we will review a
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 272
Reasoning and the Brain | 273
number of the models used to predict how people arrive at conclusions when presented
with certain evidence, research on how people deviated from these models,
followed by the newer and richer analyses of human reasoning.
This chapter will address the following questions about the way people reason: • How do they reason about situations described in conditional language
(e.g., “if–then”)? • How do they reason about situations described with quantifiers like all, some,
and none? • How do people reason from specific examples and pieces of evidence to
general conclusions?
•Reasoning and the Brain
There has been some research investigating brain areas involved in reasoning,
and it suggests that people can bring different systems to bear on different
reasoning problems. Consider an fMRI experiment by Goel, Buchel, Frith, and
Dolan (2000). They had participants solve logical syllogisms, arguments consisting
of two premises and a conclusion, of the type that will be discussed in
the second section of this chapter. Participants were presented with congruent
problems such as
All poodles are pets.
All pets have names.
All poodles have names.
and had to judge whether the third statement followed from the first two. The
content of this example is more or less consistent with what people believe
about pets and poodles. Goel et al. contrasted this type of problem with incongruent
problems whose premises and conclusions violated standard beliefs
such as
All pets are poodles.
All poodles are vicious.
All pets are vicious.
They contrasted both of these types with reasoning about abstract concepts,
such as
All P are B.
All B are C.
All P are C.
Logicians would call all three kinds of syllogism valid.
Participants achieved 84% accuracy in their ability to judge the validity of
the first congruent syllogisms and only 74% in their judgments of the second
incongruent material. They achieved 77% accuracy in their ability to judge the
content-free material. The reader might wonder about the sensibility of judging
‹
‹
‹
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 273
a participant as making a mistake in rejecting an incongruent conclusion such
as “All pets are vicious”; we will return to this matter in the second section of
the chapter. For now, of greater interest are the brain regions that were active
when participants were judging material with content and when they were
judging material without content; these areas are illustrated in Figure 10.1.
When participants were judging content-free material, parietal regions that
have been found to have roles in solving algebraic equations were active (see
Figure 1.16b).When they were judging meaningful content, left prefrontal and
temporal-parietal areas that are associated with language processing were active
(see Figure 4.1).We will find that areas such as the latter regions are frequently
active during reasoning about real problems and that they are associated both
with better performance on some problems and with what might be considered
worse performance on others. What this tells us is that people are capable of
approaching logical problems in two rather different ways.
Faced with logical problems, people can engage either brain regions associated
with the processing of meaningful content or regions associated with the processing
of more abstract information.
•Reasoning about Conditionals
The first body of research we will cover concerns deductive reasoning. Deductive
reasoning is concerned with conclusions that follow with certainty from their
premises. It is distinguished from inductive reasoning,which is concerned with
274 | Reasoning
FIGURE 10.1 Comparison of
brain regions activated when
people reason about problems
with meaningful content versus
when they reason about
material without content.
Brain Structures
Posterior parietal:
Reasoning about
content-free material
Ventral prefrontal:
Reasoning about
meaningful content
Parietal-temporal:
Reasoning about
meaningful content
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 274
conclusions that probabilistically follow from their premises. To illustrate the
distinction, suppose someone is told, “Fred is the brother of Mary” and“Mary is
the mother of Lisa.” Then, one might conclude that “Fred is the uncle of Lisa”
and that “Fred is older than Lisa.” The first conclusion, “Fred is the uncle of
Lisa,” would be a correct deductive inference given the definition of familial
relationships. On the other hand, the second conclusion, “Fred is older than Lisa,”
is a good inductive inference, because it is probably true, but not a correct
deductive inference, because it is not necessarily true.
Our first topic will concern human deductive reasoning using the conditional
connective if. A conditional statement is an assertion, such as “If you
read this chapter, you will be wiser.” The if part (you read this chapter) is called
the antecedent and the then part (you will be wiser) is called the consequent.
Table 10.1 lays out the structure of conditional statements and various valid
and invalid rules of inference.We will discuss these rules of inference below.
A particularly central rule of inference in the logic of the conditional is
known as modus ponens (loosely translates from Latin as “method for affirming”) .
It allows us to infer the consequent of a conditional if we are given the
antecedent. Thus, given both the proposition If A, then B and the proposition
A, we can infer B. So, suppose we are told the following premises and
conclusion:
Modus Ponens
If Joan understood this book, then she would get a good grade.
Joan understood this book.
Therefore, Joan got a good grade.
This example is an instance of valid deduction. By valid, we mean that, if premises
1 and 2 are true, then conclusion 3 must be true. This example also illustrates
the artificiality of applying logic to real-world situations. How is one to
really know whether Joan understands the book? One can only assign a certain
probability to her understanding. Even if Joan does understand the book, at
Reasoning about Conditionals | 275
TABLE 10.1
Analysis of a Conditional Statement and Various Valid and Invalid Rules of Inference
A conditional statement:
The antecedent The consequent
(A) (B)
If you read this chapter, you will be wiser.
Name Rule of Inference
Valid deductions Modus ponens Given A is true, infer B is true.
Modus tollens Given B is false, infer A is false.
Invalid deductions Affirmation of the consequent Given B is true, infer A is true.
Denial of the antecedent Given A is false, infer B is false.
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 275
best it is only likely—not certain—that she will get a good grade. However,
participants are asked to suspend their knowledge about such matters and treat
these facts as if they are certainties. Or, more precisely, they are asked to reason
what would follow for certain if these facts were certain.1 Participants do not
find these instructions particularly strange, but, as we will see, they are not
always able to make logically correct inferences.
Another rule of inference is known in logic as modus tollens (loosely translates
as “method of denying”) . This rule states that, if we are given the proposition
A implies B and the fact that B is false, then we can infer that A is false. The
following inference exercise requires modus tollens:
Modus Tollens
If Joan understood this book, then she would get a good grade.
Joan did not get a good grade.
Therefore, Joan did not understand this book.
This conclusion might strike the reader as less than totally compelling because,
again, in the real world such statements are not typically treated as certain.
Modus ponens allows us to infer the consequent from the antecedent; modus
tollens allows us to infer the antecedent is false if the consequent is false.
Evaluation of Conditional Arguments
There are other inference patterns that people sometimes accept but which are
invalid. One is called affirmation of the consequent and is illustrated by the
following incorrect pattern of reasoning.
Affirmation of the Consequent
If Joan understood this book, then she would get a good grade.
Joan got a good grade.
Therefore, Joan did understand this book.
The other incorrect pattern is called denial of the antecedent and is illustrated
by the following pattern of reasoning.
Denial of the Antecedent
If Joan understood this book, then she would get a good grade.
Joan did not understand this book.
Therefore, Joan did not get a grade.
In both of these cases there could be other ways that Joan could get a good
grade, such as writing a great term essay. Evans (1993) reviewed a large
number of studies that compared the frequency with which people accept the
276 | Reasoning
1 Interestingly, the mathematical theory of probability includes conditional statements. In this case, the objects
of the conditional statement are statements about probabilities, which illustrates the fact that precise
mathematics requires the formal logic of the conditional; it should not be taken as an illustration of a way
of incorporating the formal logic of the conditional into a theory of everyday reasoning.
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 276
valid modus ponens and modus tollens inferences as well as the frequency with
which they accept the invalid inferences. The average percent acceptance over
these studies is plotted in Figure 10.2. As can be seen, people rarely fail to
accept a modus ponens inference but the frequency with which they accept the
valid modus tollens is only slightly greater than the frequencies with which
they accept the invalid affirmation of the consequent or the invalid denial of
the antecedent.
People are only able to show high levels of logical reasoning with modus
ponens.
Evaluating Conditional Arguments in a Larger Context
Byrne (1989) performed an interesting variation of the typical conditional reasoning
study that illustrates that human reasoning is sensitive to things that are
ignored in a simple classification like that shown in Table 10.1. In one condition,
she presented her participants with syllogisms like these:
If she has an essay to write, she will study late in the library.
(If she has textbooks to read, she will study late in the library.)
She will stay late in the library.
Therefore, she has an essay to write.
One group of participants did not see the premise in parentheses, whereas the
other group of participants did. Without the additional premise, her participants
accepted the conclusion 71% of the time, committing the fallacy of affirmation
of the consequent. On the other hand, given the parenthetical premise
in addition to the other premises, their acceptance of the conclusion went down
Reasoning about Conditionals | 277
Modus pones Modus tollens
Percent Acceptance
Affirmation of
the consequent
Denial of the
antecedent
20%
0%
40%
60%
80%
100%
Invalid inferences
Valid inferences
FIGURE 10.2 Frequency with which various conditional syllogisms are accepted—from Evans
(1993).
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 277
to 13%. So we see people can be much more accurate in their reasoning if the
material engages them to have a richer interpretation of the situation.
These results of Byrne are even more interesting when compared with another
situation in which she used examples like the following:
If she has an essay to write, she will study late in the library.
(If the library stays open, then she will study in the library.)
She has an essay to write.
Therefore, she will study late in the library.
Without the additional statement in parentheses, participants accepted the
modus ponens inference 96% of the time. However, with the additional statement,
their acceptance rate went down to 38%. In a narrow, logical sense, the
participants are making an error in not accepting the conclusion with the additional
premise. However, in the world outside of the laboratory, they would
be viewed as making the right judgment—how could she actually study in the
library if it were not open? AI researchers would be frustrated if their programs
still made the conclusion with this additional premise. People have a
rich ability to reason about the real world, and it can intrude and cause them
to make errors in these studies where they are told to reason by the strict rules
of logic. However, it can lead them to make the right decisions in the real
world.
When people’s ability to reason about real-world situations intrudes into
logical reasoning tasks, it can result in better or worse performance.
The Wason Selection Task
A series of experiments initially begun by Peter Wason (for a review of the early
research, see Wason & Johnson-Laird, 1972, Chapters 13 and 14) have been
taken as a striking demonstration of human inability to reason correctly. In a
typical experiment in this research, four cards showing the following symbols
were placed in front of participants:
278 | Reasoning
E K 4 7
Participants were told that a letter appeared on one side of each card and a
number on the other. Their task was to judge the validity of the following rule,
which referred only to these four cards:
If a card has a vowel on one side, then it has an even number on the other side.
The participants’ task was to turn over only those cards that had to be turned
over for the correctness of the rule to be judged. This task, typically referred to
as the selection task, has received a great deal of research.
Averaging over a large number of experiments (Oaksford & Chater, 1994),
about 90% of the participants have been found to select E, which is a logically
correct choice because an odd number on the other side would disconfirm the
rule. However, about 60% of the participants also choose to turn over the 4,
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 278
which is not logically informative because neither a vowel nor a consonant on
the other side would have falsified the rule. Only 25% elect to turn over the 7,
which is a logically informative choice because a vowel behind the 7 would have
falsified the rule. Only about 15% elect to turn over the K, which would not be
an informative choice.
Thus, participants display two types of logical errors in the task. First, they
often turn over the 4, an example of the fallacy of affirming the consequent.
Even more striking is the failure to take the modus tollens step of disconfirming
the consequent and determining whether the antecedent also is disconfirmed
(in other words, to turn over the 7).
The number of people that make the right combination of choices, turning
over only the E and 7, is often only 10%, which has been taken as a damning
indictment of human reasoning. Early in the history of research on the selection
task, Wason gave a talk at the IBM Research Center in which he presented this
same problem to an audience filled with Ph.D.s, many in mathematics and
physics.He got the same poor results from this audience and, reportedly, they were
so embarrassed that they harassed Wason with complaints about how the problem
was not accurately presented or the correct answer was not really correct. This
question of what the right answer is has been recently explored but, before considering
it, we will see what happens when one puts content into these problems.
When presented with neutral material in the Wason selection task, people
have particular difficulty in recognizing the importance of exploring if the
consequent is false.
Permission Interpretation of the Conditional
A person’s performance can sometimes be greatly enhanced when the material
to be judged has meaningful content. Griggs and Cox (1982) were among the
first to demonstrate this enhancement in a paradigm that is formally equivalent
to the Wason card-selection task. Participants were instructed to imagine that
they were police officers responsible for ensuring that the following regulation
was being followed: If a person is drinking beer, then the person must be over 19.
They were presented with four cards that represented people sitting around a
table. On one side of each card was the age of the person and on the other side
was the substance that the person was drinking. The cards were labeled “Drinking
beer,”“Drinking Coke,” “16 years of age,” and “22 years of age.” The task was
to select those people (cards to turn over) from whom further information was
needed to determine whether the drinking law was being violated. In this situation,
74% of the participants selected the logically correct cards (namely,
“Drinking beer” and “16 years of age”). Interestingly, patients with damage to
the ventromedial prefrontal cortex do not show this advantage with content
(Adolphs, Tranel, Bechara, Damasio, & Damasio, 1996). We will discuss this
patient population more thoroughly in the next chapter.
It has been argued that the better performance in this task depends on the
fact that the conditional statement is being interpreted as a rule about a social
norm called the permission schema. Society has many rules about how its
Reasoning about Conditionals | 279
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 279
members should conduct themselves, and the argument is that people are good
at applying such social rules (Cheng & Holyoak, 1985). An alternate possibility
is that better performance in this task depends not on the permission semantics
but on the greater familiarity of the participants with the rule. The participants
were Florida undergraduates, and this rule about drinking was in force in
Florida at the time. Would the participants have been able to reason about a
similar but unfamiliar law? To discriminate between these two possibilities,
Cheng and Holyoak (1985) performed the following experiment. One group of
participants was asked to evaluate the following apparently senseless rule
against a set of instances: “If the form says ‘entering’ on one side, then the other
side includes cholera among the list of diseases.” Another group was given the
same rule as well as the rationale that to satisfy immigration officials upon entering
a particular country, one must have been vaccinated for cholera. This rationale
should invoke people’s ability to reason about the permission schema.
The forms indicated on one side whether the passenger was entering the country
or in transit, whereas the other side listed the names of diseases for which he
or she was vaccinated. Participants were presented with a set of forms that said
“Transit,” or “Entering,” or “cholera, typhoid, hepatitis,” or “typhoid, hepatitis.”
The performance of the group given the rationale was much better than that of
the group given just the senseless rule without explanation; that is, the former
group knew to check the “Entering” form and the “typhoid, hepatitis” form.
Because the participants were not familiar with the rule, their good performance
apparently depended on evoking the concept of permission and not on practice
in applying the specific rule.
Cosmides (1989) and Gigerenzer and Hug (1992) argued that our good performance
with such rules (which they call social contract rules) depends on the
skill with which we have learned to detect cheaters. Gigerenzer and Hug had
participants evaluate the following rule:
If a student is assigned to Grover High School, then that student must live in
Grover City.
They saw cards that stated whether the students attended Grover High or not
on one side and whether they lived in Grover city or not on the other side. As in
the original Wason experiment, they had to decide which cards to turn over. In
the cheating condition, participants were asked to take the perspective of a
member of the Grover City School Board looking for students who were illegally
attending the high school. In the non-cheating condition, participants
were asked to take the perspective of a visiting official from the German government
who just wants to find out whether this rule is in effect at Grover High
School. Gigerenzer and Hug were interested in the frequency with which participants
would choose to turn over and check just the two cards: the student
marked as going to Grover High School and the student marked as a nonresident
of Grover City, which are the logically correct choices. In the cheating
condition, where they took the perspective of a school board member, 80% of
the participants chose just these two cards, replicating other results with permission
rules. In the no-cheating condition, where they took the perspective of
a disinterested visitor, only 45% of the participants chose just these two.
280 | Reasoning
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 280
When participants take the perspective of detecting whether a social rule
has been violated, they make a large proportion of logically correct choices
in the Wason card selection task.
Probabilistic Interpretation of the Conditional
The research just reviewed demonstrates that people can show good reasoning
when they adopt what is called the permission interpretation of the conditional.
However, how are we to understand their poor performance in the original
Wason demonstrations where participants are not taking this permission
interpretation? Oaksford and Chater (1994) argued that people tend to interpret
these statements not as strict logical statements but rather as probabilistic
statements about the world. Thus, when someone says, “If A, then B,” they
mean that B will probably occur when A occurs. Even more important to the
Oaksford and Chater argument is the idea that events A and B typically have
low probabilities of occurring in the world—which is what makes such statements
informative. To illustrate their argument, suppose you visited a city and a
friend told you that the following rule held about the cars driving in that city:
If a car has a broken headlight, it will have a broken taillight.
Events A and B (broken headlights and taillights) are rare, and consequently
asserting that one implies the other is informative. Suppose you go to a large
parking lot in which there are hundreds of cars; some are parked with their
fronts exposed and others with their rears exposed. Most do not have their
headlights broken or their taillights broken, but there are one or two with
broken headlights and taillights. On which cars would you check the end not
exposed to test your friend’s claim? Let us consider the following possibilities:
1. A car with a broken headlight: If you saw such a car, like participants in all
of these experiments, you would be inclined to check its taillight. Almost
everyone sees that it is the sensible thing to do.
2. A car without a broken headlight: You would not be inclined to check this
car, like most of the participants in these experiments, and, again,
everyone agrees that you are right.
3. A car with a broken taillight: You would be sorely tempted to see whether
that car did not have a broken headlight (despite the fact that it is supposedly
unnecessary or “illogical”), and Oaksford and Chater agree with
you. The reason is that a car with a broken taillight is so rare that, if it did
have a broken headlight, you would be inclined to believe your friend’s
claim. The coincidence would be too much to shrug off.
4. A car without a broken taillight: You would be reluctant to check every car
in the lot that met this condition (despite the fact that it is supposedly the
logical thing to do), and, again, Oaksford and Chater would agree with
you. The odds of finding a broken headlight on such a car are low because
a broken headlight is rare, and so many cars would have to be checked.
Checking those hundreds of normal cars just does not seem worthwhile.
Reasoning about Conditionals | 281
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 281
Oaksford and Chater developed a mathematical analysis of the optimal
behavior that explains why the typical errors in the original Wason task can
be sensible. Their analysis predicts the frequency of choices in the Wason task.
Their analysis depends on the assumption that properties such as “broken
headlight” and “broken taillight” are rare. For this reason, it is informative
to check the car with a broken taillight as in possibility 3 and is rather
uninformative to check a car without a broken taillight as in possibility 4.
Although the properties might not always be as rare as in this example,
Oaksford and Chater argued that they generally are rare. For instance, more
things are not dogs than are dogs and more things don’t bark than do, and so
the same analysis would apply to a rule such as “If an animal is a dog, then it
will bark” and many other such rules. There is a weakness in the Oaksford
and Chater argument, however, when applied to the original Wason experiment
where the participants were reasoning about even numbers: There are
not more odd numbers than even numbers. Nonetheless, Oaksford argued
that people carry their beliefs that properties are rare into the Wason
situation. There is evidence that manipulations of the probabilities of these
properties do change people’s behavior in the expected way (Oaksford &
Wakefield, 2003).
The behavior in the Wason card selection task can be explained if we assume
that participants select cards that will be informative under a probabilistic
model.
Final Thoughts on the Connective If
The logical connective if can evoke many different interpretations, which reflect
the richness of human cognition. We have considered evidence for its probabilistic
interpretation and its permission interpretation. People are capable of
adopting the logician’s interpretation of it as well, which, not surprisingly, is the
interpretation that logicians and students of logic take of it when doing logic.
Studies of their reasoning (Lewis, 1985; Scheines & Sieg, 1994) with the connective
find it to be similar to mathematical reasoning such as in the domain of
geometry discussed in Chapter 9. Basically, they take a problem-solving approach
to formal reasoning with the connective. Qin et al. (2003) looked at participants
solving abstract logic tasks and found activation in the same parietal
regions (see Figure 10.1) that Goel et al. (2000) found active with their contentfree
material.
An amusing result is that training in logic does not necessarily result in
better behavior on the original Wason selection task. In a study by Cheng,
Holyoak, Nisbett, and Oliver (1986), college students who had just taken a semester
course in logic did only 3% better on the card selection task than those
who had no formal training in logic. It was not that they did not know the rules
of logic; rather, they did choose to apply them in the experiment. When presented
with these problems outside of the logic classroom, the students chose to
adopt some other interpretation of the word if. However, this is not necessarily
a “flaw” in human reasoning. To repeat a point made before, many researchers
282 | Reasoning
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 282
in AI wish their programs were as adaptive in how they interpret the information
they are presented.
People use different problem-solving operators, depending on their interpretation
of the logical connective if.
•Deductive Reasoning:
Reasoning about Quantifiers
Much of human knowledge is cast with logical quantifiers such as all or some.
Witness Lincoln’s famous statement: “You may fool all the people some of the
time; you can even fool some of the people all the time; but you can’t fool all of the
people all the time.” Scientific laws such as Newton’s third law, “For every action
there is always an opposite and equal reaction,” try to identify what is always the
case. It is important to understand how we reason with such quantifiers. This section
will report research on how people reason about such quantifiers when they
appear in simple sentences.As was the case for the logical connective if, we will see
that there are differences between the logician’s interpretation of quantifiers and
the way in which people frequently reason about them.
The Categorical Syllogism
Modern logic is greatly concerned with analyzing the meaning of quantifiers
such as all, no, and some, as in, for example:
All philosophers read some books.
Most of us might believe that this statement is true. The logician would then
say that we were committed to the belief that we could not find a philosopher
who did not read books, but most of us have no trouble accepting the idea that
there were philosophers in societies before there were books or that one still
might find somewhere in the world an illiterate person who professed sufficiently
profound ideas to deserve the title of “philosopher.” This example illustrates
the fact that frequently when we use all in real life, we mean “most” or
“with high probability.” Similarly, when we use no as in
No doctors are poor.
we often mean “hardly any” or “with small probability.” Logicians call both the all
and no statements universal statements because they interpret these statements
as blanket claims with no exceptions. Roger Schank, a famous AI researcher, was
once observed to make the assertion
No one uses universals.
which surely is a sign that people use these words in a richer or more complex
way than implied by the logical analysis.
By the beginning of the 20th century, the sophistication with which logicians
analyzed such quantified statements increased considerably (see Church, 1956,
Deductive Reasoning: Reasoning about Quantifiers | 283
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 283
for a historical discussion). This more advanced treatment of quantifiers is covered
in most modern logic courses. However, most of the research on quantifiers
in psychology has focused on a simpler and older kind of quantified deduction,
called the categorical syllogism. Much of Aristotle’s writing on reasoning concerned
the categorical syllogism. Extensive discussion of these types of syllogisms
can be found in old textbooks on logic such as Cohen and Nagel (1934).
Categorical syllogisms include statements containing the quantifiers some,
all, no, and some–not. Examples of such categorical statements are:
1. All doctors are rich.
2. Some lawyers are dishonest.
3. No politician is trustworthy.
4. Some actors are not handsome.
As a convenient shorthand, the categories (e.g., doctors, rich people, lawyers,
dishonest people) in such statements can be represented by letters—say, A, B, C,
and so on. Thus, the statements might be rendered in this way:
1. All A’s are B’s.
2. Some C’s are D’s.
3. No E’s are F’s.
4. Some G’s are not H’s.
Sometimes, as in the Goel et al. experiment described at the beginning of the
chapter, material is actually presented with such letters.
A categorical syllogism typically contains two premises and a conclusion.
A typical example that might be used in research follows:
1. No Pittsburgher is a Browns fan.
All Browns fans live in Cleveland.
No Pittsburgher lives in Cleveland.
Many people accept this syllogism as logically valid. To see that the conclusion
does not necessarily follow from the form of the premises, consider the following
equivalent syllogism:
2. No man is a woman.
Every woman is a human.
No man is a human.
The first example illustrates a frequent result in research on categorical syllogisms,
which is that people often accept invalid syllogisms. For instance, people accept
the invalid syllogism 1 almost asmuch as they do the following valid syllogism:
3. No Pittsburgher lives in Cleveland.
All Browns fans live in Cleveland.
‹ No Pittsburgher is a Browns fan.
‹
‹
284 | Reasoning
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 284
Research on reasoning with quantifiers has focused on trying to understand
why people accept many invalid categorical syllogisms.
The Atmosphere Hypothesis
The previous example 1 is a case where people are biased by the content of the
syllogism, but much of the research has focused on the tendency of people to
accept invalid syllogisms even when they have neutral content. People are generally
good at recognizing valid syllogisms when stated with neutral content.
For instance, almost everyone accepts
1. All A’s are B’s.
All B’s are C’s.
All A’s are C’s.
The problem is that people also accept many invalid syllogisms. For instance,
many people will accept
2. Some A’s are B’s.
Some B’s are C’s.
Some A’s are C’s.
(To see that this syllogism is invalid, consider replacing A with men, B with
humans, and C with women.) However, people are not completely indiscriminate
in what they accept as valid. For instance, even though they will accept example 2
in the preceding subsection, they will not accept example 3:
3. Some A’s are B’s.
Some B’s are C’s.
No A’s are C’s.
To account for the pattern of what participants accept and what they reject,
Woodworth and Sells (1935) proposed the atmosphere hypothesis. This
hypothesis states that the logical terms (some, all, no, and not) used in the
premises of a syllogism create an “atmosphere” that predisposes participants to
accept conclusions having the same terms. The atmosphere hypothesis consists
of two parts. One part asserts that participants tend to accept a positive conclusion
to positive premises and a negative conclusion to negative premises.When
the premises are mixed, participants tend to prefer a negative. Thus, they would
tend to accept the following invalid syllogism:
4. No A’s are B’s.
All B’s are C’s.
No A’s are C’s.
(This is an abstract form of the same invalid syllogism discussed on the previous
page.)
The other part of the atmosphere hypothesis concerns a participant’s response
to particular statements (some or some not) versus universal statements (all
or no). As example 4 illustrates, participants will accept a universal conclusion
if the premises are universal. They will tend to accept a particular conclusion if
the premises are particular, which accounts for their acceptance of syllogism 2
‹
‹
‹
‹
Deductive Reasoning: Reasoning about Quantifiers | 285
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 285
given earlier. When one premise is particular and the other universal, participants
prefer a particular conclusion. Thus they will accept the following invalid
syllogism:
5. All A’s are B’s.
Some B’s are C ’s.
Some A’s are C ’s.
(To see that this syllogism is invalid, consider replacing A with men, B with
humans, and C with women.)
The atmosphere hypothesis states that the logical terms (some, all, no,
and some not) used in the premises of a syllogism create an “atmosphere”
that predisposes participants to accept conclusions having the same terms.
Limitations of the Atmosphere Hypothesis
The atmosphere hypothesis provides a succinct characterization of participant
behavior with the various syllogisms, but it tells us little about what the participants
are actually thinking or why. It offers no explanation for why the content
of the syllogism (as in the Pittsburgh–Cleveland example) can have such a
strong effect on judgments. Its characterization of participant behavior is also
not always correct for content-free syllogisms. For example, according to the
atmosphere hypothesis, participants are just as likely to accept the atmospherefavored
conclusion when it is not valid as when it is valid. That is, it predicts
that participants would be just as likely to accept
6. All A’s are B’s.
Some B’s are C ’s.
Some A’s are C’s.
which is not valid, as they would be to accept
7. Some A’s are B’s.
All B’s are C’s.
Some A’s are C ’s.
which is valid. In fact, participants are more likely to accept the conclusion in
the valid case. Thus, participants do display some ability to evaluate a syllogism
accurately.
Another limitation of the atmosphere hypothesis is that it fails to predict
the effects that the form of a syllogism will have on participants’ validity judgments.
For instance, the hypothesis predicts that participants would be no more
likely to erroneously accept
8. Some A’s are B’s.
Some B’s are C ’s.
Some A’s are C ’s.
than they would be to erroneously accept
9. Some B’s are A’s.
Some C ’s are B’s.
‹ Some A’s are C ’s.
‹
‹
‹
‹
286 | Reasoning
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 286
In fact, participants are more willing to erroneously accept the conclusion in
the former case (Johnson-Laird & Steedman, 1978). In general, participants are
more willing to accept a conclusion from A to C if they can find a chain leading
from A to B in one premise and from B to C in the second premise.
Another problem with the atmosphere hypothesis is that it does not really
handle what participants do in the presence of two negatives. If participants are
given the following two premises,
No A’s are B’s.
No B’s are C ’s.
the atmosphere hypothesis would predict that participants should tend to accept
the invalid conclusion:
No A’s are C ’s.
Although a few participants do accept this conclusion, most refuse to accept
any conclusion when both premises are negative, which is the correct thing to
do (Dickstein, 1978).
Another problem with the atmosphere hypothesis is that it does not really
explain what people are thinking when they process such syllogisms. It just tries
to predict what conclusions they will accept. The next section will consider
some explanations of the thought processes that lead people to correct or
incorrect conclusions.
Participants only approximate the predictions of the atmosphere hypothesis
and are often more accurate than it would predict.
Process Explanations
One class of explanations is that participants choose not to do what the experimenters
think they are doing. For instance, it has been argued that it is not natural
for people to judge the logical validity of these arguments. Rather, people tend
to judge what conclusions are true. Consider the following pair of syllogisms:
All lawyers are human.
All Republicans are human.
Some lawyers are Republicans.
which has a true conclusion but is not valid (consider replacing lawyers by men
and Republicans by women) and
All bictoids are reptiles.
All bictoids are birds.
Some reptiles are birds.
which is a valid argument but has a false conclusion. People have a greater
tendency to accept the first, invalid argument having a true conclusion than
the second, valid argument having a false conclusion (Evans, Handley, &
Harper, 2001).
It is also argued that many people really do not understand what it means
for an argument to be valid and simply judge whether a conclusion is possible
‹
‹
‹
Deductive Reasoning: Reasoning about Quantifiers | 287
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Pagegiven the premises. So, for example, although the preceding syllogism concerning
lawyers and Republicans is not valid, it is certainly possible given the
premises that the conclusion is true. Evans et al. showed that there is very
little difference in the judgments that participants make when they are asked
to judge when conclusions are necessarily true given the premises (the measure
of a valid argument) and when conclusions are possibly true given the
premises.
Johnson-Laird (Johnson-Laird, 1983; Johnson-Laird & Steedman, 1978)
proposed that participants judge whether a conclusion is possible by creating a
mental model of a world that satisfies the premises of the syllogism and inspecting
that model to see whether the conclusion is satisfied. This explanation
is called mental model theory. Consider these premises:
All the squares are striped.
Some of the striped objects have bold borders.
Figure 10.3a illustrates what a participant might imagine, according to Johnson-
Laird, as an instantiation of these premises. The participant has imagined a
group of objects, some of which are square, whereas others are round; some
of which are striped, whereas others are clear; and some of which have bold
borders, whereas others do not. This world represents one possible interpretation
of these premises. When the participant is asked to judge the following
conclusion,
Some of the squares have bold borders.
the participant inspects the diagram and sees that, indeed, the conclusion is
possible given the premises. The problem is that this response establishes only
that the conclusion is possible but not that it is necessary. For the conclusion
to be necessary, it must be true in all mental models that support
the premises. Figure 10.3b illustrates a case in which the
premises are true but the conclusion does not hold. Johnson-
Laird claimed that participants have considerable difficulty
developing alternative models. Thus, the participant is building
a specific model for the premises and is inspecting it to see what
is true in that model. Johnson-Laird (1983) developed a computer
simulation of this theory that reproduces many of the
errors that participants make. Johnson-Laird (1995) also argued
that there is neurological evidence in favor of the mental model
explanation. He noted that patients with right-hemisphere
damage are more impaired in reasoning tasks than are patients
with left-hemisphere damage. He noted that the right hemisphere
tends to take part in spatial processing of mental images.
In a brain-imaging study, Kroger, Cohen, Nystrom, and
Johnson-Laird (2008) found that the right frontal cortex was
more active than the left in processing such syllogisms but that
the opposite was true when people engaged in arithmetic calculation
(this left bias for arithmetic is also illustrated in the study
‹
288 | Reasoning
(a)
(b)
FIGURE 10.3 Two possible
models that participants might
form for the premises of the
categorical syllogism dealing
with square and round objects.
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 288
described in Figure 1.16). Parsons and Osherson (2001) reported a similar finding,
with deductive reasoning being rightlocalized and probabilistic reasoning
being left localized.
Basically, Johnson-Laird’s argument is that people make errors in reasoning
because they overlook possible explanations of the premises. For example, a
participant imagines Figure 10.3a as an explanation and overlooks the possibility
of Figure 10.3b. Johnson-Laird (personal communication) argues that a great
many errors in human reasoning are produced by failures to consider possible
explanations of the data. For instance, a problem in the Chernobyl disaster was
that, for several hours, engineers failed to consider the possibility that the reactor
was no longer intact.
Evans et al. (2001) argued that people are inclined to look for only one
mental model for the premises and, if they can find such an interpretation, they
accept the argument if the conclusion is valid in that model. The effect of realworld
content is to bias that search for such a mental model. In this way, we can
explain the effect of real-world content. That is, people tend to accept the
invalid “Some lawyers are Republicans” syllogism because it naturally suggests a
model in which the conclusion holds. In contrast, they tend to reject the valid
“Some reptiles are birds” argument because, given its content, it is hard to
imagine a world where it is true.
To successfully adapt to our world, it is important to make correct inferences
about what is true in this world and not about what might be true in all
possible worlds (which is what judging logical validity is concerned with).
Given this perspective, the problem seems to lie as much with the questions
that participants are being asked to answer in these experiments as with the
participants themselves.
Errors in evaluating syllogisms can be explained by assuming that participants
fail to consider possible mental models of the syllogisms.
•Inductive Reasoning and Hypothesis Testing
In contrast to deductive reasoning where logical rules allow one to infer certain
conclusions from premises, in inductive reasoning the conclusion does not necessarily
follow from the premises. Consider the following premises:
The first number in the series is 1.
The second number in the series is 2.
The third number in the series is 4.
What conclusion follows? The numbers are doubling and so possible conclusion
is that
The fourth number is 8.
However, a better conclusion might be to state the general rule:
Each number is twice the previous number.
Inductive Reasoning and Hypothesis Testing | 289
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 289
A characteristic of a good inductive inference like the second conclusion is that it
is a statement from which one can deduce all the premises. For example, because
we know each number is twice the previous number, we can now deduce what the
original three numbers must have been. Thus, in a certain sense it is deduction
turned around. The difficulty for inductive reasoning is that there is usually not a
single conclusion that would be consistent with the premises. For instance, in the
problem above one could have concluded that the difference between successive
numbers is increasing by one and that the fourth number would be 7.
Inductive reasoning is relevant to many aspects of everyday life: a detective
trying to solve a mystery given a set of clues, a doctor trying to diagnose the cause
of a set of symptoms, someone trying to determine what is wrong with a TV, or a
researcher trying to discover a new scientific law. In all these cases, one gets a set
of specific observations from which one is trying to infer some relevant conclusion.
Many of these cases involve the sort of probabilistic reasoning that will be
discussed in the next chapter (for instance, medical symptoms are typically only
associated probabilistically with disease). In this chapter, we will focus on cases,
like the above number example, where we are looking for a hypothesis that
implies the observations with certainty.Much of the interest in such cases is how
people seek evidence relevant to formulating such a hypothesis.
Hypothesis Formation
Bruner, Goodnow, and Austin (1956) performed a classic series of experiments
on hypothesis formation. Figure 10.4 illustrates the kind of material they used.
The stimuli were all rectangular boxes containing various objects. The stimuli
varied on four dimensions: number of objects (one, two, or three); number of
borders around the boxes (one, two, or three); shape (cross, circle, or square);
290 | Reasoning
FIGURE 10.4 Material used
by Bruner et al. in one of
their studies of concept
identification. The array
consists of instances formed by
combinations of four attributes,
each exhibiting three values.
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 290
and color (green, black, or red: represented here as white, black, or
red). Participants were told that they were to discover some concept
that described a particular subset of these instances. For instance,
the concept might have been black crosses. Participants were to discover
the correct concept on the basis of information they were
given about what were and what were not instances of the concept.
Figure 10.5 contains three illustrations (the three columns) of
the information participants might have been presented. Each column
consists of a sequence of instances identified either as members
of the concept (positive, _) or not (negative, _). Each column
represents a different concept. Participants would be presented with
the instances in a column one at a time. From these instances they
would determine what the concept was. Stop reading and try to
determine the concept for each column.
The concept in the first example is two crosses. This concept is referred
to as a conjunctive concept, since the conjunction of a number
of features (in this case the features are two and cross) must be present
for the instance to be positive. People typically find conjunctive
concepts easiest to discover. In some sense conjunctive hypotheses
seem to be the most natural kind of hypotheses. They are the kind
that have been researched most extensively. The solution to the second
example is two borders or circles. This kind of concept is referred to as
a disjunctive concept because an instance is a member of the concept
if either of the features is present. In the final example, the solution is that the
number of objects must equal the number of borders. This example is a relational
concept because it specifies a relationship between two dimensions.
The problems in this series are particularly difficult because to identify the
concept, you must both determine which features are relevant and discover
the kind of rule that connects the features (e.g., conjunctive, disjunctive, or
relational). The former problem is referred to as attribute identification and
the latter as rule learning (Haygood & Bourne, 1965). In many experiments,
either the form of the rule or the relevant attributes are identified for the participant.
For instance, in the Bruner et al. (1956) experiments, participants had
to identify only the correct attributes. They knew that they would be identifying
conjunctive concepts.
Forming a hypothesis involves identifying both what features are relevant
to the hypothesis and how these features are related.
Hypothesis Testing
In the experiment illustrated in Figure 10.5, participants receive a sequence of
instances and have to figure out what the concept is. Some problems in real life
are like this—we have no control over what evidence we see but must figure out
the rules that govern it. For instance, when there is an outbreak of food poisoning
in the United States, medical health researchers check on what the victims
ate, looking for some common pattern. They have no control over what the
Inductive Reasoning and Hypothesis Testing | 291
Concept 1 Concept 2 Concept 3
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
_
FIGURE 10.5 Examples of
sequences of instances from
which a participant is to identify
concepts. Each column gives
a sequence of instances and
non-instances for a different
concept. A plus (_) signals a
positive instance and minus (_)
sign a negative instance.
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 291
victims ate. On the other hand, in other situations one can do experiments and
test certain possibilities. For instance, when medical researchers want to determine
the most effective combination of drugs to treat a disease, they will perform
clinical trials where different groups of patients receive different drug
combinations. Scientific research can reach more certain conclusions more
quickly if the researchers can choose the cases to test rather than having to take
the cases that the situation presents to them.
In their classic research, Bruner, Goodnow, and Austin (1956) also studied
situations where participants could choose which instances they get information
about. For example, in one condition, Bruner et al. gave participants a single positive
instance of concept, and then the participants could select various instances
and ask whether they were also instances of the concept. For example, if you were
told that the middle stimulus in Figure 10.4 (fifth row, fifth column) was an
instance of a conjunctive concept that you had to discover, what stimuli would
you choose to select? The approach advocated in science would be to test each
dimension, one at a time, and determine whether it was critical to the hypothesis.
For instance, you could choose to test first the dimension of number of borders
and choose a stimulus that differed from the initial stimulus only on this
dimension. If the stimulus were not an instance, you would know that the value
of the original stimulus (in this case, two borders) was relevant and if not, you
would know that it was irrelevant. Then you could try another dimension. After
four stimuli, you would have identified the conjunctive concept with certainty.
Bruner et al. called this strategy “conservative focusing,” and some of their participants
(Harvard undergraduates of the 1950s) followed it. However, many participants
practiced less well-behaved strategies. For instance, given the same initial
stimulus, they might test an instance that changed both the color and the number
of borders. If the stimulus were an instance, they would know that neither dimension
was relevant.However, if wrong, they would have learned relatively little.
A well-known case where people seem to test their hypotheses less than
optimally is the 2-4-6 task introduced by Wason (1960—the same psychologist
who introduced the card selection task that we described earlier). In this experiment,
participants are told that “2 4 6” is an instance of a triad that is consistent
with a rule and are instructed to find out what the rule is by asking whether
other triples of numbers are instances of the rule. What triples would you try?
The protocol below comes from one of Wason’s participants. The protocol gives
each triad that the participant produced and the reason for the choice, along
with the experimenter’s feedback as to whether the triad conformed to the rule.
The sequence of triads was occasionally broken when the participant decided
to announce a hypothesis. The experimenter’s feedback for each hypothesis is
given in parenthesis:
Triad Reason Given for Triad Feedback
8 10 12 2 added each time. Yes
14 16 18 Even numbers in order of magnitude. Yes
20 22 24 Same reason. Yes
1 3 5 2 added to preceding number Yes
292 | Reasoning
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 292
Announcement: The rule is that by starting with any number, 2 is added each
time to form the next number. (Incorrect)
2 6 10 The middle number is the arithmetic mean of the
other two. Yes
1 50 99 Same reason. Yes
Announcement: The rule is that the middle number is the arithmetic mean of
the other two. (Incorrect)
3 10 17 Same number, 7, added each time. Yes
0 3 6 Three added each time. Yes
Announcement: The rule is that the difference between two numbers next to
each other is the same. (Incorrect)
12 8 4 The same number is subtracted each time to form
the next number. No
Announcement: The rule is adding a number, always the same one, to form
the next number. (Incorrect)
1 4 9 Any three numbers in order of magnitude. Yes
Announcement: The rule is any three numbers in order of magnitude.
(Correct)
The important feature to note about this protocol is that the participant tested
the hypothesis by almost exclusively generating sequences consistent with it. The
better procedure in this case would have been to also try sequences that were
inconsistent. That is, the participant should have looked sooner for negative
evidence as well as positive evidence. This would have exposed the fact that
the participant had started out with a hypothesis that was too narrow and was
missing the more general correct hypothesis. The only way to discover this
error is to try examples that disconfirm the hypothesis, but this is what people
have great difficulty doing,
In another experiment, Wason (1968) asked 16 participants, after they had
announced their hypotheses, what they would do to determine whether their
hypotheses were incorrect. Nine participants said they would generate only
instances consistent with their hypotheses and wait for one to be identified as
not an instance of the rule. Only four participants said that they would generate
instances inconsistent with the hypothesis to see whether they were identified
as members of the rule. The remaining three insisted that their hypotheses
could not be incorrect.
This strategy to select only positive instances has been called the confirmation
bias. It has been argued that confirmation bias is not necessarily a
mistaken strategy (Fischhoff & Neyth-Marom, 1983; Kayman & Ha, 1987). In
many situations, selecting instances consistent with a hypothesis is an effective
way to disconfirm the hypothesis. For instance, if one did well on an exam after
drinking a glass of orange juice and entertained the hypothesis that orange
juice led to good exam performance, drinking orange juice before a couple
more exams might quickly disabuse one of that hypothesis. What made this
Inductive Reasoning and Hypothesis Testing | 293
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 293
strategy so ineffective in Wason’s experiment is simply that the correct hypothesis
was very general. It would be like finding out that consuming any drink
would improve exam performance (particularly unlikely if we include alcoholic
drinks).
In choosing instances to test a hypothesis, people often focus on instances
consistent with their hypothesis, and this can cause difficulties if their
hypothesis is too narrow.
Scientific Discovery
Whether participants are trying to infer a concept by selecting instances from a
set of options like those in Figure 10.4 or trying to infer a rule that describes a
set of examples as in the protocol we just reviewed, participants are engaged in
problem-solving searches like those we discussed in Chapter 8 (such as in
Figure 8.4 or Figure 8.8). The difference, however, is that they are searching two
problem spaces. One problem space is the space of possible hypotheses and the
other is the space of possible instances. It has been argued (e.g., Simon & Lea,
1974; Klahr & Dunbar, 1998) that this is exactly the situation that scientists face
in discovering a new theory—they search through a space of possible theories
and a space of possible experiments to test these theories.
The term “confirmation bias” has been used to describe failures in the way
people test scientific theories. In the hypothesis-testing example we described, it
just referred to a tendency to test only instances that were an example of one’s
hypothesis. However, in the broader context of testing scientific theories, it refers
to a host of behaviors that serve to protect one’s favored theory from disconfirmation.
In one study, Dunbar (1993) had undergraduates try to discover how
genes were controlled by redoing, in a highly simplified form, the research that
won Jacques Monod and Francois Jacob the 1965 Nobel Prize for medicine.
They provided the participants with computer simulations that could mimic
some of the critical experiments. They were told that their task was to determine
how one set of genes controlled another set of genes that produced an
enzyme only when lactose was present. (This enzyme serves to break down the
lactose into glucose.) All the undergraduates initially thought that there must
be a mechanism by which the first set of genes responded to the presence of
lactose and activated the second set of genes. This is the hypothesis that Monod
and Jacob had initially as well, but in fact the mechanism is an inhibitory mechanism
by which the first set of genes inhibit the enzyme-producing genes when
lactose is absent but are blocked from inhibiting when lactose is present. Showing
the confirmation bias, these undergraduates tried to find experiments that
would confirm their activation hypothesis. The majority of the participants
continued to search the experimental space for some combination of genes that
would support their activation hypothesis, but a minority began to search for
alternative hypotheses about what was in control.
Science as an institution has a way of protecting us from scientists whose confirmation
bias leads them too strongly in the wrong direction. Other scientists are
294 | Reasoning
Anderson7e_Chapter_10.qxd 8/20/09 9:50 AM Page 294
often strongly motivated to find problems with the theories of a particular scientist
(Nickerson, 1998). There is also considerable variation in how individual
scientists practice. Michael Faraday, a famous 19th-century chemist, made his
discoveries by early focusing on collecting confirmatory evidence and then
switching to focusing on disconfirmatory evidence (Tweney, 1989). Dunbar
(1997) studied scientists in three immunology laboratories and one biology
laboratory at Stanford and noted that they were quite ready to attend to unexpected
results and modify their theory to accommodate these.
Fugelsang and Dunbar (2005) performed fMRI imaging studies looking at
participants as they tried to integrate data with specific hypothesis. For instance,
participants were told that they were seeing results from a clinical trial
looking at the effect of an antidepressant on mood. They either saw patient
records that indicated the drug had an effect on mood (consistent) or that it
did not have an effect (inconsistent). Participants started out believing the drug
had an effect and so found consistent evidence more plausible. When viewing
the inconsistent evidence, participants showed greater activation in their anterior
cingulate cortex (ACC) (see Figure 3.1). As we noted in Chapter 3, the ACC
is highly active when participants are engaged in a task that requires strong cognitive
control, such as dealing with an inconsistent trial in a Stroop task. These
same basic brain mechanisms seem to be invoked when participants must deal
with inconsistent data in a scientific context, and the results suggest that scientific
reasoning evokes basic cognitive processes.
Inductive Reasoning and Hypothesis Testing | 295
dropped a rock from a 100-meter tower and timed its
fall as 1 second, it would be wise not to accept the
theory that acceleration due to gravity
was 200 meters (using the formula
distance _ 1/2 _ acceleration _ time2)
rather than the standard accepted value
of approximately 10 meters on earth.
Almost certainly something was wrong in
the measurements and the experiment
needs to be repeated. On the other hand,
the Pasteur case might seem rather extreme,
ignoring 90% of the experiments
on a theory that was much in doubt at
the time. In this case, however, he turned
out to be right.
Implications
How convincing is a 90% result?
Scientists can be subject to a confirmation bias. For
instance, Louis Pasteur was involved in a major debate
with other scientists about whether new
organisms could spontaneously generate.
It was argued that the appearance of
bacteria in apparently sterilized organic
material was evidence for spontaneous
generation of life. Pasteur performed
many experiments trying to disprove this
and 90% of these failed, but he chose
only to publish the successful experiment,
claiming that the rest were due
to experimental errors (Geison, 1995).
Scientists frequently question their
experiment if results seem to contradict
establish theory. For instance, if one
Anderson7e_Chapter_10.qxd 8/20/09 9:51 AM Page 295
1. Johnson-Laird and Goldvarg (1997) presented
Princeton undergraduates with reasoning problems
like this one:
Only one of the following premises is true about a
particular hand of cards:
There is a king in the hand or there is an ace or both.
There is a queen in the hand or there is an ace or both.
There is a jack in the hand or there is a 10, or both.
Is it possible that there is an ace in the hand?
They report that the students only were correct on 1%
of such problems.What is the correct answer for the
problem above? Why is it so hard? Johnson-Laird and
Goldvarg attribute the difficulty that people have in
creating mental models of what is not the case.
2. Johnson-Laird and Steedman (1978) presented the
following syllogisms to participants drawn from
Columbia Teachers College:
All gourmets are shopkeepers.
All bowlers are shopkeepers.
And asked them what followed. The following is
the distribution of answers:
17 agreed that no conclusion followed.
2 thought that “Some gourmets are bowlers” followed.
4 thought that “All bowlers are gourmets” followed.
Questions for Thought
In studies of scientific discovery, participants tend to focus on experiments
consistent with their favorite hypothesis and show a reluctance to search for
alternative hypotheses.
•Conclusions
Much of the research on human reasoning has found it wanting when compared
to the rules and implications of formal logic. As we just noted, this might even
be said of the process by which scientists engage in their research. However, this
dismal characterization of human reasoning fails to properly appreciate the
full context in which it occurs. In many actual reasoning situations, people do
quite well, in part because they take in the full complexity and implications of
the actual real-world content. Despite a tendency towards confirmation bias,
science as a whole has progressed with great success. To some extent, this is
because science is a social activity carried out by a community of researchers.
Competitive scientists are quick to find mistakes in each other’s approach, but
there is also a cooperative nature to science. Research takes place in teams of
researchers and they often rely on each other’s help. Okada and Simon (1997)
found that pairs of undergraduates were much more successful than individual
students at finding the inhibition mechanism in Dunbar’s (1993) genetic control
task. As Okada and Simon note, “In a collaborative situation, subjects must often
be more explicit than in an individual learning situation, to make partners
understand their ideas and to convince them. This can prompt subjects to entertain
requests for explanation and construct deeper explanations” (p. 130). The
bottom line of this chapter is that human reasoning normally takes place in
a world of complexities (both factual and social) and that what appears deficient
in the laboratory may be exquisitely tuned to that world.
296 | Reasoning
Anderson7e_Chapter_10.qxd 8/20/09 9:51 AM Page 296
Key Terms | 297
7 thought that “Some bowlers are gourmets” followed.
8 thought that “All gourmets are bowlers” followed.
Use the concepts of this chapter to help explain the
answers these participants gave and did not give.
3. Consider the third column in Figure 10.5, which was
described in the chapter as satisfying the rule that
“the number of borders is the same as the number of
objects.”An alternative rule that describes the instances
is “3 white objects or 2 black objects or 1 object with
one border.”Which is the better description of the
category and why? Is it possible to know for certain
which is the correct rule?
Key Terms
affirmation of the
consequent
antecedent
atmosphere hypothesis
attribute identification
categorical syllogism
conditional statement
confirmation bias
consequent
deductive reasoning
denial of the antecedent
inductive reasoning
logical quantifiers
mental model theory
modus ponens
modus tollens
particular statements
permission schema
rule learning
selection task
syllogisms
universal statements
Anderson7e_Chapter_10.qxd 8/20/09 9:51 AM Page 297
298
11
Judgment and
Decision Making
As we saw in Chapter 10, most of the research on human reasoning has compared
it to various prescriptive models from logic and mathematics. The prescriptive
models assume that people have access to information about which they can be
certain and that they can coolly reflect on the information. However, in the real world,
people have to make decisions in the face of incomplete and uncertain information.
Furthermore, in contrast to the relatively neutral character of the syllogisms of the
previous chapter, in real life our decisions can have important consequences for our
lives. Consider the simple task of deciding what to eat—we have all been frustrated
by the medical reports that pronounce formerly “healthy” food as “unhealthy” and vice
versa. In decision making, we must also deal with the unpleasant consequences of
what might be good decisions such as going on a diet or giving up a pleasurable
activity like smoking.
This chapter will focus on research on judgment and decision making that
comes closer to such real-life circumstances. As before, we will discuss research
showing how the performance of normal humans is wanting compared to models
that were developed for rational behavior. However, we will also see how much of
that research is incomplete, missing the complexity of everyday human decision
making. Recent research has developed a more nuanced characterization of the situations
that people face in their everyday life, and a better appreciation of the
nature of their judgments.
In this chapter, we will answer the questions: • How well do people judge the probability of uncertain events? • How do people use their past experiences to make judgments? • How do people decide among uncertain options that offer different rewards
and costs? • How does the brain support such decision making?
Anderson7e_Chapter_11.qxd 8/20/09 9:51 AM Page 298
The Brain and Decision Making | 299
•The Brain and Decision Making
In 1848, Phineas Gage, a railroad worker in Vermont, suffered a bizarre accident:
He was using an iron bar to pack gunpowder down into a hole drilled into
a rock that had to be blasted to clear a roadbed for the
railroad. The powder unexpectedly exploded and sent the
iron bar flying through his head before landing 80 feet
away. Figure 11.1 shows a reconstruction of the trajectory
of the bar through his skull (Damasio, Grabowski, Frank,
Galabruda, & Damasio, 1994). (For a more detailed reconstruction,
see Color Plate 11.1.) The bar managed to miss
any vital areas and spared most of his brain but tore
through the middle of the very front of the brain—a
region called the ventromedial prefrontal cortex. Amazingly,
he not only survived, he was even was able to talk
and walk away from the accident after being unconscious
for a few minutes. His recovery was difficult, largely because
of infections, but he eventually was able to hold jobs
such as a coach driver. Henry Jacob Bigelow, a Professor of
Surgery at Harvard University, declared him “quite recovered
in faculties of body and mind” (MacMillan, 2000).
Based on such a report, one might have thought that this
part of the brain performed no function.
However, all was not well. His personality had undergone major changes.
Before his injury he had been polite, respectful, popular, and reliable, and generally
displayed ideal behavior for an American man of that time. Afterward he
became just the opposite—as his own physician, Harlow, later described him:
fitful, irreverent, indulging at times in the grossest profanity (which was not
previously his custom), manifesting but little deference for his fellows, impatient
of restraint or advice when it conflicts with his desires, at times pertinaciously
obstinate, yet capricious and vacillating, devising many plans of future
operations, which are no sooner arranged than they are abandoned in turn
for others appearing more feasible. A child in his intellectual capacity and
manifestations, he has the animal passions of a strong man. Previous to his
injury, although untrained in the schools, he possessed a well-balanced mind,
and was looked upon by those who knew him as a shrewd, smart businessman,
very energetic and persistent in executing all his plans of operation. In
this regard his mind was radically changed, so decidedly that his friends and
acquaintances said he was “no longer Gage.” (Harlow, 1868, p. 327)
Gage is the classic case demonstrating the importance of the ventromedial
prefrontal cortex to human personality. Subsequently, a number of other similar
patients have been described and they all show the same sorts of personality
disorders. Family members and friends will describe them with phrases like
“socially incompetent,”“decides against his best interest,” and “doesn’t learn from
his mistakes” (Sanfey, Hastie, Colvin, & Grafman, 2003). Earlier in Chapter 8, we
Brain Structures
FIGURE 11.1 A representation
of the passage of the bar
through Phineas Gage’s brain.
Note that only the middle
of the frontal-most portion
has been damaged. (From Damasio
et al., 1994).
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 299
discussed the case of the patient PF who also suffered damage to his anterior
prefrontal region, like Gage. However, in his case the damage also included
lateral portions of the anterior prefrontal region, and his difficulty was more
with organizing complex problem solving than decision making. In general, it is
thought that the more medial portion of the anterior prefrontal region, where
Gage’s injury was localized, is important to motivation, emotional regulation,
and social sensitivity (Gilbert, Spengler, Simons, Frith, & Burgess, 2006).
The ventromedial prefrontal cortex plays an important role in achieving the
motivational balance and social sensitivity that is key to making successful
judgments.
•Probabilistic Judgment
There is a prescriptive model for how people should reason about probabilities
as they collect relevant evidence. This is called Bayes’s theorem, which is based
on a mathematical analysis of the nature of probability.Much of the research in
the field has been concerned with showing that human participants do not
match up with the prescriptions of Bayes’s theorem.
Bayes’s Theorem
As an example of the application of Bayes’s theorem, suppose I come home and
find the door to my house ajar. I am interested in the hypothesis that it might
be the work of a burglar. How do I evaluate this hypothesis? I might treat it as
a conditional syllogism of the following sort:
If a burglar is in the house, then the door will be ajar.
The door is ajar.
A burglar is in the house.
As a conditional syllogism, it would be judged as the erroneous affirmation of
the consequent. However, it does have a certain plausibility as an inductive
argument. Bayes’s theorem provides a way of assessing just how plausible it is
by combining what are called a prior probability and a conditional probability
to produce what is called a posterior probability, which is a measure of the
strength of the conclusion.
A prior probability is the probability that a hypothesis is true before consideration
of the evidence (e.g., the door is ajar). The less likely the hypothesis was
before the evidence, the less likely it should be after the evidence. Let us refer to
the hypothesis that my house has been burglarized as H. Suppose that I know
from police statistics that the probability of a house in my neighborhood being
burglarized on any particular day is 1 in 1,000.1 This probability is expressed as:
Prob(H) _ .001
‹
300 | Judgment and Decision Making
1 Although this makes for easy calculation, the actual number for Pittsburgh is closer to 1 burglary per
100,000 households per day.
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 300
This equation expresses the prior probability of the hypothesis, or the probability
that the hypothesis is true before the evidence is considered. The other prior
probability needed for the application of Bayes’s theorem is the probability that
the house has not been burglarized. This alternate hypothesis is denoted ~H.
This value is 1 minus Prob(H) and is expressed:
Prob(~H) _ .999
A conditional probability is the probability that a particular type of evidence
is true if a particular hypothesis is true. Let us consider what the conditional
probabilities of the evidence (door ajar) would be under the two hypotheses.
First, suppose I believe that the probability of the door’s being ajar is quite high
if I have been burglarized, for example, 4 out of 5. Let E denote the evidence, or
the event of the door being ajar. Then, we will denote this conditional probability
of E given that H is true as
Prob(E_H) _ .8
Second, we determine the probability of E if H is not true. Suppose I know that
chances are only 1 out of 100 that the door would be ajar if no burglary had taken
place (e.g., by accident, neighbors with a key).We denote this probability by
Prob(E|~H) _ .01
the probability of E given that H is not true.
The posterior probability is the probability that a hypothesis is true after
consideration of the evidence. The notation Prob(H|E) is the posterior probability
of hypothesis H given evidence E. According to Bayes’s theorem, we can
calculate the posterior probability of H, that the house has been burglarized
given the evidence, thus:
Prob(E|H) • Prob(H)
Bayes equation: Prob(H|E) _ ______________________________________
Prob(E|H) • Prob(H) _ Prob(E|~H) • Prob(~H)
Given our assumed values, we can solve for Prob(H|E) by substituting into the
preceding equation:
(.8) (.001)
Prob(H|E) _____________________ _ .074
(.8) (.001) _ (.01) (.999)
Thus, the probability that my house has been burglarized is still less than 8 in
100. Note that the posterior probability is this low even though an open door
is good evidence for a burglary and not for a
normal state of affairs: Prob(E|H) _ .8 versus
Prob(E|~H) _ .01. The posterior probability
is still quite low because the prior probability
of H—Prob(H) _ .001—was very low to begin
with. Relative to that low start, the posterior
probability of .074 is a considerable increase.
Table 11.1 offers an illustration of Bayes’s
theorem as applied to the burglary example.
It offers an analysis of 100,000 households,
Probabilistic Judgment | 301
TABLE 11.1
An Analysis of Bayes’s Theorem—100,000 Households
Burglarized Not Burglarized Sums
Door open 80 999 1,079
Door not open 20 98,901 98,921
Sums 100 99,900 100,000
Adapted from J. R. Hayes (1984).
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 301
assuming these statistics. There are four possible states of affairs, determined by
whether the burglary hypothesis is true or not and by whether there is evidence
of an open door or not. The frequency of each state of affairs is set forth in the
four cells of the table. Let’s consider the frequency in the upper-left cell, which
is the case I was worried about—the door is open and my house has been burglarized.
Because 1 in a 1,000 households are burglarized (Prob(H) is .001),
there should be 100 burglaries in the 100,000 households. This is the frequency
of both events in the left column. Because 8 times out of 10 the front door is
left open in a burglary (Prob(E|H) is .8), 80 of these 100 burglaries should leave
the door open—the number in the upper left. Similarly, in the upper-right cell,
we can calculate that of the 99,900 homes without burglary, the front door will
be left open 1 in 100 times, for 999 cases. Thus, in total there are 80_999_1,079
cases of front doors left open, and the probability of the house being burglarized
is 80/1,079 _ .074. The calculations in Bayes’s theorem perform the
same calculation as afforded by Table 11.1, but in terms of probabilities rather
than frequencies. As we will see, people find it easier to reason in terms of
frequencies.
Because Bayes’s theorem rests on a mathematical analysis of the nature of
probability, the formula can be proved to evaluate hypotheses correctly. Thus, it
enables us to precisely determine the posterior probability of a hypothesis given
the prior and conditional probabilities. The theorem serves as a prescriptive
model, or normative model, specifying the means of evaluating the probability
of a hypothesis. Such a model contrasts with a descriptive model, which specifies
what people actually do. People normally do not perform the calculations
that we have just gone through any more than they follow the steps prescribed
by formal logic. Nonetheless, they do hold various strengths of belief in assertions
such as “My house has been burglarized.” Moreover, their strength of
belief does vary with evidence such as whether the door has been found ajar.
The interesting question is whether the strength of their belief changes in accord
with Bayes’s theorem.
Bayes’s theorem specifies how to combine the prior probability of a hypothesis
with the conditional probabilities of the evidence to determine the posterior
probability of a hypothesis.
Base-Rate Neglect
Many people are surprised that the open door in the preceding example does
not provide as much evidence for a burglary as might have been expected. The
reason for the surprise is that they do not grasp the importance of the prior
probabilities. People sometimes ignore prior probabilities. In one demonstration,
Kahneman and Tversky (1973) told one group of participants that a person
had been chosen at random from a set of 100 people consisting of 70 engineers
and 30 lawyers. This group of participants was termed the engineer-high
group. A second group, the engineer-low group, was told that the person came
from a set of 30 engineers and 70 lawyers. Both groups were asked to determine
302 | Judgment and Decision Making
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 302
the probability that the person chosen at random from the group would be
an engineer, given no information about the person. Participants were able to
respond with the right prior probabilities: The engineer-high group estimated
.70 and the engineer-low group estimated .30. Then participants were told that
another person was chosen from the population and they were given the following
description:
Jack is a 45-year-old man. He is married and has four children. He is generally
conservative, careful, and ambitious. He shows no interest in political and
social issues and spends most of his free time on his many hobbies, which
include home carpentry, sailing, and mathematical puzzles.
Participants in both groups gave a .90 probability estimate to the hypothesis
that this person is an engineer. No difference was displayed between the two
groups, which had been given different prior probabilities for an engineer
hypothesis. But Bayes’s theorem prescribes that prior probability should have a
strong effect, resulting in a higher posterior probability from the engineer-high
group than from the engineer-low group.
In a second case, Kahneman and Tversky presented participants with the
following description:
Dick is a 30-year-old man. He is married with no children. A man of high
ability and high motivation, he promises to be quite successful in his field.
He is well liked by his colleagues.
This example was designed to provide no diagnostic information either way
with respect to Dick’s profession. According to Bayes’s theorem, the posterior
probability of the engineer hypothesis should be the same as the prior probability
because this description is not informative. However, both the engineer-high
and the engineer-low groups estimated that the probability was .50 that the
man described is an engineer. Thus, they allowed a completely uninformative
piece of information to change their probabilities. Once again, the participants
were shown to be completely unable to use prior probabilities in assessing the
posterior probability of a hypothesis.
The failure to take prior probabilities into account can lead people to make
some totally unwarranted conclusions. For instance, suppose you take a diagnostic
test for a cancer. Suppose also that this type of cancer, when present, results in
a positive test 95% of the time. On the other hand, if a person does not have the
cancer, the probability of a positive test result is only 5%. Suppose you are informed
that your result is positive. If you are like most people, you will assume
that your chances of dying of cancer are about 95 out of 100 (Hammerton,
1973). You would be overreacting in assuming that the cancer will be fatal, but
you would also be making a fundamental error in probability estimation.What
is the error?
You would have failed to consider the base rate (prior probability) for the
particular type of cancer in question. Suppose only 1 in 10,000 people have this
cancer. This percentage would be your prior probability. Now, with this information,
you would be able to determine the posterior probability of your
Probabilistic Judgment | 303
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 303
having the cancer. Bringing out the Bayesian formula, you would express the
problem in the following way:
Prob(H) • Prob(E|H)
Prob(H|E) _ _______________________________________
Prob(H) • Prob(E|H) _ Prob(~H) • Prob(E|~H)
where the prior probability of the cancer hypothesis is Prob(H) _ .0001, and
Prob(~H) _ .9999, Prob(E|H) _ .95, and Prob(E|~H) _ .05. Thus,
(.0001) (.95)
Prob(H|E) _ _______________________ _ .0019
(.0001) (.95) _ (.9999) (.05)
That is, the posterior probability of your having the cancer would still be less
than 1 in 500.
People often fail to take base rates into account in making probability
judgments.
Conservatism
The preceding examples show that people weigh the evidence too much and
ignore base rates. However, there are also situations in which people do not
weigh evidence enough, particularly as the evidence pointing to a conclusion
accumulates.Ward Edwards (1968) extensively investigated how people use new
information to adjust their estimates of the probabilities of various hypotheses.
In one experiment, he presented participants with two bags, each containing
100 poker chips. One of the bags contained 70 red chips and 30 blue; the other
contained 70 blue chips and 30 red. The experimenter chose one of the bags
at random and the participants’ task was to decide which bag had been chosen.
In the absence of any prior information, the probability of either bag having
been chosen was 50%. Thus,
Prob(HR) _ .50 and Prob(HB) _ .50
where HR is the hypothesis of a predominantly red bag and HB is the hypothesis
of a predominantly blue bag. To obtain further information, participants sampled
chips at random from the bag. Suppose the first chip drawn was red. The
conditional probability of a red chip drawn from each bag is
Prob(R|HR) _ .70 and Prob(R|HR) _ .30
Now, we can calculate the posterior probability of the bag’s being predominantly
red, given the red chip is drawn, by applying the Bayes equation to this situation:
Prob(R|HR) • Prob(HR)
Prob(R|HR) _ _______________________________________
Prob(R|HR) • Prob(HR) _ Prob(R|HR) • Prob(HR)
(.70) • (.50)
_ _____________________ _ .70
(.70) • (.50) _ (.30) • (.50)
304 | Judgment and Decision Making
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 304
This result seems, to both naive and sophisticated observers, to be a rather
sharp increase in probabilities. Typically, participants do not increase the probability
of a red-majority bag to .70; rather, they make a more conservative revision
to a value such as .60.
After this first drawing, the experiment continues: The poker chip is put back
in the bag and a second chip is drawn at random. Suppose this chip too is red.
Again, by applying Bayes’s theorem, we can show that the posterior probability
of a red bag is now .84. Suppose our observations continued for 10 more trials
and, after all 12 trials, we have observed eight reds and four blues. By continuing
the Bayesian analysis, we could show that the new posterior probability of the
hypothesis of a red bag is .97. Participants who see this sequence of 12 trials only
estimate subjectively a posterior probability of .75 or less for the red bag.
Edwards has used the term conservative to refer to the tendency to underestimate
full effect of available evidence. He estimates that we use between a fifth and a
half of the evidence available to us in situations like this experiment.
People frequently underestimate the cumulative force of evidence in making
probability judgments.
Correspondence to Bayes’s Theorem with Experience
All the preceding examples showed that participants can be quite far off in their
judgments of probability. One possibility is that participants really do not
understand probabilities or how to reason with respect to them. Certainly, it is
a rare participant in these experiments who could reproduce Bayes’s theorem,
let alone who would report engaging in Bayesian calculation. However, there is
evidence that, although participants cannot articulate the correct probabilities,
many aspects of their behavior are in accordance with Bayesian principles. To return
to the explicit-implicit distinction discussed in Chapter 7, people often seem
to display implicit knowledge of Bayesian principles even if they do not display
any explicit knowledge and make errors when asked to make explicit judgments.
Gluck and Bower (1988) performed an experiment that illustrates implicit
Bayesian behavior. Participants were given records of fictitious patients who
could display from one to four symptoms (bloody nose, stomach cramps, puffy
eyes, and discolored gums) and made discriminative diagnoses about which of
two hypothetical diseases the patients had. One of these diseases had a base rate
three times that of the other. Additionally, the conditional probabilities of displaying
the various symptoms, given the diseases, were varied. Participants were
not told directly about these base rates or conditional probabilities. They
merely looked at a series of 256 patient records, chose the disease they thought
the patient had, and were given feedback on the correctness of their judgments.
There are 15 possible combinations of one to four symptom patterns that
a patient might have. Gluck and Bower calculated the probability of each disease
for each pattern by using Bayes’s theorem and arranged it so that each
disease occurred with that probability when the symptoms were present. Thus,
the participants experienced the base probabilities and conditional probabilities
implicitly in terms of the frequencies of symptom–disease combinations.
Probabilistic Judgment | 305
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 305
Of interest is the probability with which they assigned
the rarer disease to various symptom combinations.
Gluck and Bower compared the participant
probabilities with the true Bayesian probabilities.
This correspondence is displayed by the scatterplot
in Figure 11.2. There we have, for each symptom
combination, the Bayesian probability (labeled objective
probability) and the proportion of times that
participants assigned the rare disease to that symptom
combination. As can be seen, these points fall
very close to a straight diagonal line with a slope of
1, which indicates that the proportion of the participants’
choices were very close to the true probabilities.
Thus, implicitly, the participants had become
quite good Bayesians in this experiment. The behavior
of choosing among alternatives in proportion to
their success is called probability matching.
After the experiment, Gluck and Bower presented
the participants with the four symptoms individually
and asked them how frequently the rare disease had
appeared with each symptom. This result is presented
in Figure 11.3 in a format similar to that of
Figure 11.2. As can be seen, participants showed
some neglect of the base rate, consistently having
overestimated the frequency of the rare disease. Still,
their judgments show some influence of base rate in
that their average estimated probability of the rare
disease is less than 50%.
Gigerenzer and Hoffrage (1995) showed that
base-rate neglect also decreases if events are stated
in terms of frequencies rather than in terms of
probabilities. Some of their participants were given
a description in terms of probabilities, such as the
one that follows:
The probability of breast cancer is 1% for
women at age 40 who participate in routine
screening. If a woman has breast cancer,
the probability is 80% that she will get a positive
mammography. If a woman does not have
breast cancer, the probability is 9.6% that she
also will get a positive mammography. A woman
in this age group had a positive mammography
in a routine screening. What is the probability
that she actually has breast cancer?
Fewer than 20 out of 100 (20%) of the participants
given such statements calculated the correct
Bayesian answer (which is about 8%). In the other
306 | Judgment and Decision Making
.2
.2
.4
.6
.8
1.0
.4 .6 .8 1.0
Proportion of choices by subjects
Objective probability
FIGURE 11.2 Participants’ proportion of choices corresponds
closely to the objective probabilities as determined by Bayes’s
theorem.
.2
.2
0
.4
.6
.8
1.0
.4 .6 .8 1.0
Estimated probability
True probability
FIGURE 11.3 Participants’ estimated probabilities systematically
overestimated the frequency of the rare disease, showing base-rate
neglect.
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 306
condition, participants were given descriptions in terms of frequencies, such as
the one that follows:
Ten out of every 1,000 women at age 40 who participate in routine screening
have breast cancer. Eight of every 10 women with breast cancer will get a
positive mammography. Ninety-five out of every 990 women without breast
cancer also will get a positive mammography. Here is a new representative
sample of women at age 40 who got a positive mammography in routine
screening. How many of these women do you expect to actually have breast
cancer?
Almost 50% of the participants given such statements calculated the correct
Bayesian answer. Gigerenzer and Hoffrage argued that we can reason better
with frequencies than with probabilities because we experience frequencies of
events, but not probabilities, in our daily lives.
There is also evidence that experience makes people more statistically
tuned. In a study of medical diagnosis,Weber, Böckenholt, Hilton, and Wallace
(1993) found that doctors were quite sensitive both to base rates and to the evidence
provided by the symptoms. Moreover, the more clinical experience the
doctors had, the more tuned were their judgments.
Although participants’ processing of abstract probabilities often does not
correspond with Bayes’s theorem, their behavior based on experience
often does.
Judgments of Probability
What are participants actually doing when they report probabilities of an event
such as the probability that someone who has bloody gums has a particular disease?
The evidence is that rather than thinking about probabilities,
they are thinking about relative frequencies. Thus they are
trying to judge the proportion of the patients that they saw with
bloody gums who had that particular disease. People are reasonably
accurate at making such proportionate judgments when they
do not have to rely on memory (Robinson, 1964; Shuford, 1961).
Consider an experiment by Shuford (1961). He presented arrays
such as that shown in Figure 11.4 to participants for 1 s. He then
asked participants to judge the proportion of vertical bars relative
to horizontal bars. The number of vertical bars varied from 10%
to 90% in different matrices. Shuford’s results are shown in Figure
11.5. As can be seen, participants’ estimates are quite close to
the true proportions.
The situation just described is one where the participants
can see the relevant information and make a judgment about
proportions.When participants cannot see the events and must
recall them from memory, they can give distorted judgments if
they recall too many of one kind from memory. A fair amount
of research has been done on the ways in which participants can
be biased in their estimation of the relative frequency of various
Probabilistic Judgment | 307
FIGURE 11.4 A random matrix presented to
participants to determine their accuracy in judging
proportions. The matrix is 90% vertical bars and
10% horizontal bars. (From Shuford, 1961. Copyright © 1961
by the American Psychological Association. Reprinted by permission.)
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 307
events in the population. Consider the following
experiment reported by Tversky and Kahneman
(1974), which demonstrates that judgments of
proportion can be biased by differential availability
of examples. These investigators asked participants
to judge the proportion of words in the language
that fit certain characteristics. For instance, they
asked participants to estimate the proportion of
English words that begin with the letter k versus
words with the letter k in the third position. How
might participants perform this task? One obvious
method is to briefly try to think of words that
satisfy the specification and words that do not and
to estimate the relative proportion of target words.
How many words can you think of that begin with
the letter k? How many words can you think of that
do not? What is your estimate of their proportion?
Now, how many words can you think of that have
the letter k in the third position? How many words
can you think of that do not? What is their relative
proportion? Participants estimated that more words begin with letter k than
have the letter k in the third position. In actual fact, three times as many words
have the letter k in the third position as begin with the letter k. Generally,
participants overestimate the frequency with which words begin with various
letters.
As in this experiment, many real-life circumstances require that we estimate
probabilities without having direct access to the population that these probabilities
describe. In such cases, we must rely on memory as the source for our estimates.
The memory factors that we studied in Chapters 6 and 7 serve to explain
how such estimates can be biased. Under the reasonable assumption that words
are more strongly associated with their first letter than with their third letter,
the bias exhibited in the experimental results can be explained by the spreadingactivation
theory (Chapter 6). With the focus of attention on the letter k, for
example, activation will spread from that letter to words beginning with it. This
process will tend to make words beginning with the letter k more available than
other words. Thus, these words will be overrepresented in the sample that participants
take from memory to estimate the true proportion in the population.
The same overestimation is not made for words with the letter k in the third
position because words are unlikely to be directly associated with the letters in
the third position. Therefore, these words cannot be associatively primed and
made more available.
Other factors besides memory lead to biases in probability estimates.
Consider another example from Tversky and Kahneman (1974).Which of the
following sequences of six tosses of a coin (where H denotes heads and T tails)
is more likely: H T H T T H or H H H H H H? Many people think the first
sequence is more probable, but both sequences are actually equally probable.
308 | Judgment and Decision Making
0
0
20
40
60
80
100
Horizontal
20
Porportion in display
Judged proportion
40 60 80 100
Vertical
FIGURE 11.5 Mean estimated
proportion as a function of the
true proportion. Participants
exhibited a fairly accurate ability
to estimate the proportions of
vertical and horizontal bars in
Figure 10.5. (From Shuford, 1961.
Copyright © 1961 by the American
Psychological Association. Reprinted
by permission.)
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 308
The probability of the first sequence is the probability of H on the first toss
(which is .50) times the probability of T on the second toss (which is .50),
times the probability of H on the third toss (which is .50), and so on. The
probability of the whole sequence is .50 * .50 * .50 * .50 * .50 * .50 _ .016.
Similarly, the probability of the second sequence is the product of the probabilities
of each coin toss, and the probability of a head on each coin toss is .50.
Thus, again, the final probability also is .50 * .50 * .50 * .50 * .50 * .50 _ .016.
Why do some people have the illusion that the first sequence is more probable?
It is because the first event seems similar to a lot of other events—for example,
H T H T H T or H T T H T H. These similar events serve to bias upward a person’s
probability estimate of the target event. On the other hand, H H H H H H,
straight heads, seems unlike any other event, and its probability will therefore
not be biased upward by other similar sequences. In conclusion, a person’s
estimate of the probability of an event will be biased by other events that are
similar to it.
A related phenomenon is what is called the gambler’s fallacy. The fallacy
is the belief that if an event has not occurred for a while, then it is more likely,
by the “law of averages,” to occur in the near future. This phenomenon can be
demonstrated in an experimental setting—for instance, one in which participants
see a sequence of coin tosses and must guess whether each toss will be a
head or a tail. If they see a string of heads, they become more and more likely to
guess that tails will come up on the next trial. Casino operators count on this
fallacy to help them make money. Players who have had a string of losses at a
table will keep playing, assuming that by the “law of averages” they will experience
a compensating string of wins. However, the game is set in favor of the
house. The dice do not know or care whether a gambler has had a string of
losses. The consequence is that players tend to lose more as they try to recoup
their losses. The “law of averages” is a fallacy.
The gambler’s fallacy can be used to advantage in certain situations—for
instance, at the racetrack. Most racetracks operate by a pari-mutuel system in
which the odds on a horse are determined by the number of people betting on
the horse. By the end of the day, if favorites have won all the races, people tend
to doubt that another favorite can win, and they switch their bets to the long
shots. As a consequence, the betting odds on the favorite deviate from what
they should be, and a person can sometimes make money by betting on the
favorite.
People can be biased in their estimates of probabilities when they must rely
on factors such as memory and similarity judgments.
The Adaptive Nature of the Recognition Heuristic
The examples in the previous section focused on cases where people came to
bad judgments relying on, for example, availability of events in memory.
Gigerenzer, Todd, and ABC Research Group (1999), in their book Simple
Heuristics That Make Us Smart, argue that such cases are the exception and not
Probabilistic Judgment | 309
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 309
the rule. They argue that people tend to identify the most valid cues for making
judgments and use these. For instance, through evolution people have
acquired a tendency to pay attention to availability of events in memory, which
is more often helpful than not.
Goldstein and Gigerenzer (1999, 2002) report studies of what they call the
recognition heuristic. This heuristic applies in cases where people can recognize
one thing and not another. This heuristic leads people to believe that the
recognized item is bigger and more important than the unrecognized item. In
one study, they looked at the ability of students at the University of Chicago
to judge the relative size of various German cities. For instance, which city is
larger—Bamberg or Heidelberg? Most of the students knew that Heidelberg is a
German city, but most do not recognize Bamberg—that is, one city is available
in memory and the other is not. Goldstein and Gigerenzer showed that when
faced with pairs like this, students almost always pick the city they can recognize.
One might think this shows another fallacy based on availability in memory.
However, Goldstein and Gigerenzer show that the students are actually
more accurate when they make their judgment for pairs of cities like this
(where they recognize one and not the other) than when they are given two
cities they can recognize and must use other bases for judging the size of the
cities (such as Munich versus Hamburg). This is because most American students
have little knowledge about the population of German cities. Thus, far
from a fallacy, this proves to be an effective basis for making judgments. Also,
American students do better at judging the relative size of German cities using
this heuristic than either American students do judging American cities or
German students do judging German cities, where this heuristic cannot be used
because almost all the cities are recognized.2 German students also do better
than American students in judging the population of American cities because
they can use the recognition heuristic and Americans cannot.
Figure 11.6 illustrates Goldstein and Gigerenzer’s explanation for why these
students were more accurate in judging the size of cities when they did not
know one. They looked at the frequency with which German cities were mentioned
in the Chicago Tribune and the frequency with which American cities
were mentioned in the German newspaper Die Zeit. It turns out that there is a
strong correlation between the actual size of the city and the frequency of mention
in these newspapers. Not surprisingly, people read about the larger cities in
other countries more frequently. Gigerenzer and Goldstein also show that there
is a strong correlation between the frequency of mention in the newspapers
(and the media more generally) and the probability that these students will
recognize the name. This is just the basic effect of frequency on memory. As a
consequence of these two strong correlations, there will be a strong correlation
between availability in memory and the actual size of the city.
310 | Judgment and Decision Making
2My German informant (Angela Brunstein) tells me that almost all Germans would recognize Bamberg and
Heidelberg, but many would be puzzled by which is larger. Interestingly, Google search on English texts
reports 37 million hits on Heidelberg and 3.5 million on Bamberg. Google search on German texts reports
30 million hits on Heidelberg and 12 million on Bamberg—a much closer ratio and many more hits
on Bamberg.
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 310
Goldstein and Gigerenzer argue that the recognition heuristic is useful in
many but not all domains. In some domains, researchers have shown that
people intelligently combine it with other information. For instance, Richter
and Späth (2006) had participants judge which of two animals has the larger
population size. For example, consider the following questions:
Are there more hainan partridges or arctic hares?
Are there more giant pandas or mottled umbers?
In the first case, most people have heard of the arctic hares and not hainan
partridges and would correctly choose arctic hares using the recognition
heuristic. In the second case, most people would recognize giant pandas and
not mottled umbers (a moth). Nonetheless, they also know giant pandas are
an endangered species and therefore correctly choose mottled umbers. This is
an example of how people can adaptively choose what aspects of information
to pay attention to.
People can use their ability to recognize an item, and combine this with other
information, to make good judgments.
Probabilistic Judgment | 311
Ecological correlation
Mediator
Recognition correlation
.66/.60
.72/.70
.86/.79
Criterion Recognition
Surrogate correlation
FIGURE 11.6 Ecological correlation (correlation between frequency of mention in newspapers
and population size), surrogate correlation (correlation between frequency of mention in
newspapers and probability of recognition), and recognition correlation (correlation between
probability of recognition and population size). The first value is for American cities and
the German newspaper Die Zeit as mediator, and the second value is for German cities and
the Chicago Tribune as mediator. (From Goldstein and Gigerenzer, 2002.)
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 311
312 | Judgment and Decision Making
•Decision Making
An extension of the research on probabilistic reasoning is research on decision
making, which is concerned with the way in which people make choices. Sometimes,
the choices that we have to make are easy. If we are offered the choice between
$400 and $1000, most of us would not have much difficulty in figuring
out which to accept. However, if we were faced with the choice of a certainty of
$400 but only a 50% chance of $1000, which would we select then? Something
like it might happen if we inherited a risky stock that we could cash in for $400
or that we could hold on to and see whether the company would take off or
fold. A great deal of research on decision making under uncertainty requires
participants to make choices among gambles. For instance, a participant might
be asked to choose between the following two gambles:
A. $8 with a probability of 1/3
B. $3 with a probability of 5/6
In some cases, participants are just asked for their opinions; in other cases, they
actually play the gamble that they choose. As an example of the latter possibility,
a participant might roll a die and win in case A if he gets a 5 or 6 and win in
case B if he gets a number other than 1.Which gamble would you choose?
As in the other domains of reasoning, decision making has its own standard
prescriptive theory for the way that people should behave in such situations
(von Neumann & Morgenstern, 1944). This theory says that they should choose
the alternative with highest expected value. The expected value of an alternative
is to be calculated by multiplying the probability by the value. Thus, the expected
value of alternative A is $8 _ 1_3 _ $2.67, whereas the expected value of
alternative B is $3 _ 5_6 _ $2.50. Thus, the normative theory says that participants
should select gamble A. However, most participants will select gamble B.
As a perhaps more extreme example of the same result, suppose you are
given a choice between
A. 1 million dollars with a probability of 1
B. 2.5 million dollars with a probability of 1/2
Maybe, in this case, you are on a game show and are offered a choice between
this great wealth with certainty or the opportunity to toss a coin and get even
more. I (and I assume you) would take the money (1 million)
and run, but in fact, if we do the utility calculations, we should
prefer the second choice because its utility is .5 _ 2.5 million _
1.25 million. Are we really behaving irrationally?
Most people, when asked to justify their behavior in such
situations, will argue that there comes a point when one has
enough money (if we could only convince CEOs of this notion!)
and that there really isn’t much difference for them
between 1 million dollars and 2.5 million dollars. This idea has
been formalized in the terms of what is referred to as subjective
utility—the value that we place on money is not linear with the
face value of the money. Figure 11.7 shows a typical function
proposed for the relation of subjective utility to money. It has
Value
Gains Losses
FIGURE 11.7 A function that
relates subjective value to
magnitude of gain and loss.
(From Kahneman & Tversky, 1984.
Reprinted by permission from the
American Psychological Association.)
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 312
a couple of properties. The first property is that it is curvilinear
such that it takes more than a doubling in the amount of money
to double its utility. Thus, in the preceding example, we may value
2.5 million only 20% more than 1 million. Let us say that the utility
of 1 million is U. The utility of 2.5 million can then be expressed
as 1.2U. The expected value of gamble A is 1 _ U _ U,
and the expected value of gamble B is 1_2 _ 1.2U _ .6U. Thus, in
terms of subjective utility, gamble A is more valuable and is to be
preferred.
The second property of this utility function is that it is steeper
in the loss region than in the gain region. Thus, participants given
the following choice of gambles
A: Gain $10 with 1/2 probability and lose $10 with 1/2 probability
B: Nothing with certainty
prefer B because they weight the loss of $10 more strongly than the gain of $10.
Kahneman and Tversky (1984) also argued that, as with subjective utility,
people associate a subjective probability with an event that is not identical
with the objective probability. They proposed the function in Figure 11.8 to
relate subjective probability to objective probability. According to this function,
very low probabilities are overweighted relative to high probabilities, producing
a bowing in the function. Thus, a participant might prefer a 1% chance of $400
to a 2% chance of $200 because 1% is not represented as half of 2%. Kahneman
and Tversky (1979) showed that a great deal of human decision making can be
explained by assuming that participants are responding in terms of these subjective
utilities and subjective probabilities.
An interesting question is whether the subjective functions in Figures 11.7
and 11.8 represent irrational tendencies. Generally, the utility function in Figure
11.7 is thought to be reasonable. As we get more money, getting even more
seems less and less important. Certainly, the amount of happiness that a billion
dollars can buy is not 1,000 times the amount of happiness that a million
dollars can buy. It should be noted that not all the utility functions of different
people are like that shown in Figure 11.7, which represents a sort of average.
One can imagine someone needing $10,000 for an important medical procedure.
Then, all sums less than $10,000 would be rather useless, and all sums
greater than $10,000 would be relatively equally good. Thus, such a person
might have a step in the utility function at $10,000.
There is less agreement about how we should assess the subjective probability
function in Figure 11.8. I (Anderson, 1990) have argued that it might actually
make sense to discount the extremity of low probabilities in the way that function
does. The argument is that, sometimes when we are told that probabilities
are extreme, we are being misinformed (see the third Question for Thought at
the end of the chapter). However, there is little consensus in the field about how
to evaluate the subjective probability function.
People make decisions under uncertainty in terms of subjective utilities and
subjective probabilities.
Decision Making | 313
Subjective probability
Objective probability
0
.5
.5
1.0
1.0
FIGURE 11.8 A function that
relates subjective probability
to objective probability.
(From Kahneman & Tversky, 1984.
Reprinted by permission from the
American Psychological Association.)
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 313
Framing Effects
Although one might view the functions in Figures 11.7 and 11.8 as reasonable,
there is evidence that they can lead people to do rather strange things. These
demonstrations deal with framing effects. These effects refer to the fact that
people’s decisions vary, depending on where they perceive themselves to be on
the utility curve in Figure 11.7. In an example from Kahneman and Tversky
(1984), someone must purchase a $15 item versus a $125 item. If another store
offers a $5 discount off the $15, the person is likely to make an effort to go to
the other store, whereas he is not likely to do so if the same $5 discount is
offered on the $125 item. However, in both cases, it is the same $5 savings, and
the question is simply whether one’s time is worth the $5. However, the two
contexts place the person on different points of the utility curve, which is negatively
accelerated. According to that curve, the difference between $15 and $10
is larger than the difference between $125 and $120. Thus, in the first case, the
saving seems worth it, but in the second case, it does not.
Another example has to do with betting behavior. Consider someone who
has lost $140 at the racetrack and has an opportunity to bet $10 on a horse that
will pay 15 to 1. The bettor can view this choice in one of two ways. In one way,
it becomes this choice:
A. Refuse the bet and accept a certainty of losing $140.
B. Make the bet and face a good chance of losing $150 and a poor chance of breaking
even.
Because the subjective difference between losing $140 and $150 is small, the
person will likely choose B and make the bet. On the other hand, the bettor
could view it as the following choice:
C. Refuse the bet and face the certainty of having nothing change.
D. Make the bet and face a good chance of losing an additional $10 and a poor chance
of gaining $140.
In this case, because of the greater weight on losses than on gains and because
of the negatively accelerated utility function, the bettor is likely to avoid the bet.
The only difference is whether one places oneself at the –$140 point or the 0
point on the curve in Figure 11.7. However, one gets a different evaluation of
the two outcomes, depending on where one places oneself.
As an example that appears to be more consequential, consider this situation
described by Kahneman and Tversky (1984):
Problem 1: Imagine that the U.S. is preparing for the outbreak of an unusual
Asian disease, which is expected to kill 600 people. Two alternative programs
to combat the disease have been proposed. Assume that the exact scientific
estimates of the consequences of the programs are as follows:
If Program A is adopted, 200 people will be saved.
If Program B is adopted, there is a one-third probability that 600 people will be saved
and a two-thirds probability that no people will be saved.
Which of the two programs would you favor?
Seventy-two percent of the participants preferred program A, which guarantees
lives, to dealing with the risk of program B. However, consider what
314 | Judgment and Decision Making
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 314
happens when, rather than describing the two programs in regard to saving
lives, the two programs are described as follows:
If Program C is adopted, 400 people will die.
If Program D is adopted, there is a one-third probability nobody will die and a twothirds
probability that 600 people will die.
With this description, only 22% preferred program C, which the reader will
recognize as equivalent to A (and D is equivalent to B). Both of these choices
can be understood in terms of a negatively accelerated utility function for
lives. In the first case, the subjective value of 600 lives saved is less than three
times the subjective value of 200 lives saved, whereas in the second case, the
subjective value of 400 deaths is more than two-thirds the subjective value of
600 deaths. McNeil, Pauker, Cox, and Tversky (1982) found that this tendency
extended to actual medical treatment. What treatment a doctor will choose
depends on whether the treatment is described in terms of odds of living or
odds of dying.
Situations in which framing effects are most prevalent tend to have one
thing in common—no clear basis for choice. This commonality is true of the
three examples that we have reviewed. In the case in which the shopper has an
opportunity for a saving, whether $5 is worth going to another store is unclear.
In the gambling example, there is no clear basis for making a decision.3 The
stakes are very high in the third case, but it is, unfortunately, one of those social
policy decisions that defy a clear analysis. Thus, these cases are hard to decide
on their merits alone.
Shafir (1993) suggested that, in such situations, we may make a decision not
on the basis of which decision is actually the best one but on the basis of which
will be easiest to justify (to ourselves or to others). Different framings make it
easier or harder to justify an action. In the disease example, the first framing
focuses one on saving lives and the second framing focuses one on avoiding
deaths. In the first case, one would justify the action by pointing to the people
whose lives have been saved (therefore it is critical that there be some people to
point to). In the second case, a justification would have to explain why people
died (and it would be better if there were no such people).
This need to justify one’s action can lead one to pick the same alternative
whether asked to pick something to accept or something to reject. Consider the
example in Table 11.2 in which two parents are described in a divorce case and
participants are asked to play the role of a judge who must decide to which parent
to award custody of the child. In the award condition, participants are asked
to decide who is to be awarded custody; in the deny condition, they are asked to
decide who is to be denied custody. The parents are overall rather equivalent, but
parent B has rather more extreme positive and negative factors. Asked to make
an award decision, more participants choose to award custody to parent B; asked
to make a deny decision, they tend to deny custody, again, to parent B. The
reason, Shafir argued, is that parent B offers reasons, such as a close relation with
Decision Making | 315
3 That is, there is no basis for making the gambling decision that would not have rejected gambling as
irrational in the first place.
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 315
the child, that can be used to justify the awarding of custody, but parent B also
has reasons, such as time away from home, to justify denying custody of the
child to that parent.
An interesting study in framing was performed by Greene, Sommerville,
Nystrom, Darley, and Cohen (2001). They compared ethical dilemmas such as
the following pair. In the first dilemma, a runaway trolley is headed for five
people who will be killed if it proceeds on its current course. The only way to
save them is to hit a switch that will turn the trolley onto an alternate set of
tracks where it will kill one person instead of five. The second dilemma is like the
first, except that you are standing next to a large stranger on a footbridge
that spans the tracks in between the oncoming trolley and the five people. In
this scenario, the only way to save the five people is to push the stranger off the
bridge onto the tracks below. He will die, but his body will stop the trolley from
reaching the others. In the first case, most people are willing to sacrifice one
person to save five, but in the second case, they are not.
In an fMRI study, Greene et al. compared the brain areas activated when
people considered an impersonal dilemma such as the first case, with the brain
areas activated when people considered a personal dilemma such as the second.
In the impersonal case, the regions of the parietal cortex that are associated
with cold calculation were active. On the other hand, when they judged the personal
case, regions of the brain associated with emotion (such as the ventromedial
prefrontal cortex that we discussed in the beginning of the chapter) were
active. Thus, part of what can be involved in the different framing of problems
seems which brain regions are engaged.
316 | Judgment and Decision Making
TABLE 11.2
Imagine that you serve on the jury of an only-child sole-custody case following a
relatively messy divorce. The facts of the case are complicated by ambiguous
economic, social, and emotional considerations, and you decide to base your
decision entirely on the following few observations.
(To which parent would you award sole custody of the child?/To which parent would
you deny sole custody of the child?)
Decisions
Award Deny
Parent A Average income 36% 45%
Average health
Average working hours
Reasonable rapport with the child
Relatively stable social life
Parent B Above-average income 64% 55%
Very close relation with the child
Extremely active social life
Lots of work-related travel
Minor health problems
Adapted from Shafir, 1993. Reprinted by permission from the Psychonomic Society, Inc.
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 316
When there is no clear basis for making a decision, people are influenced by
the way in which the problem is framed.
Neural Representation of Subjective Utility and Probability
The subjective utility of an outcome appears to be related to the activity of
dopamine neurons in the basal ganglia. The importance of this region to motivation
has been known since the 1950s, when Olds and Milner (1954) discovered
that rats would press a lever to the point of exhaustion to receive electrical
Decision Making | 317
risk. Reyna and Farley argue that adults don’t think
through the potential costs and benefits of a risky
behavior, but rather they simply
recognize the risk and avoid the situation—
just as the chess masters
discussed in Chapter 9 could recognize
the risk of a potential chess
position. In contrast, adolescents
often have to try to reason through
the consequences of a situation,
much as a chess duffer does, and
can make errors in reasoning.
2. Different values and situations. Risky behavior has
benefits such as immediate pleasure, and adolescents
value these benefits more. Adolescents are
particularly likely to weight the benefits of risky
behavior heavily in the context of their peers,
where social acceptance is at stake. Thus their
utilities in computing expected value are different.
Reyna and Farley speculate that this is related
to the fact that brain regions like the
ventromedial prefrontal cortex continue to mature
into the early 20s. Fischhoff also notes that risky
behavior often arises when adolescents attempt
to establish independence and personal competence,
which are important to achieve. However,
this can put adolescents in situations where older
adults seldom find themselves. If adults found
themselves in similar situations, they might find
themselves also acting in a more risky manner.
Implications
Why are adolescents more likely to make bad decisions?
One of society’s great concerns is risk taking in adolescents.
Compared to older adults, adolescents are more
likely to engage in risky sexual behavior,
abuse drugs and alcohol, and
drive recklessly. Such poor adolescent
choices are the leading cause
of death in adolescence and can
lead to a lifetime of suffering due to
such things as failed education, destroyed
personal relationships, and
addiction to cigarettes, alcohol, and
other drugs. This has been a subject
of a great deal of research (e.g., Fischhoff, 2008; Reyna &
Farley, 2006) and the results are a bit surprising. Contrary
to common belief, adolescents do not perceive themselves
to be any more invulnerable than older adults do
and often perceive greater danger from risky behavior than
do older adults. Also in many laboratory studies, late adolescents
often show as good or better performance as
older adults on abstract tasks of reasoning and decision
making (this will be discussed further in Chapter 14). Thus,
it does not appear that adolescents are poorer thinkers
about risk than older adults. Rather, it appears that the
explanation is involves two classes of factors:
1. Knowledge and experience. Adolescents lack some
of the information that adults have. For instance,
adolescents may know it is important to “practice
safe sex” but not know all that they should about
how to practice safe sex. Also, through experience
adults have become experts on reasoning about
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 317
stimulation from electrodes near this region. This stimulation caused release of
dopamine in a region of the basal ganglia called the nucleus accumbens. Drugs
like heroin and cocaine have their effect by producing increased levels of
dopamine from this region. These dopamine neurons show increased activity
for all sorts of positive rewards including basic rewards like food and sex, but
also social rewards like money or sports cars (Camerer, Loewenstein, & Prelec,
2005). Thus they appear to be the neural equivalent of subjective utility.
Although one can record activity in dopamine neurons while animals do
things like touch morsels of food, much of the knowledge of the human system
comes from neural imaging studies. In one fMRI study, Knutson, Taylor,
Kaufman, Peterson, and Glover (2005) presented participants with various uncertain
outcomes. For instance, on one trial participants might be told that they
had a 50% chance of winning $5; on another trial that they had a 50% chance
of winning $1. Knutson et al. imaged the activity while participants contemplated
the gamble and before they actually saw the results of the gamble. The
magnitude of the fMRI response in the nucleus accumbens reflected the differential
magnitude of these rewards. However, this region does not respond differently
to probability of reward. For instance, it did not respond differentially
when participants were told on one trial that they had an 80% probability of a
reward versus a 20% probability on another trial. In this case, the ventromedial
prefrontal cortex responded to probability of the reward although it had not responded
to the magnitude of the reward. Figure 11.9 illustrates the contrasting
response of these regions to reward magnitude and reward probability.
Although the Knutson et al. study found the ventromedial prefrontal region
only responding to probabilities, other research suggests it is involved in integrating
probabilities and utilities. The ventromedial region is that portion that
was destroyed in Phineas Gage (see Figure 11.1) and his problems went beyond
318 | Judgment and Decision Making
(a)
_0.20
0.20
NAcc
cue ant rsp
MPFC
**
*
0.15
0.10
0.05
0
_0.05
_0.10
_0.15
0 2 4 6 8 10 12 14
Seconds
% Signal change (SEM)
(b)
_0.20
0.20
0.15
0.10
0.05
0
_0.05
_0.10
_0.15
0 2 4 6 8 10 12 14
Seconds
% Signal change (SEM)
*
_$5.00/50%
_$1.00/50%
_$5.00/80%
_$5.00/20%
FIGURE 11.9 (A) The magnitude of a reward is represented in the activity of the nucleus
accumbens (A); (B) the probability of a reward is represented in the activity of the
ventromedial prefrontal cortex. (From Knutson et al., 2005.)
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 318
judging probabilities. Subsequent research has confirmed that people who have
damage to this region do have difficulty in responding adaptively in situations
where they experience good and bad outcomes with different probabilities. For
instance, this has been studied extensively in a task known as the Iowa gambling
task (Bechara, Damasio, Damasio, & Anderson, 1994; Bechara, Damasio,
Tranel, & Damasio, 2005), illustrated in Figure 11.10. The participants choose
cards from four decks. In this version of the problem, decks A and B are equivalent
and decks C and D are equivalent. Every time one selects from deck A
or B, the participant will gain $100 dollars but 1 time out of 10 will also lose
$1250 dollars. So, applying our formula for expected value, the expected value
of selecting a card from one of these decks is
$100 _ 0.1 _ $1250 __$25
or equivalently if participants play these decks 10 trials they can expect to lose
$250. Every time they select a card from decks C and D, they get only $50, but
they also only lose $250 on that 1 out of every 10 draws. The expected value of
selecting from one of these desks is
$50 _ 0.1 _ $250 __$25
and so choosing from these decks, participants can expect to make $250 every
10 trials. Players are initially attracted to decks A and B because of their higher
payoff, but normal participants eventually learn to avoid them. In contrast,
patients with ventromedial damage keep coming back to the high-paying decks.
Also, unlike normal participants, they do not show measures of emotional
engagement (such as increased galvanic skin response) when they choose from
these desks.
Decision Making | 319
“Bad” decks
A B C D
The Iowa gambling task
Gain per card $100
$1250
_$250
$100
$1250
_$250
$50
$250
_$250
$50
$250
_$250
Loss per 10 cards
Net per 10 cards
“Good” decks
FIGURE 11.10 A schematic diagram of the Iowa Gambling Task. The participants are given four
decks of cards, a loan of $2000 facsimile U.S. bills, and asked to play so as to win the most
money. Turning each card carries an immediate reward ($100 in decks A and B and $50 in
decks C and D). Unpredictably, however, the turning of some cards also carries a penalty
(which is large in decks A and B and small in decks C and D). Playing mostly from decks A
and B leads to an overall loss. Playing mostly from decks C and D leads to an overall gain.
(From Bechara et al., 2005.)
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 319
Dopamine activityin the nucleus accumbens reflects the magnitude of
reward, whereas the human ventromedial cortex is involved in integrating
probabilities with reward.
•Conclusions
Decision making deals with choosing actions that can have real consequences
in the presence of real uncertainty. All mammals have the dopamine system
that we just described, which gives them a basic ability to seek things that are
rewarding and avoid things that are harmful. However, humans, by virtue of
their greatly expanded prefrontal cortex, have the capacity to reflect on their
circumstances and select actions other than what their more primitive systems
might urge. Research suggests that the ventromedial portion of the human prefrontal
cortex, which is greatly expanded in size even in comparison to the
genetically similar apes, might play a particularly important role in such regulation.
Humans attempt acts of self-regulation—for example, diet plans—that are
far beyond the reach of any other species. However, we live in an uncertain
world, as witnessed by all the contradictory claims made for various diet plans.
Perhaps if we understood better how people responded to such uncertainty and
contradiction, we would also be in a better position to understand why there
are so many failures of our good resolutions.
320 | Judgment and Decision Making
1. Consider the Monte Hall problem:
Suppose you’re on a game show, and you’re given the choice of
three doors: Behind one door is a car; behind the others, goats.
You pick a door—for example, door 1—and the host, who
knows what’s behind the doors, opens another door—for example,
door 3—that has a goat. He then says to you, “Do you
want to pick door 2?” Is it to your advantage to switch your
choice? (Whitaker, 1990, p. 16)
This can be analyzed using the following form of
Bayes’s theorem:
P(H2)P(E3|H2)
P(H2|E3) _ ___________________________________________
P(H1)P(E3|H1) _ P(H2)P(E3|H2)+P(H3)P(E3|H3)
where P(H2|E3) is the probability that the car is behind
door 2 given that the host has opened door 3. P(H1),
P(H2), and P(H3) are the prior probabilities that the
car is behind each door and all three are 1_3. P(E3|H1),
P(E3|H2), and P(E3|H3) are the conditional probabilities
that the host opens each door given each hypothesis.
In calculating these probabilities, keep in mind that
the host cannot open the door you chose and must
open a door that has a goat.
2. Conservatism and Bayes rate neglect seem to be in conflict
(Fischhoff & Beyth-Marom, 1983; Gigerenzer et al.,
1989). Conservatism says that people pay too little
attention to data, whereas Bayes rate neglect says they
only pay attention to evidence and ignore base rates.
Could the contradiction be explained by differences
between studies like Edwards’s that show conservatism
and those like Kahneman and Tversky’s that demonstrate
base-rate neglect?
3. Consult the Web site
http://www.rense.com/general81/dw.htm for a list of
things that people said would never happen.What does
this imply about what our subjective probability should
be when someone informs us that the objective
probability is 0?
4. In the 1980s, it used to be recommended that a
pregnant woman 35 years or older be tested to find
Questions for Thought
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 320
Key Terms | 321
out whether the fetus had Down syndrome. The logic
behind this recommendation was that the probability
of having a Down syndrome baby increases with age
and is about 1_250 for when the expectant mother is
age 35, whereas the probability of the procedure
resulting in a miscarriage was also 1_250. Analyze the
assumptions behind this decision-making criterion
used in the 1980s in terms of the expected-value
calculations described in this chapter. Do you agree
with the recommendation?
Key Terms
Bayes’s theorem
conditional probability
descriptive model
framing effects
gambler’s fallacy
posterior probability
prescriptive model
prior probability
probability matching
recognition heuristic
subjective probability
subjective utility
ventromedial prefrontal
cortex
Anderson7e_Chapter_11.qxd 8/20/09 9:52 AM Page 321
322