Cognitive Psychology and Its Implications, Ch. 3
3Attention and Performance
Chapter 2 described how the human visual system and other perceptual systems
simultaneously process information from all over their sensory fields. However, there
are limits on how much we can do in parallel. There are points at which we can attend
to only one spoken message or one visual object at a time. This chapter addresses how
higher level cognition determines what to attend to
In this chapter, we will address these questions: • In a busy world filled with sounds, how do we select what to listen to? • How do we find meaningful information within a complex visual scene? • What role does attention play in putting visual patterns together as
recognizable objects? • How do we coordinate parallel activities like driving a car and holding
a conversation?
•Serial Bottlenecks
Psychologists have proposed that there are serial bottlenecks in human information
processing, points at which it is no longer possible to continue processing
everything in parallel. For example, it is generally accepted that there are limits to
parallelism in the motor systems. Although most of us can perform separate
actions simultaneously when the actions involve different motor systems (such as
walking and chewing gum), we have difficulty in getting one motor system to do
two things at once. Thus, even though we have two hands, we have only one
system for moving our hands, so it is hard to get our two hands to move in different
ways at the same time. Think of the familiar problem of trying to pat your
head while rubbing your stomach. It is hard to prevent one of the movements
from dominating—one tends to wind up rubbing one’s head or patting one’s
63
Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 63
64 | Attention and Performance
stomach.1 The many human motor systems—one for moving feet, one for moving
hands, one for moving eyes, and so on—can and do work independently and
separately, but it is difficult to get any one of these systems to do two things at
the same time.
One question that has occupied psychologists is how early do the bottlenecks
occur: before we perceive the stimulus, after we perceive the stimulus but before
we think about it, or only just before motor action is required? Common sense
suggests that there must be a certain amount of serial thinking before motor
action can occur. For instance, we find it basically impossible to add two digits
and multiply them simultaneously. Still, there remains the question of just where
the bottlenecks in information processing lie. Various theories about when they
happen are referred to as early-selection theories or late-selection theories,
depending on when they propose that bottlenecks take place. This question is one
of those examined by psychologists who are interested in studying attention.
Wherever there is a bottleneck, our cognitive processes must select which pieces
of information to attend to and which to ignore.A related research question concerns
how we select what to attend to.
In addition to the question of where information selection takes place in the
information-processing stream, researchers have asked the question of what factors
determine what we attend to. A major distinction is made
between goal-directed factors (sometimes called endogenous
control) and stimulus-driven factors (sometimes
called exogenous control). To illustrate the distinction,
Corbetta and Shulman (2002) ask us to imagine ourselves
at the El Prado Museum in Madrid, looking at Bosch’s
painting The Garden of Earthly Delights (see Color
Plate 3.1). Initially, our eyes will be probably be drawn
to large, salient objects like the instrument in the center
of picture. This would be an instance of stimulus-driven
attention—it is not that we wanted to attend to this; it
just grabbed our attention. However, our guide may
start to comment on a “small animal playing a musical
instrument.” Now we have a goal and will direct our
attention over the picture to find the object being described.
Continuing their story, Corbetta and Shulman
ask us to imagine that we hear an alarm system starting to
ring in the next room. Now a stimulus-driven factor has
intervened, and our attention will be drawn away fromour
goal of understanding our guide and the picture and switch
to the adjacent room. Corbetta and Shulman argue that somewhat different brain
systems control goal-directed attention versus stimulus-driven attention. For instance,
neural imaging evidence suggests that the goal-directed attentional system is
more left lateralized, whereas the stimulus-driven system is more right lateralized.
The brain regions that select information to process can be distinguished (to
an approximation) from those that process the information selected. Figure 3.1
1 Drummers (including my son) are particularly good at doing this—I definitely am not a drummer. This
suggests that the real problem might be motor timing.
Brain Structures
Parietal cortex: attends
to locations and objects
Motor cortex:
controls hands
Dorsolateral prefrontal
cortex: directs central
cognition
Auditory cortex:
processes auditory
information
Extrastriate cortex:
processes visual
information
Anterior cingulate:
(midline structure)
monitors conflict
FIGURE 3.1 A representation of
some of the brain areas involved
in attention and some to the
perceptual and motor regions
they control. The parietal regions
are particularly important in
directing perceptual resources.
The prefrontal regions (dorsolateral
prefrontal cortex, anterior
cingulate) are particularly
important in executive control.
Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 64
Auditory Attention | 65
highlights the parietal cortex, which influences information processing in regions
such as the visual cortex and auditory cortex. It also highlights prefrontal
regions that influence processing in the motor area and more posterior regions.
These prefrontal regions include the dorsolateral prefrontal cortex and, well
below the surface, the anterior cingulate cortex. The regions that do the selecting
and direct the processing are responsible for attentional control. As this chapter
proceeds, it will elaborate on the research involving the various brain regions in
Figure 3.1.
Attentional systems select information to process at serial bottlenecks where
it is no longer possible to do things in parallel.
•Auditory Attention
Some of the early research on attention was concerned with auditory attention.
Much of this research centered on the dichotic listening task. In a typical dichotic
listening experiment, illustrated in Figure 3.2, participants wear a set of
headphones. They hear two messages at the same time, one entering each ear,
and are asked to “shadow” one of the two messages (i.e., repeat back the words
from one message only). Most participants are able to attend to one message
and tune out the other.
Psychologists (e.g., Cherry, 1953; Moray, 1959) have discovered that very
little information about the unattended message is processed in a dichotic
listening task. After hearing the messages, participants report that they can
tell whether the unattended message was a human voice or a noise; whether
... and then John turned rapidly toward ...
ran house ox cat
and, um, John turned . . .
FIGURE 3.2 A typical dichotic listening task. Different messages are presented to the left and
right ears, and the participant attempts to “shadow” the message entering one ear. (From Lindsay &
Norman, 1977. Reprinted by permission of the publisher. © 1977 by Academic Press.)
Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 65
66 | Attention and Performance
the human voice was male or female; and whether the sex of the speaker
changed during the test. This limited information is often all that they can
report. They cannot tell what language was spoken or remember any of the
words, even if the same word was repeated over and over again. An analogy is
often made between performing this task and being at a cocktail party, where a
guest tunes in to one message (a conversation) and filters out others. This is an
example of goal-directed processing—the listener selects the message to be
processed. However, to return to the distinction between goal-directed and
stimulus-driven processing, important stimulus information can disrupt our
goals.We have probably all experienced the situation in which we are listening
intently to one person and hear our name mentioned by someone else. It is
very hard in this situation to keep your attention on what the original speaker
is saying.
The Filter Theory
Broadbent (1958) proposed a particular early-selection theory called the filter
theory to account for these results. His basic assumption was that sensory
information comes through the system until some bottleneck is reached. At that
point, a person chooses which message to process on the basis of some physical
characteristic. The person is said to filter out the other information. In a
dichotic listening task, it was assumed that the message to each ear was registered
but that at some point the participant selected one ear to listen with. At
a cocktail party, we pick which speaker to follow on the
basis of physical characteristics, such as the pitch of the
speaker’s voice.
A crucial feature of Broadbent’s original filter model is
its proposal that we select a message to process on the basis
of physical characteristics such as ear or pitch. This hypothesis
made a certain amount of neurophysiological
sense. Messages entering each ear arrive on different
nerves. Nerves also vary in which frequencies they carry
from each ear. Thus, we might imagine that the brain, in
some way, selects certain nerves to “pay attention to.”
People can certainly choose to attend to a message on
the basis of its physical characteristics. There is evidence,
however, that we can also select messages to process on the
basis of their semantic content. In one study, Gray and
Wedderburn (1960), who at the time were undergraduate
students at Oxford University, demonstrated that participants
were successful in following a message that jumped
back and forth between the ears. Figure 3.3 illustrates the participants’ task
in their experiment. Suppose part of the meaningful message that participants
were to shadow was dogs scratch fleas. The message to one ear might be
dogs six fleas, whereas the message to the other might be eight scratch two.
Instructed to shadow the meaningful message, participants would report
dogs scratch fleas. Thus, participants can shadow a message on the basis of
dogs six fleas . . .
. . . eight scratch two
dogs scratch fleas . . .
FIGURE 3.3 An illustration
of the shadowing task in the
Gray and Wedderburn (1960)
experiment. The participant follows
the meaningful message as
it moves from ear to ear. (After
Klatzky, 1975. Adapted by permission of
the publisher. © 1975 by W. H. Freeman.)
Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 66
Auditory Attention | 67
meaning rather than on the basis of what each ear
physically hears.
Treisman (1960) looked at a situation in which
participants were instructed to shadow a particular
ear (Figure 3.4). The message in the ear to be
shadowed was meaningful up to a certain point;
then it turned into a random sequence of words.
Simultaneously, the meaningful message switched
to the other ear—the one to which the participant
had not been attending. Some participants switched
ears, against instructions, and continued to follow
the meaningful message. Others continued to follow
the shadowed ear. Thus, it seems that sometimes
people use the physical ear to select which message
to follow, and sometimes they choose semantic
content.
Broadbent’s filter model proposes that we use physical features, such as ear or
pitch, to select one message to process, but it has been shown that people can
also use the meaning of the message as the basis for selection.
The Attenuation Theory and the Late-Selection Theory
To account for these kinds of results, Treisman (1964) proposed a modification
of the Broadbent model that has come to be known as the attenuation theory.
This model hypothesized that certain messages would be weakened but not
filtered out entirely on the basis of their physical properties. Thus, in a dichotic
listening task, participants would minimize the signal from the unattended ear
but not eliminate it. Semantic selection criteria could apply to all messages,
whether they were attenuated or not. If the message were attenuated, it would
be harder to apply these selection criteria, but it would still be possible, as in the
Gray and Wedderburn (1960) experiment. Treisman (personal communication,
1978) emphasized that in her 1960 experiment, most participants actually
continued to shadow the prescribed ear. It was easier to follow the message that
is not being attenuated than to apply semantic criteria to switch attention to the
attenuated message.
An alternative explanation was offered by J. A. Deutsch and D. Deutsch
(1963) in their late-selection theory. They proposed that all the information is
processed completely without attentuation. Their hypothesis was that the
capacity limitation is in the response system, not the perceptual system. They
claimed that people can perceive multiple messages but that they can shadow
only one message at a time. Thus, people need some basis for selecting which
message to shadow. If they use meaning as the criterion (either according to or
in contradiction to instructions), they will switch ears to follow the message. If
they use the ear of origin in deciding what to attend to, they will shadow the
chosen ear.
I SAW THE GIRL/Song was wishing . . .
The to-be-shadowed ear
I SAW THE GIRL JUMPING . . .
. . . me that bird
JUMPING IN THE STREET.
FIGURE 3.4 An illustration of
the Treisman (1960) experiment.
The meaningful message moves
to the other ear, and the
participant sometimes
continues to shadow it against
instructions. (After Klatzky, 1975.
Adapted by permission of the publisher.
© 1975 by W. H. Freeman.)
Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 67
The difference between the two theories is illustrated in Figure 3.5. Both
models assume that there is some filter or bottleneck in processing. Treisman’s
theory (Figure 3.5a) assumes that the filter selects which message to attend to,
whereas Deutsch and Deutsch’s theory (Figure 3.5b) assumes that the filter
occurs after the perceptual stimulus has been analyzed for verbal content.
Treisman and Geffen (1967) tried to address the difference between these two
theories. They used a dichotic listening task in which participants had to
shadow one message and also had to process both messages for a target word. If
they heard the target word, they were to signal by tapping. According to the
Deutsch and Deutsch late-selection theory, messages from both ears would get
through and participants should have been able to detect the critical word
equally well in either ear. In contrast, the attenuation theory predicted much
less detection in the unshadowed ear because the message would be attenuated.
In the experiment, participants detected 87% of the target words in the shadowed
ear and only 8% percent in the unshadowed ear. Other evidence consistent
with the attenuation theory was reported by A. M. Treisman and Riley
(1969) and by Johnston and Heinz (1978).
There is neural evidence for a version of the attenuation theory that asserts
that there is both enhancement of the signal coming from the attended ear and
attenuation of the signal coming from the unattended ear. The primary auditory
area of the cortex (see Figure 3.1) shows an enhanced response to auditory signals
coming from the ear the listener is attending to and a decreased response to
signals coming from the other ear. Through ERP recording,Woldorff et al. (1993)
showed that these responses occur between 20 and 50 ms after stimulus onset.
68 | Attention and Performance
FIGURE 3.5 Treisman and
Geffen’s illustration of attentional
limitations produced by
(a) Treisman’s (1964) attenuation
theory and (b) Deutsch and
Deutsch’s (1963) late-selection
theory. (From Treisman & Geffen, 1967.
Reprinted by permission of the publisher.
© 1967 by the Quarterly Journal of
Experimental Psychology.)
(a) (b)
Responses
Selection and
organization of
responses
Analysis of
verbal content
Perceptual
filter
Input messages
1 2
1 2
Responses
Selection and
organization of
responses
Analysis of
verbal content
Input messages
Response filter
1 2
1 2
Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 68
Visual Attention | 69
The enhanced responses occurmuch sooner in auditory processing than the point
at which the meaning of the message can be identified. There is also evidence for
enhancement of the message in the auditory cortex on the basis of features other
than location. For instance, Zatorre, Mondor, and Evans (1999) found in a PET
study that when people attend to a message on the basis of pitch, there is similar
enhancement (registered as increased activation) in the auditory cortex. This
study also found increased activation in the parietal areas that direct attention.
Although auditory attention can enhance processing in the primary auditory
cortex, there is no evidence of reliable effects of attention on earlier portions
of auditory processing, such as in the auditory nerve or in brain-stem
processing (Picton & Hillyard, 1974). The various results we have reviewed
suggest that the primary auditory cortex is the earliest area to be influenced by
attention. It should be stressed that the effects at the auditory cortex are a
matter of attenuation and enhancement. Messages are not completely filtered
out and so it is still possible to select them at later points of processing.
Attention can enhance or reduce the magnitude of response to an auditory
signal in the primary auditory cortex.
•Visual Attention
The bottleneck in visual information processing is even more apparent than the
one in auditory information processing. As we saw in Chapter 2, the retina
varies in acuity, with the greatest acuity in a very small area called the fovea.
Although the human eye registers a large part of the visual field, the fovea registers
only a small fraction of that field. Thus, in choosing where to focus our
vision, we also choose to devote most of our visual processing resources to a
particular part of the visual field, and we attenuate the resources allocated to
processing other parts of the field. Usually, we are attending to that part of the
visual field on which we are focusing. For instance, as we read, we move our
eyes so that we are fixating the words we are attending to.
The focus of visual attention is not always identical with the part of the
visual field being processed by the fovea, however. People can be instructed to
fixate on one part of the visual field (making that part the focus of the fovea)
and attend to another, nonfoveal region of the visual field.2 In one experiment,
Posner, Nissen, and Ogden (1978) had participants focus on a constant point
and then presented them with a stimulus 7° to the left or the right of the
fixation point. In some trials, participants were told on which side the stimulus
was likely to occur; in other trials, there was no such warning.When there was a
warning, it was correct 80% of the time—but 20% of the time the stimulus
appeared on the unexpected side. The researchers monitored eye movements
and included only those trials in which the eyes had stayed on the fixation
2 This is what quarterbacks are supposed to do when they pass the football, so that they don’t “give away” the
position of the intended receiver.
Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 69
point. Figure 3.6 shows the time required to judge the
stimulus if it appeared in the expected location (80% of
the time), if the participant had not been given a cue
(50% of the time), and if it appeared in the unexpected
location (20% of the time). Participants were able to
shift their attention from where their eyes were fixated:
Their responses to the stimuli were faster when the stimulus
appeared in the expected location and slower when
it appeared in the unexpected location.
Posner, Snyder, and Davidson (1980) found that people
can attend to regions of the visual field as far as 24° from
the fovea. Although visual attention can be moved without
accompanying eye movements, people usually do move
their eyes, so that the fovea processes the portion of the
visual field to which they are attending. Posner (1988)
pointed out that successful control of eye movements
requires us to attend to places outside the fovea. That is, we
must attend to and identify an interesting nonfoveal region
so that we can guide our eyes to fixate on that region to
achieve the greatest acuity in processing it. Thus, a shift of
attention often precedes the corresponding eye movement.
To process a complex visual scene, we must move our attention around in
the visual field to track the visual information. This process is like shadowing a
conversation. Neisser and Becklen (1975) performed the visual analog of the
auditory shadowing task. They had participants observe two videotapes superimposed
over each other. One was of two people playing a hand-slapping game,
the other of some people playing a basketball game. Figure 3.7 shows how the
situation appeared to the participants. They were instructed to pay attention
to one of the two films and to watch for odd events such as the two players in
the hand-slapping game pausing and shaking hands. Participants were able to
monitor one film successfully and reported filtering out the other.When asked
to monitor both films for odd events, the participants experienced great difficulty
and missed many of the critical events.
70 | Attention and Performance
Reaction time (ms)
Unexpected No expectation Expected
Condition
320
300
280
260
240
220
FIGURE 3.6 The results of an
experiment to determine how
people react to a stimulus that
occurs 7° to the left or right of
the fixation point. The graph
shows participants’ reaction
times to expected, unexpected,
and neutral (no expectation)
signals. (From Posner et al., 1978.
Reprinted by permission of the publisher.
© 1978 by Erlbaum.)
(a) (b) (c)
FIGURE 3.7 Frames from the two films used by Neisser and Becklen in their visual analog of the
auditory shadowing task. (a) The “hand-game” film; (b) the basketball film; and (c) the two figures
superimposed. (From Neisser & Becklen, 1975. Reprinted by permission of the publisher. © 1975 by Academic Press.)
Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 70
Visual Attention | 71
As Neisser and Becklen (1975) noted, this situation
involved an interesting combination of the use
of physical cues and the use of content cues. Participants
moved their eyes and focused their attention
in such a way that the critical aspects of the monitored
event fell on their fovea and the center of their
attentive spotlight. On the other hand, the only way
they knew how to move their eyes to follow an event
was by making reference to the content of the event
they were processing. Thus, physical cues facilitated
their processing of the critical film, which in turn
facilitated extracting content so they would know
where to move their eyes.
Figure 3.8 shows examples of the overlapping
stimuli used in an experiment by O’Craven,
Downing, and Kanwisher (1999) to study the neural
consequences of attending to one object or the other.
Participants in their experiment saw a series of pictures
that consisted of faces superimposed on houses.
They were instructed to either look for repetition of
the same face in the series or repetition of the same house. Recall from Chapter 2
that there is a region of the fusiformgyrus, the fusiformface area, which becomes
active when people are observing faces. There is another area within the temporal
cortex, the parahippocampal place area, which becomes more active when people
are observing places.What is special about these pictures is that they consisted of
both places and locations.Which region would become active—the fusiform face
area or the parahippocampal place area? As the reader might suspect, the answer
depended on what the participant was attending to.When participants were looking
for repetition of faces, the fusiform face area became more active; when they
were looking for repetition of places, the parahippocampal place area became
more active. Attention was able to select which region of the temporal cortex was
engaged in the processing of the stimulus.
People can focus their attention on parts of the visual field and move their
focus of attention to process what they are interested in
The Neural Basis of Visual Attention
It appears that the neural mechanisms underlying visual attention are very
similar to those underlying auditory attention. Just as auditory attention directed
to one ear enhances the cortical signal from that ear, visual attention directed to
a spatial location appears to enhance the cortical signal from that location. If
a person attends to a particular spatial location, a distinct neural response
(detected using ERP records) in the visual cortex occurs within 70 to 90 ms after
the onset of the stimulus. On the other hand, when a person is attending to a
particular object (attending to chairs and not tables, say), rather than to a particular
location in space, we do not see a response for more than 200 ms. Thus, it
FIGURE 3.8 An example of a
picture used in the study of
O’Craven et al. (1999). When
the face is attended, there is
activation in the fusiform face
area, and when the house is
attended, there is activation
in the parahippocampal
place area.
Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 71
appears to take more effort to direct visual attention on the basis of content than
on the basis of physical features, just as is the case with auditory attention.
Mangun, Hillyard, and Luck (1993) had participants fixate on the center of a
computer screen, then judge the lengths of bars presented in positions different
from the fixation location (upper left, lower left, upper right, and lower right).
They documented how much an ERP recording was affected at various locations
across the back of the scalp. Figure 3.9 shows the distribution of scalp activity
when a participant was attending to one of the four different regions of the visual
array (while fixating on the center of the screen). Consistent with the topographic
organization of the visual cortex, there was greatest activity over the side
of the scalp opposite the side of the visual field where the object appeared. Recall
from Chapters 1 and 2 (see Figure 2.6) that the visual cortex (at the back of the
head) is topographically organized, with each visual field (left or right) represented
in the opposite hemisphere. Thus, it appears that there is enhanced neural
processing in the portion of the visual
cortex corresponding to the location of
visual attention.
A study by Roelfsema, Lamme, and
Spekrejse (1998) illustrates the impact of
visual attention of information processing
in the primary visual area of the
macaque monkey. They trained monkeys
to perform the rather complex task
illustrated in Figure 3.10. A trial in the
experiment would begin with the monkey
fixating on a particular stimulus in
the visual field, as in part (a) of the
figure. Then two curves would appear, as
72 | Attention and Performance
P1 attention effect
(current density)
P1 P1 P1 P1
Stimulus
+ + + +
FIGURE 3.9 Results from an experiment by Mangun, Hillyard, and Luck. Distribution of scalp
activity recorded by ERP when a participant was attending to one of the four different regions
of the visual array depicted in the right-hand column while fixating on the center of the screen.
The greatest activity was recorded over the side of the scalp opposite the side of the visual
field where the object appeared, confirming that there is enhanced neural processing in portions
of the visual cortex corresponding to the location of visual attention. (After Mangun et al., 1993.
Adapted by permission of the publisher. © 1993 by MIT Press.)
(a)
Fixation (300 ms)
(b)
Stimulus (600 ms)
Receptive field
(c)
Saccade
Fixation point
FIGURE 3.10 The experimental
procedure in Roelfsema et al.
(1998): (a) The monkey fixates
the start point. (b) Two curves
are presented, one of which
links the start point to a target
point. (c) The monkey saccades
to the target point. The box
represents the receptive field in
the primary visual cortex V1.
Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 72
Visual Attention | 73
TWLN
XJBU
UDXI
HSFP
XSCQ
SDJU
PODC
ZVBP
PEVZ
SLRA
JCEN
ZLRD
XBOD
PHMU
ZHFK
PNJW
CQXT
GHNR
IXYD
QSVB
GUCH
OWBN
BVQN
FOAS
ITZN
in part (b), only one of which connected the original object to the fixation
point. The monkey had to remain fixated on this point for 600 ms and then
perform a saccade (an eye movement) to the end of the curve that connected
the original point (part c).While monkeys performed this task, Roelfsema et al.
recorded from cells in the monkey’s primary visual cortex (where cells with
receptive fields like those in Figure 2.8 are found). Indicated by the square in
Figure 3.10 is a receptive field of one of these cells. It shows increased response
when a line falls on that part of the visual field and so responds when the curve
appears that crosses it. The cell responds more during the 600-ms waiting
period when the receptive is on the curve that connects the fixation point to the
destination than when it is on the other curve. The monkey must shift its
attention along this curve to determine where to saccade to, and this shift of
attention causes the line detectors in V1 to respond more strongly.
When people attend to a particular spatial location, there is greater neural
processing in portions of the visual cortex corresponding to that location.
Visual Search
People are able to select stimuli to attend to, either in the visual or auditory
domain, on the basis of physical properties and, in particular, on the basis of
location. Although selection based on simple features can occur early and quickly
in the visual system, not everything people look for can be defined in terms of
simple features. How do they find an object with particular higher order properties,
such as the face of a friend in a crowd? In such cases, it seems that they must
search through the faces in the crowd, looking for one that has the desired
properties. Much of the research on visual attention has focused on how people
perform such searches. Rather than study how people find faces in a crowd, however,
researchers have tended to use simpler material.
Figure 3.11, for instance, shows a portion of the display
that Neisser (1964) used in one of the early studies. Try
to find the first K in the set of letters displayed.
Presumably, you tried to find the K by going
through the letters row by row, looking for the target.
Figure 3.12 graphs the average time it took participants
in Neisser’s experiment to find the letter as a
function of which row it appeared in. The slope of
the best-fitting function in the graph is about 0.6,
which implies that participants took about 0.6 s to
scan each line. When people engage in such searches,
they appear to be allocating their attention intensely
to the search process. For instance, brain-imaging
experiments have found strong activation in the
parietal cortex during such searches (see Kanwisher &
Wojciulik, 2000, for a review).
Although a search can be intense and difficult, it is
not always that way. Sometimes we can find what we
are looking for without much effort. If we know that
0
0
10
20
30
40
10 20 30 40 50
Time (s)
Position of critical item (line number)
FIGURE 3.12 The time required to find a target letter in the
array shown in Figure 3.9 as a function of the line number in
which it appears. (After Neisser, 1964. Adapted by permission of the publisher.
© 1964 by Scientific American.)
FIGURE 3.11 A representation
of lines 7–31 of the letter array
used in Neisser’s search experiment.
(After Neisser, 1964. Adapted by
permission of the publisher. © 1964 by
Scientific American.)
Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 73
our friend is wearing a bright red jacket, it can be relatively
easy to find him or her in the crowd, provided
that no one else is wearing a bright red jacket. Our
friend will just pop out of the crowd. Indeed, if there
were just one red jacket in a sea of white jackets it
would probably pop out even if we were not looking
for it—an instance of stimulus-driven attention. It
seems that if there is some distinctive feature in an
array, we can find it without a search.
Treisman studied this sort of pop-out. For instance,
Treisman and Gelade (1980) instructed participants
to try to detect a T in an array of 30 I’s and Y’s
(Figure 3.13a). They reasoned that participants could
do this simply by looking for the crossbar feature of
the T that distinguishes it from all I’s and Y’s. Participants
took an average of about 400 ms to perform
this task. Treisman and Gelade also asked participants
to detect a T in an array of I’s and Z’s (Figure 3.13b).
In this task, they could not use just the vertical bar
or just the horizontal bar of the T; they would have
to look for the conjunction of these features and
perform the feature combination required in pattern
recognition. It took participants more than 800 ms,
on average, to find the letter in this case. Thus, a task
requiring them to recognize the conjunction of features
took about 400 ms longer than one in which
perception of a single feature was sufficient. Moreover,
when Treisman and Gelade varied the number
of letters in the array, they found that participants were much more
affected by array size in the task that required recognition of the
conjunction of features. Figure 3.14 shows these results.
It is necessary to search through a visual array for an object only
when a unique visual feature does not distinguish that object.
The Binding Problem
As discussed in Chapter 2, there are different types of neurons in
the visual system that respond to various features, such as colors,
lines at various orientations, and objects in motion. A single object
in our visual field will involve a number of features; for instance, a
red vertical line combines the vertical feature and the red feature.
The fact that different features of the same object are represented
by different neurons gives rise to a logical question: How are these
features put back together to produce perception of the object?
This would not be much of a problem if there were a single object
in the visual field. We could assume that all the features belonged
74 | Attention and Performance
Array size (number of items)
T in I, Z
Reaction time (ms)
1
0
400
800
1200
5 15 30
T in I, Y
(a)
(b)
FIGURE 3.13 Stimuli used by Treisman and Gelade to determine
how people identify objects in the visual field. They found that it is
easier to pick out a target letter (T ) from a group of distracter letters
(I ’s and Y ’s) if (a) the target letter has a feature that makes it
easily distinguishable from the distracters than if (b) the same target
letter is in an array of distracters (I ’s and Z ’s) that offer no obvious
distinctive features. (After Treisman & Gelade, 1980. Adapted by permission of the
publisher. © 1980 by Cognitive Psychology.)
FIGURE 3.14 Results from the Treisman and
Gelade experiment. The graph plots the average
reaction times required to detect a target letter as
a function of the number of distracters and whether
the distracters contain separately all the features
of the target. (After Treisman & Gelade, 1980. Adapted
by permission of the publisher. © 1980 by Cognitive Psychology.)
Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 74
Visual Attention | 75
to that object. But what if there are multiple objects in the field? For instance,
suppose there were just two objects: a red vertical bar and a green horizontal bar.
These two objects might result in the firing of neurons for red, neurons for
green, neurons for vertical lines, and neurons for horizontal lines. If these firings
were all that occurred, though, how would the visual system know it saw a red
vertical bar and a green horizontal bar rather than a red horizontal bar and a
green vertical bar? The question of how the brain puts together various features
in the visual field is referred to as the binding problem.
Treisman (e.g., Treisman & Gelade, 1980) developed her feature-integration
theory as an answer to the binding problem. She proposed that people must focus
their attention on a stimulus before they can synthesize its features into a pattern.
For instance, in the example just given, the visual system can first direct its attention
to the location of the red vertical bar and synthesize that object, then direct
its attention to the green horizontal bar and synthesize that object. According to
Treisman, people must search through an array when they need to synthesize
features to recognize an object (for instance, when trying to identify a K, which
consists of a vertical line and two diagonal lines). When there is a single unique
feature, such as a red jacket or a line at a particular orientation, we can move our
attention directly to the object and recognize it, thus avoiding the need to search.
The binding problem is not just a hypothetical dilemma—it is something
that humans actually suffer from. One source of evidence comes from studies
of illusory conjunctions in which people report combinations of features
that did not occur. For instance, Treisman and Schmidt (1982) looked at what
happens to feature combinations when the stimuli are out of the focus of attention.
Participants were asked to report the identity of two black digits flashed in
one part of the visual field. This was their primary task, and it was where their
attention was focused. In another part of the visual field, letters in various
colors were presented. Thus, participants might be presented with a pink T,
a yellow S, and a blue N in the unattended portion of the field. After they
reported the numbers, participants were asked to report any letters they had
seen and the colors of these letters. They reported seeing illusory conjunctions
of features (e.g., a pink S) almost as often as they reported seeing correct combinations.
Thus, it appears that we are able to combine features into an accurate
perception only when our attention is focused on an object. Otherwise, we
perceive the features but may well combine them into a perception of objects
that were never there. Although rather special circumstances are required to
produce illusory conjunctions in an ordinary person, there are certain patients
with damage to the parietal cortex who are particularly prone to such illusions.
For instance, one patient studied by Friedman-Hill, Robertson, and Treisman
(1995) confused which letters were presented in which colors even when shown
the letters for as long as 10 seconds.
A number of studies have been conducted on the neural mechanisms
involved in binding together the features of a single object. The neurons in
visual area V4 have large receptive fields (several degrees of visual angle), and
multiple objects in a display may be within the visual field of a single neuron.
Luck, Chelazzi, Hiillyard, and Desimone (1997) trained macaque monkeys to
fixate on a certain part in the visual field and recorded neurons around this
Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 75
region. They found neurons that were specific to particular types of
objects. For instance, they found a cell that responded to a blue
vertical bar.What happens when a blue vertical bar and a green horizontal
bar are presented both within the receptive field of this cell?
If the monkey attended to the blue vertical bar, the rate of response
of the cell will remain at the same level as it would have had there
been only a blue vertical bar alone. On the other hand, if the monkey
focused on the green horizontal bar, the rate of firing of this
same cell will be greatly depressed. Thus, the same stimulus (blue
vertical bar plus green horizontal bar) can evoke different responses
depending on which object is attended to. It is speculated that this
phenomenon occurs because attention suppresses responses to all
features in the receptive field except those at the attended location.
Similar results have been obtained in fMRI experiments with humans.
Kastner, DeWeerd, Desimone, and Ungerleider (1998) measured the fMRI
signal in visual areas that responded to stimuli presented in one region of the
visual field. They found that when attention was directed away from that region,
the fMRI response to stimuli in that region decreased; but when attention was
focused on that region, the fMRI response was maintained. These experiments
indicate enhanced neural processing of attended objects and locations.
A striking demonstration of the effects of sustained attention was reported
by Simons and Chabris (1999). They asked participants to watch a video in
which a team dressed in black tossed a basketball back and forth and a team
dressed in white did the same (Figure 3.15). Participants were instructed to
count either the number of times the team in black tossed the ball or the number
of times the team in white did so. Presumably, in one condition participants
were looking for events involving the team in black and in the other for events
involving the team in white. Because the players were intermixed, the task was
difficult and required sustained attention. In the middle of the game, a person in
a black gorilla suit walked through the room. When participants were tracking
the team in white, they noticed the black gorilla only 8% of the time; when they
were tracking the team in black, they noticed it 67% of the time. They were so
fixed on searching the video for events involving team members dressed in white
that they completely missed an event involving a black object. People passively
watching the video never miss the black gorilla. The actual video is currently
available from the demonstrations page of Simons’s Visual Cognition Lab:
http://viscog.beckman.uiuc.edu/djs_lab/demos.html
For feature information to be synthesized into a pattern, it must be in the
focus of attention.
Neglect of the Visual Field
We have discussed the evidence that visual attention to a spatial location results
in enhanced activation in the appropriate portion of the primary visual cortex.
The neural structures that control this shift of attention, however, appear to
be located elsewhere. Three areas of the monkey brain have been shown to be
76 | Attention and Performance
FIGURE 3.15 Single frame from
the movie used by Simons and
Chabris to demonstrate the
effects of sustained attention.
When participants were intent
on tracking the ball passed
among the players dressed in
T-shirts, they tended not to
notice the black gorilla walking
through the room. (From Simons &
Chabris, 1999. Reprinted by permission of
the publisher. © 1999 by Perception.)
Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 76
Visual Attention | 77
involved in controlling attention (S. E. Peterson, Robinson, &
Morris, 1987;Wurtz, Goldberg, & Robinson, 1980). These areas
are the superior colliculus, the posterior parietal lobe, and a
midbrain area known as the pulvinar. Damage to these areas in
human patients, particularly to the parietal lobe (see Figure 3.1),
has been shown to result in deficits in visual attention. For
instance, Posner, Walker, Friedrich, and Rafal (1984) showed
that patients with parietal lobe injuries have difficulty in disengaging
attention from one side of the visual field.
Damage to right parietal regions produces distinctive
patterns of deficit. Posner, Cohen, and Rafal (1982) studied the
attention deficit in one such patient. The patient was cued to
expect a stimulus to the left or right of the fixation point, and
80% of the time that is where the stimulus was present. However,
20% of the time the stimulus appeared in the unexpected
field. Figure 3.16 shows the time required to detect the stimulus
as a function of which visual field it was presented in and
which field had been cued. When the stimulus was presented
in the right field, the patient showed only a little disadvantage
if inappropriately cued. If the stimulus appeared in the left
field, however, the patient showed a large deficit if inappropriately cued. Because
the right parietal lobe processes the left visual field, damage to the right lobe
impairs its ability to draw attention back to the left visual field once attention is
focused on the right visual field. This sort of one-sided attentional deficit can be
temporarily created in normal individuals by presenting TMS to the parietal
cortex (Pascaul-Leone et al., 1994—see Chapter 1 for discussion of TMS).
A more extreme version of this attentional disorder
is called unilateral visual neglect. Patients
with damage to the right hemisphere completely
ignore the left side of the visual field, and patients
with damage to the left hemisphere ignore the
right side of the field. Figure 3.17 shows the performance
of a patient with damage to the right
hemisphere, which caused her to neglect the left
visual field (Albert, 1973). She had been instructed
to put slashes through all the circles. As can be
seen, she ignored the circles in the left part of her
visual field. Such patients will often behave peculiarly.
For instance, one patient failed to shave half
of his face (Sacks, 1985).
It seems that the right parietal lobes are involved
in allocating spatial attention in many
modalities, not just the visual (Zatorre et al.,
1999). For instance, when one attends to the
location of auditory or visual stimuli, there is
increased activation in the right parietal region.
It also appears that the right parietal lobes are
Latency (ms)
1400
1200
1000
800
600
400
Left
Field of presentation
Right
Cued for right field
Cued for left field
FIGURE 3.16 The attention
deficit shown by a patient with
right parietal lobe damage when
switching attention to the left
visual field. (From Posner, Cohen, & Rafal,
1982. Reprinted by permission of the
publisher. © by the Royal Society of London.)
FIGURE 3.17 The performance of a patient with damage to the right
hemisphere who had been asked to put slashes through all the circles.
Because of the damage to the right hemisphere, she ignored the
circles in the left part of her visual field. (From Ellis & Young, 1988. Reprinted by
permission of the publisher. © 1988 by Erlbaum.)
Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 77
more responsible for the spatial allocation of attention and that this is why right
parietal damage tends to produce such dramatic effects. Left parietal damage
tends to produce a subtler pattern of deficits. Robertson and Rafal (2000) argue
that the right parietal lobe is responsible for attention to such global features as
spatial location, whereas the left parietal region is responsible for directing
attention to local aspects of objects. Figure 3.18 is a striking illustration of the
different types of deficits associated with left and right parietal damage. Patients
were asked to draw the objects in Figure 3.18a. Patients with right parietal
damage (Figure 3.18b) were able to reproduce the specific components of the
picture but were not able to reproduce their spatial configuration. In contrast,
patients with left parietal damage (Figure 3.18c) were able to reproduce
the overall configuration, but not the detail. Similarly, brain-imaging studies
78 | Attention and Performance
(a) (b) (c)
FIGURE 3.18 (a) The pictures presented to patients with parietal damage. (b) Examples of
drawings made by patients with right-hemisphere damage. These patients could reproduce
the specific components of the picture but not their spatial configuration. (c) Examples of
drawings made by patients with left-hemisphere damage. These patients could reproduce
the overall configuration but not the detail. (After Robertson & Lamb, 1991. Adapted by permission of
the publisher. © 1991 by Cognitive Psychology.)
Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 78
Visual Attention | 79
have found more activation of the right parietal region when a person is
responding to global patterns and more activation of the left hemisphere when
a person is attending to local patterns (Fink et al., 1996;Martinez et al., 1997).
Parietal regions are responsible for the allocation of attention, with the right
hemisphere more concerned with global features and the left hemisphere
with local features.
Object-Based Attention
So far we have talked about space-based attention, where people allocate their
attention to a region of space. There is also evidence, though, for object-based
attention, where people focus their attention on particular objects rather than
regions of space. An experiment by Behrmann, Zemel, and Mozer (1998) is an
example of the research that shows people sometimes find it easier to
attend to an object than to a location. Figure 3.19 illustrates some of
the stimuli used in the experiment. Participants were asked to judge
whether the numbers of bumps on the two ends of objects were the same.
The left column shows instances in which the numbers of bumps were
the same, the right column instances in which the numbers were not the
same. Participants made these judgments faster when the bumps were on
the same object (top and bottom rows in Figure 3.19) than when they
were on different objects (middle row). This result occurred despite the
fact that when the bumps were on different objects, they were located
closer together, which should have facilitated judgment. Behrmann et al.
argue that participants can shift attention to one object at a time, but not
one location at a time. Therefore, when the bumps were all on the same
object, participants did not need to shift their attention between objects.
Other evidence for object-centered attention involves a phenomenon
called inhibition of return. If we have looked at a particular region
of space, we find it harder to return our attention to that region. This
phenomenon also makes sense. If we are searching for something and have
already looked at a location, we would prefer our visual system to find other
locations to look at rather than return to an already searched location. If we
move our eyes to location A and then to location B, we are slower to return our
eyes to location A than to some new location C. This is also true when we move
our attention without moving our eyes (Posner, Rafal, Chaote, & Vaughn, 1985).
Tipper, Driver, and Weaver (1991) performed one demonstration of the
inhibition of return that also provided evidence for object-based attention. In
their experiments, participants viewed three squares in a frame, similar to what
is shown in each part of Figure 3.20. In one condition, the squares did not
move (unlike the moving condition illustrated in Figure 3.20, which we will
discuss in the next paragraph). The participants’ attention was drawn to one of
the outer squares by making it flicker. Attention was drawn back to the center
square 200 ms later by making that square flicker. A probe was then presented
in one of the two outer positions, and participants were instructed to press a
key indicating that they had seen the probe. On average, they took 420 ms to
(a) (d)
(e)
(c) (f)
(b)
FIGURE 3.19 Stimuli used in an
experiment by Behrmann, Zemel,
and Mozer to demonstrate that
it is sometimes easier to attend
to an object than to a location.
The left and right columns
indicate same and different
judgments, respectively; and
the rows from top to bottom
indicate the single-object,
two-object, and occluded
conditions, respectively. (From
Behrmann, Zemel, & Mozer, 1998. Reprinted
by permission of the publisher. © 1998
by the Journal of Experimental Psychology:
Human Perception and Performance.)
Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 79
see the probe when it occurred at the outer square that had not flickered and
460 ms when it occurred at the outer square that had. This 40-ms advantage is
an example of a spatially defined inhibition of return. People are slower to
move their attention to a location where it has already been.
Figure 3.20 illustrates the other condition of their experiment, in which the
objects were rotated around the screen after the flicker. By the end of the motion,
the object that had flickered on one side was now on the other side—the
two outer objects had traded positions. The question of interest was whether
participants would be slower to detect a target on the right (where the flickering
had been—which would indicate location-based inhibition) or on the left
(where the flickered object had ended up—which would indicate object-based
inhibition). The results showed that they were about 20 ms slower to detect an
object in the location that had not flickered but that contained the object that
80 | Attention and Performance
(a)
(b)
(c)
(d)
(e)
FIGURE 3.20 Examples of frames used in an experiment by Tipper, Driver, and
Weaver to determine whether inhibition of return would attach to a particular object
or to its location. Arrows represent motion. (a) Display onset, with no motion for
500 ms. After two moving frames, the three filled squares were horizontally aligned
(b), whereupon the cue appeared (one of the boxes flickered). Clockwise motion then
continued, with cueing in the center for the initial three frames (c–e). The outer
squares continued to rotate clockwise (d) until they were horizontally aligned (e),
at which point a probe was presented, as before. (From Tipper, Driver, & Weaver, 1991. Reprinted
by permission of the publisher. © 1991 by the Quarterly Journal of Experimental Psychology.)
Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 80
Central Attention: Selecting Lines of Thought to Pursue | 81
had flickered. Thus, their visual systems displayed an inhibition of return to the
same object, not the same location.
Another example of object-based attention comes from studies of visual
neglect. Earlier, we noted that some patients with damage to the right parietal
lobe have difficulty detecting information in the left side of the visual field
(see Figure 3.17). Researchers have identified a number of patients who neglect
the left side of objects regardless of which visual field these objects occur in
(Behrmann & Moscovitch, 1994; Driver, Baylis, Goodrich, & Rafal, 1994).
It seems that the visual system can direct attention either to locations in space
or to objects. Experiments like those just described indicate that the visual system
can track objects. On the other hand, there are many experiments in which people
direct their attention to regions of space where there are no objects (see Figure 3.6
for the results of such an experiment). It is interesting that the left parietal regions
seem to be more involved in object-based attention and the right parietal regions
in location-based attention. Patients with left parietal damage appear to have
deficits in focusing attention on objects (Egly, Driver, & Rafal, 1994), unlike the
location-based deficits that I have described in patients with right parietal damage.
Also, when participants without brain damage attend to objects rather than locations,
there is greater left parietal activation, as revealed by fMRI (Arrington, Carr,
Mayer, & Rao, 2000). This seems consistent with the earlier research we reviewed
(see Figure 3.18) showing that the right parietal region is responsible for attention
to global features and the left for attention to local features.
Visual attention can be directed either toward objects independent of their
location or toward locations independent of what objects are present.
•Central Attention: Selecting Lines
of Thought to Pursue
So far, this chapter has considered how people allocate their attention to process
stimuli in the visual and auditory modalities. What about cognition after the
stimuli are attended to and encoded? How do we select which lines of thought to
pursue? Suppose we are driving down a highway and encode the fact that a dog is
sitting in the middle of the road.We might want to figure out why the dog is sitting
there, we might want to consider whether there is something we should do
to help the dog, and we certainly want to decide how best to steer the car to avoid
an accident. Can we do all these things at once? If not, how do we select the most
important problem of deciding how to steer and save the rest for later? It appears
that people allocate central attention to competing lines of thought in much the
same way they allocate perceptual attention to competing objects.
In many (but not all) circumstances, people are able to pursue only one line
of thought at a time. This section will describe two laboratory tasks: one in
which it appears that people have no ability to overlap two tasks and another
pair in which they appear to have almost total ability to do so. Then we will
address how people can develop the ability to overlap tasks and how they select
among tasks when they cannot or do not want to overlap them.
Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 81
The first experiment, which Mike Byrne and I did (Byrne & Anderson,
2001), illustrates the claim made at the beginning of the chapter about it being
impossible to multiply and add two numbers at the same time. Participants in
this experiment saw a string of three digits, such as “3 4 7.” There were two tasks
they might be asked to do:
• Task 1: Judge whether the first two digits add up to the third and press a
key with the right index finger if they do and another key with the left
index finger if they do not. • Task 2: Report verbally the product of the first and third numbers. In this
case, the answer is 21, because 3 _ 7 _ 21.
Participants either performed these two tasks individually or tried to do
both at once. Figure 3.21 compares the time required to do each task in the
single-task condition versus the time required for each task in the dual-task
condition. Participants took almost twice as long to do either task when they
had to perform the other as well. The illustration also displays the time participants
took to complete both tasks in the dual-task condition. (They answered
the multiplication problem first 59% of the time and the addition problem first
41% of the time.) The horizontal black line near the top of Figure 3.21 represents
the time they took to give the second answer, whichever it was. This line
reflects the mean time required to complete both tasks (1.99 s). Note that this
time is greater than the sum of the time for the verification task by itself (0.88 s)
82 | Attention and Performance
Verify
addition
Latency (ms)
Generate
multiplication
250
0
500
750
1000
1250
1500
1750
2000
2250
Stimulus: 3 4 7
Mean time to complete both tasks
Single task
Dual task
FIGURE 3.21 The results of
an experiment by Byrne and
Anderson to see whether people
can overlap two tasks. The bars
show the response times
required to solve two problems—
one of addition and one of
multiplication—when done by
themselves and when done
together. The results indicate
that the participants were not
able to overlap the addition
and multiplication computations.
(From Byrne & Anderson, 2001. Reprinted
by permission of the publisher. © 2001
by Psychological Review.)
Anderson7e_Chapter_03.qxd 8/20/09 9:42 AM Page 82
Central Attention: Selecting Lines of Thought to Pursue | 83
and the time for the multiplication task by itself (1.05 s). The extra time probably
reflects the cost of shifting between tasks (for a review, see Monsell, 2003). In
any case, it appears that the participants were not able to overlap the addition
and multiplication computations at all.
The second experiment, reported by Schumacher et al. (2001), illustrates
what is referred to as perfect time-sharing. The task was much simpler than
the Byrne and Anderson (2001) experiment. Participants simultaneously saw a
single letter on a screen and heard a tone. As in the first experiment, they had to
perform two tasks:
• Task 1: Press a left, middle, or right key according to whether the letter
occurred on the left, in the middle, or the right. • Task 2: Report “One,” “two,” or “three” according to whether the tone was
low, middle, or high in frequency.
Figure 3.22 compares the times required to do each task in the single-task
condition and the dual-task condition. As can be seen, these times are nearly
unaffected by the requirement to do the two tasks at once. There are many
differences between this task and the Byrne and Anderson task. Perhaps the most
apparent is the complexity of the tasks. Participants were able to do the individual
tasks in the second experiment in a few hundred milliseconds, whereas the
individual tasks in the first experiment took around a second. Thus, there was
significantly more thought required in the first experiment, and it is apparently
harder for people to engage in both streams of thought simultaneously. Also,
Location
discrimination
Response time (ms)
0
50
100
150
200
250
300
350
400
450
500
Tone
discrimination
Single task
Dual task
FIGURE 3.22 The results of
an experiment by Schumacher
et al. illustrating perfect
time-sharing. The bars show
the times required to perform
two tasks—a simple location
discrimination task and a tone
discrimination task—when done
by themselves and when done
together. The times were nearly
unaffected by the requirement
to do the two tasks at once,
indicating that the participants
achieved almost perfect
time-sharing. (From Schumacher et al.,
2001. Reprinted by permission of the
publisher. © 2001 by Psychological
Science.)
Anderson7e_Chapter_03.qxd 8/20/09 9:42 AM Page 83
participants in the second experiment achieved perfect time-sharing only after
five sessions of practice. There was no such practice in the first experiment.
Figure 3.23 presents an analysis of what occurred in the Schumacher et al.
(2001) experiment. It shows what was happening at various points in time in
five streams of processing: (1) perceiving the visual location of a letter, (2) generating
manual actions, (3) central cognition, (4) perceiving auditory stimuli,
and (5) generating speech. Task 1 involved visually encoding the location of the
letter, using central cognition to select which finger to press, and then performing
the actual finger movement. Task 2 involved detecting and encoding the tone,
using central cognition to select which word to say (“one,” “two,” or “three”), and
then generating the word and speaking it. The lengths of the boxes in Figure 3.23
represent estimates of the duration of each component based on human performance
studies. Each of these streams can go on in parallel with the others.
Thus, for instance, during the time the tone is being detected and encoded, the
location of the letter is being encoded (which happens much faster), a finger is
being selected by central cognition, and the motor system is starting to program
the action. Although all these streams can go on in parallel, within each stream
only one thing can happen at a time. This is a potential problem in the case of
the central cognition stream because central cognition must direct all activities.
In this case, it must serve both task 1 and task 2. In this experiment, however, the
length of time devoted to central cognition was so brief that the two tasks did
84 | Attention and Performance
Encode
letter
location
Vision
Manul action
Central cognition
Speech
Time (ms)
Streams of processing:
Task 2
Task 1
Audition
Select
action
Program key press
Detect and
encode tone
0 100 200 300 400
Generate speech
Select
action
FIGURE 3.23 An analysis of the timing of events in five streams of processing during execution
of the dual task in the Schumacher et al. (2001) experiment: (1) vision, (2) manual action,
(3) central cognition, (4) speech, and (5) audition.
Anderson7e_Chapter_03.qxd 8/20/09 9:42 AM Page 84
Central Attention: Selecting Lines of Thought to Pursue | 85
not contend for the resource. The five days of practice in this experiment played
a critical role in reducing the amount of time devoted to central cognition.
Although the discussion here has focused on bottlenecks in central cognition,
there can be bottlenecks in any of these modalities. People cannot attend
to two locations at once. Earlier, we reviewed evidence that they must shift their
attention across locations in the visual array serially if they must attend to more
than one location. Similarly, they can process only one speech stream at a time,
move their hands in one way at a time, or say one thing at a time. Even though
all these peripheral processes have bottlenecks, it is generally thought that bottlenecks
in central cognition can have the most significant effects, and they are
the reason we seldom find ourselves thinking about two things at once. This
bottleneck in central cognition is referred to as the central bottleneck.
People can process multiple perceptual modalities at once or execute actions
in multiple motor systems at once, but they cannot process multiple things
in a single system including central cognition.
Automaticity: Expertise Through Practice
In Chapter 9, we will discuss at some length how people become expert with
practice. The general effect of practice is to reduce the central cognitive component
of information processing. When one has practiced the central cognitive
component of a task so much that the task requires little or no thought, we say
phones. In contrast, listening to a radio or books on tape
does not interfere with driving. Strayer and Drews
suggest that the demands
of participating in a conversation
place more requirements
on central cognition. When
someone says something on
the cell phone, they expect an
answer and are unaware of
the driving conditions. Strayer
and Drews note that participating
in a conversation with
a passenger in the car is not
as distracting because the
passenger will adjust the conversation
to driving demands and even point out things
like exits to the driver.
Implications
Why is cell phone use and driving a dangerous combination?
Bottlenecks in information processing can have important
practical implications. A study by the Harvard
Center for Risk Analysis
(Cohen & Graham, 2003)
estimates that cell phone
distraction results in
2,600 deaths, 330,000
injuries, and 1.5 million instances
of property damage
in the United States
each year. Strayer and
Drews (2007) review the
evidence that people are
more likely to miss traffic
lights and other critical
information while talking on a cell phone. Moreover,
these problems are not any better with hands-free
Anderson7e_Chapter_03.qxd 8/20/09 9:42 AM Page 85
that doing the task is automatic. Automaticity is a matter of degree. A nice
example is driving. For experienced drivers in unchallenging conditions, driving
has become so automatic that they can carry on a conversation while driving
with little difficulty. Experienced drivers are much more successful at doing
secondary tasks like changing the radio (Wikman,Nieminen, & Summala, 1998).
Experienced drivers also often have the experience of traveling long stretches of
highway with no memory of what they did.
There have been a number of dramatic demonstrations in the psychological
literature of how practice can enable parallel processing. One demonstration
of the way practice affects attentional limitations is the study reported by
Underwood (1974) on the psychologist Neville Moray, who had spent many
years studying shadowing. During that time, Moray practiced shadowing a
great deal. Unlike most participants in experiments, he was very good at
reporting what was contained in the unattended channel. Through a great deal
of practice, the process of shadowing had become partially automatic for
Moray, and he had capacity left over to attend to the unshadowed channel.
Spelke, Hirst, and Neisser (1976) provided an interesting demonstration of
how a highly practiced skill ceases to interfere with other ongoing behaviors.
(This was a follow-up of a demonstration pioneered by the writer Gertrude
Stein when she was at Harvard University.) Their participants had to perform
two tasks: read a text silently for comprehension while simultaneously copying
words dictated by the experimenter. At first, these tasks were extremely difficult
to do simultaneously. Participants read much more slowly than normal. After
six weeks of practice, however, the participants were reading at normal speed.
They had become so skilled that their comprehension scores were the same as
for normal reading. For these participants, reading while copying had become
no more difficult than reading while walking. It is of interest that participants
reported no awareness of what it was they were copying. Much as with driving,
the participants lost their awareness of the automated activity.3
Another example of automaticity is transcription typing. The typist is
simultaneously reading the text and executing the finger strokes for typing. In this
case, we have three systems operating in parallel: perception of the text to be
typed, central translation of the earlier perceived letters into keystrokes, and the
actual typing of still earlier letters. Skilled transcription typists often report little
awareness of what they are typing, because this task has become so automated.
Skilled typists also find it impossible to stop typing instantaneously. If suddenly
told to stop, they will hit a few more letters before quitting (Salthouse, 1985, 1986).
As tasks become practiced, they become more automatic and require less and
less central cognition to execute.
The Stroop Effect
Automatic processes not only require little or no central cognition to execute but
also appear to be difficult to prevent. A good example is word recognition for
86 | Attention and Performance
3 When given further training with the intention of remembering what they were transcribing, participants
were also able to recall this information.
Anderson7e_Chapter_03.qxd 8/20/09 9:42 AM Page 86
Central Attention: Selecting Lines of Thought to Pursue | 87
practiced readers. It is virtually impossible to look
at a common word and not read it. This strong
tendency for words to be recognized automatically
has been studied in a phenomenon known as the
Stroop effect, after the psychologist who first
demonstrated it, J. Ridley Stroop (1935). The task
requires participants to say the ink color in which
words are printed. Color Plate 3.2 provides an illustration
of such a task. Try naming the colors of the
words in each column as fast as you can.Which column
was easiest to read? Which was hardest?
The three columns illustrate three of the conditions
in which the Stroop effect is studied. The first
column illustrates a neutral, or control, condition
in which the words are not color words. The second
column illustrates the congruent condition in which
the words are the same as the color. The third column
illustrates the conflict condition in which there are
color words but they are different from their colors. A
typical modern experiment, rather than having participants
read a whole column, will present a single word
at a time and measure the time to name that word. Figure
3.24 shows the results from an experiment on the Stroop effect by Dunbar and
MacLeod (1984). Compared to the control condition of a neutral word, participants
could name the ink color somewhat faster in the congruent condition—
when the word was the name of the ink color. In the conflict condition, when the
word was the name of a different color, they named the ink color much more
slowly. For instance, they had great difficulty in saying that the ink color of the word
red is green. Figure 3.24 also shows the results when the task is switched and participants
are asked to read the word and not name the color. The effects are asymmetrical;
that is, individual participants experienced very little interference in reading a
word as a function of its ink color. This reflects the highly automatic character of
reading.Additional evidence for its automaticity is that participants could also read
a word much faster than they could name its ink color. Reading is such an automatic
process that not only is it unaffected by the color, but participants are unable
to inhibit reading the word, and that reading can interfere with the color naming.
MacLeod and Dunbar (1988) looked at the effect of practice on performance
in a Stroop task. They used an experiment in which the participants learned
the color names for random shapes. Part (a) of Color Plate 3.3 illustrates the
shape-color associations they might learn. The experimenters then presented the
participants with test geometric shapes and asked them to say either the color
name associated with the shape or the actual ink color of the shape. As in the
original Stroop experiment, there were three conditions, and these are illustrated
in part (b) of Color Plate 3.3:
1. Congruent: The random shape was in the same ink color as its name.
2. Control:White shapes were presented when participants were to say the
color name for the shape; colored squares were presented when they were
Reaction time (ms)
900
800
700
600
500
400
Congruent Control Conflict
Color naming
Word reading
Condition
FIGURE 3.24 Performance data
for the standard Stroop task.
The curves plot the average
reaction time of the participants
as a function of the condition
tested: congruent (the word
was the name of the ink color);
control (the word was not
related to color at all); and
conflict (the word was the name
of a color different from the
ink color). (From Dunbar & MacLeod,
1984. Reprinted by permission of the
publisher. © 1984 by the Journal of
Experimental Psychology: Human
Perception and Performance.)
Anderson7e_Chapter_03.qxd 8/20/09 9:42 AM Page 87
to name the ink color of the shape. (The square shape was not associated
with any color.)
3. Conflict: The random shape was in a different ink color from its name.
Figure 3.25 shows the results from this experiment. Color naming was much
more automatic than shape naming and was relatively unaffected by congruence
with the shape, whereas shape naming was affected by congruence with the ink
color (Figure 3.25a). Then MacLeod and Dunbar gave the participants 20 days
of practice at naming the shapes. Participants became much faster at naming
shapes, and now shape naming interefered with color naming rather than vice
versa (Figure 3.25b). Thus, the consequence of the training was to make shape
naming automatic, like word reading, so that it affected color naming.
Reading a word is such an automatic process that it is difficult to inhibit,
and it will interfere with processing other information about the word.
Prefrontal Sites of Executive Control
We have seen that the parietal cortex is important in the exercise of attention in
the perceptual domain. There is evidence that the prefrontal regions are particularly
important in direction of central cognition, often known as executive
control. The prefrontal cortex is that portion of the frontal cortex anterior to the
premotor region (the premotor region is area 6 in Color Plate 1.1). Just as damage
to parietal regions results in deficits in the deployment of perceptual attention,
damage to prefrontal regions results in deficits of executive control. Patients with
88 | Attention and Performance
Reaction time (ms)
750
700
650
600
550
500
450
Congruent Control Conflict
(a) Condition (b)
Reaction time (ms)
750
700
650
600
550
500
450
Congruent Control Conflict
Condition
Color naming
Shape naming
Color naming
Shape naming
FIGURE 3.25 Results from the experiment created by MacLeod and Dunbar (1988) to evaluate
the effect of practice on the performance of a Stroop task. The data reported are the average
times required to name shapes and colors as a function of color-shape congruence: (a) initial
performance (note correction); (b) after 20 days of practice. The practice made shape naming
automatic, like word reading, so that it affected color naming.
Anderson7e_Chapter_03.qxd 8/20/09 9:42 AM Page 88
Central Attention: Selecting Lines of Thought to Pursue | 89
such damage often seem totally driven by stimulus and fail to control their behavior
according to their intentions. A patient who sees a comb on the table may
simply pick it up and begin combing her hair; another who sees a pair of glasses
will put them on even if he already has a pair on his face. Patients with damage to
prefrontal regions show marked deficits in the Stroop task and often cannot refrain
from reading the word rather than naming the color (Janer & Pardo, 1991).
Two prefrontal regions seem particularly important in executive control. One
is the dorsolateral prefrontal cortex (DLPFC), which is the upper portion of
the prefrontal cortex (see Figure 3.1). It is called dorsolateral because it is high
(dorsal) and to the side (lateral). The second region is the anterior cingulate
cortex (ACC), which is a structure below the surface of the brain along the midline.
It is part of the cortex that is folded under its visible surface. The DLPFC
seems particularly important in the setting of intentions and the control of
behavior. For instance, it is highly active during the simultaneous performance
of dual tasks such as those whose results are reported in Figures 3.19 and 3.20
(Szameitat, Schubert,Muller, & von Cramon, 2002). The ACC seems particularly
active when people must monitor conflict between competing tendencies.
For instance, brain-imaging studies show that it is highly active
in Stroop trials when one must name the color of a word printed in an
ink of conflicting color (J.V. Pardo, P. J. Pardo, Janer, & Raichle, 1990).
There is a strong relationship between the ACC and cognitive control
in many tasks. For instance, it appears that children develop more cognitive
control as their ACC develops. The amount of activation in the ACC
appears to be correlated with the performance by children in tasks
requiring cognitive control (Casey et al., 1997a). Developmentally, there
also appears to be a positive correlation between performance and sheer
volume of the ACC (Casey et al., 1997b). A nice paradigm for demonstrating
the development of cognitive control in children is the “Simon
says” task. In one study, Jones, Rothbart, and Posner (2003) had children
receive instructions from two dolls—a bear and an elephant. The instructions
were things like “Elephant says, ‘Touch your nose.’” The children
were to follow the instructions from one doll (the act doll) and ignore the
instructions from another (the inhibit doll). All children successfully followed
the act doll but many had difficulty ignoring the inhibit doll. From
the age of 36 to 48 months children progressed from 22% success to 91%
success in ignoring the inhibit doll. A few children used regulatory selfspeech
to control their behavior, as famously proposed by Luria (1961),
but they also used strategies such as sitting on their hands or distorting
their actions—pointing to their ear rather than their nose.
Another way to appreciate the importance of prefrontal regions to
cognitive control is to compare performance of humans with other
primates. As reviewed in Chapter 1, a major dimension of the evolution
from primates to humans has been the increase in the size of prefrontal
regions. Primates can be trained to do many tasks that humans do and
so permit careful comparison. One such task involving a variant of the Stroop task
presents participants with a display of numerals (e.g., five 3’s) and pits naming the
number of objects against indicating the identity of the numerals. Figure 3.26
5 5 5
1 1 1 1
2
3 3 3 3 3
4 4
5 5 5
4 4 4 4 4
5 5 5 5
3
4 4 4
2 2 2 2
3 3
4 4 4
1 1 1 1
3
2 2 2
FIGURE 3.26 A numerical
Stroop task comparable to the
color Stroop task (see Color
Plate 3.2).
Anderson7e_Chapter_03.qxd 8/20/09 9:42 AM Page 89
provides an example of this task in the same form as the original Stroop task
(Color Plate 3.2): trying to count the number of numerals in each line versus
trying to name the numerals in each line. The stronger interference in this case
is from the numeral naming to the counting (Windes, 1968). This paradigm has
been used to compare Stroop interference in humans versus rhesus monkeys
who had been trained to use the numerals (Washburn, 1994; see Table 3.1).
Both groups of participants were shown two arrays and were required to indicate
which had more numerals independent of the identity of the numerals.
Compared to a baseline where they had to judge which array of letters had more
objects, both humans and monkeys performed better when the numerals agreed
with the difference in cardinality and performed worse when the numerals disagreed
(as they do in Figure 3.26). Both populations showed similar reaction
time effects, but whereas the humans made 3% errors in the incongruent condition,
the monkeys made 27% errors. The level of performance observed of the
monkeys was like the level of performance observed with patients with damage
to their frontal lobes.
Prefrontal regions, particularly DLPFC and ACC, play a major role in
executive control.
•Conclusions
There has been a gradual shift in the way cognitive psychology has perceived
the issue of attention. For a long time, the implicit assumption was captured
by this famous quote from William James (1890) over a century ago:
Everyone knows what attention is. It is the taking possession by the mind, in a
clear and vivid form, of one out of what seem several simultaneously possible
objects or trains of thought. Focalization, concentration of consciousness are
of its essence. It implies withdrawal from some things in order to deal effectively
with others. (pp. 403–404)
90 | Attention and Performance
TABLE 3.1
Mean Response Times and Accuracy Levels as a Function of Species and Condition
Condition Accuracy (%) Response Time (ms)
Rhesus Monkeys (N _ 6)
Congruent numerals 92 676
Baseline (letters) 86 735
Incongruent numerals 73 829
Human Subjects (N _ 28)
Congruent numerals 99 584
Baseline (letters) 99 613
Incongruent numerals 97 661
Anderson7e_Chapter_03.qxd 8/20/09 9:42 AM Page 90
Key Terms | 91
Two features of this quote reflect conceptions once held about attention. The
first is that attention is strongly related to consciousness—we cannot attend
to one thing unless we are conscious of it. The second is that attention, like
consciousness, is a unitary system.More and more, cognitive psychology is coming
to recognize that attention operates at an unconscious level. For instance,
people often are not conscious of where they have moved their eyes. Along with
this recognition has come the realization that attention is multifaceted (e.g.,
Pashler, 1995). We have seen that it makes sense to separate auditory attention
from visual attention and attention in perceptual processing from attention in
executive control from attention in response generation. The brain consists of a
number of parallel processing systems for the various perceptual systems,
motor systems, and central cognition. Each of these parallel systems seems to
suffer bottlenecks—points at which it must focus its processing on a single
thing. Attention is best conceived as the processes by which each of these
systems is allocated to potentially competing information-processing demands.
The amount of interference that occurs among tasks is a function of the overlap
in the demands that these tasks make on the same systems.
1. The chapter discussed how listening to one spoken
message makes it difficult to process a second spoken
message. Do you think that listening to a conversation
on a cell phone while driving makes it harder to process
other sounds like a car horn honking?
2. Which search should produce greater parietal activation:
searching Figure 3.13a for a T or searching Figure 3.13b
for a T?
3. Describe circumstances where it would be advantageous
to focus one’s attention on an object rather than
a region of space, and describe circumstances where the
opposite would be true.
4. We have discussed how automatic behaviors can
intrude on other behaviors and discussed how some
aspects of driving have become automatic. Consider
the situation in which a passenger in the car is a skilled
driver and has automatic aspects of driving evoked by
the driving experience. Can you think of examples
where automatic aspects of driving seem to affect a
passenger’s behavior in a car? Might this help explain
why having a conversation with a passenger in a car
is not as distracting as having a conversation over a
cell phone?
Questions for Thought
Key Terms
anterior cingulate cortex
(ACC)
attention
attenuation theory
automaticity
binding problem
central bottleneck
dichotic listening task
dorsolateral prefrontal
cortex (DLPFC)
early-selection theories
executive control
feature-integration theory
filter theory
goal-directed attention
illusory conjunction
inhibition of return
late-selection theories
object-based attention
perfect time-sharing
serial bottleneck
space-based attention
stimulus-driven attention