Cognitive Psychology and Its Implications, Ch. 3

profilebenisd
chapter_3_attention_and_performance.docx

3Attention and Performance

Chapter 2 described how the human visual system and other perceptual systems

simultaneously process information from all over their sensory fields. However, there

are limits on how much we can do in parallel. There are points at which we can attend

to only one spoken message or one visual object at a time. This chapter addresses how

higher level cognition determines what to attend to

In this chapter, we will address these questions: • In a busy world filled with sounds, how do we select what to listen to? • How do we find meaningful information within a complex visual scene? • What role does attention play in putting visual patterns together as

recognizable objects? • How do we coordinate parallel activities like driving a car and holding

a conversation?

Serial Bottlenecks

Psychologists have proposed that there are serial bottlenecks in human information

processing, points at which it is no longer possible to continue processing

everything in parallel. For example, it is generally accepted that there are limits to

parallelism in the motor systems. Although most of us can perform separate

actions simultaneously when the actions involve different motor systems (such as

walking and chewing gum), we have difficulty in getting one motor system to do

two things at once. Thus, even though we have two hands, we have only one

system for moving our hands, so it is hard to get our two hands to move in different

ways at the same time. Think of the familiar problem of trying to pat your

head while rubbing your stomach. It is hard to prevent one of the movements

from dominating—one tends to wind up rubbing one’s head or patting one’s

63

Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 63

64 | Attention and Performance

stomach.1 The many human motor systems—one for moving feet, one for moving

hands, one for moving eyes, and so on—can and do work independently and

separately, but it is difficult to get any one of these systems to do two things at

the same time.

One question that has occupied psychologists is how early do the bottlenecks

occur: before we perceive the stimulus, after we perceive the stimulus but before

we think about it, or only just before motor action is required? Common sense

suggests that there must be a certain amount of serial thinking before motor

action can occur. For instance, we find it basically impossible to add two digits

and multiply them simultaneously. Still, there remains the question of just where

the bottlenecks in information processing lie. Various theories about when they

happen are referred to as early-selection theories or late-selection theories,

depending on when they propose that bottlenecks take place. This question is one

of those examined by psychologists who are interested in studying attention.

Wherever there is a bottleneck, our cognitive processes must select which pieces

of information to attend to and which to ignore.A related research question concerns

how we select what to attend to.

In addition to the question of where information selection takes place in the

information-processing stream, researchers have asked the question of what factors

determine what we attend to. A major distinction is made

between goal-directed factors (sometimes called endogenous

control) and stimulus-driven factors (sometimes

called exogenous control). To illustrate the distinction,

Corbetta and Shulman (2002) ask us to imagine ourselves

at the El Prado Museum in Madrid, looking at Bosch’s

painting The Garden of Earthly Delights (see Color

Plate 3.1). Initially, our eyes will be probably be drawn

to large, salient objects like the instrument in the center

of picture. This would be an instance of stimulus-driven

attention—it is not that we wanted to attend to this; it

just grabbed our attention. However, our guide may

start to comment on a “small animal playing a musical

instrument.” Now we have a goal and will direct our

attention over the picture to find the object being described.

Continuing their story, Corbetta and Shulman

ask us to imagine that we hear an alarm system starting to

ring in the next room. Now a stimulus-driven factor has

intervened, and our attention will be drawn away fromour

goal of understanding our guide and the picture and switch

to the adjacent room. Corbetta and Shulman argue that somewhat different brain

systems control goal-directed attention versus stimulus-driven attention. For instance,

neural imaging evidence suggests that the goal-directed attentional system is

more left lateralized, whereas the stimulus-driven system is more right lateralized.

The brain regions that select information to process can be distinguished (to

an approximation) from those that process the information selected. Figure 3.1

1 Drummers (including my son) are particularly good at doing this—I definitely am not a drummer. This

suggests that the real problem might be motor timing.

Brain Structures

Parietal cortex: attends

to locations and objects

Motor cortex:

controls hands

Dorsolateral prefrontal

cortex: directs central

cognition

Auditory cortex:

processes auditory

information

Extrastriate cortex:

processes visual

information

Anterior cingulate:

(midline structure)

monitors conflict

FIGURE 3.1 A representation of

some of the brain areas involved

in attention and some to the

perceptual and motor regions

they control. The parietal regions

are particularly important in

directing perceptual resources.

The prefrontal regions (dorsolateral

prefrontal cortex, anterior

cingulate) are particularly

important in executive control.

Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 64

Auditory Attention | 65

highlights the parietal cortex, which influences information processing in regions

such as the visual cortex and auditory cortex. It also highlights prefrontal

regions that influence processing in the motor area and more posterior regions.

These prefrontal regions include the dorsolateral prefrontal cortex and, well

below the surface, the anterior cingulate cortex. The regions that do the selecting

and direct the processing are responsible for attentional control. As this chapter

proceeds, it will elaborate on the research involving the various brain regions in

Figure 3.1.

Attentional systems select information to process at serial bottlenecks where

it is no longer possible to do things in parallel.

Auditory Attention

Some of the early research on attention was concerned with auditory attention.

Much of this research centered on the dichotic listening task. In a typical dichotic

listening experiment, illustrated in Figure 3.2, participants wear a set of

headphones. They hear two messages at the same time, one entering each ear,

and are asked to “shadow” one of the two messages (i.e., repeat back the words

from one message only). Most participants are able to attend to one message

and tune out the other.

Psychologists (e.g., Cherry, 1953; Moray, 1959) have discovered that very

little information about the unattended message is processed in a dichotic

listening task. After hearing the messages, participants report that they can

tell whether the unattended message was a human voice or a noise; whether

... and then John turned rapidly toward ...

ran house ox cat

and, um, John turned . . .

FIGURE 3.2 A typical dichotic listening task. Different messages are presented to the left and

right ears, and the participant attempts to “shadow” the message entering one ear. (From Lindsay &

Norman, 1977. Reprinted by permission of the publisher. © 1977 by Academic Press.)

Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 65

66 | Attention and Performance

the human voice was male or female; and whether the sex of the speaker

changed during the test. This limited information is often all that they can

report. They cannot tell what language was spoken or remember any of the

words, even if the same word was repeated over and over again. An analogy is

often made between performing this task and being at a cocktail party, where a

guest tunes in to one message (a conversation) and filters out others. This is an

example of goal-directed processing—the listener selects the message to be

processed. However, to return to the distinction between goal-directed and

stimulus-driven processing, important stimulus information can disrupt our

goals.We have probably all experienced the situation in which we are listening

intently to one person and hear our name mentioned by someone else. It is

very hard in this situation to keep your attention on what the original speaker

is saying.

The Filter Theory

Broadbent (1958) proposed a particular early-selection theory called the filter

theory to account for these results. His basic assumption was that sensory

information comes through the system until some bottleneck is reached. At that

point, a person chooses which message to process on the basis of some physical

characteristic. The person is said to filter out the other information. In a

dichotic listening task, it was assumed that the message to each ear was registered

but that at some point the participant selected one ear to listen with. At

a cocktail party, we pick which speaker to follow on the

basis of physical characteristics, such as the pitch of the

speaker’s voice.

A crucial feature of Broadbent’s original filter model is

its proposal that we select a message to process on the basis

of physical characteristics such as ear or pitch. This hypothesis

made a certain amount of neurophysiological

sense. Messages entering each ear arrive on different

nerves. Nerves also vary in which frequencies they carry

from each ear. Thus, we might imagine that the brain, in

some way, selects certain nerves to “pay attention to.”

People can certainly choose to attend to a message on

the basis of its physical characteristics. There is evidence,

however, that we can also select messages to process on the

basis of their semantic content. In one study, Gray and

Wedderburn (1960), who at the time were undergraduate

students at Oxford University, demonstrated that participants

were successful in following a message that jumped

back and forth between the ears. Figure 3.3 illustrates the participants’ task

in their experiment. Suppose part of the meaningful message that participants

were to shadow was dogs scratch fleas. The message to one ear might be

dogs six fleas, whereas the message to the other might be eight scratch two.

Instructed to shadow the meaningful message, participants would report

dogs scratch fleas. Thus, participants can shadow a message on the basis of

dogs six fleas . . .

. . . eight scratch two

dogs scratch fleas . . .

FIGURE 3.3 An illustration

of the shadowing task in the

Gray and Wedderburn (1960)

experiment. The participant follows

the meaningful message as

it moves from ear to ear. (After

Klatzky, 1975. Adapted by permission of

the publisher. © 1975 by W. H. Freeman.)

Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 66

Auditory Attention | 67

meaning rather than on the basis of what each ear

physically hears.

Treisman (1960) looked at a situation in which

participants were instructed to shadow a particular

ear (Figure 3.4). The message in the ear to be

shadowed was meaningful up to a certain point;

then it turned into a random sequence of words.

Simultaneously, the meaningful message switched

to the other ear—the one to which the participant

had not been attending. Some participants switched

ears, against instructions, and continued to follow

the meaningful message. Others continued to follow

the shadowed ear. Thus, it seems that sometimes

people use the physical ear to select which message

to follow, and sometimes they choose semantic

content.

Broadbent’s filter model proposes that we use physical features, such as ear or

pitch, to select one message to process, but it has been shown that people can

also use the meaning of the message as the basis for selection.

The Attenuation Theory and the Late-Selection Theory

To account for these kinds of results, Treisman (1964) proposed a modification

of the Broadbent model that has come to be known as the attenuation theory.

This model hypothesized that certain messages would be weakened but not

filtered out entirely on the basis of their physical properties. Thus, in a dichotic

listening task, participants would minimize the signal from the unattended ear

but not eliminate it. Semantic selection criteria could apply to all messages,

whether they were attenuated or not. If the message were attenuated, it would

be harder to apply these selection criteria, but it would still be possible, as in the

Gray and Wedderburn (1960) experiment. Treisman (personal communication,

1978) emphasized that in her 1960 experiment, most participants actually

continued to shadow the prescribed ear. It was easier to follow the message that

is not being attenuated than to apply semantic criteria to switch attention to the

attenuated message.

An alternative explanation was offered by J. A. Deutsch and D. Deutsch

(1963) in their late-selection theory. They proposed that all the information is

processed completely without attentuation. Their hypothesis was that the

capacity limitation is in the response system, not the perceptual system. They

claimed that people can perceive multiple messages but that they can shadow

only one message at a time. Thus, people need some basis for selecting which

message to shadow. If they use meaning as the criterion (either according to or

in contradiction to instructions), they will switch ears to follow the message. If

they use the ear of origin in deciding what to attend to, they will shadow the

chosen ear.

I SAW THE GIRL/Song was wishing . . .

The to-be-shadowed ear

I SAW THE GIRL JUMPING . . .

. . . me that bird

JUMPING IN THE STREET.

FIGURE 3.4 An illustration of

the Treisman (1960) experiment.

The meaningful message moves

to the other ear, and the

participant sometimes

continues to shadow it against

instructions. (After Klatzky, 1975.

Adapted by permission of the publisher.

© 1975 by W. H. Freeman.)

Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 67

The difference between the two theories is illustrated in Figure 3.5. Both

models assume that there is some filter or bottleneck in processing. Treisman’s

theory (Figure 3.5a) assumes that the filter selects which message to attend to,

whereas Deutsch and Deutsch’s theory (Figure 3.5b) assumes that the filter

occurs after the perceptual stimulus has been analyzed for verbal content.

Treisman and Geffen (1967) tried to address the difference between these two

theories. They used a dichotic listening task in which participants had to

shadow one message and also had to process both messages for a target word. If

they heard the target word, they were to signal by tapping. According to the

Deutsch and Deutsch late-selection theory, messages from both ears would get

through and participants should have been able to detect the critical word

equally well in either ear. In contrast, the attenuation theory predicted much

less detection in the unshadowed ear because the message would be attenuated.

In the experiment, participants detected 87% of the target words in the shadowed

ear and only 8% percent in the unshadowed ear. Other evidence consistent

with the attenuation theory was reported by A. M. Treisman and Riley

(1969) and by Johnston and Heinz (1978).

There is neural evidence for a version of the attenuation theory that asserts

that there is both enhancement of the signal coming from the attended ear and

attenuation of the signal coming from the unattended ear. The primary auditory

area of the cortex (see Figure 3.1) shows an enhanced response to auditory signals

coming from the ear the listener is attending to and a decreased response to

signals coming from the other ear. Through ERP recording,Woldorff et al. (1993)

showed that these responses occur between 20 and 50 ms after stimulus onset.

68 | Attention and Performance

FIGURE 3.5 Treisman and

Geffen’s illustration of attentional

limitations produced by

(a) Treisman’s (1964) attenuation

theory and (b) Deutsch and

Deutsch’s (1963) late-selection

theory. (From Treisman & Geffen, 1967.

Reprinted by permission of the publisher.

© 1967 by the Quarterly Journal of

Experimental Psychology.)

(a) (b)

Responses

Selection and

organization of

responses

Analysis of

verbal content

Perceptual

filter

Input messages

1 2

1 2

Responses

Selection and

organization of

responses

Analysis of

verbal content

Input messages

Response filter

1 2

1 2

Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 68

Visual Attention | 69

The enhanced responses occurmuch sooner in auditory processing than the point

at which the meaning of the message can be identified. There is also evidence for

enhancement of the message in the auditory cortex on the basis of features other

than location. For instance, Zatorre, Mondor, and Evans (1999) found in a PET

study that when people attend to a message on the basis of pitch, there is similar

enhancement (registered as increased activation) in the auditory cortex. This

study also found increased activation in the parietal areas that direct attention.

Although auditory attention can enhance processing in the primary auditory

cortex, there is no evidence of reliable effects of attention on earlier portions

of auditory processing, such as in the auditory nerve or in brain-stem

processing (Picton & Hillyard, 1974). The various results we have reviewed

suggest that the primary auditory cortex is the earliest area to be influenced by

attention. It should be stressed that the effects at the auditory cortex are a

matter of attenuation and enhancement. Messages are not completely filtered

out and so it is still possible to select them at later points of processing.

Attention can enhance or reduce the magnitude of response to an auditory

signal in the primary auditory cortex.

Visual Attention

The bottleneck in visual information processing is even more apparent than the

one in auditory information processing. As we saw in Chapter 2, the retina

varies in acuity, with the greatest acuity in a very small area called the fovea.

Although the human eye registers a large part of the visual field, the fovea registers

only a small fraction of that field. Thus, in choosing where to focus our

vision, we also choose to devote most of our visual processing resources to a

particular part of the visual field, and we attenuate the resources allocated to

processing other parts of the field. Usually, we are attending to that part of the

visual field on which we are focusing. For instance, as we read, we move our

eyes so that we are fixating the words we are attending to.

The focus of visual attention is not always identical with the part of the

visual field being processed by the fovea, however. People can be instructed to

fixate on one part of the visual field (making that part the focus of the fovea)

and attend to another, nonfoveal region of the visual field.2 In one experiment,

Posner, Nissen, and Ogden (1978) had participants focus on a constant point

and then presented them with a stimulus 7° to the left or the right of the

fixation point. In some trials, participants were told on which side the stimulus

was likely to occur; in other trials, there was no such warning.When there was a

warning, it was correct 80% of the time—but 20% of the time the stimulus

appeared on the unexpected side. The researchers monitored eye movements

and included only those trials in which the eyes had stayed on the fixation

2 This is what quarterbacks are supposed to do when they pass the football, so that they don’t “give away” the

position of the intended receiver.

Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 69

point. Figure 3.6 shows the time required to judge the

stimulus if it appeared in the expected location (80% of

the time), if the participant had not been given a cue

(50% of the time), and if it appeared in the unexpected

location (20% of the time). Participants were able to

shift their attention from where their eyes were fixated:

Their responses to the stimuli were faster when the stimulus

appeared in the expected location and slower when

it appeared in the unexpected location.

Posner, Snyder, and Davidson (1980) found that people

can attend to regions of the visual field as far as 24° from

the fovea. Although visual attention can be moved without

accompanying eye movements, people usually do move

their eyes, so that the fovea processes the portion of the

visual field to which they are attending. Posner (1988)

pointed out that successful control of eye movements

requires us to attend to places outside the fovea. That is, we

must attend to and identify an interesting nonfoveal region

so that we can guide our eyes to fixate on that region to

achieve the greatest acuity in processing it. Thus, a shift of

attention often precedes the corresponding eye movement.

To process a complex visual scene, we must move our attention around in

the visual field to track the visual information. This process is like shadowing a

conversation. Neisser and Becklen (1975) performed the visual analog of the

auditory shadowing task. They had participants observe two videotapes superimposed

over each other. One was of two people playing a hand-slapping game,

the other of some people playing a basketball game. Figure 3.7 shows how the

situation appeared to the participants. They were instructed to pay attention

to one of the two films and to watch for odd events such as the two players in

the hand-slapping game pausing and shaking hands. Participants were able to

monitor one film successfully and reported filtering out the other.When asked

to monitor both films for odd events, the participants experienced great difficulty

and missed many of the critical events.

70 | Attention and Performance

Reaction time (ms)

Unexpected No expectation Expected

Condition

320

300

280

260

240

220

FIGURE 3.6 The results of an

experiment to determine how

people react to a stimulus that

occurs 7° to the left or right of

the fixation point. The graph

shows participants’ reaction

times to expected, unexpected,

and neutral (no expectation)

signals. (From Posner et al., 1978.

Reprinted by permission of the publisher.

© 1978 by Erlbaum.)

(a) (b) (c)

FIGURE 3.7 Frames from the two films used by Neisser and Becklen in their visual analog of the

auditory shadowing task. (a) The “hand-game” film; (b) the basketball film; and (c) the two figures

superimposed. (From Neisser & Becklen, 1975. Reprinted by permission of the publisher. © 1975 by Academic Press.)

Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 70

Visual Attention | 71

As Neisser and Becklen (1975) noted, this situation

involved an interesting combination of the use

of physical cues and the use of content cues. Participants

moved their eyes and focused their attention

in such a way that the critical aspects of the monitored

event fell on their fovea and the center of their

attentive spotlight. On the other hand, the only way

they knew how to move their eyes to follow an event

was by making reference to the content of the event

they were processing. Thus, physical cues facilitated

their processing of the critical film, which in turn

facilitated extracting content so they would know

where to move their eyes.

Figure 3.8 shows examples of the overlapping

stimuli used in an experiment by O’Craven,

Downing, and Kanwisher (1999) to study the neural

consequences of attending to one object or the other.

Participants in their experiment saw a series of pictures

that consisted of faces superimposed on houses.

They were instructed to either look for repetition of

the same face in the series or repetition of the same house. Recall from Chapter 2

that there is a region of the fusiformgyrus, the fusiformface area, which becomes

active when people are observing faces. There is another area within the temporal

cortex, the parahippocampal place area, which becomes more active when people

are observing places.What is special about these pictures is that they consisted of

both places and locations.Which region would become active—the fusiform face

area or the parahippocampal place area? As the reader might suspect, the answer

depended on what the participant was attending to.When participants were looking

for repetition of faces, the fusiform face area became more active; when they

were looking for repetition of places, the parahippocampal place area became

more active. Attention was able to select which region of the temporal cortex was

engaged in the processing of the stimulus.

People can focus their attention on parts of the visual field and move their

focus of attention to process what they are interested in

The Neural Basis of Visual Attention

It appears that the neural mechanisms underlying visual attention are very

similar to those underlying auditory attention. Just as auditory attention directed

to one ear enhances the cortical signal from that ear, visual attention directed to

a spatial location appears to enhance the cortical signal from that location. If

a person attends to a particular spatial location, a distinct neural response

(detected using ERP records) in the visual cortex occurs within 70 to 90 ms after

the onset of the stimulus. On the other hand, when a person is attending to a

particular object (attending to chairs and not tables, say), rather than to a particular

location in space, we do not see a response for more than 200 ms. Thus, it

FIGURE 3.8 An example of a

picture used in the study of

O’Craven et al. (1999). When

the face is attended, there is

activation in the fusiform face

area, and when the house is

attended, there is activation

in the parahippocampal

place area.

Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 71

appears to take more effort to direct visual attention on the basis of content than

on the basis of physical features, just as is the case with auditory attention.

Mangun, Hillyard, and Luck (1993) had participants fixate on the center of a

computer screen, then judge the lengths of bars presented in positions different

from the fixation location (upper left, lower left, upper right, and lower right).

They documented how much an ERP recording was affected at various locations

across the back of the scalp. Figure 3.9 shows the distribution of scalp activity

when a participant was attending to one of the four different regions of the visual

array (while fixating on the center of the screen). Consistent with the topographic

organization of the visual cortex, there was greatest activity over the side

of the scalp opposite the side of the visual field where the object appeared. Recall

from Chapters 1 and 2 (see Figure 2.6) that the visual cortex (at the back of the

head) is topographically organized, with each visual field (left or right) represented

in the opposite hemisphere. Thus, it appears that there is enhanced neural

processing in the portion of the visual

cortex corresponding to the location of

visual attention.

A study by Roelfsema, Lamme, and

Spekrejse (1998) illustrates the impact of

visual attention of information processing

in the primary visual area of the

macaque monkey. They trained monkeys

to perform the rather complex task

illustrated in Figure 3.10. A trial in the

experiment would begin with the monkey

fixating on a particular stimulus in

the visual field, as in part (a) of the

figure. Then two curves would appear, as

72 | Attention and Performance

P1 attention effect

(current density)

P1 P1 P1 P1

Stimulus

+ + + +

FIGURE 3.9 Results from an experiment by Mangun, Hillyard, and Luck. Distribution of scalp

activity recorded by ERP when a participant was attending to one of the four different regions

of the visual array depicted in the right-hand column while fixating on the center of the screen.

The greatest activity was recorded over the side of the scalp opposite the side of the visual

field where the object appeared, confirming that there is enhanced neural processing in portions

of the visual cortex corresponding to the location of visual attention. (After Mangun et al., 1993.

Adapted by permission of the publisher. © 1993 by MIT Press.)

(a)

Fixation (300 ms)

(b)

Stimulus (600 ms)

Receptive field

(c)

Saccade

Fixation point

FIGURE 3.10 The experimental

procedure in Roelfsema et al.

(1998): (a) The monkey fixates

the start point. (b) Two curves

are presented, one of which

links the start point to a target

point. (c) The monkey saccades

to the target point. The box

represents the receptive field in

the primary visual cortex V1.

Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 72

Visual Attention | 73

TWLN

XJBU

UDXI

HSFP

XSCQ

SDJU

PODC

ZVBP

PEVZ

SLRA

JCEN

ZLRD

XBOD

PHMU

ZHFK

PNJW

CQXT

GHNR

IXYD

QSVB

GUCH

OWBN

BVQN

FOAS

ITZN

in part (b), only one of which connected the original object to the fixation

point. The monkey had to remain fixated on this point for 600 ms and then

perform a saccade (an eye movement) to the end of the curve that connected

the original point (part c).While monkeys performed this task, Roelfsema et al.

recorded from cells in the monkey’s primary visual cortex (where cells with

receptive fields like those in Figure 2.8 are found). Indicated by the square in

Figure 3.10 is a receptive field of one of these cells. It shows increased response

when a line falls on that part of the visual field and so responds when the curve

appears that crosses it. The cell responds more during the 600-ms waiting

period when the receptive is on the curve that connects the fixation point to the

destination than when it is on the other curve. The monkey must shift its

attention along this curve to determine where to saccade to, and this shift of

attention causes the line detectors in V1 to respond more strongly.

When people attend to a particular spatial location, there is greater neural

processing in portions of the visual cortex corresponding to that location.

Visual Search

People are able to select stimuli to attend to, either in the visual or auditory

domain, on the basis of physical properties and, in particular, on the basis of

location. Although selection based on simple features can occur early and quickly

in the visual system, not everything people look for can be defined in terms of

simple features. How do they find an object with particular higher order properties,

such as the face of a friend in a crowd? In such cases, it seems that they must

search through the faces in the crowd, looking for one that has the desired

properties. Much of the research on visual attention has focused on how people

perform such searches. Rather than study how people find faces in a crowd, however,

researchers have tended to use simpler material.

Figure 3.11, for instance, shows a portion of the display

that Neisser (1964) used in one of the early studies. Try

to find the first K in the set of letters displayed.

Presumably, you tried to find the K by going

through the letters row by row, looking for the target.

Figure 3.12 graphs the average time it took participants

in Neisser’s experiment to find the letter as a

function of which row it appeared in. The slope of

the best-fitting function in the graph is about 0.6,

which implies that participants took about 0.6 s to

scan each line. When people engage in such searches,

they appear to be allocating their attention intensely

to the search process. For instance, brain-imaging

experiments have found strong activation in the

parietal cortex during such searches (see Kanwisher &

Wojciulik, 2000, for a review).

Although a search can be intense and difficult, it is

not always that way. Sometimes we can find what we

are looking for without much effort. If we know that

0

0

10

20

30

40

10 20 30 40 50

Time (s)

Position of critical item (line number)

FIGURE 3.12 The time required to find a target letter in the

array shown in Figure 3.9 as a function of the line number in

which it appears. (After Neisser, 1964. Adapted by permission of the publisher.

© 1964 by Scientific American.)

FIGURE 3.11 A representation

of lines 7–31 of the letter array

used in Neisser’s search experiment.

(After Neisser, 1964. Adapted by

permission of the publisher. © 1964 by

Scientific American.)

Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 73

our friend is wearing a bright red jacket, it can be relatively

easy to find him or her in the crowd, provided

that no one else is wearing a bright red jacket. Our

friend will just pop out of the crowd. Indeed, if there

were just one red jacket in a sea of white jackets it

would probably pop out even if we were not looking

for it—an instance of stimulus-driven attention. It

seems that if there is some distinctive feature in an

array, we can find it without a search.

Treisman studied this sort of pop-out. For instance,

Treisman and Gelade (1980) instructed participants

to try to detect a T in an array of 30 I’s and Y’s

(Figure 3.13a). They reasoned that participants could

do this simply by looking for the crossbar feature of

the T that distinguishes it from all I’s and Y’s. Participants

took an average of about 400 ms to perform

this task. Treisman and Gelade also asked participants

to detect a T in an array of I’s and Z’s (Figure 3.13b).

In this task, they could not use just the vertical bar

or just the horizontal bar of the T; they would have

to look for the conjunction of these features and

perform the feature combination required in pattern

recognition. It took participants more than 800 ms,

on average, to find the letter in this case. Thus, a task

requiring them to recognize the conjunction of features

took about 400 ms longer than one in which

perception of a single feature was sufficient. Moreover,

when Treisman and Gelade varied the number

of letters in the array, they found that participants were much more

affected by array size in the task that required recognition of the

conjunction of features. Figure 3.14 shows these results.

It is necessary to search through a visual array for an object only

when a unique visual feature does not distinguish that object.

The Binding Problem

As discussed in Chapter 2, there are different types of neurons in

the visual system that respond to various features, such as colors,

lines at various orientations, and objects in motion. A single object

in our visual field will involve a number of features; for instance, a

red vertical line combines the vertical feature and the red feature.

The fact that different features of the same object are represented

by different neurons gives rise to a logical question: How are these

features put back together to produce perception of the object?

This would not be much of a problem if there were a single object

in the visual field. We could assume that all the features belonged

74 | Attention and Performance

Array size (number of items)

T in I, Z

Reaction time (ms)

1

0

400

800

1200

5 15 30

T in I, Y

(a)

(b)

FIGURE 3.13 Stimuli used by Treisman and Gelade to determine

how people identify objects in the visual field. They found that it is

easier to pick out a target letter (T ) from a group of distracter letters

(I ’s and Y ’s) if (a) the target letter has a feature that makes it

easily distinguishable from the distracters than if (b) the same target

letter is in an array of distracters (I ’s and Z ’s) that offer no obvious

distinctive features. (After Treisman & Gelade, 1980. Adapted by permission of the

publisher. © 1980 by Cognitive Psychology.)

FIGURE 3.14 Results from the Treisman and

Gelade experiment. The graph plots the average

reaction times required to detect a target letter as

a function of the number of distracters and whether

the distracters contain separately all the features

of the target. (After Treisman & Gelade, 1980. Adapted

by permission of the publisher. © 1980 by Cognitive Psychology.)

Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 74

Visual Attention | 75

to that object. But what if there are multiple objects in the field? For instance,

suppose there were just two objects: a red vertical bar and a green horizontal bar.

These two objects might result in the firing of neurons for red, neurons for

green, neurons for vertical lines, and neurons for horizontal lines. If these firings

were all that occurred, though, how would the visual system know it saw a red

vertical bar and a green horizontal bar rather than a red horizontal bar and a

green vertical bar? The question of how the brain puts together various features

in the visual field is referred to as the binding problem.

Treisman (e.g., Treisman & Gelade, 1980) developed her feature-integration

theory as an answer to the binding problem. She proposed that people must focus

their attention on a stimulus before they can synthesize its features into a pattern.

For instance, in the example just given, the visual system can first direct its attention

to the location of the red vertical bar and synthesize that object, then direct

its attention to the green horizontal bar and synthesize that object. According to

Treisman, people must search through an array when they need to synthesize

features to recognize an object (for instance, when trying to identify a K, which

consists of a vertical line and two diagonal lines). When there is a single unique

feature, such as a red jacket or a line at a particular orientation, we can move our

attention directly to the object and recognize it, thus avoiding the need to search.

The binding problem is not just a hypothetical dilemma—it is something

that humans actually suffer from. One source of evidence comes from studies

of illusory conjunctions in which people report combinations of features

that did not occur. For instance, Treisman and Schmidt (1982) looked at what

happens to feature combinations when the stimuli are out of the focus of attention.

Participants were asked to report the identity of two black digits flashed in

one part of the visual field. This was their primary task, and it was where their

attention was focused. In another part of the visual field, letters in various

colors were presented. Thus, participants might be presented with a pink T,

a yellow S, and a blue N in the unattended portion of the field. After they

reported the numbers, participants were asked to report any letters they had

seen and the colors of these letters. They reported seeing illusory conjunctions

of features (e.g., a pink S) almost as often as they reported seeing correct combinations.

Thus, it appears that we are able to combine features into an accurate

perception only when our attention is focused on an object. Otherwise, we

perceive the features but may well combine them into a perception of objects

that were never there. Although rather special circumstances are required to

produce illusory conjunctions in an ordinary person, there are certain patients

with damage to the parietal cortex who are particularly prone to such illusions.

For instance, one patient studied by Friedman-Hill, Robertson, and Treisman

(1995) confused which letters were presented in which colors even when shown

the letters for as long as 10 seconds.

A number of studies have been conducted on the neural mechanisms

involved in binding together the features of a single object. The neurons in

visual area V4 have large receptive fields (several degrees of visual angle), and

multiple objects in a display may be within the visual field of a single neuron.

Luck, Chelazzi, Hiillyard, and Desimone (1997) trained macaque monkeys to

fixate on a certain part in the visual field and recorded neurons around this

Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 75

region. They found neurons that were specific to particular types of

objects. For instance, they found a cell that responded to a blue

vertical bar.What happens when a blue vertical bar and a green horizontal

bar are presented both within the receptive field of this cell?

If the monkey attended to the blue vertical bar, the rate of response

of the cell will remain at the same level as it would have had there

been only a blue vertical bar alone. On the other hand, if the monkey

focused on the green horizontal bar, the rate of firing of this

same cell will be greatly depressed. Thus, the same stimulus (blue

vertical bar plus green horizontal bar) can evoke different responses

depending on which object is attended to. It is speculated that this

phenomenon occurs because attention suppresses responses to all

features in the receptive field except those at the attended location.

Similar results have been obtained in fMRI experiments with humans.

Kastner, DeWeerd, Desimone, and Ungerleider (1998) measured the fMRI

signal in visual areas that responded to stimuli presented in one region of the

visual field. They found that when attention was directed away from that region,

the fMRI response to stimuli in that region decreased; but when attention was

focused on that region, the fMRI response was maintained. These experiments

indicate enhanced neural processing of attended objects and locations.

A striking demonstration of the effects of sustained attention was reported

by Simons and Chabris (1999). They asked participants to watch a video in

which a team dressed in black tossed a basketball back and forth and a team

dressed in white did the same (Figure 3.15). Participants were instructed to

count either the number of times the team in black tossed the ball or the number

of times the team in white did so. Presumably, in one condition participants

were looking for events involving the team in black and in the other for events

involving the team in white. Because the players were intermixed, the task was

difficult and required sustained attention. In the middle of the game, a person in

a black gorilla suit walked through the room. When participants were tracking

the team in white, they noticed the black gorilla only 8% of the time; when they

were tracking the team in black, they noticed it 67% of the time. They were so

fixed on searching the video for events involving team members dressed in white

that they completely missed an event involving a black object. People passively

watching the video never miss the black gorilla. The actual video is currently

available from the demonstrations page of Simons’s Visual Cognition Lab:

http://viscog.beckman.uiuc.edu/djs_lab/demos.html

For feature information to be synthesized into a pattern, it must be in the

focus of attention.

Neglect of the Visual Field

We have discussed the evidence that visual attention to a spatial location results

in enhanced activation in the appropriate portion of the primary visual cortex.

The neural structures that control this shift of attention, however, appear to

be located elsewhere. Three areas of the monkey brain have been shown to be

76 | Attention and Performance

FIGURE 3.15 Single frame from

the movie used by Simons and

Chabris to demonstrate the

effects of sustained attention.

When participants were intent

on tracking the ball passed

among the players dressed in

T-shirts, they tended not to

notice the black gorilla walking

through the room. (From Simons &

Chabris, 1999. Reprinted by permission of

the publisher. © 1999 by Perception.)

Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 76

Visual Attention | 77

involved in controlling attention (S. E. Peterson, Robinson, &

Morris, 1987;Wurtz, Goldberg, & Robinson, 1980). These areas

are the superior colliculus, the posterior parietal lobe, and a

midbrain area known as the pulvinar. Damage to these areas in

human patients, particularly to the parietal lobe (see Figure 3.1),

has been shown to result in deficits in visual attention. For

instance, Posner, Walker, Friedrich, and Rafal (1984) showed

that patients with parietal lobe injuries have difficulty in disengaging

attention from one side of the visual field.

Damage to right parietal regions produces distinctive

patterns of deficit. Posner, Cohen, and Rafal (1982) studied the

attention deficit in one such patient. The patient was cued to

expect a stimulus to the left or right of the fixation point, and

80% of the time that is where the stimulus was present. However,

20% of the time the stimulus appeared in the unexpected

field. Figure 3.16 shows the time required to detect the stimulus

as a function of which visual field it was presented in and

which field had been cued. When the stimulus was presented

in the right field, the patient showed only a little disadvantage

if inappropriately cued. If the stimulus appeared in the left

field, however, the patient showed a large deficit if inappropriately cued. Because

the right parietal lobe processes the left visual field, damage to the right lobe

impairs its ability to draw attention back to the left visual field once attention is

focused on the right visual field. This sort of one-sided attentional deficit can be

temporarily created in normal individuals by presenting TMS to the parietal

cortex (Pascaul-Leone et al., 1994—see Chapter 1 for discussion of TMS).

A more extreme version of this attentional disorder

is called unilateral visual neglect. Patients

with damage to the right hemisphere completely

ignore the left side of the visual field, and patients

with damage to the left hemisphere ignore the

right side of the field. Figure 3.17 shows the performance

of a patient with damage to the right

hemisphere, which caused her to neglect the left

visual field (Albert, 1973). She had been instructed

to put slashes through all the circles. As can be

seen, she ignored the circles in the left part of her

visual field. Such patients will often behave peculiarly.

For instance, one patient failed to shave half

of his face (Sacks, 1985).

It seems that the right parietal lobes are involved

in allocating spatial attention in many

modalities, not just the visual (Zatorre et al.,

1999). For instance, when one attends to the

location of auditory or visual stimuli, there is

increased activation in the right parietal region.

It also appears that the right parietal lobes are

Latency (ms)

1400

1200

1000

800

600

400

Left

Field of presentation

Right

Cued for right field

Cued for left field

FIGURE 3.16 The attention

deficit shown by a patient with

right parietal lobe damage when

switching attention to the left

visual field. (From Posner, Cohen, & Rafal,

1982. Reprinted by permission of the

publisher. © by the Royal Society of London.)

FIGURE 3.17 The performance of a patient with damage to the right

hemisphere who had been asked to put slashes through all the circles.

Because of the damage to the right hemisphere, she ignored the

circles in the left part of her visual field. (From Ellis & Young, 1988. Reprinted by

permission of the publisher. © 1988 by Erlbaum.)

Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 77

more responsible for the spatial allocation of attention and that this is why right

parietal damage tends to produce such dramatic effects. Left parietal damage

tends to produce a subtler pattern of deficits. Robertson and Rafal (2000) argue

that the right parietal lobe is responsible for attention to such global features as

spatial location, whereas the left parietal region is responsible for directing

attention to local aspects of objects. Figure 3.18 is a striking illustration of the

different types of deficits associated with left and right parietal damage. Patients

were asked to draw the objects in Figure 3.18a. Patients with right parietal

damage (Figure 3.18b) were able to reproduce the specific components of the

picture but were not able to reproduce their spatial configuration. In contrast,

patients with left parietal damage (Figure 3.18c) were able to reproduce

the overall configuration, but not the detail. Similarly, brain-imaging studies

78 | Attention and Performance

(a) (b) (c)

FIGURE 3.18 (a) The pictures presented to patients with parietal damage. (b) Examples of

drawings made by patients with right-hemisphere damage. These patients could reproduce

the specific components of the picture but not their spatial configuration. (c) Examples of

drawings made by patients with left-hemisphere damage. These patients could reproduce

the overall configuration but not the detail. (After Robertson & Lamb, 1991. Adapted by permission of

the publisher. © 1991 by Cognitive Psychology.)

Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 78

Visual Attention | 79

have found more activation of the right parietal region when a person is

responding to global patterns and more activation of the left hemisphere when

a person is attending to local patterns (Fink et al., 1996;Martinez et al., 1997).

Parietal regions are responsible for the allocation of attention, with the right

hemisphere more concerned with global features and the left hemisphere

with local features.

Object-Based Attention

So far we have talked about space-based attention, where people allocate their

attention to a region of space. There is also evidence, though, for object-based

attention, where people focus their attention on particular objects rather than

regions of space. An experiment by Behrmann, Zemel, and Mozer (1998) is an

example of the research that shows people sometimes find it easier to

attend to an object than to a location. Figure 3.19 illustrates some of

the stimuli used in the experiment. Participants were asked to judge

whether the numbers of bumps on the two ends of objects were the same.

The left column shows instances in which the numbers of bumps were

the same, the right column instances in which the numbers were not the

same. Participants made these judgments faster when the bumps were on

the same object (top and bottom rows in Figure 3.19) than when they

were on different objects (middle row). This result occurred despite the

fact that when the bumps were on different objects, they were located

closer together, which should have facilitated judgment. Behrmann et al.

argue that participants can shift attention to one object at a time, but not

one location at a time. Therefore, when the bumps were all on the same

object, participants did not need to shift their attention between objects.

Other evidence for object-centered attention involves a phenomenon

called inhibition of return. If we have looked at a particular region

of space, we find it harder to return our attention to that region. This

phenomenon also makes sense. If we are searching for something and have

already looked at a location, we would prefer our visual system to find other

locations to look at rather than return to an already searched location. If we

move our eyes to location A and then to location B, we are slower to return our

eyes to location A than to some new location C. This is also true when we move

our attention without moving our eyes (Posner, Rafal, Chaote, & Vaughn, 1985).

Tipper, Driver, and Weaver (1991) performed one demonstration of the

inhibition of return that also provided evidence for object-based attention. In

their experiments, participants viewed three squares in a frame, similar to what

is shown in each part of Figure 3.20. In one condition, the squares did not

move (unlike the moving condition illustrated in Figure 3.20, which we will

discuss in the next paragraph). The participants’ attention was drawn to one of

the outer squares by making it flicker. Attention was drawn back to the center

square 200 ms later by making that square flicker. A probe was then presented

in one of the two outer positions, and participants were instructed to press a

key indicating that they had seen the probe. On average, they took 420 ms to

(a) (d)

(e)

(c) (f)

(b)

FIGURE 3.19 Stimuli used in an

experiment by Behrmann, Zemel,

and Mozer to demonstrate that

it is sometimes easier to attend

to an object than to a location.

The left and right columns

indicate same and different

judgments, respectively; and

the rows from top to bottom

indicate the single-object,

two-object, and occluded

conditions, respectively. (From

Behrmann, Zemel, & Mozer, 1998. Reprinted

by permission of the publisher. © 1998

by the Journal of Experimental Psychology:

Human Perception and Performance.)

Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 79

see the probe when it occurred at the outer square that had not flickered and

460 ms when it occurred at the outer square that had. This 40-ms advantage is

an example of a spatially defined inhibition of return. People are slower to

move their attention to a location where it has already been.

Figure 3.20 illustrates the other condition of their experiment, in which the

objects were rotated around the screen after the flicker. By the end of the motion,

the object that had flickered on one side was now on the other side—the

two outer objects had traded positions. The question of interest was whether

participants would be slower to detect a target on the right (where the flickering

had been—which would indicate location-based inhibition) or on the left

(where the flickered object had ended up—which would indicate object-based

inhibition). The results showed that they were about 20 ms slower to detect an

object in the location that had not flickered but that contained the object that

80 | Attention and Performance

(a)

(b)

(c)

(d)

(e)

FIGURE 3.20 Examples of frames used in an experiment by Tipper, Driver, and

Weaver to determine whether inhibition of return would attach to a particular object

or to its location. Arrows represent motion. (a) Display onset, with no motion for

500 ms. After two moving frames, the three filled squares were horizontally aligned

(b), whereupon the cue appeared (one of the boxes flickered). Clockwise motion then

continued, with cueing in the center for the initial three frames (c–e). The outer

squares continued to rotate clockwise (d) until they were horizontally aligned (e),

at which point a probe was presented, as before. (From Tipper, Driver, & Weaver, 1991. Reprinted

by permission of the publisher. © 1991 by the Quarterly Journal of Experimental Psychology.)

Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 80

Central Attention: Selecting Lines of Thought to Pursue | 81

had flickered. Thus, their visual systems displayed an inhibition of return to the

same object, not the same location.

Another example of object-based attention comes from studies of visual

neglect. Earlier, we noted that some patients with damage to the right parietal

lobe have difficulty detecting information in the left side of the visual field

(see Figure 3.17). Researchers have identified a number of patients who neglect

the left side of objects regardless of which visual field these objects occur in

(Behrmann & Moscovitch, 1994; Driver, Baylis, Goodrich, & Rafal, 1994).

It seems that the visual system can direct attention either to locations in space

or to objects. Experiments like those just described indicate that the visual system

can track objects. On the other hand, there are many experiments in which people

direct their attention to regions of space where there are no objects (see Figure 3.6

for the results of such an experiment). It is interesting that the left parietal regions

seem to be more involved in object-based attention and the right parietal regions

in location-based attention. Patients with left parietal damage appear to have

deficits in focusing attention on objects (Egly, Driver, & Rafal, 1994), unlike the

location-based deficits that I have described in patients with right parietal damage.

Also, when participants without brain damage attend to objects rather than locations,

there is greater left parietal activation, as revealed by fMRI (Arrington, Carr,

Mayer, & Rao, 2000). This seems consistent with the earlier research we reviewed

(see Figure 3.18) showing that the right parietal region is responsible for attention

to global features and the left for attention to local features.

Visual attention can be directed either toward objects independent of their

location or toward locations independent of what objects are present.

Central Attention: Selecting Lines

of Thought to Pursue

So far, this chapter has considered how people allocate their attention to process

stimuli in the visual and auditory modalities. What about cognition after the

stimuli are attended to and encoded? How do we select which lines of thought to

pursue? Suppose we are driving down a highway and encode the fact that a dog is

sitting in the middle of the road.We might want to figure out why the dog is sitting

there, we might want to consider whether there is something we should do

to help the dog, and we certainly want to decide how best to steer the car to avoid

an accident. Can we do all these things at once? If not, how do we select the most

important problem of deciding how to steer and save the rest for later? It appears

that people allocate central attention to competing lines of thought in much the

same way they allocate perceptual attention to competing objects.

In many (but not all) circumstances, people are able to pursue only one line

of thought at a time. This section will describe two laboratory tasks: one in

which it appears that people have no ability to overlap two tasks and another

pair in which they appear to have almost total ability to do so. Then we will

address how people can develop the ability to overlap tasks and how they select

among tasks when they cannot or do not want to overlap them.

Anderson7e_Chapter_03.qxd 8/20/09 9:41 AM Page 81

The first experiment, which Mike Byrne and I did (Byrne & Anderson,

2001), illustrates the claim made at the beginning of the chapter about it being

impossible to multiply and add two numbers at the same time. Participants in

this experiment saw a string of three digits, such as “3 4 7.” There were two tasks

they might be asked to do:

• Task 1: Judge whether the first two digits add up to the third and press a

key with the right index finger if they do and another key with the left

index finger if they do not. • Task 2: Report verbally the product of the first and third numbers. In this

case, the answer is 21, because 3 _ 7 _ 21.

Participants either performed these two tasks individually or tried to do

both at once. Figure 3.21 compares the time required to do each task in the

single-task condition versus the time required for each task in the dual-task

condition. Participants took almost twice as long to do either task when they

had to perform the other as well. The illustration also displays the time participants

took to complete both tasks in the dual-task condition. (They answered

the multiplication problem first 59% of the time and the addition problem first

41% of the time.) The horizontal black line near the top of Figure 3.21 represents

the time they took to give the second answer, whichever it was. This line

reflects the mean time required to complete both tasks (1.99 s). Note that this

time is greater than the sum of the time for the verification task by itself (0.88 s)

82 | Attention and Performance

Verify

addition

Latency (ms)

Generate

multiplication

250

0

500

750

1000

1250

1500

1750

2000

2250

Stimulus: 3 4 7

Mean time to complete both tasks

Single task

Dual task

FIGURE 3.21 The results of

an experiment by Byrne and

Anderson to see whether people

can overlap two tasks. The bars

show the response times

required to solve two problems—

one of addition and one of

multiplication—when done by

themselves and when done

together. The results indicate

that the participants were not

able to overlap the addition

and multiplication computations.

(From Byrne & Anderson, 2001. Reprinted

by permission of the publisher. © 2001

by Psychological Review.)

Anderson7e_Chapter_03.qxd 8/20/09 9:42 AM Page 82

Central Attention: Selecting Lines of Thought to Pursue | 83

and the time for the multiplication task by itself (1.05 s). The extra time probably

reflects the cost of shifting between tasks (for a review, see Monsell, 2003). In

any case, it appears that the participants were not able to overlap the addition

and multiplication computations at all.

The second experiment, reported by Schumacher et al. (2001), illustrates

what is referred to as perfect time-sharing. The task was much simpler than

the Byrne and Anderson (2001) experiment. Participants simultaneously saw a

single letter on a screen and heard a tone. As in the first experiment, they had to

perform two tasks:

• Task 1: Press a left, middle, or right key according to whether the letter

occurred on the left, in the middle, or the right. • Task 2: Report “One,” “two,” or “three” according to whether the tone was

low, middle, or high in frequency.

Figure 3.22 compares the times required to do each task in the single-task

condition and the dual-task condition. As can be seen, these times are nearly

unaffected by the requirement to do the two tasks at once. There are many

differences between this task and the Byrne and Anderson task. Perhaps the most

apparent is the complexity of the tasks. Participants were able to do the individual

tasks in the second experiment in a few hundred milliseconds, whereas the

individual tasks in the first experiment took around a second. Thus, there was

significantly more thought required in the first experiment, and it is apparently

harder for people to engage in both streams of thought simultaneously. Also,

Location

discrimination

Response time (ms)

0

50

100

150

200

250

300

350

400

450

500

Tone

discrimination

Single task

Dual task

FIGURE 3.22 The results of

an experiment by Schumacher

et al. illustrating perfect

time-sharing. The bars show

the times required to perform

two tasks—a simple location

discrimination task and a tone

discrimination task—when done

by themselves and when done

together. The times were nearly

unaffected by the requirement

to do the two tasks at once,

indicating that the participants

achieved almost perfect

time-sharing. (From Schumacher et al.,

2001. Reprinted by permission of the

publisher. © 2001 by Psychological

Science.)

Anderson7e_Chapter_03.qxd 8/20/09 9:42 AM Page 83

participants in the second experiment achieved perfect time-sharing only after

five sessions of practice. There was no such practice in the first experiment.

Figure 3.23 presents an analysis of what occurred in the Schumacher et al.

(2001) experiment. It shows what was happening at various points in time in

five streams of processing: (1) perceiving the visual location of a letter, (2) generating

manual actions, (3) central cognition, (4) perceiving auditory stimuli,

and (5) generating speech. Task 1 involved visually encoding the location of the

letter, using central cognition to select which finger to press, and then performing

the actual finger movement. Task 2 involved detecting and encoding the tone,

using central cognition to select which word to say (“one,” “two,” or “three”), and

then generating the word and speaking it. The lengths of the boxes in Figure 3.23

represent estimates of the duration of each component based on human performance

studies. Each of these streams can go on in parallel with the others.

Thus, for instance, during the time the tone is being detected and encoded, the

location of the letter is being encoded (which happens much faster), a finger is

being selected by central cognition, and the motor system is starting to program

the action. Although all these streams can go on in parallel, within each stream

only one thing can happen at a time. This is a potential problem in the case of

the central cognition stream because central cognition must direct all activities.

In this case, it must serve both task 1 and task 2. In this experiment, however, the

length of time devoted to central cognition was so brief that the two tasks did

84 | Attention and Performance

Encode

letter

location

Vision

Manul action

Central cognition

Speech

Time (ms)

Streams of processing:

Task 2

Task 1

Audition

Select

action

Program key press

Detect and

encode tone

0 100 200 300 400

Generate speech

Select

action

FIGURE 3.23 An analysis of the timing of events in five streams of processing during execution

of the dual task in the Schumacher et al. (2001) experiment: (1) vision, (2) manual action,

(3) central cognition, (4) speech, and (5) audition.

Anderson7e_Chapter_03.qxd 8/20/09 9:42 AM Page 84

Central Attention: Selecting Lines of Thought to Pursue | 85

not contend for the resource. The five days of practice in this experiment played

a critical role in reducing the amount of time devoted to central cognition.

Although the discussion here has focused on bottlenecks in central cognition,

there can be bottlenecks in any of these modalities. People cannot attend

to two locations at once. Earlier, we reviewed evidence that they must shift their

attention across locations in the visual array serially if they must attend to more

than one location. Similarly, they can process only one speech stream at a time,

move their hands in one way at a time, or say one thing at a time. Even though

all these peripheral processes have bottlenecks, it is generally thought that bottlenecks

in central cognition can have the most significant effects, and they are

the reason we seldom find ourselves thinking about two things at once. This

bottleneck in central cognition is referred to as the central bottleneck.

People can process multiple perceptual modalities at once or execute actions

in multiple motor systems at once, but they cannot process multiple things

in a single system including central cognition.

Automaticity: Expertise Through Practice

In Chapter 9, we will discuss at some length how people become expert with

practice. The general effect of practice is to reduce the central cognitive component

of information processing. When one has practiced the central cognitive

component of a task so much that the task requires little or no thought, we say

phones. In contrast, listening to a radio or books on tape

does not interfere with driving. Strayer and Drews

suggest that the demands

of participating in a conversation

place more requirements

on central cognition. When

someone says something on

the cell phone, they expect an

answer and are unaware of

the driving conditions. Strayer

and Drews note that participating

in a conversation with

a passenger in the car is not

as distracting because the

passenger will adjust the conversation

to driving demands and even point out things

like exits to the driver.

Implications

Why is cell phone use and driving a dangerous combination?

Bottlenecks in information processing can have important

practical implications. A study by the Harvard

Center for Risk Analysis

(Cohen & Graham, 2003)

estimates that cell phone

distraction results in

2,600 deaths, 330,000

injuries, and 1.5 million instances

of property damage

in the United States

each year. Strayer and

Drews (2007) review the

evidence that people are

more likely to miss traffic

lights and other critical

information while talking on a cell phone. Moreover,

these problems are not any better with hands-free

Anderson7e_Chapter_03.qxd 8/20/09 9:42 AM Page 85

that doing the task is automatic. Automaticity is a matter of degree. A nice

example is driving. For experienced drivers in unchallenging conditions, driving

has become so automatic that they can carry on a conversation while driving

with little difficulty. Experienced drivers are much more successful at doing

secondary tasks like changing the radio (Wikman,Nieminen, & Summala, 1998).

Experienced drivers also often have the experience of traveling long stretches of

highway with no memory of what they did.

There have been a number of dramatic demonstrations in the psychological

literature of how practice can enable parallel processing. One demonstration

of the way practice affects attentional limitations is the study reported by

Underwood (1974) on the psychologist Neville Moray, who had spent many

years studying shadowing. During that time, Moray practiced shadowing a

great deal. Unlike most participants in experiments, he was very good at

reporting what was contained in the unattended channel. Through a great deal

of practice, the process of shadowing had become partially automatic for

Moray, and he had capacity left over to attend to the unshadowed channel.

Spelke, Hirst, and Neisser (1976) provided an interesting demonstration of

how a highly practiced skill ceases to interfere with other ongoing behaviors.

(This was a follow-up of a demonstration pioneered by the writer Gertrude

Stein when she was at Harvard University.) Their participants had to perform

two tasks: read a text silently for comprehension while simultaneously copying

words dictated by the experimenter. At first, these tasks were extremely difficult

to do simultaneously. Participants read much more slowly than normal. After

six weeks of practice, however, the participants were reading at normal speed.

They had become so skilled that their comprehension scores were the same as

for normal reading. For these participants, reading while copying had become

no more difficult than reading while walking. It is of interest that participants

reported no awareness of what it was they were copying. Much as with driving,

the participants lost their awareness of the automated activity.3

Another example of automaticity is transcription typing. The typist is

simultaneously reading the text and executing the finger strokes for typing. In this

case, we have three systems operating in parallel: perception of the text to be

typed, central translation of the earlier perceived letters into keystrokes, and the

actual typing of still earlier letters. Skilled transcription typists often report little

awareness of what they are typing, because this task has become so automated.

Skilled typists also find it impossible to stop typing instantaneously. If suddenly

told to stop, they will hit a few more letters before quitting (Salthouse, 1985, 1986).

As tasks become practiced, they become more automatic and require less and

less central cognition to execute.

The Stroop Effect

Automatic processes not only require little or no central cognition to execute but

also appear to be difficult to prevent. A good example is word recognition for

86 | Attention and Performance

3 When given further training with the intention of remembering what they were transcribing, participants

were also able to recall this information.

Anderson7e_Chapter_03.qxd 8/20/09 9:42 AM Page 86

Central Attention: Selecting Lines of Thought to Pursue | 87

practiced readers. It is virtually impossible to look

at a common word and not read it. This strong

tendency for words to be recognized automatically

has been studied in a phenomenon known as the

Stroop effect, after the psychologist who first

demonstrated it, J. Ridley Stroop (1935). The task

requires participants to say the ink color in which

words are printed. Color Plate 3.2 provides an illustration

of such a task. Try naming the colors of the

words in each column as fast as you can.Which column

was easiest to read? Which was hardest?

The three columns illustrate three of the conditions

in which the Stroop effect is studied. The first

column illustrates a neutral, or control, condition

in which the words are not color words. The second

column illustrates the congruent condition in which

the words are the same as the color. The third column

illustrates the conflict condition in which there are

color words but they are different from their colors. A

typical modern experiment, rather than having participants

read a whole column, will present a single word

at a time and measure the time to name that word. Figure

3.24 shows the results from an experiment on the Stroop effect by Dunbar and

MacLeod (1984). Compared to the control condition of a neutral word, participants

could name the ink color somewhat faster in the congruent condition—

when the word was the name of the ink color. In the conflict condition, when the

word was the name of a different color, they named the ink color much more

slowly. For instance, they had great difficulty in saying that the ink color of the word

red is green. Figure 3.24 also shows the results when the task is switched and participants

are asked to read the word and not name the color. The effects are asymmetrical;

that is, individual participants experienced very little interference in reading a

word as a function of its ink color. This reflects the highly automatic character of

reading.Additional evidence for its automaticity is that participants could also read

a word much faster than they could name its ink color. Reading is such an automatic

process that not only is it unaffected by the color, but participants are unable

to inhibit reading the word, and that reading can interfere with the color naming.

MacLeod and Dunbar (1988) looked at the effect of practice on performance

in a Stroop task. They used an experiment in which the participants learned

the color names for random shapes. Part (a) of Color Plate 3.3 illustrates the

shape-color associations they might learn. The experimenters then presented the

participants with test geometric shapes and asked them to say either the color

name associated with the shape or the actual ink color of the shape. As in the

original Stroop experiment, there were three conditions, and these are illustrated

in part (b) of Color Plate 3.3:

1. Congruent: The random shape was in the same ink color as its name.

2. Control:White shapes were presented when participants were to say the

color name for the shape; colored squares were presented when they were

Reaction time (ms)

900

800

700

600

500

400

Congruent Control Conflict

Color naming

Word reading

Condition

FIGURE 3.24 Performance data

for the standard Stroop task.

The curves plot the average

reaction time of the participants

as a function of the condition

tested: congruent (the word

was the name of the ink color);

control (the word was not

related to color at all); and

conflict (the word was the name

of a color different from the

ink color). (From Dunbar & MacLeod,

1984. Reprinted by permission of the

publisher. © 1984 by the Journal of

Experimental Psychology: Human

Perception and Performance.)

Anderson7e_Chapter_03.qxd 8/20/09 9:42 AM Page 87

to name the ink color of the shape. (The square shape was not associated

with any color.)

3. Conflict: The random shape was in a different ink color from its name.

Figure 3.25 shows the results from this experiment. Color naming was much

more automatic than shape naming and was relatively unaffected by congruence

with the shape, whereas shape naming was affected by congruence with the ink

color (Figure 3.25a). Then MacLeod and Dunbar gave the participants 20 days

of practice at naming the shapes. Participants became much faster at naming

shapes, and now shape naming interefered with color naming rather than vice

versa (Figure 3.25b). Thus, the consequence of the training was to make shape

naming automatic, like word reading, so that it affected color naming.

Reading a word is such an automatic process that it is difficult to inhibit,

and it will interfere with processing other information about the word.

Prefrontal Sites of Executive Control

We have seen that the parietal cortex is important in the exercise of attention in

the perceptual domain. There is evidence that the prefrontal regions are particularly

important in direction of central cognition, often known as executive

control. The prefrontal cortex is that portion of the frontal cortex anterior to the

premotor region (the premotor region is area 6 in Color Plate 1.1). Just as damage

to parietal regions results in deficits in the deployment of perceptual attention,

damage to prefrontal regions results in deficits of executive control. Patients with

88 | Attention and Performance

Reaction time (ms)

750

700

650

600

550

500

450

Congruent Control Conflict

(a) Condition (b)

Reaction time (ms)

750

700

650

600

550

500

450

Congruent Control Conflict

Condition

Color naming

Shape naming

Color naming

Shape naming

FIGURE 3.25 Results from the experiment created by MacLeod and Dunbar (1988) to evaluate

the effect of practice on the performance of a Stroop task. The data reported are the average

times required to name shapes and colors as a function of color-shape congruence: (a) initial

performance (note correction); (b) after 20 days of practice. The practice made shape naming

automatic, like word reading, so that it affected color naming.

Anderson7e_Chapter_03.qxd 8/20/09 9:42 AM Page 88

Central Attention: Selecting Lines of Thought to Pursue | 89

such damage often seem totally driven by stimulus and fail to control their behavior

according to their intentions. A patient who sees a comb on the table may

simply pick it up and begin combing her hair; another who sees a pair of glasses

will put them on even if he already has a pair on his face. Patients with damage to

prefrontal regions show marked deficits in the Stroop task and often cannot refrain

from reading the word rather than naming the color (Janer & Pardo, 1991).

Two prefrontal regions seem particularly important in executive control. One

is the dorsolateral prefrontal cortex (DLPFC), which is the upper portion of

the prefrontal cortex (see Figure 3.1). It is called dorsolateral because it is high

(dorsal) and to the side (lateral). The second region is the anterior cingulate

cortex (ACC), which is a structure below the surface of the brain along the midline.

It is part of the cortex that is folded under its visible surface. The DLPFC

seems particularly important in the setting of intentions and the control of

behavior. For instance, it is highly active during the simultaneous performance

of dual tasks such as those whose results are reported in Figures 3.19 and 3.20

(Szameitat, Schubert,Muller, & von Cramon, 2002). The ACC seems particularly

active when people must monitor conflict between competing tendencies.

For instance, brain-imaging studies show that it is highly active

in Stroop trials when one must name the color of a word printed in an

ink of conflicting color (J.V. Pardo, P. J. Pardo, Janer, & Raichle, 1990).

There is a strong relationship between the ACC and cognitive control

in many tasks. For instance, it appears that children develop more cognitive

control as their ACC develops. The amount of activation in the ACC

appears to be correlated with the performance by children in tasks

requiring cognitive control (Casey et al., 1997a). Developmentally, there

also appears to be a positive correlation between performance and sheer

volume of the ACC (Casey et al., 1997b). A nice paradigm for demonstrating

the development of cognitive control in children is the “Simon

says” task. In one study, Jones, Rothbart, and Posner (2003) had children

receive instructions from two dolls—a bear and an elephant. The instructions

were things like “Elephant says, ‘Touch your nose.’” The children

were to follow the instructions from one doll (the act doll) and ignore the

instructions from another (the inhibit doll). All children successfully followed

the act doll but many had difficulty ignoring the inhibit doll. From

the age of 36 to 48 months children progressed from 22% success to 91%

success in ignoring the inhibit doll. A few children used regulatory selfspeech

to control their behavior, as famously proposed by Luria (1961),

but they also used strategies such as sitting on their hands or distorting

their actions—pointing to their ear rather than their nose.

Another way to appreciate the importance of prefrontal regions to

cognitive control is to compare performance of humans with other

primates. As reviewed in Chapter 1, a major dimension of the evolution

from primates to humans has been the increase in the size of prefrontal

regions. Primates can be trained to do many tasks that humans do and

so permit careful comparison. One such task involving a variant of the Stroop task

presents participants with a display of numerals (e.g., five 3’s) and pits naming the

number of objects against indicating the identity of the numerals. Figure 3.26

5 5 5

1 1 1 1

2

3 3 3 3 3

4 4

5 5 5

4 4 4 4 4

5 5 5 5

3

4 4 4

2 2 2 2

3 3

4 4 4

1 1 1 1

3

2 2 2

FIGURE 3.26 A numerical

Stroop task comparable to the

color Stroop task (see Color

Plate 3.2).

Anderson7e_Chapter_03.qxd 8/20/09 9:42 AM Page 89

provides an example of this task in the same form as the original Stroop task

(Color Plate 3.2): trying to count the number of numerals in each line versus

trying to name the numerals in each line. The stronger interference in this case

is from the numeral naming to the counting (Windes, 1968). This paradigm has

been used to compare Stroop interference in humans versus rhesus monkeys

who had been trained to use the numerals (Washburn, 1994; see Table 3.1).

Both groups of participants were shown two arrays and were required to indicate

which had more numerals independent of the identity of the numerals.

Compared to a baseline where they had to judge which array of letters had more

objects, both humans and monkeys performed better when the numerals agreed

with the difference in cardinality and performed worse when the numerals disagreed

(as they do in Figure 3.26). Both populations showed similar reaction

time effects, but whereas the humans made 3% errors in the incongruent condition,

the monkeys made 27% errors. The level of performance observed of the

monkeys was like the level of performance observed with patients with damage

to their frontal lobes.

Prefrontal regions, particularly DLPFC and ACC, play a major role in

executive control.

Conclusions

There has been a gradual shift in the way cognitive psychology has perceived

the issue of attention. For a long time, the implicit assumption was captured

by this famous quote from William James (1890) over a century ago:

Everyone knows what attention is. It is the taking possession by the mind, in a

clear and vivid form, of one out of what seem several simultaneously possible

objects or trains of thought. Focalization, concentration of consciousness are

of its essence. It implies withdrawal from some things in order to deal effectively

with others. (pp. 403–404)

90 | Attention and Performance

TABLE 3.1

Mean Response Times and Accuracy Levels as a Function of Species and Condition

Condition Accuracy (%) Response Time (ms)

Rhesus Monkeys (N _ 6)

Congruent numerals 92 676

Baseline (letters) 86 735

Incongruent numerals 73 829

Human Subjects (N _ 28)

Congruent numerals 99 584

Baseline (letters) 99 613

Incongruent numerals 97 661

Anderson7e_Chapter_03.qxd 8/20/09 9:42 AM Page 90

Key Terms | 91

Two features of this quote reflect conceptions once held about attention. The

first is that attention is strongly related to consciousness—we cannot attend

to one thing unless we are conscious of it. The second is that attention, like

consciousness, is a unitary system.More and more, cognitive psychology is coming

to recognize that attention operates at an unconscious level. For instance,

people often are not conscious of where they have moved their eyes. Along with

this recognition has come the realization that attention is multifaceted (e.g.,

Pashler, 1995). We have seen that it makes sense to separate auditory attention

from visual attention and attention in perceptual processing from attention in

executive control from attention in response generation. The brain consists of a

number of parallel processing systems for the various perceptual systems,

motor systems, and central cognition. Each of these parallel systems seems to

suffer bottlenecks—points at which it must focus its processing on a single

thing. Attention is best conceived as the processes by which each of these

systems is allocated to potentially competing information-processing demands.

The amount of interference that occurs among tasks is a function of the overlap

in the demands that these tasks make on the same systems.

1. The chapter discussed how listening to one spoken

message makes it difficult to process a second spoken

message. Do you think that listening to a conversation

on a cell phone while driving makes it harder to process

other sounds like a car horn honking?

2. Which search should produce greater parietal activation:

searching Figure 3.13a for a T or searching Figure 3.13b

for a T?

3. Describe circumstances where it would be advantageous

to focus one’s attention on an object rather than

a region of space, and describe circumstances where the

opposite would be true.

4. We have discussed how automatic behaviors can

intrude on other behaviors and discussed how some

aspects of driving have become automatic. Consider

the situation in which a passenger in the car is a skilled

driver and has automatic aspects of driving evoked by

the driving experience. Can you think of examples

where automatic aspects of driving seem to affect a

passenger’s behavior in a car? Might this help explain

why having a conversation with a passenger in a car

is not as distracting as having a conversation over a

cell phone?

Questions for Thought

Key Terms

anterior cingulate cortex

(ACC)

attention

attenuation theory

automaticity

binding problem

central bottleneck

dichotic listening task

dorsolateral prefrontal

cortex (DLPFC)

early-selection theories

executive control

feature-integration theory

filter theory

goal-directed attention

illusory conjunction

inhibition of return

late-selection theories

object-based attention

perfect time-sharing

serial bottleneck

space-based attention

stimulus-driven attention

Stroop effect