KEY TERMS categorical perception
fixation
Ganong effect
hemispherical specialisation
McGurk effect
phoneme restoration
right ear advantage
saccade
segmentation
variability
PREVIEW
This chapter introduces some key issues in both visual and
auditory perception that relate to language processing. By
the end of the chapter you will know that:
• there are similarities as well as crucial differences in the
perceptual processes involved in spoken and visual
language comprehension;
• there are perceptual skills that are particularly important
for efficient language processing;
• the language system influences the interpretation of
perceptual cues.
Perception for language
7 CHAPTER
https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at
100 INTRODUCING PSYCHOLINGUISTICS
7.1 Introduction
In this chapter, we consider some significant findings in both visual and
auditory perception studies. In particular we consider those skills that are
specifically language-oriented and which act as a pre-requisite for success-
ful language processing. The chapter will also consider some of the find-
ings from language acquisition that indicate the early stage at which
these perceptual skills develop.
There are some issues that are common to both visual and spoken lan-
guage processing, but – as we will see – there are also some issues that are
unique to either modality. The common questions concern the nature of
the information extracted in the early stages of the recognition process
and how this information is extracted, the mapping from the input to the
lexicon and the nature of any intervening units of analysis, and the time-
course of the flow of information throughout the system.
7.2 Basic issues in perception for language
Some basic tasks for successful language comprehension are that lan-
guage users must recognise the signals that reach the brain (from eye or
from ear, or even from the fingers in the context of Braille) as being lan-
guage rather than non-language, they must recognise them as being in a
language that they understand, and they must interpret them as mean-
ingful. In the comprehension of written and spoken language, these tasks
involve knowledge about how letters and sounds are used, but also knowl-
edge about writers and speakers, about the processes of writing and speak-
ing, and about the structures and units of language. In this section we will
focus on some of the general issues that exist for perception for lan-
guage.
Hemispherical specialisation It is clear that as a species, humans have become specially adapted for
language. Our upright posture, the position of our larynx (voice box) in
the throat and the shape and dimensions of our vocal tract all contribute
to our ability to produce a rich and well-controlled range of speech
sounds. Our hearing for language is helped by the fact that these speech
sounds have sound frequencies and amplitudes to which our auditory
system is especially sensitive. There is also neurophysiological evidence
that humans have perceptual specialisation for language. This includes
hemispherical specialisation , where the two halves of the brain have dif-
ferent specialisations. It is typically (though not always) the case that
language faculties are predominantly in the left hemisphere of the brain.
Interestingly, and following the general pattern that the left hemisphere
is responsive to and responsible for the right-hand side of the body, this is
linked to a right ear advantage (REA) for speech for most people. This
was demonstrated in the 1960s and 1970s in dichotic listening experi-
ments (e.g. Kimura, 1961 ; Studdert-Kennedy, Shankweiler & Pisoni, 1972 ;
Hemispherical specialisation Various tasks are
under the control
of certain brain
areas, and there is
considerable
evidence that one
hemisphere (half)
of the brain is
responsible for
some tasks, and
the other for
other tasks. It is
quite well known
for instance that
there is a cross-
over whereby
motor commands
to the left body
side – e.g. mental
instructions to
move the left arm –
are controlled by
the right brain
hemisphere, and
vice versa.
https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at
Perception for language 101
Studdert-Kennedy & Shankweiler, 1970 ). In these experiments, partici-
pants hear competing sequences of words presented over headphones to
each ear. More accurate identification occurs for words presented to the
right ear, as long as participants have no basic hearing imbalance between
the ears.
An interesting question is whether the REA arises because we hear bet-
ter with the right ear, or because we hear speech sounds better with the
right ear, or because we process language better when we receive it
through the right ear. That is, at what processing level does the left hemi-
sphere get its advantage? To test this, dichotic listening experiments have
been carried out with a number of different kinds of stimuli. Musical
stimuli fail to show the REA, and indeed have been found to give a left ear
advantage (Bryden, 1988 ). This shows that the REA is not a reflection of
auditory processing per se, as this would predict an REA for any kind of
auditory input. In addition, neurophysiological studies using speech and
equivalently complex non-speech sounds found differences in left brain
hemisphere activation for the speech and non-speech, but equal activa-
tion levels for the two types in the right hemisphere (Parviainen, Helenius
& Salmelin, 2005 ).
The REA is also clearly not phonetic, i.e. not an advantage specifically
for speech sounds, because Morse code signals (sequences of short and
long tones acting as a code for letters of the alphabet) also show the REA
(Papcun, Krashen, Terbeek, Remington & Harshman, 1974 ). It is most likely
therefore that the REA reflects the linguistic processing that takes place
in the left brain hemisphere, and which will apply to both speech and
Morse code input. It also turns out that speech-like but unintelligible
stimuli also show an REA (i.e. more can be remembered about these
stimuli when presented to the right ear). Even though these stimuli are
unintelligible, our linguistic processing system attempts to make sense of
them, and so we can recall more about them as a consequence.
If participants are instructed to pay greater attention to one ear than
the other, then this can enhance or decrease the REA, suggesting again
that the advantage is not an automatic peripheral effect (Hugdahl &
Andersson, 1986 ). In addition, the REA for simple syllables in dichotic lis-
tening tasks is affected by the nature of a preceding prime stimulus pre-
sented to both ears simultaneously (Sætrevik & Hugdahl, 2007 ). If the
prime differs from both of two test items, one presented to each ear, then
the REA persists, i.e. the item presented to the right ear is recalled more
accurately. If the prime is identical to the left-channel test item, then the
REA increases, and if the prime is identical to the right-channel test item,
the REA decreases, i.e. there is inhibition of the previously presented
prime item. Inhibition has been shown independently for primes that
have to be ignored (i.e. where no response is expected for the prime item,
as in this task). Sætrevik & Hugdahl argue that after the prime item is
presented, cognitive control inhibits it because it is a potential interfering
factor, and this leads to a recognition advantage for the novel item.
Interestingly, this effect is found in this dichotic listening task even when
the primes are visually displayed on a computer screen (e.g. <ga> before
In priming tasks
researchers are
interested in how
quickly and/or
accurately
participants
respond to a
stimulus (the
probe) that has
been preceded by
another stimulus
(the prime) that
might be related
to it in some way.
Priming can
include for
example identity
priming (the same
stimulus is
presented twice)
or semantic
priming (the
prime word is
related in
meaning to the
probe, e.g.
DOCTOR then
NURSE).
https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at
102 INTRODUCING PSYCHOLINGUISTICS
dichotic auditory stimuli consisting of /ga/ to one ear and /ba/ to the
other).
Mapping from the input to the linguistic system An important part of perceptual processing for language is how the lis-
tener or reader gets from the input signal to the linguistic system. In
Chapters 8 and 9 we discuss this process more closely in the context of
spoken and visual word recognition respectively, building on assumptions
that words provide a significant linguistic building block and that recog-
nising words in the input stream is therefore an important objective of
the perceptual system.
An issue that is common to both visual and auditory processing, though
different in its specific workings, is the nature of pre-lexical processing,
i.e. what kinds of units need to be identified before words can be accessed.
We are so used to a particular way of thinking about how written words
are made up that it might seem obvious that word recognition would
involve the recognition of a word’s component letters. However, there is
evidence that practised readers recognise individual letters only in the
case of relatively uncommon words, and that a lot of word recognition is
based on overall word shape. While this suggests recognition units larger
than the letter, other approaches to visual word recognition argue that
there are smaller recognition units than letters, and that letter features
(such as horizontal, vertical or diagonal lines at various heights on a text
line) form an important part of the recognition process.
Similarly, models of spoken word recognition argue for different types
of intervening representation between the input and the word. The most
obvious is the phoneme , or distinct speech sound, as the nearest equiva-
lent in speech to the letter in visually presented words. But smaller units,
e.g. phonetic features have also been claimed to have perceptual validity.
In addition, pre-lexical units larger than the phoneme have also been
proposed, such as the syllable or the diphone .
In both the visual and auditory domains there are also peripheral per-
ceptual processes that must take place before linguistically relevant infor-
mation can be extracted from the input. These more automatic processes
do not generally form part of the subject matter of psycholinguistics,
except insofar as they may be relevant to the extraction of linguistically
relevant features.
Variability A significant issue in perception for language is variability . That is, the
input that we receive can be highly variable in its detail. This variability
adds to the difficulty of identifying the units of writing or of speech.
Writing styles and legibility vary from one person to the next. The care
taken over writing depends on the nature of the writing task and the
intended reader of the material (notes written by a student during a lec-
ture will differ from a scholarship application letter from the same stu-
dent). The choice of font in a typed document will affect letter shape, as
Figure 7.1 illustrates.
Phonetic
features usually
involve the
presence or
absence of a
feature, such as
[±voice], i.e.
whether or not
the sound is
voiced, as in the
contrast between
/z/ (voiced) and /s/
(voiceless).
A diphone is a
sequence of two
sounds, capturing
the important
transitions from
one sound to the
next. Diphones
are often used to
create more
natural sounding
speech synthesis.
https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at
Perception for language 103
Variability is very obviously present in speech too. Speakers have differ-
ent vocal tract shapes and sizes and different chest cavity sizes. These and
other physical factors will contribute to variation in the sounds produced
by different speakers. Even the same speaker will produce qualitatively
different versions of the same sound on separate occasions, depending on
a range of factors such as health, emotional state, the situation of speak-
ing, the phonetic and linguistic context in which a sound is found, and
so on.
Variability is a potential problem for perception, as too much variabil-
ity will result in difficulty identifying the intended letter or sound or
word or message. Researchers in speech perception and automatic
speech recognition have often struggled to identify invariant cues to the
identity of individual speech sounds. Of course, some variability is pre-
dictable and therefore potentially useful. Predictable variation in speech
can indicate differences between speakers in terms of their age, sex, size,
social class, place of origin, and many other demographic, social and
personal factors. Some variation, both in writing and in speech, results
from the effects of the context in which a letter or sound is found, and
can therefore provide information about that context which can actu-
ally help perception. For instance, nasalisation on an English vowel may
make that vowel different from other instances of that vowel, and there-
fore increase the vowel’s variability, but at the same time this nasalisa-
tion may be informative, because it may tell you that the following
consonant is nasal.
Exemplars Rather than assume that the input is matched against a single template
for a phoneme, word or other recognition unit, recent approaches to lan-
guage perception and comprehension have argued that our memory sys-
tems allow us to store multiple representations for a given unit. These are
known as exemplars , and are assumed to be rich in information that
relates to the actual utterances on which they are based. For example, it
has been argued that we have exemplar representations for words which
include information about the speaker who uttered the word, such as
their age, sex, social grouping, dialect, etc., as well as possibly about the
time and place of the utterance, and so on (Hay, Warren & Drager, 2006 ;
Johnson, 1997 ; Pierrehumbert, 2001 ; Strand, 1999 ). Exemplars provide a
possible mechanism for coping with variation, because the latter becomes
part of the richness of the set of representations rather than a problem to
Figure 7.1 Variability in the input.
Nasalisation
occurs when the
velum, or soft
palate, is lowered
during speaking,
allowing air to
flow through the
nose. English has
no nasal vowels,
though it does
have nasal
consonants (e.g.
the final
consonants in
sum , sun and sung ).
When a vowel in
English precedes a
nasal consonant,
then the vowel
often has a nasal
quality as a result
of assimilation to
(i.e. becoming
more like) that
consonant. In this
case we say the
vowel is nasal ised ,
but not that it is a
nasal vowel.
https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at
104 INTRODUCING PSYCHOLINGUISTICS
be overcome. Since social information is also associated with the exem-
plars, speech perception has a built-in mechanism for recognising social
variation and for normalising for it. As new exemplars are encountered,
they are added to our exemplar sets, and older exemplars that are not re-
activated fade over time.
Segmentation If language comprehension involves the recognition of basic units of writ-
ing or of speech, then these units need to be separable from adjacent
units. But segmentation of the input is not always straightforward.
(7.1) Some joined up writing
Take for instance the example in (7.1). Although this is a highly regular-
ised version of connected letters, using a computer font rather than
actual handwriting, there are areas of ambiguity and uncertainty con-
cerning where one letter finishes and the next begins. For instance, the
beginning of the final word could be a <u>, the second letter in the first
word might be <a>.
Likewise, speech sounds run into one another as the articulators move
from the position for one sound to that for the next. Figure 7.2 gives a
visual representation – a spectrogram (see sidebar) – of speech, for the
utterance Pete is keen to lead the team . An approximate segmentation into
words is shown below the spectrogram. Note that there are seldom any
clear ‘boundaries’ between the words in the spectrogram, i.e. segmenta-
tion of the speech into words is difficult, let alone into individual sounds
within these words. In this respect, speech is different from most instanc-
es of writing, in that writing – even joined up writing as in (7.1) – usually
places spaces between words. Note also that Figure 7.2 includes a good
example of variability resulting from the context in which a sound is
uttered – there are four instances of the /i/ sound in this utterance, and
the portions of the spectrogram corresponding to those instances are not
identical, as shown in Figure 7.3 . They differ both in their duration and in
the shape of the darker bands showing how the sound energy is distrib-
uted, though there are also some common features.
This section has highlighted some of the common issues for the percep-
tion of written and spoken language. These issues relate to the fact that
the perceiver has to extract linguistic information from the input signal,
and that this can be made difficult by two major problems. One is the lack
of invariance in how the ‘same’ letter or sound is produced on different
occasions or by different people. The other is how to segment a piece of
text or an utterance into its constituent parts. We have noted above that
the segmentation of speech into words is more problematic than the seg-
mentation of written or printed text. In addition, the transitory nature of
speech means the initial re-coding of the input into linguistic units is
likely to be more critical with speech than with writing, where the reader
can go back and look again at the input. In the next section we will look
at further issues for speech perception.
Spectrograms
are based on the
analysis of the
sound energy
present in speech
at different
frequencies. Time
is represented on
the horizontal
axis, frequency on
the vertical axis,
and the darkness
of the shading
shows how much
energy is present
at any frequency
at a given point in
time during the
utterance.
The darker
bands during the
vowels shown in
the spectrogram
show the resonant
frequencies,
known as
formants, of the
vocal tract. These
differ for different
vowels because of
the position of the
tongue, the shape
of the lips, the
height of the
lower jaw, etc.
They also differ
for what seems to
be the same
vowel, depending
on what sounds
precede and
follow that vowel.
https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at
Perception for language 105
7.3 Basic issues in speech perception
As already mentioned, human auditory perception is especially well
tuned to speech sounds. Our hearing is most sensitive to sounds in the
frequency range in which most speech sounds are found, i.e. between
600 Hz and 4000 Hz.
It is also the case that the human perceptual system streams language
and non-language signals, i.e. treats them as separate inputs, thereby
reducing the distracting effect of non-speech signals on speech percep-
tion. This has been shown through phoneme restoration effects (Samuel,
1990 ). When listeners hear words in which a speech sound (phoneme) has
been replaced by a non-speech sound such as a cough, they are highly
likely to report the word as intact, i.e. the cough is treated as part of a
separate stream. There is a clear linguistic influence here, as the restora-
tion effect is stronger with real words than with nonsense words. It has
also been shown that when the word-level information is ambiguous,
then the word that is restored is one which matches the sentence context.
For example, the sequence /#il/, where # indicates a non-speech sound
replacing or overlaid on a consonant, could represent many possible
words ( deal, feel, heal , etc.). In the different conditions shown in (7.2), a
Figure 7.2 Spectrogram for the utterance Pete is keen to lead the team .
pete is keen to lead the team
Figure 7.3 Four /i/ sounds from the utterance in Figure 7.2 .
Hz = Hertz, or
cycles per second.
It is the rate at
which a sound-
wave repeats
itself, and is the
standard measure
of frequency for
sounds.
https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at
106 INTRODUCING PSYCHOLINGUISTICS
word will be reported that is appropriate to the context shown by the final
word in the utterance (originally demonstrated by Warren & Warren,
1970 ). So if that word is orange , then peel is reported, if it is table , then meal ,
and so on. Linguistic effects on perception are powerful, and will be dis-
cussed in more detail later in this chapter.
(7.2) It was found that the /#il/ was on the { shoe | orange | table }
The language-specific nature of streaming effects is indicated by anecdo-
tal evidence from students who are asked to listen to recordings from a
language with a very different sound inventory from their own, and who
experience some of the sounds as non-speech sounds external to the
speech stream. A good example of this is when English-speaking students
first listen to recordings of a click language, with many students reporting
the click consonants as a tapping or knocking sound happening sepa-
rately from the speech.
Despite the evidence that listeners segregate speech and non-speech
signals, it is also clear that the perceptual system will integrate these if at
all plausible. That is, if a non-speech sound could be part of the simultane-
ous speech signal, we will generally perceive it as such. For example, if the
final portion of the /s/ sound in the word slit is replaced by silence, then
this silence is interpreted as a /p/ sound, resulting in the word split being
heard (see the exercises at the end of this chapter). A stretch of silence is
one of several cues to a voiceless plosive such as /p/ (the silence results
from the closure of the lips with no simultaneous voicing noise), and is
sufficient in this context to result in the percept of a speech sound.
Frequently there are multiple cues to a speech sound, or to the distinc-
tion between that speech sound and a very similar one. The voiceless
bilabial plosive /p/ sound is cued not just by the silence during the lip-
closure portion of that consonant, but also by changes that take place in
the formant structure of any preceding vowel as the lips come together to
make the closure, by the duration of a preceding vowel (voiceless stops in
English tend to be preceded by shorter variants of a vowel than voiced
stops such as /b/), by several properties of the burst noise as the lips are
opened, and so on. While some of these cues may be more important or
more reliable than others, it is clear that the perception of an individual
sound depends on cue integration , involving a range of cues that distin-
guish this sound from others in the sound inventory of the language.
A fascinating instance of cue integration comes from studies of speech
perception that involve visual cues. We are often able to see the people we
are listening to, and their faces tell us much about what they are saying.
A particular set of cues comes from the shape and movements of the
mouth. For a bilabial plosive (/b/ or /p/) there will be a visible lip closure
gesture; for an alveolar plosive (/d/ or /t/) it might be possible to see the
tongue making a closure at the front of the mouth, just behind the top
teeth; for a velar plosive (/DZ/ or /k/) the closure towards the back of the
mouth will be visually less evident. Normally, these visual cues will be
compatible with the auditory cues from the speech signal, and therefore
will supplement them. If however, the visual cues and the auditory cues
Click consonants
are found in a
range of languages
in southern and
eastern Africa.
Speakers of
English use click
sounds, such as
the ‘tut-tut’ noise
of disapproval, or
the ‘gee-up’
clicking sound
made to
encourage a horse.
In click languages,
such sounds are
used as parts of
words.
https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at
Perception for language 107
have been experimentally manipulated so that they are no longer compat-
ible, then they can merge on a percept that is different from that sig-
nalled by either set of cues on their own. This is known as the McGurk
effect , after one of the early researchers to identify the phenomenon
(McGurk & MacDonald, 1976 ). For instance, if the auditory information
indicates a /ba/ syllable, but the visual information is from a /DZa/ syllable,
showing no lip closure, then the interpretation is that the speaker has
said /da/. Examples of this effect are available on the website for this book
(and see also the exercises at the end of this chapter).
Another and at first glance somewhat bizarre cue integration effect has
been reported in what have been referred to as the “puff of air” experi-
ments. In these experiments participants listen to stimulus syllables that
are ambiguous between, say, /ba/ and /pa/. For speakers of English and
many other languages, one of several characteristics that distinguish the
/b/ and /p/ sounds in these syllables is that there is a stronger puff of air
that accompanies the /p/ than is found with the /b/. In phonetics terminol-
ogy, the /p/ is aspirated and the /b/ is unaspirated. In the experiments, it
was found that participants were more likely to report the ambiguous
stimulus as /pa/ if they also felt a puff of air that was presented simultane-
ously with the speech signal. The effect was found whether the puff of air
was directed at the hand or at the neck (Gick & Derrick, 2009 ), or even at
the ankle (Derrick & Gick, 2010 ).
It has also been shown that cue trading is involved in speech percep-
tion. For instance, if the release burst of a /p/ is unclear, perhaps because
of some non-speech sound that happened at the same time, then the lis-
tener may assign greater perceptual significance to other cues such as the
relative duration of the preceding vowel and movements in the formants
at the end of that vowel.
These cues in the formant movements are a result of coarticulation –
the articulation of one sound is influenced by the articulation of a neigh-
bouring sound. It appears that our perceptual system is so used to the
phenomenon of coarticulation that it will compensate for it in the percep-
tion of sounds. For instance, Elman and McClelland ( 1988 ) asked partici-
pants to identify a word as capes or tapes . Their experiment hinged on the
fact that a /k/ is pronounced further forward in the mouth, so closer to a
/t/, when it follows /s/ (as in Christmas capes ) than when it follows /ʃ/ (as in
foolish capes ). This is because /s/ is itself further forward in the mouth than
/ʃ/ and there is coarticulation of the following /k/ towards the place of
articulation of the /s/. Elman and McClelland manipulated the first speech
sound in capes or tapes to make it sound more /k/-like or more /t/-like. One
of the cues to the difference between the /k/ and /t/ sounds is the height
in the frequency scale of the burst of noise that is emitted when the plo-
sive is released – it is higher for front sounds like /t/ than it is for back
sounds like /k/. In the experiment, the noise burst of the initial consonant
in tapes or capes was manipulated to produce a range of values that were
intermediate between the target values for /t/ and /k/. Participants heard
tokens from this range of tapes/capes stimuli after either Christmas or fool-
ish , and had to report whether they heard the word as tapes or capes . The
Aspiration is
the puff of air
that accompanies
the release of
certain stop or
plosive consonants
such as /p/ in
English. You can
demonstrate
aspiration to
yourself by
dangling a piece
of paper loosely in
front of your
mouth and saying
/pa/ and /ba/. The
paper should
move more with
/pa/. It should also
move more with
/pa/ than with
/spa/, since another
characteristic of
/p/ in English is
that it is aspirated
at the beginning
of a syllable, but
not when it
follows /s/.
Burst noise as a
cue to place of
articulation – The
spectrogram in
Figure 7.2 contains
/k/ in keen and
/t/ in team . Notice
that the dark
band of acoustic
energy for /k/
covers lower parts
of the frequency
range than the
corresponding
band of energy
for /t/.
https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at
108 INTRODUCING PSYCHOLINGUISTICS
results were very clear – after the word Christmas , tokens on the /t/--/k/
continuum were more likely to be heard as /k/ than when the same tokens
followed foolish (see Figure 7.4 ). That is, the participants expected the coar-
ticulation effect to lead to a ‘fronted’ /k/ after /s/, and compensated for this
in their interpretation of the frequency level of the burst noise.
Our perception and comprehension of speech is also affected by signal
continuity . That is, listeners are better able to follow a stream of speech if
it sounds like it comes in a continuous fashion from one source. This lies
behind the cocktail party effect , where we are able to follow one speaker
in a crowded room full of conversation despite other talk around us
(Arons, 1992 ). This effect can be demonstrated in various ways. In one task,
participants hear two voices over stereo headphones, and are asked to
focus on what is being said on just one of the headphone channels, the
left channel for example. If the voice on the left channel switches to the
right part way through the recording, then participants find that their
attention follows the voice to the right channel. They then report at least
some of what is then said on the right channel, despite the instructions
to focus on the left channel (Treisman, 1960 ). The strength of this effect is
reduced if the utterance prosody is disrupted at the switch. The impor-
tance of signal continuity is also demonstrated in the relative unnatural-
ness of some computer-generated or concatenated speech, such as is
found for instance in the automated speech of some phone-in banking
systems.
Active and passive speech perception There have been numerous attempts to frame aspects of speech percep-
tion in models or theories (some of these are reviewed by Klatt, 1989 ). One
distinction that has been made between different models concerns the
degree of involvement of the listener as speaker, characterised as a differ-
ence between active and passive perception processes.
Passive models of speech perception assume that we have a stored sys-
tem of patterns or recognition units, against which we match the speech
Figure 7.4 Compensation for coarticulation (based on Elman and McClelland, 1988 ).
https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at
Perception for language 109
that we hear. Depending on the specific claims of the model, these stored
patterns might be phonetic features or perhaps templates for phonemes
or diphone sequences, and so on. A phoneme-based perception model
might for example include a template for the /i/ phoneme that shares
some of the common characteristics of the spectrogram slices shown in
Figure 7.3 . A feature-based model might include a voicing detector that
examines the input for the presence of the regular repetition of speech
waves that corresponds to vocal cord vibration, and would have similar
detectors for other features that define a speech sound.
Incoming speech data is matched against the templates, and a score
given for how well the data matches the templates. These scores are evalu-
ated and a best match determined. Many automatic speech recognition
systems operate like this – they have templates for each recognition unit
and match slices of the input speech data against these templates. Such
systems perform best when they have had some training, usually requir-
ing the user to repeat some standard phrases so that the speech process-
ing system can develop appropriate templates.
Active models of speech perception argue that our perception is influ-
enced by or depends on our capabilities as producers of speech. One
model of this type involves analysis-by-synthesis. Here, the listener matches
the incoming speech data not against a stored template for input units of
speech, but against the patterns that would result from the listener’s own
speech production, i.e. synthesises an output (or a series of alternative
outputs) and matches that against the analysis of the input.
7.4 Basic issues in visual perception for language
We have seen that speech perception involves some pre-linguistic analysis
of the auditory input, the precise details of which vary according to the
model of perception. In visual perception for language, some sort of pre-
liminary analysis of the input is also assumed to take place. The whole
word-shape is probably important, as shown by the fact that we can recog-
nise words containing misspellings or ordering errors, jsut as long as the
overoll shape of the word is relatively unaffected. This can even result in
errors remaining undetected. For instance, if you have been reading this
paragraph quickly, for meaning rather than to identify the detail of every
word, you might not have noticed the misspellings of just and overall two
sentences ago.
It seems that the visual input is not recognised on a straightforward
letter-by-letter basis. This is shown by the word superiority effect . This is
the finding that individual letters (e.g. D) are recognised more rapidly and
reliably when they occur in words (e.g. in WORD) than when they are
either in nonwords (legitimate sequences of letters that happen not to
make a word, as in WROD) or in jumbled letter strings (WLOD). Neverthe-
less, individual letters and letter shapes are important in visual percep-
tion for language, and form an important part of approaches to visual
Several years ago
a passage was
published in the
press and on the
internet that
began “Aoccdring
to rscheearch at
Cmabrigde
Uinervtisy it
deosn’t mttaer in
waht oredr the
ltteers in a wrod
are”. Crucially,
though, the
disruptions in that
text involve local
reorderings within
the word, keep the
overall word shape
reasonably intact,
and have no effect
on shorter words.
Information from
the context (the
overall meaning)
further reduces
the disruptive
effect.
https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at
110 INTRODUCING PSYCHOLINGUISTICS
word recognition (see Chapter 9 ). It has also been argued that letter fea-
tures are significant. Letter features might include horizontal and vertical
lines (as in <H>), sloping lines (in <M>), various loops (<P, p, B>) and so on.
Evidence for the importance of letter features comes from studies that
show that it is more difficult to find a letter if it is embedded in a set of
letters with similar features.
The specific nature and level of detail that is claimed to be important
in the initial visual analysis varies according to the particular theory of
perception. It may also depend on the automaticity of visual word recogni-
tion, which varies with reading experience.
Stages during visual perception It is generally assumed that there are stages in the visual perception of
language. The initial visual analysis transfers the input into some sort of
buffer. This is followed by further analysis in working memory, and then
by integration of the analysed input with the linguistic and cognitive
interpretation of the text.
Evidence for these stages comes from a number of sources. As we will
see below, reading normally proceeds by means of a series of eye fixations,
where different portions of text are in the visual field. Although these
fixations are relatively brief, at around 250 msec, they are actually longer
than is needed to recognise a word. We know this because studies that use
very brief presentations of words show that we only need to see a word for
about 50 msec in order to be able to recall it quite well. It seems that we
fixate on a stretch of text for longer than this because it takes longer to
transfer information from a visual buffer into working memory . This was
demonstrated in some early experiments where participants were asked
to report which letter in a series of letters had been marked by having
a vertical line placed next to it (Sperling, 1960 ). The letters were presented
very briefly, and could be recognised from such a brief presentation. The
vertical mark was not presented simultaneously with the letters, but after
a very short delay. The experience for the participants, though, was that
the mark appeared at the same time as the letters, and so they were able
to report accurately which letter was marked. But this only happened
if the interval between the letters and the vertical mark was not too long.
The interpretation of these results is that it takes a short while to transfer
information from the visual buffer to working memory, where further
processing can take place. If further visual input is received too soon, then
the first lot of visual information is added to or overwritten, resulting in
the marking of a letter. It follows therefore that if the eyes skip ahead too
soon during reading, then the visual information relating to one set of
words will be merged with or replaced by the information relating to the
next fixation, before it can be shuffled into working memory.
Eye movements during reading It is clear that the initial visual analysis of text is different from other
kinds of visual perception. There are some obvious differences between
text and much other visual input, such as the fact that text is generally
The letter recall
experiments
worked like this:
First presentation:
G D J P S K W A
< short delay>
Second presentation:
|
Percept:
G D J|P S K W A
If the delay was too
long (over half a
second), then the
two presentations
were not integrated
in this way.
https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at
Perception for language 111
stationary and can easily be looked at again, text is usually sharply
defined but two-dimensional, and so on. There are some important differ-
ences in how text is read, compared with other visual processing. An obvi-
ous external indication of this can be gained by comparing people’s eye
movements when reading text and when watching someone walking
across a room. The following summary of eye movements during reading
is based on Rayner and Balota ( 1989 ). (See also Balota, Yap, & Cortese,
2006 .)
During reading, eye movements are characteristically not smooth, but
demonstrate sequences of fixations and saccades (jumps). Most of the
time readers’ eyes are not in fact moving, since the fixations last about
250 msec while the saccades are very quick, taking between 10 and
20 msec. Each jump moves the eyes forward by approximately 8--9 charac-
ters, as shown by the example in Figure 7.5 , although the size of the jump
can depend on the complexity of the text. The example in the figure is
based on data from reading studies using precision cameras that mea-
sure the movements of the eyes. It shows three consecutive fixations on
a piece of text, and illustrates the information that is available during
each fixation.
As indicated by the key in the figure, WI shows that word identification
can take place for the word(s) in question, i.e. those at or near the fixation
point. BL indicates that the beginning letters of following words can be
identified, and LF shows that some of the letter features of the next few
letters are discernible. WL means that the reader can judge the relative
lengths of words to the right of the fixation point. The solid vertical lines
show the total perceptual span of the fixation, which is clearly greater
than the size of the jump between fixations, and includes more words
than the one(s) that can be readily identified. The span also takes in more
information to the right of the fixation point than to the left. These
aspects of the perceptual span during reading have been measured in
Figure 7.5 Eye movements during silent reading (adapted from Rayner & Balota, 1989 : 268).
Key:
• = fixation point; WI = word identification; BL = beginning letters;
LF = letter features; WL = word length
The perceptual
span asymmetry,
where more
information is
taken in from one
side of a fixation
than the other,
depends on the
direction of
reading. Readers
who deal with
scripts that are
read from right-to-
left, such as
Hebrew, show the
opposite
asymmetry to
those shown by
English readers. In
other words, the
fixation is
‘looking ahead’ in
terms of what is
to be read.
https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at
112 INTRODUCING PSYCHOLINGUISTICS
tasks controlling the amount and nature of information available on-
screen as a reader progresses through a text, and assessing the impact on
the speed and success of reading that results from changing the available
information.
The look-ahead evident in Figure 7.5 plays an important part in reading.
Being able to tell the difference between shorter and longer words (using
the word length information marked by WL in the figure) helps the
reader to distinguish function and content words (the former tend to be
quite short) and to determine which words should be the landing point
for the next fixation. Identification of the beginning letters of words (BL)
and some of the letter features in the rest of the word (LF) allows the lexi-
cal search to be started before the word itself is looked at more closely
during a subsequent fixation.
This preview effect is one of many factors that influence fixation times
and jump distances during reading. Others include the difficulty of the
text, the predictability of the fixated words from the prior context, how
frequent the words are and how recently they have been encountered,
whether the words are ambiguous, and whether there is any priming of
the words by previous mention of related words or other material (see
Chapters 8 and 9 for further discussion of some of these parameters in the
context of word recognition).
The nature of the input It is important to note at this point that languages do not all use the same
writing system as English, and that many of the differences between writ-
ing systems will have consequences for the visual perception of words or
for their recognition. The following are some of the differences. First,
there is a range of orthographic writing systems, i.e. of writing systems
that represent some aspect of the sounds of words. As we will see in
Chapter 9 , pronunciation plays a role during reading, and so the relation-
ship between letters and sounds is important. English uses an alphabetic
system, but one in which there are many irregularities in the correspond-
ences between letters and sounds. For instance <c> has a /s/ pronunciation
in cease , but a /k/ pronunciation in cat . Languages with a high degree of
irregularity are said to have a deep orthography . They contrast with lan-
guages in which there is a reasonably direct letter-sound correspondence,
and therefore a shallow orthography , such as Italian, Serbo-Croatian or
Ma – ori, as well as with languages in which the pronunciation of a particu-
lar letter string is reasonably predictable but where there may be many
letter strings with the same pronunciation (e.g. French, in which the spell-
ings <o>, <au>, <aux>, <eau>, <eaux> can represent the same sound).
Other orthographic writing systems include consonantal systems, such
as Hebrew and Arabic, in which the letters often represent the conso-
nants only, and syllabic systems, such as Kannada and the Japanese kana
writing system, in which each orthographic symbol represents a syllable.
Looking beyond orthographic systems, we find ideographic systems, such
as in Chinese, where a symbol corresponds to a word (or a portion of a
word in a morphologically complex form); with the exception of diacritic
https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at
Perception for language 113
markers, the symbol does not bear a direct representation of the sounds
of the word.
The range of different writing systems and of the relationships of these
writing systems to pronunciation adds to the complexity of the study of
visual perception and visual word recognition.
7.5 Influence of the linguistic system on perception
There are many ways in which our perception is influenced by the linguis-
tic system, i.e. by the fact that the sounds or visual shapes we are process-
ing might be considered part of language. We saw evidence of this earlier
(p. 105) where phoneme restoration effects depended on the sentence
context in which the replaced phoneme occurred. We see it also in the
perceptual advantages that language-relevant stimuli have. For example,
studies in which sounds are embedded in noise have shown that speech
sounds are more reliably identified or recalled than non-speech sounds.
On top of this, if the speech sounds make up real words rather than
nonsense, then they have an even greater processing advantage.
Neurophysiological studies, using EEG and MEG to measure electrical
activity in the brain, show that this processing advantage for real words
over nonsense words becomes distinct as early as 150 msec after the point
at which a word can be recognised (Pulvermüller et al. , 2001 ). The fact that
this is a linguistic effect rather than a consequence of differences in the
sounds involved is shown by the fact this is found for native speakers of
the language being tested (Finnish in this case) but not for foreigners who
do not know the language.
A very significant area in which a linguistic influence on speech percep-
tion has been demonstrated is categorical perception . Work in this area
dates back to the 1950s (Liberman, Harris, Hoffman & Griffith, 1957 ), but
our understanding of what is involved in categorical perception has been
refined over the intervening decades. As the name suggests, categorical
perception relates to the finding that we hear speech sounds as belonging
to categories. That is, if we are presented with a series of speech stimuli
that vary along some linguistically relevant phonetic dimension, then we
will classify some of them as belonging to category X and some as belong-
ing to category Y, rather than hearing a particular stimulus as being X-ish,
another as being a little more Y-like, and so on.
Although categorical perception has been demonstrated for a range of
contrasts, the one most widely cited is the distinction between voiced and
voiceless plosive consonants, e.g. between /b/ and /p/. This distinction is
actually signalled by a range of different parameters, but the one that has
usually been explored is voice onset time (VOT). Voice onset time is the lag
between the release of the closure for the plosive consonant (in this case
the opening of the lips) and the beginning of voicing for a following vowel.
This is shorter for a voiced plosive like /b/ than for a voiceless one like
/p/ – in English typically less than and more than 30 msec respectively.
https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at
114 INTRODUCING PSYCHOLINGUISTICS
In experiments, participants are given synthesised /ba/-like or /pa/-like
syllables to respond to. The response can either be an identification one
(is this /ba/ or /pa/?) or a discrimination one (are two stimuli the same or
different?). The identification responses show that rather than there
being a gradual change of responses from /ba/ to /pa/ as VOT increases,
there is a dramatic switch of response preference part-way along the con-
tinuum. In terms of the hypothetical responses shown in Figure 7.6 ,
responses to VOT continua follow a categorical function – also referred to
as an S-curve because of the shape of the line in such a figure – rather
than a smooth or straight-line function.
Responses in the discrimination tasks show that participants treat as
‘same’ two stimuli that fall on the same side of the dramatic shift in the
identification response, e.g. in terms of Figure 7.6 , stimuli 2 and 3, or
3 and 4, or 5 and 6, but will label as ‘different’ two stimuli that fall on
either side of the shift, i.e. stimuli 4 and 5 in the figure. They do this even
if in terms of VOT the difference between 4 and 5 is of the same magni-
tude as the difference between 3 and 4 or between 5 and 6. This is evi-
dence that sounds are placed into perceptual categories, with rather
sharp boundaries between the categories. These findings have been sup-
ported by neurophysiological studies that have demonstrated different
brain responses (measured through changes in event-related magnetic
fields, or ERFs) to tokens either side of a category boundary, but similar
responses to stimuli within a VOT category (e.g. Simos et al., 1998 ).
Categorical perception makes sense from the point of view of language
perception and comprehension. This is because we expect spoken sound
to indicate linguistic entities as part of the comprehension process.
Studies with infants have shown that they are able to discriminate
categories, e.g. between /b/ and /p/, at a very young age (around 3 months
old), before they can speak (Eimas, Siqueland, Jusczyk & Vigorito, 1971 ).
While an initial reaction to this finding was to conjecture that phonetic
Figure 7.6 Smooth and categorical response functions.
https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at
Perception for language 115
distinctions are a specially evolved and innate human characteristic that
serves as a precursor to speech, a number of factors indicate that this is
not the case. First, it turns out that there is nothing special about
speech – other stimuli, including non-speech sounds, can also be distin-
guished categorically, with appropriate training (as demonstrated in
early work by Lane, 1965 ). Second, it was found that humans are not
alone in making categorical perceptual distinctions between speech
sounds – chinchillas for instance can also do this, even though they have
not evolved to speak (Kuhl, 1987 ). Third, there are cross-linguistic differ-
ences in where the VOT category boundary between, say, /b/ and/p/ is
found. This implies that the categories themselves, and the boundaries
between the categories, have to be learned (Abramson & Lisker, 1973 ),
although some studies have demonstrated that infants may initially be
more sensitive to certain category boundaries than to others (Aslin,
Pisoni, Hennessey & Perey, 1981 ).
In further categorical perception studies with adults, it has been
shown that category boundaries are not fixed, even within a language,
but can be affected by the linguistic context. We saw this earlier in this
chapter in the discussion of listeners’ compensation for coarticulation
(see Figure 7.4 ), where the /k/ identification function for tapes vs capes was
shifted depending on the nature of the consonant at the end of the pre-
ceding word.
The Ganong effect (named after the author of the first famous study of
this phonemenon) is a good demonstration of how the boundaries
between phonetic categories can be affected by linguistic information
(Ganong, 1980 ). In this case, it is the lexical status of the word containing
the manipulated segment that is important. For a /d/--/t/ continuum, for
instance, the category boundary falls nearer the /d/ end of the continuum
if the stimulus is /?ask/ (where ? indicates the manipulated segment). This
is because task is a word and dask is not, so more of the stimuli towards
the /d/ end of the continuum are accepted as falling within the /t/ category.
Conversely, the boundary falls nearer to the /t/ end if the stimulus is /?esk/,
i.e. more desk responses are given overall, because desk is a word and tesk
is not.
The perceptual
abilities of infants
are measured
using methods
such as High
Amplitude
Sucking. An
infant sucks on a
dummy connected
to equipment that
measures how
hard she is
sucking and
controls the
presentation of
stimuli. When the
infant gets bored
with the same
stimulus, her
sucking intensity
reduces, at which
point a new
stimulus is
presented. If she
hears this as
different, then she
becomes excited
and her sucking
increases.
Summary
We have seen that there are some clear commonalities in language
perception in the auditory and visual modalities:
• the input has to be segmented and analysed into linguistically useful units;
• variability in the input has to be overcome.
https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at
116 INTRODUCING PSYCHOLINGUISTICS
But there are also some differences in how speech and writing are
processed, which result from the obvious differences between the
modalities:
• speech is transitory; • speech lacks many clear boundaries between the linguistic units it
contains;
• writing has greater permanence and can be re-inspected with rela- tive ease;
• writing contains clearer boundaries between words and within words between letters.
We have also seen some key issues relevant to perception for language:
• the human brain is specialised for language processing; • the perceptual system streams language and non-language stimuli; • but will integrate stimuli if this makes linguistic sense; • cue integration includes visual and tactile cues as well as auditory
ones;
• the linguistic system can have an influence on perception, leading to effects such as real word advantages and/or biases in perception;
• infants begin to tune-in to the relevant cues for the language they are learning at a very early age.
Exercises Exercise 7.1 If you replaced some of the end of the /s/ in the words soil , suit and seep with silence, would you expect to hear a /p/ in each case? Why? If not, why not? (If you want to hear this for yourself, the website has examples of these words in the original and manipulated forms. Alternatively, use a speech editing package such as Praat ( www.praat.org ) to record and manipulate words like this yourself.)
Exercise 7.2 The sound file represented in Figure 7.2 shows two nasal sounds (at the end of keen and of team ). Can you identify a characteristic of nasal sounds from the spectrogram? Notice that the vowel portions of the words to and the look very similar to one another in the spectrogram. Why do you think this is the case? Figure 7.3 showed four examples of the /i/ vowel from this sound file. Can you explain any of the differences between these examples? What do they have in common that might tell you that the same vowel is involved?
Exercise 7.3 Work with two fellow students, friends or relatives for this exercise, forming a group of three – A, B and C. A watches B’s eyes as B either reads a text or watches C walking across the room. A should note any differences in the pat- terns of B’s eye movements. What differences are there? Swap roles and do this again.
https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at
Perception for language 117
Further reading
Samuel ( 1990 ) provides a review of phoneme restoration effects. The McGurk effect has been reported in many places, and a short and accessible overview is given by McGurk & MacDonald ( 1976 ). An excellent review of eye movements during reading, with particular reference to word recognition, is provided by Rayner and Balota ( 1989 ). A recent review of work in speech perception, including discussion of the recent interest in episodic exemplar representations, can be found in Pisoni and Levi ( 2007 ).
Exercise 7.4 The perceptual integration of visual and auditory inputs known as the McGurk effect is a powerful effect. Would you expect it to happen if the face you see is clearly not the origin of the voice you hear (e.g. if the face is female and the voice is male)? Look at the example of this on the website.
https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at
https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at