chapter7.pdf

KEY TERMS categorical perception

fixation

Ganong effect

hemispherical specialisation

McGurk effect

phoneme restoration

right ear advantage

saccade

segmentation

variability

PREVIEW

This chapter introduces some key issues in both visual and

auditory perception that relate to language processing. By

the end of the chapter you will know that:

• there are similarities as well as crucial differences in the

perceptual processes involved in spoken and visual

language comprehension;

• there are perceptual skills that are particularly important

for efficient language processing;

• the language system influences the interpretation of

perceptual cues.

Perception for language

7 CHAPTER

https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at

100 INTRODUCING PSYCHOLINGUISTICS

7.1 Introduction

In this chapter, we consider some significant findings in both visual and

auditory perception studies. In particular we consider those skills that are

specifically language-oriented and which act as a pre-requisite for success-

ful language processing. The chapter will also consider some of the find-

ings from language acquisition that indicate the early stage at which

these perceptual skills develop.

There are some issues that are common to both visual and spoken lan-

guage processing, but – as we will see – there are also some issues that are

unique to either modality. The common questions concern the nature of

the information extracted in the early stages of the recognition process

and how this information is extracted, the mapping from the input to the

lexicon and the nature of any intervening units of analysis, and the time-

course of the flow of information throughout the system.

7.2 Basic issues in perception for language

Some basic tasks for successful language comprehension are that lan-

guage users must recognise the signals that reach the brain (from eye or

from ear, or even from the fingers in the context of Braille) as being lan-

guage rather than non-language, they must recognise them as being in a

language that they understand, and they must interpret them as mean-

ingful. In the comprehension of written and spoken language, these tasks

involve knowledge about how letters and sounds are used, but also knowl-

edge about writers and speakers, about the processes of writing and speak-

ing, and about the structures and units of language. In this section we will

focus on some of the general issues that exist for perception for lan-

guage.

Hemispherical specialisation It is clear that as a species, humans have become specially adapted for

language. Our upright posture, the position of our larynx (voice box) in

the throat and the shape and dimensions of our vocal tract all contribute

to our ability to produce a rich and well-controlled range of speech

sounds. Our hearing for language is helped by the fact that these speech

sounds have sound frequencies and amplitudes to which our auditory

system is especially sensitive. There is also neurophysiological evidence

that humans have perceptual specialisation for language. This includes

hemispherical specialisation , where the two halves of the brain have dif-

ferent specialisations. It is typically (though not always) the case that

language faculties are predominantly in the left hemisphere of the brain.

Interestingly, and following the general pattern that the left hemisphere

is responsive to and responsible for the right-hand side of the body, this is

linked to a right ear advantage (REA) for speech for most people. This

was demonstrated in the 1960s and 1970s in dichotic listening experi-

ments (e.g. Kimura, 1961 ; Studdert-Kennedy, Shankweiler & Pisoni, 1972 ;

Hemispherical specialisation Various tasks are

under the control

of certain brain

areas, and there is

considerable

evidence that one

hemisphere (half)

of the brain is

responsible for

some tasks, and

the other for

other tasks. It is

quite well known

for instance that

there is a cross-

over whereby

motor commands

to the left body

side – e.g. mental

instructions to

move the left arm –

are controlled by

the right brain

hemisphere, and

vice versa.

https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at

Perception for language 101

Studdert-Kennedy & Shankweiler, 1970 ). In these experiments, partici-

pants hear competing sequences of words presented over headphones to

each ear. More accurate identification occurs for words presented to the

right ear, as long as participants have no basic hearing imbalance between

the ears.

An interesting question is whether the REA arises because we hear bet-

ter with the right ear, or because we hear speech sounds better with the

right ear, or because we process language better when we receive it

through the right ear. That is, at what processing level does the left hemi-

sphere get its advantage? To test this, dichotic listening experiments have

been carried out with a number of different kinds of stimuli. Musical

stimuli fail to show the REA, and indeed have been found to give a left ear

advantage (Bryden, 1988 ). This shows that the REA is not a reflection of

auditory processing per se, as this would predict an REA for any kind of

auditory input. In addition, neurophysiological studies using speech and

equivalently complex non-speech sounds found differences in left brain

hemisphere activation for the speech and non-speech, but equal activa-

tion levels for the two types in the right hemisphere (Parviainen, Helenius

& Salmelin, 2005 ).

The REA is also clearly not phonetic, i.e. not an advantage specifically

for speech sounds, because Morse code signals (sequences of short and

long tones acting as a code for letters of the alphabet) also show the REA

(Papcun, Krashen, Terbeek, Remington & Harshman, 1974 ). It is most likely

therefore that the REA reflects the linguistic processing that takes place

in the left brain hemisphere, and which will apply to both speech and

Morse code input. It also turns out that speech-like but unintelligible

stimuli also show an REA (i.e. more can be remembered about these

stimuli when presented to the right ear). Even though these stimuli are

unintelligible, our linguistic processing system attempts to make sense of

them, and so we can recall more about them as a consequence.

If participants are instructed to pay greater attention to one ear than

the other, then this can enhance or decrease the REA, suggesting again

that the advantage is not an automatic peripheral effect (Hugdahl &

Andersson, 1986 ). In addition, the REA for simple syllables in dichotic lis-

tening tasks is affected by the nature of a preceding prime stimulus pre-

sented to both ears simultaneously (Sætrevik & Hugdahl, 2007 ). If the

prime differs from both of two test items, one presented to each ear, then

the REA persists, i.e. the item presented to the right ear is recalled more

accurately. If the prime is identical to the left-channel test item, then the

REA increases, and if the prime is identical to the right-channel test item,

the REA decreases, i.e. there is inhibition of the previously presented

prime item. Inhibition has been shown independently for primes that

have to be ignored (i.e. where no response is expected for the prime item,

as in this task). Sætrevik & Hugdahl argue that after the prime item is

presented, cognitive control inhibits it because it is a potential interfering

factor, and this leads to a recognition advantage for the novel item.

Interestingly, this effect is found in this dichotic listening task even when

the primes are visually displayed on a computer screen (e.g. <ga> before

In priming tasks

researchers are

interested in how

quickly and/or

accurately

participants

respond to a

stimulus (the

probe) that has

been preceded by

another stimulus

(the prime) that

might be related

to it in some way.

Priming can

include for

example identity

priming (the same

stimulus is

presented twice)

or semantic

priming (the

prime word is

related in

meaning to the

probe, e.g.

DOCTOR then

NURSE).

https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at

102 INTRODUCING PSYCHOLINGUISTICS

dichotic auditory stimuli consisting of /ga/ to one ear and /ba/ to the

other).

Mapping from the input to the linguistic system An important part of perceptual processing for language is how the lis-

tener or reader gets from the input signal to the linguistic system. In

Chapters 8 and 9 we discuss this process more closely in the context of

spoken and visual word recognition respectively, building on assumptions

that words provide a significant linguistic building block and that recog-

nising words in the input stream is therefore an important objective of

the perceptual system.

An issue that is common to both visual and auditory processing, though

different in its specific workings, is the nature of pre-lexical processing,

i.e. what kinds of units need to be identified before words can be accessed.

We are so used to a particular way of thinking about how written words

are made up that it might seem obvious that word recognition would

involve the recognition of a word’s component letters. However, there is

evidence that practised readers recognise individual letters only in the

case of relatively uncommon words, and that a lot of word recognition is

based on overall word shape. While this suggests recognition units larger

than the letter, other approaches to visual word recognition argue that

there are smaller recognition units than letters, and that letter features

(such as horizontal, vertical or diagonal lines at various heights on a text

line) form an important part of the recognition process.

Similarly, models of spoken word recognition argue for different types

of intervening representation between the input and the word. The most

obvious is the phoneme , or distinct speech sound, as the nearest equiva-

lent in speech to the letter in visually presented words. But smaller units,

e.g. phonetic features have also been claimed to have perceptual validity.

In addition, pre-lexical units larger than the phoneme have also been

proposed, such as the syllable or the diphone .

In both the visual and auditory domains there are also peripheral per-

ceptual processes that must take place before linguistically relevant infor-

mation can be extracted from the input. These more automatic processes

do not generally form part of the subject matter of psycholinguistics,

except insofar as they may be relevant to the extraction of linguistically

relevant features.

Variability A significant issue in perception for language is variability . That is, the

input that we receive can be highly variable in its detail. This variability

adds to the difficulty of identifying the units of writing or of speech.

Writing styles and legibility vary from one person to the next. The care

taken over writing depends on the nature of the writing task and the

intended reader of the material (notes written by a student during a lec-

ture will differ from a scholarship application letter from the same stu-

dent). The choice of font in a typed document will affect letter shape, as

Figure 7.1 illustrates.

Phonetic

features usually

involve the

presence or

absence of a

feature, such as

[±voice], i.e.

whether or not

the sound is

voiced, as in the

contrast between

/z/ (voiced) and /s/

(voiceless).

A diphone is a

sequence of two

sounds, capturing

the important

transitions from

one sound to the

next. Diphones

are often used to

create more

natural sounding

speech synthesis.

https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at

Perception for language 103

Variability is very obviously present in speech too. Speakers have differ-

ent vocal tract shapes and sizes and different chest cavity sizes. These and

other physical factors will contribute to variation in the sounds produced

by different speakers. Even the same speaker will produce qualitatively

different versions of the same sound on separate occasions, depending on

a range of factors such as health, emotional state, the situation of speak-

ing, the phonetic and linguistic context in which a sound is found, and

so on.

Variability is a potential problem for perception, as too much variabil-

ity will result in difficulty identifying the intended letter or sound or

word or message. Researchers in speech perception and automatic

speech recognition have often struggled to identify invariant cues to the

identity of individual speech sounds. Of course, some variability is pre-

dictable and therefore potentially useful. Predictable variation in speech

can indicate differences between speakers in terms of their age, sex, size,

social class, place of origin, and many other demographic, social and

personal factors. Some variation, both in writing and in speech, results

from the effects of the context in which a letter or sound is found, and

can therefore provide information about that context which can actu-

ally help perception. For instance, nasalisation on an English vowel may

make that vowel different from other instances of that vowel, and there-

fore increase the vowel’s variability, but at the same time this nasalisa-

tion may be informative, because it may tell you that the following

consonant is nasal.

Exemplars Rather than assume that the input is matched against a single template

for a phoneme, word or other recognition unit, recent approaches to lan-

guage perception and comprehension have argued that our memory sys-

tems allow us to store multiple representations for a given unit. These are

known as exemplars , and are assumed to be rich in information that

relates to the actual utterances on which they are based. For example, it

has been argued that we have exemplar representations for words which

include information about the speaker who uttered the word, such as

their age, sex, social grouping, dialect, etc., as well as possibly about the

time and place of the utterance, and so on (Hay, Warren & Drager, 2006 ;

Johnson, 1997 ; Pierrehumbert, 2001 ; Strand, 1999 ). Exemplars provide a

possible mechanism for coping with variation, because the latter becomes

part of the richness of the set of representations rather than a problem to

Figure 7.1 Variability in the input.

Nasalisation

occurs when the

velum, or soft

palate, is lowered

during speaking,

allowing air to

flow through the

nose. English has

no nasal vowels,

though it does

have nasal

consonants (e.g.

the final

consonants in

sum , sun and sung ).

When a vowel in

English precedes a

nasal consonant,

then the vowel

often has a nasal

quality as a result

of assimilation to

(i.e. becoming

more like) that

consonant. In this

case we say the

vowel is nasal ised ,

but not that it is a

nasal vowel.

https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at

104 INTRODUCING PSYCHOLINGUISTICS

be overcome. Since social information is also associated with the exem-

plars, speech perception has a built-in mechanism for recognising social

variation and for normalising for it. As new exemplars are encountered,

they are added to our exemplar sets, and older exemplars that are not re-

activated fade over time.

Segmentation If language comprehension involves the recognition of basic units of writ-

ing or of speech, then these units need to be separable from adjacent

units. But segmentation of the input is not always straightforward.

(7.1) Some joined up writing

Take for instance the example in (7.1). Although this is a highly regular-

ised version of connected letters, using a computer font rather than

actual handwriting, there are areas of ambiguity and uncertainty con-

cerning where one letter finishes and the next begins. For instance, the

beginning of the final word could be a <u>, the second letter in the first

word might be <a>.

Likewise, speech sounds run into one another as the articulators move

from the position for one sound to that for the next. Figure 7.2 gives a

visual representation – a spectrogram (see sidebar) – of speech, for the

utterance Pete is keen to lead the team . An approximate segmentation into

words is shown below the spectrogram. Note that there are seldom any

clear ‘boundaries’ between the words in the spectrogram, i.e. segmenta-

tion of the speech into words is difficult, let alone into individual sounds

within these words. In this respect, speech is different from most instanc-

es of writing, in that writing – even joined up writing as in (7.1) – usually

places spaces between words. Note also that Figure 7.2 includes a good

example of variability resulting from the context in which a sound is

uttered – there are four instances of the /i/ sound in this utterance, and

the portions of the spectrogram corresponding to those instances are not

identical, as shown in Figure 7.3 . They differ both in their duration and in

the shape of the darker bands showing how the sound energy is distrib-

uted, though there are also some common features.

This section has highlighted some of the common issues for the percep-

tion of written and spoken language. These issues relate to the fact that

the perceiver has to extract linguistic information from the input signal,

and that this can be made difficult by two major problems. One is the lack

of invariance in how the ‘same’ letter or sound is produced on different

occasions or by different people. The other is how to segment a piece of

text or an utterance into its constituent parts. We have noted above that

the segmentation of speech into words is more problematic than the seg-

mentation of written or printed text. In addition, the transitory nature of

speech means the initial re-coding of the input into linguistic units is

likely to be more critical with speech than with writing, where the reader

can go back and look again at the input. In the next section we will look

at further issues for speech perception.

Spectrograms

are based on the

analysis of the

sound energy

present in speech

at different

frequencies. Time

is represented on

the horizontal

axis, frequency on

the vertical axis,

and the darkness

of the shading

shows how much

energy is present

at any frequency

at a given point in

time during the

utterance.

The darker

bands during the

vowels shown in

the spectrogram

show the resonant

frequencies,

known as

formants, of the

vocal tract. These

differ for different

vowels because of

the position of the

tongue, the shape

of the lips, the

height of the

lower jaw, etc.

They also differ

for what seems to

be the same

vowel, depending

on what sounds

precede and

follow that vowel.

https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at

Perception for language 105

7.3 Basic issues in speech perception

As already mentioned, human auditory perception is especially well

tuned to speech sounds. Our hearing is most sensitive to sounds in the

frequency range in which most speech sounds are found, i.e. between

600 Hz and 4000 Hz.

It is also the case that the human perceptual system streams language

and non-language signals, i.e. treats them as separate inputs, thereby

reducing the distracting effect of non-speech signals on speech percep-

tion. This has been shown through phoneme restoration effects (Samuel,

1990 ). When listeners hear words in which a speech sound (phoneme) has

been replaced by a non-speech sound such as a cough, they are highly

likely to report the word as intact, i.e. the cough is treated as part of a

separate stream. There is a clear linguistic influence here, as the restora-

tion effect is stronger with real words than with nonsense words. It has

also been shown that when the word-level information is ambiguous,

then the word that is restored is one which matches the sentence context.

For example, the sequence /#il/, where # indicates a non-speech sound

replacing or overlaid on a consonant, could represent many possible

words ( deal, feel, heal , etc.). In the different conditions shown in (7.2), a

Figure 7.2 Spectrogram for the utterance Pete is keen to lead the team .

pete is keen to lead the team

Figure 7.3 Four /i/ sounds from the utterance in Figure 7.2 .

Hz = Hertz, or

cycles per second.

It is the rate at

which a sound-

wave repeats

itself, and is the

standard measure

of frequency for

sounds.

https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at

106 INTRODUCING PSYCHOLINGUISTICS

word will be reported that is appropriate to the context shown by the final

word in the utterance (originally demonstrated by Warren & Warren,

1970 ). So if that word is orange , then peel is reported, if it is table , then meal ,

and so on. Linguistic effects on perception are powerful, and will be dis-

cussed in more detail later in this chapter.

(7.2) It was found that the /#il/ was on the { shoe | orange | table }

The language-specific nature of streaming effects is indicated by anecdo-

tal evidence from students who are asked to listen to recordings from a

language with a very different sound inventory from their own, and who

experience some of the sounds as non-speech sounds external to the

speech stream. A good example of this is when English-speaking students

first listen to recordings of a click language, with many students reporting

the click consonants as a tapping or knocking sound happening sepa-

rately from the speech.

Despite the evidence that listeners segregate speech and non-speech

signals, it is also clear that the perceptual system will integrate these if at

all plausible. That is, if a non-speech sound could be part of the simultane-

ous speech signal, we will generally perceive it as such. For example, if the

final portion of the /s/ sound in the word slit is replaced by silence, then

this silence is interpreted as a /p/ sound, resulting in the word split being

heard (see the exercises at the end of this chapter). A stretch of silence is

one of several cues to a voiceless plosive such as /p/ (the silence results

from the closure of the lips with no simultaneous voicing noise), and is

sufficient in this context to result in the percept of a speech sound.

Frequently there are multiple cues to a speech sound, or to the distinc-

tion between that speech sound and a very similar one. The voiceless

bilabial plosive /p/ sound is cued not just by the silence during the lip-

closure portion of that consonant, but also by changes that take place in

the formant structure of any preceding vowel as the lips come together to

make the closure, by the duration of a preceding vowel (voiceless stops in

English tend to be preceded by shorter variants of a vowel than voiced

stops such as /b/), by several properties of the burst noise as the lips are

opened, and so on. While some of these cues may be more important or

more reliable than others, it is clear that the perception of an individual

sound depends on cue integration , involving a range of cues that distin-

guish this sound from others in the sound inventory of the language.

A fascinating instance of cue integration comes from studies of speech

perception that involve visual cues. We are often able to see the people we

are listening to, and their faces tell us much about what they are saying.

A particular set of cues comes from the shape and movements of the

mouth. For a bilabial plosive (/b/ or /p/) there will be a visible lip closure

gesture; for an alveolar plosive (/d/ or /t/) it might be possible to see the

tongue making a closure at the front of the mouth, just behind the top

teeth; for a velar plosive (/DZ/ or /k/) the closure towards the back of the

mouth will be visually less evident. Normally, these visual cues will be

compatible with the auditory cues from the speech signal, and therefore

will supplement them. If however, the visual cues and the auditory cues

Click consonants

are found in a

range of languages

in southern and

eastern Africa.

Speakers of

English use click

sounds, such as

the ‘tut-tut’ noise

of disapproval, or

the ‘gee-up’

clicking sound

made to

encourage a horse.

In click languages,

such sounds are

used as parts of

words.

https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at

Perception for language 107

have been experimentally manipulated so that they are no longer compat-

ible, then they can merge on a percept that is different from that sig-

nalled by either set of cues on their own. This is known as the McGurk

effect , after one of the early researchers to identify the phenomenon

(McGurk & MacDonald, 1976 ). For instance, if the auditory information

indicates a /ba/ syllable, but the visual information is from a /DZa/ syllable,

showing no lip closure, then the interpretation is that the speaker has

said /da/. Examples of this effect are available on the website for this book

(and see also the exercises at the end of this chapter).

Another and at first glance somewhat bizarre cue integration effect has

been reported in what have been referred to as the “puff of air” experi-

ments. In these experiments participants listen to stimulus syllables that

are ambiguous between, say, /ba/ and /pa/. For speakers of English and

many other languages, one of several characteristics that distinguish the

/b/ and /p/ sounds in these syllables is that there is a stronger puff of air

that accompanies the /p/ than is found with the /b/. In phonetics terminol-

ogy, the /p/ is aspirated and the /b/ is unaspirated. In the experiments, it

was found that participants were more likely to report the ambiguous

stimulus as /pa/ if they also felt a puff of air that was presented simultane-

ously with the speech signal. The effect was found whether the puff of air

was directed at the hand or at the neck (Gick & Derrick, 2009 ), or even at

the ankle (Derrick & Gick, 2010 ).

It has also been shown that cue trading is involved in speech percep-

tion. For instance, if the release burst of a /p/ is unclear, perhaps because

of some non-speech sound that happened at the same time, then the lis-

tener may assign greater perceptual significance to other cues such as the

relative duration of the preceding vowel and movements in the formants

at the end of that vowel.

These cues in the formant movements are a result of coarticulation –

the articulation of one sound is influenced by the articulation of a neigh-

bouring sound. It appears that our perceptual system is so used to the

phenomenon of coarticulation that it will compensate for it in the percep-

tion of sounds. For instance, Elman and McClelland ( 1988 ) asked partici-

pants to identify a word as capes or tapes . Their experiment hinged on the

fact that a /k/ is pronounced further forward in the mouth, so closer to a

/t/, when it follows /s/ (as in Christmas capes ) than when it follows /ʃ/ (as in

foolish capes ). This is because /s/ is itself further forward in the mouth than

/ʃ/ and there is coarticulation of the following /k/ towards the place of

articulation of the /s/. Elman and McClelland manipulated the first speech

sound in capes or tapes to make it sound more /k/-like or more /t/-like. One

of the cues to the difference between the /k/ and /t/ sounds is the height

in the frequency scale of the burst of noise that is emitted when the plo-

sive is released – it is higher for front sounds like /t/ than it is for back

sounds like /k/. In the experiment, the noise burst of the initial consonant

in tapes or capes was manipulated to produce a range of values that were

intermediate between the target values for /t/ and /k/. Participants heard

tokens from this range of tapes/capes stimuli after either Christmas or fool-

ish , and had to report whether they heard the word as tapes or capes . The

Aspiration is

the puff of air

that accompanies

the release of

certain stop or

plosive consonants

such as /p/ in

English. You can

demonstrate

aspiration to

yourself by

dangling a piece

of paper loosely in

front of your

mouth and saying

/pa/ and /ba/. The

paper should

move more with

/pa/. It should also

move more with

/pa/ than with

/spa/, since another

characteristic of

/p/ in English is

that it is aspirated

at the beginning

of a syllable, but

not when it

follows /s/.

Burst noise as a

cue to place of

articulation – The

spectrogram in

Figure 7.2 contains

/k/ in keen and

/t/ in team . Notice

that the dark

band of acoustic

energy for /k/

covers lower parts

of the frequency

range than the

corresponding

band of energy

for /t/.

https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at

108 INTRODUCING PSYCHOLINGUISTICS

results were very clear – after the word Christmas , tokens on the /t/--/k/

continuum were more likely to be heard as /k/ than when the same tokens

followed foolish (see Figure 7.4 ). That is, the participants expected the coar-

ticulation effect to lead to a ‘fronted’ /k/ after /s/, and compensated for this

in their interpretation of the frequency level of the burst noise.

Our perception and comprehension of speech is also affected by signal

continuity . That is, listeners are better able to follow a stream of speech if

it sounds like it comes in a continuous fashion from one source. This lies

behind the cocktail party effect , where we are able to follow one speaker

in a crowded room full of conversation despite other talk around us

(Arons, 1992 ). This effect can be demonstrated in various ways. In one task,

participants hear two voices over stereo headphones, and are asked to

focus on what is being said on just one of the headphone channels, the

left channel for example. If the voice on the left channel switches to the

right part way through the recording, then participants find that their

attention follows the voice to the right channel. They then report at least

some of what is then said on the right channel, despite the instructions

to focus on the left channel (Treisman, 1960 ). The strength of this effect is

reduced if the utterance prosody is disrupted at the switch. The impor-

tance of signal continuity is also demonstrated in the relative unnatural-

ness of some computer-generated or concatenated speech, such as is

found for instance in the automated speech of some phone-in banking

systems.

Active and passive speech perception There have been numerous attempts to frame aspects of speech percep-

tion in models or theories (some of these are reviewed by Klatt, 1989 ). One

distinction that has been made between different models concerns the

degree of involvement of the listener as speaker, characterised as a differ-

ence between active and passive perception processes.

Passive models of speech perception assume that we have a stored sys-

tem of patterns or recognition units, against which we match the speech

Figure 7.4 Compensation for coarticulation (based on Elman and McClelland, 1988 ).

https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at

Perception for language 109

that we hear. Depending on the specific claims of the model, these stored

patterns might be phonetic features or perhaps templates for phonemes

or diphone sequences, and so on. A phoneme-based perception model

might for example include a template for the /i/ phoneme that shares

some of the common characteristics of the spectrogram slices shown in

Figure 7.3 . A feature-based model might include a voicing detector that

examines the input for the presence of the regular repetition of speech

waves that corresponds to vocal cord vibration, and would have similar

detectors for other features that define a speech sound.

Incoming speech data is matched against the templates, and a score

given for how well the data matches the templates. These scores are evalu-

ated and a best match determined. Many automatic speech recognition

systems operate like this – they have templates for each recognition unit

and match slices of the input speech data against these templates. Such

systems perform best when they have had some training, usually requir-

ing the user to repeat some standard phrases so that the speech process-

ing system can develop appropriate templates.

Active models of speech perception argue that our perception is influ-

enced by or depends on our capabilities as producers of speech. One

model of this type involves analysis-by-synthesis. Here, the listener matches

the incoming speech data not against a stored template for input units of

speech, but against the patterns that would result from the listener’s own

speech production, i.e. synthesises an output (or a series of alternative

outputs) and matches that against the analysis of the input.

7.4 Basic issues in visual perception for language

We have seen that speech perception involves some pre-linguistic analysis

of the auditory input, the precise details of which vary according to the

model of perception. In visual perception for language, some sort of pre-

liminary analysis of the input is also assumed to take place. The whole

word-shape is probably important, as shown by the fact that we can recog-

nise words containing misspellings or ordering errors, jsut as long as the

overoll shape of the word is relatively unaffected. This can even result in

errors remaining undetected. For instance, if you have been reading this

paragraph quickly, for meaning rather than to identify the detail of every

word, you might not have noticed the misspellings of just and overall two

sentences ago.

It seems that the visual input is not recognised on a straightforward

letter-by-letter basis. This is shown by the word superiority effect . This is

the finding that individual letters (e.g. D) are recognised more rapidly and

reliably when they occur in words (e.g. in WORD) than when they are

either in nonwords (legitimate sequences of letters that happen not to

make a word, as in WROD) or in jumbled letter strings (WLOD). Neverthe-

less, individual letters and letter shapes are important in visual percep-

tion for language, and form an important part of approaches to visual

Several years ago

a passage was

published in the

press and on the

internet that

began “Aoccdring

to rscheearch at

Cmabrigde

Uinervtisy it

deosn’t mttaer in

waht oredr the

ltteers in a wrod

are”. Crucially,

though, the

disruptions in that

text involve local

reorderings within

the word, keep the

overall word shape

reasonably intact,

and have no effect

on shorter words.

Information from

the context (the

overall meaning)

further reduces

the disruptive

effect.

https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at

110 INTRODUCING PSYCHOLINGUISTICS

word recognition (see Chapter 9 ). It has also been argued that letter fea-

tures are significant. Letter features might include horizontal and vertical

lines (as in <H>), sloping lines (in <M>), various loops (<P, p, B>) and so on.

Evidence for the importance of letter features comes from studies that

show that it is more difficult to find a letter if it is embedded in a set of

letters with similar features.

The specific nature and level of detail that is claimed to be important

in the initial visual analysis varies according to the particular theory of

perception. It may also depend on the automaticity of visual word recogni-

tion, which varies with reading experience.

Stages during visual perception It is generally assumed that there are stages in the visual perception of

language. The initial visual analysis transfers the input into some sort of

buffer. This is followed by further analysis in working memory, and then

by integration of the analysed input with the linguistic and cognitive

interpretation of the text.

Evidence for these stages comes from a number of sources. As we will

see below, reading normally proceeds by means of a series of eye fixations,

where different portions of text are in the visual field. Although these

fixations are relatively brief, at around 250 msec, they are actually longer

than is needed to recognise a word. We know this because studies that use

very brief presentations of words show that we only need to see a word for

about 50 msec in order to be able to recall it quite well. It seems that we

fixate on a stretch of text for longer than this because it takes longer to

transfer information from a visual buffer into working memory . This was

demonstrated in some early experiments where participants were asked

to report which letter in a series of letters had been marked by having

a vertical line placed next to it (Sperling, 1960 ). The letters were presented

very briefly, and could be recognised from such a brief presentation. The

vertical mark was not presented simultaneously with the letters, but after

a very short delay. The experience for the participants, though, was that

the mark appeared at the same time as the letters, and so they were able

to report accurately which letter was marked. But this only happened

if the interval between the letters and the vertical mark was not too long.

The interpretation of these results is that it takes a short while to transfer

information from the visual buffer to working memory, where further

processing can take place. If further visual input is received too soon, then

the first lot of visual information is added to or overwritten, resulting in

the marking of a letter. It follows therefore that if the eyes skip ahead too

soon during reading, then the visual information relating to one set of

words will be merged with or replaced by the information relating to the

next fixation, before it can be shuffled into working memory.

Eye movements during reading It is clear that the initial visual analysis of text is different from other

kinds of visual perception. There are some obvious differences between

text and much other visual input, such as the fact that text is generally

The letter recall

experiments

worked like this:

First presentation:

G D J P S K W A

< short delay>

Second presentation:

|

Percept:

G D J|P S K W A

If the delay was too

long (over half a

second), then the

two presentations

were not integrated

in this way.

https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at

Perception for language 111

stationary and can easily be looked at again, text is usually sharply

defined but two-dimensional, and so on. There are some important differ-

ences in how text is read, compared with other visual processing. An obvi-

ous external indication of this can be gained by comparing people’s eye

movements when reading text and when watching someone walking

across a room. The following summary of eye movements during reading

is based on Rayner and Balota ( 1989 ). (See also Balota, Yap, & Cortese,

2006 .)

During reading, eye movements are characteristically not smooth, but

demonstrate sequences of fixations and saccades (jumps). Most of the

time readers’ eyes are not in fact moving, since the fixations last about

250 msec while the saccades are very quick, taking between 10 and

20 msec. Each jump moves the eyes forward by approximately 8--9 charac-

ters, as shown by the example in Figure 7.5 , although the size of the jump

can depend on the complexity of the text. The example in the figure is

based on data from reading studies using precision cameras that mea-

sure the movements of the eyes. It shows three consecutive fixations on

a piece of text, and illustrates the information that is available during

each fixation.

As indicated by the key in the figure, WI shows that word identification

can take place for the word(s) in question, i.e. those at or near the fixation

point. BL indicates that the beginning letters of following words can be

identified, and LF shows that some of the letter features of the next few

letters are discernible. WL means that the reader can judge the relative

lengths of words to the right of the fixation point. The solid vertical lines

show the total perceptual span of the fixation, which is clearly greater

than the size of the jump between fixations, and includes more words

than the one(s) that can be readily identified. The span also takes in more

information to the right of the fixation point than to the left. These

aspects of the perceptual span during reading have been measured in

Figure 7.5 Eye movements during silent reading (adapted from Rayner & Balota, 1989 : 268).

Key:

• = fixation point; WI = word identification; BL = beginning letters;

LF = letter features; WL = word length

The perceptual

span asymmetry,

where more

information is

taken in from one

side of a fixation

than the other,

depends on the

direction of

reading. Readers

who deal with

scripts that are

read from right-to-

left, such as

Hebrew, show the

opposite

asymmetry to

those shown by

English readers. In

other words, the

fixation is

‘looking ahead’ in

terms of what is

to be read.

https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at

112 INTRODUCING PSYCHOLINGUISTICS

tasks controlling the amount and nature of information available on-

screen as a reader progresses through a text, and assessing the impact on

the speed and success of reading that results from changing the available

information.

The look-ahead evident in Figure 7.5 plays an important part in reading.

Being able to tell the difference between shorter and longer words (using

the word length information marked by WL in the figure) helps the

reader to distinguish function and content words (the former tend to be

quite short) and to determine which words should be the landing point

for the next fixation. Identification of the beginning letters of words (BL)

and some of the letter features in the rest of the word (LF) allows the lexi-

cal search to be started before the word itself is looked at more closely

during a subsequent fixation.

This preview effect is one of many factors that influence fixation times

and jump distances during reading. Others include the difficulty of the

text, the predictability of the fixated words from the prior context, how

frequent the words are and how recently they have been encountered,

whether the words are ambiguous, and whether there is any priming of

the words by previous mention of related words or other material (see

Chapters 8 and 9 for further discussion of some of these parameters in the

context of word recognition).

The nature of the input It is important to note at this point that languages do not all use the same

writing system as English, and that many of the differences between writ-

ing systems will have consequences for the visual perception of words or

for their recognition. The following are some of the differences. First,

there is a range of orthographic writing systems, i.e. of writing systems

that represent some aspect of the sounds of words. As we will see in

Chapter 9 , pronunciation plays a role during reading, and so the relation-

ship between letters and sounds is important. English uses an alphabetic

system, but one in which there are many irregularities in the correspond-

ences between letters and sounds. For instance <c> has a /s/ pronunciation

in cease , but a /k/ pronunciation in cat . Languages with a high degree of

irregularity are said to have a deep orthography . They contrast with lan-

guages in which there is a reasonably direct letter-sound correspondence,

and therefore a shallow orthography , such as Italian, Serbo-Croatian or

Ma – ori, as well as with languages in which the pronunciation of a particu-

lar letter string is reasonably predictable but where there may be many

letter strings with the same pronunciation (e.g. French, in which the spell-

ings <o>, <au>, <aux>, <eau>, <eaux> can represent the same sound).

Other orthographic writing systems include consonantal systems, such

as Hebrew and Arabic, in which the letters often represent the conso-

nants only, and syllabic systems, such as Kannada and the Japanese kana

writing system, in which each orthographic symbol represents a syllable.

Looking beyond orthographic systems, we find ideographic systems, such

as in Chinese, where a symbol corresponds to a word (or a portion of a

word in a morphologically complex form); with the exception of diacritic

https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at

Perception for language 113

markers, the symbol does not bear a direct representation of the sounds

of the word.

The range of different writing systems and of the relationships of these

writing systems to pronunciation adds to the complexity of the study of

visual perception and visual word recognition.

7.5 Influence of the linguistic system on perception

There are many ways in which our perception is influenced by the linguis-

tic system, i.e. by the fact that the sounds or visual shapes we are process-

ing might be considered part of language. We saw evidence of this earlier

(p. 105) where phoneme restoration effects depended on the sentence

context in which the replaced phoneme occurred. We see it also in the

perceptual advantages that language-relevant stimuli have. For example,

studies in which sounds are embedded in noise have shown that speech

sounds are more reliably identified or recalled than non-speech sounds.

On top of this, if the speech sounds make up real words rather than

nonsense, then they have an even greater processing advantage.

Neurophysiological studies, using EEG and MEG to measure electrical

activity in the brain, show that this processing advantage for real words

over nonsense words becomes distinct as early as 150 msec after the point

at which a word can be recognised (Pulvermüller et al. , 2001 ). The fact that

this is a linguistic effect rather than a consequence of differences in the

sounds involved is shown by the fact this is found for native speakers of

the language being tested (Finnish in this case) but not for foreigners who

do not know the language.

A very significant area in which a linguistic influence on speech percep-

tion has been demonstrated is categorical perception . Work in this area

dates back to the 1950s (Liberman, Harris, Hoffman & Griffith, 1957 ), but

our understanding of what is involved in categorical perception has been

refined over the intervening decades. As the name suggests, categorical

perception relates to the finding that we hear speech sounds as belonging

to categories. That is, if we are presented with a series of speech stimuli

that vary along some linguistically relevant phonetic dimension, then we

will classify some of them as belonging to category X and some as belong-

ing to category Y, rather than hearing a particular stimulus as being X-ish,

another as being a little more Y-like, and so on.

Although categorical perception has been demonstrated for a range of

contrasts, the one most widely cited is the distinction between voiced and

voiceless plosive consonants, e.g. between /b/ and /p/. This distinction is

actually signalled by a range of different parameters, but the one that has

usually been explored is voice onset time (VOT). Voice onset time is the lag

between the release of the closure for the plosive consonant (in this case

the opening of the lips) and the beginning of voicing for a following vowel.

This is shorter for a voiced plosive like /b/ than for a voiceless one like

/p/ – in English typically less than and more than 30 msec respectively.

https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at

114 INTRODUCING PSYCHOLINGUISTICS

In experiments, participants are given synthesised /ba/-like or /pa/-like

syllables to respond to. The response can either be an identification one

(is this /ba/ or /pa/?) or a discrimination one (are two stimuli the same or

different?). The identification responses show that rather than there

being a gradual change of responses from /ba/ to /pa/ as VOT increases,

there is a dramatic switch of response preference part-way along the con-

tinuum. In terms of the hypothetical responses shown in Figure 7.6 ,

responses to VOT continua follow a categorical function – also referred to

as an S-curve because of the shape of the line in such a figure – rather

than a smooth or straight-line function.

Responses in the discrimination tasks show that participants treat as

‘same’ two stimuli that fall on the same side of the dramatic shift in the

identification response, e.g. in terms of Figure 7.6 , stimuli 2 and 3, or

3 and 4, or 5 and 6, but will label as ‘different’ two stimuli that fall on

either side of the shift, i.e. stimuli 4 and 5 in the figure. They do this even

if in terms of VOT the difference between 4 and 5 is of the same magni-

tude as the difference between 3 and 4 or between 5 and 6. This is evi-

dence that sounds are placed into perceptual categories, with rather

sharp boundaries between the categories. These findings have been sup-

ported by neurophysiological studies that have demonstrated different

brain responses (measured through changes in event-related magnetic

fields, or ERFs) to tokens either side of a category boundary, but similar

responses to stimuli within a VOT category (e.g. Simos et al., 1998 ).

Categorical perception makes sense from the point of view of language

perception and comprehension. This is because we expect spoken sound

to indicate linguistic entities as part of the comprehension process.

Studies with infants have shown that they are able to discriminate

categories, e.g. between /b/ and /p/, at a very young age (around 3 months

old), before they can speak (Eimas, Siqueland, Jusczyk & Vigorito, 1971 ).

While an initial reaction to this finding was to conjecture that phonetic

Figure 7.6 Smooth and categorical response functions.

https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at

Perception for language 115

distinctions are a specially evolved and innate human characteristic that

serves as a precursor to speech, a number of factors indicate that this is

not the case. First, it turns out that there is nothing special about

speech – other stimuli, including non-speech sounds, can also be distin-

guished categorically, with appropriate training (as demonstrated in

early work by Lane, 1965 ). Second, it was found that humans are not

alone in making categorical perceptual distinctions between speech

sounds – chinchillas for instance can also do this, even though they have

not evolved to speak (Kuhl, 1987 ). Third, there are cross-linguistic differ-

ences in where the VOT category boundary between, say, /b/ and/p/ is

found. This implies that the categories themselves, and the boundaries

between the categories, have to be learned (Abramson & Lisker, 1973 ),

although some studies have demonstrated that infants may initially be

more sensitive to certain category boundaries than to others (Aslin,

Pisoni, Hennessey & Perey, 1981 ).

In further categorical perception studies with adults, it has been

shown that category boundaries are not fixed, even within a language,

but can be affected by the linguistic context. We saw this earlier in this

chapter in the discussion of listeners’ compensation for coarticulation

(see Figure 7.4 ), where the /k/ identification function for tapes vs capes was

shifted depending on the nature of the consonant at the end of the pre-

ceding word.

The Ganong effect (named after the author of the first famous study of

this phonemenon) is a good demonstration of how the boundaries

between phonetic categories can be affected by linguistic information

(Ganong, 1980 ). In this case, it is the lexical status of the word containing

the manipulated segment that is important. For a /d/--/t/ continuum, for

instance, the category boundary falls nearer the /d/ end of the continuum

if the stimulus is /?ask/ (where ? indicates the manipulated segment). This

is because task is a word and dask is not, so more of the stimuli towards

the /d/ end of the continuum are accepted as falling within the /t/ category.

Conversely, the boundary falls nearer to the /t/ end if the stimulus is /?esk/,

i.e. more desk responses are given overall, because desk is a word and tesk

is not.

The perceptual

abilities of infants

are measured

using methods

such as High

Amplitude

Sucking. An

infant sucks on a

dummy connected

to equipment that

measures how

hard she is

sucking and

controls the

presentation of

stimuli. When the

infant gets bored

with the same

stimulus, her

sucking intensity

reduces, at which

point a new

stimulus is

presented. If she

hears this as

different, then she

becomes excited

and her sucking

increases.

Summary

We have seen that there are some clear commonalities in language

perception in the auditory and visual modalities:

• the input has to be segmented and analysed into linguistically useful units;

• variability in the input has to be overcome.

https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at

116 INTRODUCING PSYCHOLINGUISTICS

But there are also some differences in how speech and writing are

processed, which result from the obvious differences between the

modalities:

• speech is transitory; • speech lacks many clear boundaries between the linguistic units it

contains;

• writing has greater permanence and can be re-inspected with rela- tive ease;

• writing contains clearer boundaries between words and within words between letters.

We have also seen some key issues relevant to perception for language:

• the human brain is specialised for language processing; • the perceptual system streams language and non-language stimuli; • but will integrate stimuli if this makes linguistic sense; • cue integration includes visual and tactile cues as well as auditory

ones;

• the linguistic system can have an influence on perception, leading to effects such as real word advantages and/or biases in perception;

• infants begin to tune-in to the relevant cues for the language they are learning at a very early age.

Exercises Exercise 7.1 If you replaced some of the end of the /s/ in the words soil , suit and seep with silence, would you expect to hear a /p/ in each case? Why? If not, why not? (If you want to hear this for yourself, the website has examples of these words in the original and manipulated forms. Alternatively, use a speech editing package such as Praat ( www.praat.org ) to record and manipulate words like this yourself.)

Exercise 7.2 The sound file represented in Figure 7.2 shows two nasal sounds (at the end of keen and of team ). Can you identify a characteristic of nasal sounds from the spectrogram? Notice that the vowel portions of the words to and the look very similar to one another in the spectrogram. Why do you think this is the case? Figure 7.3 showed four examples of the /i/ vowel from this sound file. Can you explain any of the differences between these examples? What do they have in common that might tell you that the same vowel is involved?

Exercise 7.3 Work with two fellow students, friends or relatives for this exercise, forming a group of three – A, B and C. A watches B’s eyes as B either reads a text or watches C walking across the room. A should note any differences in the pat- terns of B’s eye movements. What differences are there? Swap roles and do this again.

https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at

Perception for language 117

Further reading

Samuel ( 1990 ) provides a review of phoneme restoration effects. The McGurk effect has been reported in many places, and a short and accessible overview is given by McGurk & MacDonald ( 1976 ). An excellent review of eye movements during reading, with particular reference to word recognition, is provided by Rayner and Balota ( 1989 ). A recent review of work in speech perception, including discussion of the recent interest in episodic exemplar representations, can be found in Pisoni and Levi ( 2007 ).

Exercise 7.4 The perceptual integration of visual and auditory inputs known as the McGurk effect is a powerful effect. Would you expect it to happen if the face you see is clearly not the origin of the voice you hear (e.g. if the face is female and the voice is male)? Look at the example of this on the website.

https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at

https://www.cambridge.org/core/terms. https://doi.org/10.1017/CBO9780511978531.008 Downloaded from https://www.cambridge.org/core. University of California, Santa Cruz, on 29 Jul 2020 at 09:59:46, subject to the Cambridge Core terms of use, available at