Seyed Masoud Hosseini · Overview · Study log · Ideas · Transcript · RSS feed
The Human Brain · Lecture 15 of 17 · 1:18:30
Lecture 15: Hearing and Speech
Study guide
What this lecture covers
This lecture asks how the auditory system extracts so much from a single, simple input: air pressure compressions arriving at the ear. It works through the computational theory of hearing (what problem the brain is solving and why it's hard), then moves to what is known about the ear and auditory cortex, ending with recent human neuroimaging work on speech-selective cortex. It sits alongside the course's earlier treatment of vision, repeatedly drawing analogies between the two senses.
After watching, you should be able to name the three classic computational challenges of hearing, explain why reverb and multi-talker speech are ill-posed problems, describe the tonotopic organization of primary auditory cortex, and explain how spectrotemporal receptive field models were tested in humans using synthetic "model-matched" sounds.
Key ideas
- Spectrogram: a plot of sound energy across frequency (vertical axis) and time (horizontal axis), produced by a Fourier analysis of the sound wave; vowels show up as stacked horizontal harmonics, consonants as vertical smears.
- Cocktail party effect: the ability to selectively attend to one sound source when multiple sources are superimposed in the incoming signal, a classic ill-posed problem because the mixture alone doesn't determine the separate sources.
- Reverb as an ill-posed problem: the sound reaching your ears is the source sound convolved with the room's impulse response; recovering the source requires the auditory system to use built-in knowledge of how real-world reverb decays with frequency.
- Formant: a band of concentrated frequency energy in a vowel sound; the timing of formant transitions (for example, before or after a consonant burst) distinguishes sounds like
bafrompa. - Phoneme: a sound distinction that changes word meaning in a given language; which sounds count as phonemes differs across languages (English has no phonemic click consonants, though it uses clicks non-linguistically).
- Talker variability: the same vowel produced by different speakers lands at very different points in formant space, so recognizing a word and recognizing a voice are mutually confounded, each partly solved using knowledge of the other.
- Tonotopic map: primary auditory cortex is organized by frequency (a high-low-high gradient) rather than by spatial location, reflecting the physical frequency separation already performed by the cochlea.
- Spectrotemporal receptive field (STRF): a linear-filter model of how an auditory cortex neuron responds to changes in frequency over time; used to generate synthetic sounds that let researchers test whether human A1 works the same way.
Walkthrough
What hearing can do, and where the field starts (0:10)
The lecture opens with demonstrations of everyday auditory abilities: localizing a speaker with your eyes closed, recognizing environmental sounds and materials from a single impact, and identifying what an object is made of just from the sound of it hitting a table. These demonstrations motivate the computational approach used throughout the course: define the inputs and outputs, then ask what information in the stimulus makes the task possible.
What sound is, and reading a spectrogram (5:11)
Sound is described physically as traveling compressions and rarefactions of air, illustrated with schlieren photography and a speaker vibrating paint. The lecture then introduces spectrograms, showing how whistling produces a single frequency band, a trombone produces multiple harmonic bands (a pitched sound), and speech produces a mix of harmonic vowel stripes and vertical consonant bursts.
Three challenges: invariance, the cocktail party problem, and reverb (9:13)
The lecture lays out why extracting meaning from sound is computationally hard. Invariance problems mean the same word sounds different across speakers while the same speaker's different words also look very different, so recognizing words and recognizing voices require separate kinds of invariance. The cocktail party problem is the challenge of separating simultaneous overlapping sound sources, framed as an ill-posed problem with infinite possible decompositions. Reverb compounds this: real-world sound is a mix of the direct signal and many delayed reflections. A study by Josh McDermott and colleagues is described in which researchers measured impulse response functions across many real locations and showed that listeners use built-in knowledge of natural reverb decay to recover the original source; synthetic, non-physical reverb breaks this ability.
Speech perception: formants, phonemes, and talker variability (24:28)
The lecture examines spectrograms of vowels and consonants, showing how formant bands distinguish vowels and how the timing of a formant transition distinguishes consonants such as ba and pa. It introduces phonemes as language-specific sound distinctions, illustrated with click consonants from Xhosa that are phonemic in some southern African languages but not in English. The talker variability problem is shown with an overlapping-vowel-space plot: because different speakers' vowels overlap heavily in formant space, understanding speech and recognizing a voice mutually constrain each other. A study on dyslexic readers is used as evidence: typical listeners recognize voices more accurately in a language they know, but this advantage disappears in dyslexic listeners, consistent with dyslexia involving impaired speech-sound processing.
From cochlea to cortex: the tonotopic map (42:43)
The lecture traces the physical path of sound from the ear canal through the tympanic membrane and middle-ear bones to the cochlea, which performs a physical frequency separation along its length, with nerve endings at each position tuned to a different frequency band. It contrasts the many subcortical relay stops in the auditory pathway with the single relay (retina to LGN to cortex) in vision, then shows how this frequency map is preserved as a tonotopic high-low-high gradient in primary auditory cortex.
Testing receptive field models in humans, and speech-selective cortex (57:51)
Animal recordings show primary auditory cortex neurons behave like spectrotemporal filters (STRFs) tuned to particular frequency changes over time. The lecture describes a study by McDermott and Norman-Haignere that generated synthetic "model-matched" sounds designed to produce identical STRF responses to real recordings, then played both to human listeners in an MRI scanner. Primary auditory cortex responded almost identically to original and synthetic versions, supporting the STRF model, but regions just outside primary auditory cortex responded very differently, especially for speech and music, whose synthetic versions became unrecognizable. This led to identifying a band of cortex selective for speech sounds, responding similarly to speech in known and unknown languages, showing the selectivity is for phonemes rather than language meaning, and shown to be robust to matching for low-level acoustic properties and backward speech.
Before you watch
- Review the vision lectures on ill-posed problems and color constancy in this course; the lecture repeatedly compares reverb and voice/word confounds to those problems.
- Be comfortable reading a basic spectrogram (time versus frequency, with color for intensity).
- Recall the distinction between primary sensory cortex and downstream, more selective cortical regions from the vision lectures.
Check your understanding
- Why are the cocktail party problem and the reverb problem both described as "ill-posed," and what kind of information does the auditory system bring in to constrain them?
- How does formant transition timing distinguish a
basound from apasound? - Explain why recognizing a talker's voice and recognizing the words they say are mutually dependent problems, and how the dyslexia study supports this.
- What does the tonotopic organization of primary auditory cortex represent, and how does the cochlea produce it?
- How did researchers use "model-matched" synthetic sounds to test whether spectrotemporal receptive fields explain responses in human primary auditory cortex, and what did the result outside primary auditory cortex suggest?
Chapters
- 0:00 <Untitled Chapter 1>
- 1:50 Environmental Sounds
- 2:31 The Cocktail Party Effect
- 4:16 How Hearing Works
- 5:26 What Is Sound
- 7:23 Spectrograms of Different Sounds
- 8:33 Pitched Sounds
- 9:55 Invariance Problems
- 12:44 Ill-Posed Problem
- 14:59 Vision
- 15:32 Dry Speech with no Reverb
- 18:26 Impulse Response Function
- 19:30 Color Vision
- 23:15 Measuring Reverb
- 24:22 Speech Perception
- 25:27 Intonation of Speech
- 26:14 Formant
- 27:56 Consonants
- 33:15 Phonemes
- 34:13 Click Consonant
- 35:59 Click Consonants
- 36:50 Why Is Speech Perception Challenging
- 37:31 Differences across Speakers in the Language
- 37:50 Talker Variability
- 42:23 Accuracy at Voice Recognition
- 42:47 Dyslexics
- 43:34 The Ear
- 43:53 The Cochlea
- 44:32 Low Frequency Sound Waves
- 48:15 Cortex
- 48:23 Primary Auditory Cortex
- 49:35 Functional Mri
- 53:40 Spectrotemporal Receptive Fields
- 55:33 Design of the Experiment
- 58:10 Model Match Stimuli
- 1:00:38 Control Stimulus
- 1:07:19 Selective Responses to Speech
- 1:10:58 Non-Speech Vocalizations
- 1:11:22 Is Instrumental Music Perceived as Speech
From the YouTube description
MIT 9.13 The Human Brain, Spring 2019
Instructor: Nancy Kanwisher
View the complete course: https://ocw.mit.edu/9-13S19
YouTube Playlist: https://www.youtube.com/playlist?list=PLUl4u3cNGP60IKRN_pFptIBxeiMc0MCJP
Humans use hearing in species-specific ways, for speech and music. Ongoing research is working out the functional organization of these and other human auditory skills.
* NOTE: Lecture 14: New Methods Applied to Number (student breakout groups—video not recorded)
License: Creative Commons BY-NC-SA
More information at https://ocw.mit.edu/terms
More courses at https://ocw.mit.edu
Support OCW at http://ow.ly/a1If50zVRlQ
We encourage constructive comments and discussion on OCW’s YouTube and other social media channels. Personal attacks, hate speech, trolling, and inappropriate comments are not allowed and may be removed. More details at https://ocw.mit.edu/comments.
