Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

5. Psychoacoustics

How we hear, perceive, and are fooled by sound

This week we turn to psychoacoustics, the study of how we perceive and interpret sound. It sits between the physical properties of a sound wave and the experience of hearing it.

In terms of the four levels of description, this is the step from the physical signal of the past two weeks to perception. Watch for the places where the two come apart: equal-loudness contours, masking, and the missing fundamental are all cases where the ear reports something the microphone does not. Why does one sound seem louder than another of the same intensity? How do we tell a violin from a piano playing the same note, and why do some sounds please us while others grate? Questions like these set up the more detailed look at “vertical” and “horizontal” sound perception in the coming weeks.

Introduction to psychoacoustics

Psychoacoustics is a branch of psychophysics, the study of how physical stimuli relate to the sensations and perceptions they produce. Psychophysics covers all the senses; psychoacoustics narrows the focus to hearing. Both use experiments to measure how changes in physical properties such as frequency, intensity, or duration map onto changes in perception such as pitch, loudness, or timbre.

These insights have practical uses. They inform audio engineering, music production, and hearing aid design, and they underpin audio codecs such as MP3, which exploit the limits of perception to compress audio without an audible loss in quality. The standard reference works for the field are Moore (1997) and Fastl & Zwicker (2007); Cook (1999) covers much of the same ground with audio demonstrations.

The roots of the field go back to the 19th century. Hermann von Helmholtz (1821–1894) studied the sensations of tone and the physical basis of music, and later S. S. Stevens (1906–1973) developed methods for psychophysical scaling that quantify the relationship between a stimulus and how it is perceived (for example, Stevens’ power law for loudness). Since then the field has grown to draw on physics, psychology, neuroscience, and engineering.

Auditory perception covers the processes by which the ear and brain detect, analyse, and make sense of acoustic signals, turning vibrations in the air into speech, music, or environmental sounds. Psychoacoustics explains the perceptual side of this: why a sound is heard as louder, higher, or more pleasant, and how we separate different sources in a complex scene.

Cognition adds the higher-level functions: attention, memory, learning, and decision-making. Where psychoacoustics deals with the sensory and perceptual mechanics of hearing, cognition shapes how we interpret, remember, and respond to what we hear. Cognitive processes let us follow a single voice in a noisy room (the “cocktail party effect”), recognise a familiar melody, or tie a sound to an emotion or memory.

The human auditory system

The human ear lets us detect sounds and keep our sense of balance. It has three main sections: the outer ear, middle ear, and inner ear, each with a distinct job in the hearing process.

The outer ear

The outer ear consists of the visible part called the pinna and the ear canal. Its job is to collect sound waves from the environment and funnel them towards the eardrum (sometimes called the tympanic membrane). The shape of the pinna helps us locate the direction a sound comes from. When sound waves reach the eardrum, they make it vibrate.

Outer Ear

Image Source: Wikipedia - Outer Ear

The eardrum works much like the membrane in a microphone, familiar from the previous chapter. Both are sensitive barriers that vibrate in response to incoming sound waves. In the ear, vibrations of the eardrum pass through the ossicles to the inner ear, where they become electrical signals for the brain to interpret. In a microphone, sound waves hit a thin, flexible diaphragm; its vibrations are converted into electrical signals that can be amplified, recorded, or transmitted. In both cases, air pressure variations become mechanical vibrations, which is the first step in turning acoustic energy into a form that can be processed further, biologically in the ear and electronically in the microphone.

The middle ear

Beyond the eardrum lies the middle ear, which contains three tiny bones known as the ossicles: the malleus, incus, and stapes. These bones act as a mechanical lever system, amplifying the vibrations from the eardrum and passing them to the oval window, a membrane-covered opening to the inner ear. The amplification matters because sound energy has to transfer efficiently from air into the fluid-filled inner ear.

Middle Ear

Image Source: Wikipedia - Middle Ear

This is much like a microphone running into a preamplifier. The signal from a microphone membrane is weak, so it is usually sent through a preamplifier to bring it up to a level suitable for recording or further processing. In the same way, the ossicles boost the mechanical vibrations from the eardrum so the signal is strong enough to drive the inner ear.

The inner ear

The inner ear is where mechanical vibrations become electrical signals that the brain can interpret as sound. The main structure for this is the cochlea, a spiral-shaped, fluid-filled organ lined with thousands of tiny hair cells. As vibrations travel through the cochlear fluid, they move the hair cells, generating nerve impulses that travel to the brain along the auditory nerve.

Inner Ear

Image Source: Wikipedia - Inner Ear

The hair cells are laid out in an orderly map. Those near the base of the cochlea respond best to high frequencies, those near the apex to low ones. This frequency map is called tonotopic organisation, and it is preserved all the way up to the auditory cortex (see the brain). It means the ear delivers something close to a spectrum to the brain, rather than a raw waveform.

The cochlea works rather like the analogue-to-digital converter (ADC) from the previous chapter. An ADC turns continuous analogue sound waves into discrete digital signals that a computer can process; the cochlea turns mechanical vibrations into electrical nerve impulses that the brain can read. Inside the cochlea, different hair cells respond to different frequencies, so the organ performs a biological frequency analysis while also encoding the intensity and timing of sounds, much as an ADC samples and quantises an audio signal.

The inner ear also contains the vestibular system, which keeps us balanced and spatially oriented. We will return to this when we get to music-related body motion.

Human vs machine perception

The human auditory system and machine-based audio systems (microphones and ADCs) are built on completely different biological and technological foundations, but the analogy between them is useful. Each is a chain of components that transforms and processes sound, turning physical vibrations in the air into meaningful information: neural signals in the brain, digital data in a computer.

The different parts are as follows:

  • Transduction: In humans, the eardrum and ossicles convert air pressure variations into mechanical vibrations, and the cochlea converts these into electrical nerve impulses. In machines, a microphone membrane converts sound waves into electrical signals, which are then amplified and digitised.
  • Frequency Analysis: The cochlea performs a kind of real-time frequency analysis, with different regions responding to different frequencies (a biological “filter bank”). Digital systems do something similar with mathematical transforms such as the Fourier transform.
  • Encoding and Transmission: The auditory nerve encodes and transmits information about sound to the brain, where it is processed and interpreted. In digital systems, the ADC encodes the analogue signal as binary data that can be stored, transmitted, and processed by computers.

There are also important differences:

  • The human auditory system is adaptive and context-sensitive, shaped by attention, learning, and memory. It can pick out specific sounds in noisy surroundings (the “cocktail party effect”) and fill in missing information from prior experience.
  • Machine perception is bounded by hardware (microphone quality, sampling rate) and by the algorithms used for analysis. Modern systems can do impressive work, but they lack the flexibility and subjective interpretation of human hearing.

These parallels and differences matter for audio engineering, hearing aid design, and music information retrieval, where the aim is often to connect physical sound with perceptual experience.

Loudness

As we saw in the acoustics chapter, loudness is not the same as sound pressure level (SPL). Sound pressure level (SPL) is an objective, physical measurement of the intensity of a sound wave, expressed in decibels (dB). Loudness is the subjective side: how “loud” a sound feels to a listener. It is shaped by context, such as background noise and recent exposure to other sounds, and the gap between the two runs through much of audio engineering, hearing science, and music production.

Measuring loudness

The most common unit for perceived loudness is the phon, which is based on equal-loudness contours (more on these below). The phon scale is anchored to a 1,000 Hz pure tone. A sound judged as loud as a 40 dB SPL, 1,000 Hz tone has a loudness of 40 phons, whatever its own frequency or SPL.

Several standards exist for estimating loudness in audio signals. ITU-R BS.1770 is widely used in broadcasting; it defines algorithms for measuring loudness and true-peak levels, and it forms the basis for loudness normalisation in radio, TV, and streaming.

Modern audio software and digital audio workstations (DAWs) often include loudness meters that display values in LUFS: loudness units relative to full scale. These tools help engineers and producers meet broadcast standards and keep listening comfortable. When mastering music for streaming, engineers typically aim for an integrated loudness of around -14 LUFS, as recommended by services like Spotify and YouTube. In film and TV, loudness normalisation keeps dialogue, music, and effects balanced so viewers do not have to keep reaching for the volume control.

Threshold of hearing

The threshold of hearing is the quietest sound the average ear can detect in a silent environment. It is usually defined as 0 decibels (dB SPL) at 1,000 Hz for a healthy young adult, but it varies with frequency and with individual hearing. At very low or very high frequencies, a sound has to be much louder before we hear it.

The threshold is not fixed. It shifts with age, with exposure to loud noise, and even with temporary conditions such as an ear infection or fatigue. As people age, sensitivity to high frequencies usually drops, raising the threshold there.

The graph below shows average hearing thresholds for different frequencies and age groups, and how sensitivity changes across the audible spectrum.

Threshold of Hearing

Image Source: Wikipedia - Hearing Thresholds

Safe listening levels

Here is a list of some different sound levels:

Sound SourceApproximate Loudness (dB SPL)Example / Context
Threshold of hearing0Quietest sound a healthy ear can detect
Rustling leaves10Very quiet, barely audible
Whisper20–30Soft whisper at close distance
Quiet library30–40Typical background noise
Normal conversation60At 1 metre distance
Busy street traffic70–85Inside a car with windows closed
Subway train95–100Inside the train
Rock concert110–120Near speakers
Threshold of pain130Jet engine at 30 metres
Fireworks / Gunshot140–150Close range

As a rule of thumb, prolonged exposure to sounds above 85 dB can damage hearing. Typical conversation is around 60 dB, city traffic can reach 85 dB, and concerts or clubs often exceed 100 dB. The louder the sound, the shorter the safe exposure time. At 100 dB, damage can set in within about 15 minutes.

The World Health Organization (WHO) recommends keeping personal listening devices below 80 dB for adults (75 dB for children) and limiting time in loud environments. Ear protection in noisy settings and regular listening breaks both help.

Sounds above 120–130 dB, such as a jet engine at close range or fireworks, can cause immediate pain and permanent hearing loss. Even brief exposure to extremely loud sounds can irreversibly damage the hair cells in the inner ear.

Tinnitus and hearing health

Tinnitus is the perception of sound, typically a ringing, hissing, or buzzing, when no external sound is present. Many people notice a temporary version after a loud concert, often together with a muffled feeling called a temporary threshold shift, a short-lived drop in hearing sensitivity while the ear recovers. With repeated overexposure the shift can become permanent, because the hair cells of the inner ear do not grow back once destroyed (see noise-induced hearing loss). Musicians are a risk group. Surveys of rock, pop, and orchestral musicians alike find more tinnitus and hearing loss than in the general population, simply because rehearsals and concerts add up to many hours above 85 dB.

The practical response is hearing protection, and it does not have to ruin the music. Ordinary foam earplugs dampen high frequencies much more than low ones, which makes music sound dull and distant. Musicians’ earplugs instead offer flat attenuation, reducing all frequencies by roughly the same amount, typically 10–20 dB, so the mix keeps its balance and simply becomes quieter. Filtered plugs cost little, and custom-moulded versions are a sensible investment for anyone who rehearses or performs regularly.

Some ringing after a loud night is common and usually fades. If tinnitus lasts more than a day or two, or if hearing still feels muffled once the sound has stopped, book a hearing test with a doctor or audiologist. Persistent tinnitus cannot yet be cured, but it can be assessed and managed, and an early check protects the hearing you still have.

Equal-loudness contours

SPL and loudness are related but not the same. The same SPL can be heard as louder or softer depending on the frequency of the sound and the listener’s hearing sensitivity.

Harvey Fletcher and Wilden A. Munson studied this systematically in the 1930s. Their experiments Fletcher & Munson, 1933 produced the Fletcher–Munson curves, now usually called equal-loudness contours. The curves show that the ear is not equally sensitive across frequencies. We hear best between roughly 2,000 and 5,000 Hz, and are less sensitive to very low or very high frequencies. So a low-frequency sound has to be played at a higher SPL than a mid-frequency sound to seem equally loud.

Equal-loudness contour

Image Source: Wikipedia - Equal-loudness contour

Our heightened sensitivity between 2,000 and 5,000 Hz may have evolutionary roots. This range matches the spectral content of human speech, especially the consonants that carry much of the meaning. Picking up small differences in these frequencies would have helped with communication and with detecting important sounds such as a baby’s cry or a warning call. Over time, natural selection may have favoured hearing tuned to this range.

Equal-loudness contours matter in audio engineering, music production, and hearing science. They explain why music or speech sounds different at low and high volumes, and why audio equipment often includes a “loudness” control that compensates for these perceptual differences.

Just noticeable differences for loudness

The just noticeable difference (JND), also called the difference limen, is the smallest change in a physical stimulus that a listener can reliably detect. For loudness, it is the smallest change in sound intensity that we hear as a change in loudness.

The JND for loudness is often given as a percentage of the original level. For mid-level sounds it is typically about 1 dB, so most people need at least a 1 dB change to notice a difference. The exact value depends on the frequency, the absolute loudness, and the listening environment.

Source
import numpy as np
import matplotlib.pyplot as plt

# Demonstrate equal loudness contours conceptually
# Equal loudness at different frequencies requires different SPL values

frequencies = np.array([20, 50, 100, 200, 500, 1000, 2000, 5000, 10000, 20000])
loudness_40_phon = np.array([60, 45, 35, 28, 20, 20, 18, 12, 10, 14])  # Approximate values
loudness_80_phon = np.array([85, 68, 55, 43, 33, 30, 28, 26, 26, 30])  # Approximate values

plt.figure(figsize=(10, 6))
plt.semilogx(frequencies, loudness_40_phon, 'o-', label='40 phons', linewidth=2, markersize=6)
plt.semilogx(frequencies, loudness_80_phon, 's-', label='80 phons', linewidth=2, markersize=6)
plt.xlabel('Frequency (Hz)')
plt.ylabel('Sound Pressure Level (dB SPL)')
plt.title('Equal-Loudness Contours (Approximate)')
plt.grid(True, which='both', alpha=0.3)
plt.legend()
plt.xlim(10, 20000)
plt.ylim(10, 90)
plt.tight_layout()
plt.show()

# Demonstrate just noticeable difference (JND) for loudness
spl_values = np.linspace(20, 100, 50)
jnd_db = 1  # Approximately 1 dB for mid-level sounds

plt.figure(figsize=(10, 5))
plt.fill_between(spl_values, spl_values - jnd_db, spl_values + jnd_db, alpha=0.3, label='JND range (~1 dB)')
plt.plot(spl_values, spl_values, 'b-', linewidth=2, label='Reference SPL')
plt.xlabel('Sound Pressure Level (dB SPL)')
plt.ylabel('Perceived Loudness (dB SPL)')
plt.title('Just Noticeable Difference for Loudness')
plt.legend()
plt.grid(True, alpha=0.3)
plt.tight_layout()
plt.show()
<Figure size 1000x600 with 1 Axes>
<Figure size 1000x500 with 1 Axes>

Pitch

When we hear tonal sounds—musical instruments, but also voices—we usually hear a pitch. Pitch is closely tied to frequency, the rate at which a sound wave vibrates, but it is ultimately a subjective experience shaped by the auditory system.

Pitch range

The audible range for humans runs from roughly 20 Hz to 20,000 Hz, covering the frequencies we perceive as pitch. Sounds below 20 Hz are infrasound and generally beyond our hearing, while those above 20,000 Hz are ultrasound and also lie outside our range. Within the audible spectrum, our sensitivity to pitch varies, and is sharpest in the mid-frequency range that matters most for speech and music.

human hearing

Dogs are well known for hearing better than humans, and the graphs below show that human hearing sits in the middle of the animal range. Different species have evolved to detect the frequencies most relevant to their survival and communication. Dogs can hear up to around 45,000 Hz, picking up high-pitched sounds we cannot. Bats and dolphins perceive even higher frequencies, well into the ultrasonic range, which they use for echolocation. Elephants and some whales, on the other hand, hear infrasound—very low frequencies below our threshold—which helps them communicate over long distances.

The chart below compares the hearing ranges of various animals and shows where human hearing fits in the wider spectrum. The differences reflect the ecological needs and evolutionary pressures each species has faced.

animal hearing

Just noticeable differences for pitch

There is a just noticeable difference (JND) for pitch too: the smallest change in frequency a listener can reliably detect. The ear is highly sensitive to small frequency changes, especially in the mid range (about 500–4000 Hz, where speech and music are most prominent). At 1000 Hz, the JND for pitch is typically around 3 Hz for trained listeners, about 0.3% of the frequency. In other words, play two tones at 1000 Hz and 1003 Hz and a trained listener can usually tell them apart. At very low or very high frequencies the JND grows larger, so small pitch differences become harder to hear.

The JND for pitch varies with factors such as the listener’s age, hearing ability, and training, the loudness of the tones, and whether the tones are heard in isolation or in a complex sound environment.

Timbre

Timbre is the quality of a sound that lets us tell different sources apart even when they share the same pitch and loudness. It is what distinguishes a piano from a violin playing the same note at the same volume. Timbre comes from the spectral content (the mix of the fundamental and the overtones), the temporal envelope (attack, decay, sustain, release), and other features such as vibrato or noise components. In music and audio, it is how we identify instruments, voices, and textures.

This is the definition the rest of the course uses. Timbre appeared in the previous chapter as something you shape deliberately in a signal chain, and it returns in harmony and melody as something that carries musical structure by holding voices apart or fusing them together.

Overtones

When an instrument or voice produces a note, it does not generate a single frequency (the fundamental) but also a series of higher frequencies called overtones or harmonics. These are integer multiples of the fundamental and give the sound its colour.

  • Fundamental: The lowest frequency of a sound, typically perceived as its pitch.
  • Overtones/Harmonics: Frequencies above the fundamental. The first overtone is the second harmonic (2× the fundamental), the second overtone is the third harmonic (3× the fundamental), and so on.

A pure sine wave has no overtones and sounds plain, while most musical sounds are complex because of their overtone content. The pattern and strength of the overtones set the timbre of an instrument. A clarinet and a violin playing the same note have different overtone structures, which is why they sound distinct.

Source
import numpy as np
import matplotlib.pyplot as plt

# Generate synthetic sounds for spectral comparison
sr = 22050  # Sample rate
duration = 1.0  # Duration in seconds
t = np.linspace(0, duration, int(sr * duration), endpoint=False)

# sound1: Pure sine tone (fundamental frequency: 440 Hz)
f0 = 440  # A4 note
sound1 = 0.5 * np.sin(2 * np.pi * f0 * t)

# sound2: Complex tone with overtones (fundamental + harmonics)
# Mix of fundamental + 2nd, 3rd, and 4th harmonics
sound2 = (0.5 * np.sin(2 * np.pi * f0 * t) +        # Fundamental
          0.3 * np.sin(2 * np.pi * 2 * f0 * t) +    # 2nd harmonic
          0.2 * np.sin(2 * np.pi * 3 * f0 * t) +    # 3rd harmonic
          0.1 * np.sin(2 * np.pi * 4 * f0 * t))     # 4th harmonic
sound2 = sound2 / np.max(np.abs(sound2)) * 0.5  # Normalize
Source
import librosa
import librosa.display
import matplotlib.pyplot as plt

plt.figure(figsize=(10, 6))

# Compute STFT for spectrograms
S1 = librosa.amplitude_to_db(np.abs(librosa.stft(sound1, n_fft=128)), ref=np.max)
S2 = librosa.amplitude_to_db(np.abs(librosa.stft(sound2, n_fft=128)), ref=np.max)

# Spectrogram for sound1 (log scale, limited to 0-5000 Hz)
plt.subplot(2, 1, 1)
librosa.display.specshow(S1, sr=sr, x_axis='time', y_axis='hz', cmap='viridis')
plt.ylim(0, 5000)
plt.title('Spectrogram (0-5000 Hz): Pure Tone')
plt.xlabel('Time (s)')
plt.ylabel('Frequency (Hz)')
plt.colorbar(format='%+2.0f dB')

# Spectrogram for sound2 (log scale, limited to 0-5000 Hz)
plt.subplot(2, 1, 2)
librosa.display.specshow(S2, sr=sr, x_axis='time', y_axis='hz', cmap='viridis')
plt.ylim(0, 5000)
plt.title('Spectrogram (0-5000 Hz): Tone with Overtones')
plt.xlabel('Time (s)')
plt.ylabel('Frequency (Hz)')
plt.colorbar(format='%+2.0f dB')

plt.tight_layout()
plt.show()
<Figure size 1000x600 with 4 Axes>

Envelope

The envelope of a sound describes how its amplitude changes over time, shaping the sound’s overall character and dynamics. It governs how a sound starts, develops, and ends, and it helps us tell different instruments and sound types apart.

In audio synthesis, the most common model is the ADSR envelope, which stands for Attack, Decay, Sustain, and Release:

  • Attack: The time it takes for the sound to reach its maximum amplitude after being triggered.
  • Decay: The time it takes for the amplitude to decrease from the peak level to the sustain level.
  • Sustain: The level at which the sound holds while the note is sustained.
  • Release: The time it takes for the sound to fade to silence after the note is released.
ADSR

Image source: ADSR envelope on Wikipedia.

Just noticeable differences of timbre

How many partials do you need to recognise the timbre of a sound? Below you can find examples of a Sonny Rollins saxophone tone that has been spectrally decomposed and then resynthesised with an increasing number of overtones. The code counts how many harmonics of the fundamental actually rise above the noise floor in each version, and labels the players with that number. Notice that the count stops growing near the end: the last few resyntheses add partials so weak that they barely change what you hear.

Source
import librosa
import numpy as np
import matplotlib.pyplot as plt
from pathlib import Path

# The nine resyntheses of one saxophone tone, each adding more of the
# overtone series on top of the fundamental.
examples = sorted(Path("audio").glob("rollins_ex16?.wav"))

f0 = 210.0          # fundamental of the analysed tone, in Hz
n_fft = 16384

spectra, partial_counts = [], []
for path in examples:
    y, sr = librosa.load(path, sr=None, mono=True)
    mag = np.abs(librosa.stft(y, n_fft=n_fft)).mean(axis=1)
    freqs = librosa.fft_frequencies(sr=sr, n_fft=n_fft)
    db = librosa.amplitude_to_db(mag, ref=mag.max())
    spectra.append((freqs, db))
    # Count how many harmonics of the fundamental rise above -45 dB.
    harmonics = [n for n in range(1, 45)
                 if db[(freqs > n * f0 - 25) & (freqs < n * f0 + 25)].max() > -45]
    partial_counts.append(len(harmonics))

fig, axes = plt.subplots(3, 3, figsize=(11, 7), sharex=True, sharey=True)
for ax, (freqs, db), n, path in zip(axes.ravel(), spectra, partial_counts, examples):
    ax.plot(freqs, db, linewidth=0.8)
    ax.set_xlim(0, 6000)
    ax.set_ylim(-80, 5)
    ax.set_title(f"{n} partial{'s' if n > 1 else ''}", fontsize=10)
    ax.grid(alpha=0.3)
for ax in axes[-1]:
    ax.set_xlabel("Frequency (Hz)")
for ax in axes[:, 0]:
    ax.set_ylabel("Level (dB)")
fig.suptitle("One saxophone tone, resynthesised with a growing overtone series")
fig.tight_layout()
plt.show()
<Figure size 1100x700 with 9 Axes>
Source
from IPython.display import display, Audio, Markdown

for path, n in zip(examples, partial_counts):
    display(Markdown(f"**{n} partial{'s' if n > 1 else ''}** ({path.name})"))
    y, sr = librosa.load(path, sr=None, mono=True)
    display(Audio(y, rate=sr))
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...

Spatial perception

Spatial hearing is our ability to perceive where sound sources are and how they move. It lets us judge whether a sound is in front, behind, above, below, or to the side. We rely on it to navigate, to follow speech in busy settings, and to enjoy immersive audio in music and virtual reality.

Binaural hearing is the use of both ears together to pick up spatial cues and locate sounds. By comparing differences in timing (ITD) and intensity (ILD) between the two ears, the brain builds a three-dimensional auditory scene. This is what gives us depth, localisation, and the ability to separate sources in a complex environment.

Interaural time difference

The interaural time difference (ITD) is the difference in the arrival time of a sound at each ear. When a source is off-centre, the sound reaches one ear slightly before the other. The brain uses this tiny gap, especially for low-frequency sounds, to locate the direction of the source on the horizontal plane. ITD is a primary cue for the azimuth (left-right position) of a sound.

ITD

Interaural level difference

The interaural level difference (ILD) is the difference in sound pressure level reaching each ear. When a sound comes from one side, the head acts as a barrier, so the sound is louder in the near ear and quieter in the far ear. ILD works best for high-frequency sounds, where this head-shadow effect is more pronounced. The brain combines ILD with ITD to locate sounds accurately. With headphones on, you can hear the two cues separately in the Spatial hearing app, which moves a source around a schematic head with ITD and ILD switchable on and off.

ILD

The head-related transfer function (HRTF) describes how an ear receives a sound from a particular point in space, taking into account the effects of the listener’s head, torso, and outer ear (pinna). HRTFs are unique to each individual and are key to perceiving elevation and front-back differences. They are widely used in 3D audio and virtual reality to simulate realistic spatial sound.

HRTF

The precedence effect is a phenomenon where the first-arriving sound dominates our sense of where the sound is, even when reflections or echoes follow shortly after. It helps us lock onto the direct source in reverberant spaces, such as a speaker’s voice in a large hall, rather than getting confused by reflections.

Cocktail party effect

The cocktail party effect is our ability to focus on a single sound source, such as a conversation partner, in a noisy environment full of competing sounds. This selective attention draws on spatial hearing cues as well as cognitive processes to filter out background noise and bring out the target sound. It is a central part of auditory scene analysis and of how we manage to communicate in social settings.

Source
import numpy as np
import matplotlib.pyplot as plt

# Visualize interaural time difference (ITD) for sound localization
# ITD helps localize sound on the horizontal plane

angles = np.linspace(-180, 180, 100)  # Azimuth angles in degrees
head_radius = 0.087  # Head radius in meters (approximately 8.7 cm)
sound_speed = 343  # Speed of sound in m/s at 20°C

# Calculate ITD based on angle (simplified model)
itd = (head_radius / sound_speed) * np.sin(np.deg2rad(angles)) * 1000  # in milliseconds

plt.figure(figsize=(12, 4))

# ITD plot
plt.subplot(1, 2, 1)
plt.plot(angles, itd, 'b-', linewidth=2)
plt.axhline(y=0, color='k', linestyle='--', alpha=0.3)
plt.xlabel('Azimuth Angle (degrees)')
plt.ylabel('Interaural Time Difference (ms)')
plt.title('Interaural Time Difference (ITD)')
plt.grid(True, alpha=0.3)
plt.xlim(-180, 180)

# Illustration of ITD concept
plt.subplot(1, 2, 2)
angle_example = 45  # Example angle
itd_example = (head_radius / sound_speed) * np.sin(np.deg2rad(angle_example)) * 1000

# Draw head top view
head_circle = plt.Circle((0, 0), head_radius * 100, fill=False, color='brown', linewidth=2)
plt.gca().add_patch(head_circle)

# Draw ears
left_ear_y = head_radius * 100 * 0.7
right_ear_y = -head_radius * 100 * 0.7
plt.plot(-head_radius * 100, left_ear_y, 'rs', markersize=10, label='Left ear')
plt.plot(head_radius * 100, right_ear_y, 'bs', markersize=10, label='Right ear')

# Draw sound source
sound_dist = 3
source_x = sound_dist * np.cos(np.deg2rad(angle_example))
source_y = sound_dist * np.sin(np.deg2rad(angle_example))
plt.plot(source_x, source_y, 'g*', markersize=20, label='Sound source')

# Draw lines from source to ears
plt.plot([source_x, -head_radius * 100], [source_y, left_ear_y], 'g--', alpha=0.5)
plt.plot([source_x, head_radius * 100], [source_y, right_ear_y], 'g--', alpha=0.5)

plt.xlim(-3.5, 3.5)
plt.ylim(-3.5, 3.5)
plt.axis('equal')
plt.title(f'Sound Source at {angle_example}° (ITD ≈ {itd_example:.2f} ms)')
plt.legend(loc='upper right')
plt.grid(True, alpha=0.3)

plt.tight_layout()
plt.show()
<Figure size 1200x400 with 2 Axes>

Auditory scene analysis

Everything so far has concerned single sounds: one tone’s loudness, one tone’s pitch, one source’s location. Real listening is almost never like that. At any moment the eardrum receives one pressure wave that is the sum of every sound source in the room. Somehow we hear a singer, a guitar, a passing car and a fan, not one blended smear.

The Canadian psychologist Albert Bregman (1936–2023) named this problem auditory scene analysis (ASA) in his 1990 book of the same title Bregman, 1990. The auditory system has to solve two grouping problems at once:

  • Simultaneous grouping decides which frequency components belong to the same sound right now. Partials that start together, share a harmonic relationship, and change together tend to fuse into one tone, which is why we hear a single trumpet note rather than a dozen separate partials.
  • Sequential grouping decides which sounds belong to the same source over time. Sounds that are close in pitch, timbre, and location, and follow each other at a moderate rate, tend to bind into a single stream.

Streaming

The classic demonstration alternates two tones in a gallop: A–B–A–silence–A–B–A. When A and B are close in frequency, you hear one bouncing melody. Pull them apart, and the sequence splits into two streams—a fast one and a slow one—and the galloping rhythm disappears, even though not a single note has moved in time. This is stream segregation, and it shows that rhythm is a property of a stream, not of the signal.

Streaming builds up over a few seconds and depends on rate. The faster the alternation and the wider the frequency separation, the more likely the split. Composers have exploited it for centuries. In compound melody—a solo Bach partita, say—a single instrument plays one line of notes that listeners hear as two interleaved voices.

The continuity illusion

If a tone is interrupted by a short silence, you hear a gap. Fill that gap with a loud burst of noise instead, and the tone seems to continue right through it: the continuity illusion. The auditory system treats the noise as a plausible masker and reconstructs what was probably there.

This is not a curiosity but the everyday case. Conversations survive passing traffic, and melodies survive coughs in the concert hall, because the auditory system fills in the parts that could have been masked. It is the same logic as the missing fundamental below. Perception is constructive, not a passive copy of the waveform.

Grouping principles and their musical consequences—Gestalt cues, polyphonic listening, and Diana Deutsch’s illusions—are picked up in harmony and melody. The engineering counterpart, getting a machine to separate sources, appears in time and rhythm and machine listening.

Demo: streaming and the continuity illusion

The first pair of clips is the galloping A–B–A sequence, first with a small frequency separation (one stream) and then with a large one (two streams). The second pair interrupts a steady tone with silence, then with noise. Listen for the tone appearing to continue through the noise.

Source
import numpy as np
from IPython.display import Audio, display

sr = 22050

def gallop(f_a, f_b, n_cycles=8, tone_ms=100):
    """A-B-A-silence gallop used in streaming experiments."""
    t = np.linspace(0, tone_ms / 1000, int(sr * tone_ms / 1000), endpoint=False)
    env = np.minimum(1, np.minimum(t, t[-1] - t) / 0.01)  # 10 ms ramps
    tone_a = np.sin(2 * np.pi * f_a * t) * env
    tone_b = np.sin(2 * np.pi * f_b * t) * env
    gap = np.zeros_like(tone_a)
    return np.tile(np.concatenate([tone_a, tone_b, tone_a, gap]), n_cycles) * 0.3

print("Small separation (400 / 480 Hz) — one galloping stream:")
display(Audio(gallop(400, 480), rate=sr))
print("Large separation (400 / 1600 Hz) — two separate streams:")
display(Audio(gallop(400, 1600), rate=sr))

# Continuity illusion: the same tone interrupted by silence, then by noise
dur, gap_ms = 2.0, 250
t = np.linspace(0, dur, int(sr * dur), endpoint=False)
tone = 0.25 * np.sin(2 * np.pi * 500 * t)
mask = np.ones_like(tone)
start = int(sr * (dur - gap_ms / 1000) / 2)
mask[start:start + int(sr * gap_ms / 1000)] = 0

rng = np.random.default_rng(0)
noise = np.zeros_like(tone)
noise[start:start + int(sr * gap_ms / 1000)] = 0.6 * rng.normal(0, 1, int(sr * gap_ms / 1000))

print("Tone interrupted by silence — you hear a gap:")
display(Audio(tone * mask, rate=sr))
print("Same gap filled with noise — the tone seems to continue through it:")
display(Audio(tone * mask + noise, rate=sr))
Small separation (400 / 480 Hz) — one galloping stream:
Loading...
Large separation (400 / 1600 Hz) — two separate streams:
Loading...
Tone interrupted by silence — you hear a gap:
Loading...
Same gap filled with noise — the tone seems to continue through it:
Loading...

Auditory illusions

The continuity illusion above is one case of a wider phenomenon. Auditory illusions show how the brain interprets sound, often departing from the raw physical signal. Studying them tells us how we organise, prioritise, and sometimes misinterpret acoustic information. That knowledge feeds into music production, audio engineering, hearing aid design, and the design of efficient audio codecs.

Hysteresis

Hysteresis is when a system’s response depends not only on its current state but also on its past states. In psychoacoustics, it turns up in loudness perception. How loud a sound seems can depend on the sounds that came before it. After a loud sound, a softer one may seem even quieter than it would in isolation. A gradual rise in volume can be perceived differently from a sudden jump, even when the final sound pressure level is the same.

Hysteresis

Image Source: Wikipedia - Hysteresis

Masking effects

Masking is when one sound (the masker) makes another sound (the maskee) harder to hear. The auditory system has limited frequency and temporal resolution, so a strong or similar sound can “cover up” a weaker or nearby one, making it less perceptible or even inaudible.

There are several types. In simultaneous masking, two sounds play at the same time and a louder one can render a softer one inaudible, even when both are within the listener’s hearing range. The effect is strongest when the two are close in frequency. In music production, for example, a loud bass drum can mask a softer bass guitar note if they occur together and share similar frequencies. The same principle drives audio compression formats such as MP3, which discard masked sounds to shrink file size without a noticeable loss in quality.

Frequency masking is strongest when masker and maskee are close in frequency, typically within the same critical band. The ear divides the spectrum into critical bands Fastl & Zwicker, 2007, and sounds in the same band are more likely to mask each other, which is why a high-pitched sound is unlikely to be masked by a low-pitched one, and vice versa. Masking is generally stronger for frequencies above the masker (upward spread) than below it, so a loud low-frequency sound masks higher frequencies more effectively than the reverse.

Temporal masking is when a loud sound masks a softer one that comes immediately before (pre-masking) or after (post-masking) it, even though the two do not overlap in time. Pre-masking can last up to about 20 ms before the masker, and post-masking up to about 100 ms after it. This reflects the temporal resolution limits of human hearing and is also exploited in perceptual audio coding.

The graph below shows how a masker tone at a given frequency and intensity raises the threshold of hearing for nearby frequencies, making them inaudible unless they rise above the masking threshold. This is a key idea in psychoacoustics and audio engineering.

Audio Masking Graph

Image Source: Wikipedia - Audio Masking

Masking turns up everywhere. A passing truck can mask a conversation until it has gone by. In music, masking can be used on purpose to blend instruments or to hide imperfections in a recording. In hearing aids, understanding it helps in designing algorithms that bring out speech while suppressing background noise.

Shepard and Risset tones

One of the best-known auditory illusions is the Shepard tone, named after Roger Shepard (1929–2022), who first described it Shepard, 1964. It is a series of tones that seem to climb or fall endlessly without ever actually getting higher or lower, the auditory equivalent of a barber pole.

A Shepard tone is built by layering several sine waves spaced an octave apart. As the sequence runs, all the sine waves move up (or down) together. The highest one fades out as it reaches the top of the range while a new, low one fades in at the bottom. The overall amplitude envelope is shaped so the listener always hears a similar blend of frequencies, with no clear start or end. This constant overlap and cross-fading make the brain hear a never-ending rise (or fall), even though the actual frequencies cycle within a fixed range. The illusion plays on the way our hearing groups harmonically related tones, which makes it hard to pin down when the scale “resets”.

Risset tones, named after the French composer and scientist Jean-Claude Risset (1938–2016), are often confused with Shepard tones because they are so similar. The difference is that Shepard tones are discrete, while Risset showed the same illusion can be made with continuous sounds, a pitch that seems to rise or fall without end. The effect comes from overlapping sine waves that fade in and out at different frequencies, creating the impression of an endless scale. Risset tones show how our perception of pitch can be steered by careful control of spectral content and amplitude envelopes. The Shepard tones app plays both versions endlessly, rising or falling, stepped or gliding, while drawing the octave components and their loudness envelope.

Both Shepard and Risset tones have been used in sound design and music. One of many examples is the video Audio Illusion: ascending melody (Advanced Shepard tones) by Smashed Transistors:

(If the video is unavailable, try the archived page.)

Demo: frequency masking

A soft tone can become inaudible when a louder sound sits close to it in frequency. Below, a quiet 1 kHz tone is played alone, then together with a band of noise centred on 1 kHz. The tone is still physically present in the second clip, but it is hard to hear. This masking is exactly what perceptual audio codecs like MP3 exploit.

Source
import numpy as np
from numpy.fft import rfft, irfft, rfftfreq
from IPython.display import Audio, display

sr = 22050
dur = 2.0
t = np.linspace(0, dur, int(sr * dur), endpoint=False)

tone = 0.05 * np.sin(2 * np.pi * 1000 * t)          # soft 1 kHz probe tone

rng = np.random.default_rng(0)                       # noise band around 1 kHz
spec = rfft(rng.standard_normal(len(t)))
freqs = rfftfreq(len(t), 1 / sr)
spec[(freqs < 700) | (freqs > 1300)] = 0
band = irfft(spec, n=len(t))
band = 0.5 * band / np.max(np.abs(band))

print("1) Soft 1 kHz tone alone:")
display(Audio(tone, rate=sr))
print("2) The same tone plus louder noise around 1 kHz (the tone is masked):")
display(Audio(tone + band, rate=sr))
print("3) The masking noise on its own:")
display(Audio(band, rate=sr))
1) Soft 1 kHz tone alone:
Loading...
2) The same tone plus louder noise around 1 kHz (the tone is masked):
Loading...
3) The masking noise on its own:
Loading...
Source
import numpy as np
import matplotlib.pyplot as plt
from IPython.display import Audio

# Generate Shepard tone - ascending pitch illusion
sr = 22050
duration = 3  # seconds
t = np.linspace(0, duration, int(sr * duration), endpoint=False)

# Shepard tone: layers of sine waves spaced one octave apart
shepard = np.zeros_like(t)
start_freq = 110  # A2
num_octaves = 4

for octave in range(num_octaves):
    freq = start_freq * (2 ** octave)
    # Frequency sweep from 110 Hz to 220 Hz (changes from octave to octave)
    freq_sweep = start_freq + (start_freq * 2) * (t / duration) + freq
    
    # Amplitude envelope: fade in and out at different octaves
    if octave == 0:
        env = np.sin(np.pi * (t / duration)) ** 2 * 0.3  # Fades in
    elif octave == num_octaves - 1:
        env = np.cos(np.pi * (t / duration)) ** 2 * 0.3  # Fades out
    else:
        env = np.ones_like(t) * 0.3
    
    shepard += env * np.sin(2 * np.pi * freq_sweep * t)

# Normalize
shepard = shepard / np.max(np.abs(shepard)) * 0.7

# Visualize the Shepard tone
plt.figure(figsize=(12, 8))

# Time-frequency representation
plt.subplot(2, 1, 1)
S = librosa.amplitude_to_db(np.abs(librosa.stft(shepard, n_fft=2048)), ref=np.max)
librosa.display.specshow(S, sr=sr, x_axis='time', y_axis='log', cmap='viridis')
plt.title('Shepard Tone Spectrogram: Ascending Pitch Illusion')
plt.ylabel('Frequency (Hz)')
plt.ylim(80, 1000)
plt.colorbar(format='%+2.0f dB')

# Waveform
plt.subplot(2, 1, 2)
plt.plot(t[:10000], shepard[:10000])  # Show first 0.5 seconds
plt.xlabel('Time (s)')
plt.ylabel('Amplitude')
plt.title('Shepard Tone Waveform')
plt.grid(True, alpha=0.3)

plt.tight_layout()
plt.show()

# Play the Shepard tone
print("Playing Shepard tone (3 seconds) - Notice how it seems to continuously rise!")
# Audio(shepard, rate=sr)  # Uncomment to listen
<Figure size 1200x800 with 3 Axes>
Playing Shepard tone (3 seconds) - Notice how it seems to continuously rise!

Missing fundamental

Another striking auditory illusion is the missing fundamental. When a complex tone lacks its fundamental frequency but keeps its harmonics, listeners still hear the pitch of the missing fundamental. Pitch perception, in other words, is based on the pattern of overtones, not just on the lowest frequency present.

Play a sound with frequencies at 200 Hz, 300 Hz, and 400 Hz (the 2nd, 3rd, and 4th harmonics of 100 Hz) but leave out the 100 Hz fundamental, and most listeners still hear the pitch as if the 100 Hz tone were there. The auditory system reads the spacing between the harmonics and infers the fundamental, even when it is physically absent.

This matters in music and audio technology. Small speakers, like those in smartphones, often cannot reproduce very low frequencies, yet listeners still perceive the intended bass notes from the higher harmonics. The effect is also used in telephony and audio compression to suggest full-range sound from a limited frequency range.

The missing fundamental shows that pitch perception is constructive. The brain detects regularities and patterns in the spectrum rather than simply responding to the frequencies that are physically present. Harmony and melody returns to the same effect under its other name, virtual pitch, and uses it to build melodies out of partials alone.

Binaural beats

When two slightly different frequencies are played separately to each ear (using headphones), the listener perceives a rhythmic beating at the frequency difference. This illusion arises from the brain’s processing of phase differences between the ears.

A closer look: binaural beats as a brain hack?

Search any streaming service for “binaural beats” and you will find playlists promising focus, deep sleep, or reduced anxiety, each tuned to a specific frequency. The illusion itself is real, as described above. The claims built on top of it deserve a closer look.

  • The claim is that the beating percept entrains brain rhythms at the beat frequency, so that a 6 Hz beat induces a relaxed “theta state” and a 40 Hz beat sharpens attention.
  • The evidence is a pile of small experiments with mixed results. A meta-analysis of twenty-two studies found a modest average benefit for anxiety, attention, and pain Garcia-Argibay et al., 2019, but the studies differed in beat frequency, duration, masking noise, and outcome measures, and small positive studies reach print more easily than null ones, a problem known as publication bias.
  • The method in most studies cannot separate the supposed entrainment from ordinary relaxation, expectation, or simply resting with a steady sound in the ears.
  • The limits are therefore sharp. Listening to binaural beats is harmless and may well be pleasant, but the specific frequency-prescription story is not established.

An audible illusion is not evidence of a neural mechanism. That distinction, between what you hear and what it proves, is the core skill this chapter has been practising.

Demo: the Shepard tone illusion

A Shepard tone seems to rise forever. It stacks octave-spaced components under a fixed bell-shaped spectral envelope. As each component glides upward and fades out at the top, a new one fades in at the bottom, so the pitch class keeps climbing without ever leaving the range.

Source
import numpy as np
from IPython.display import Audio, display

sr = 22050
dur = 8.0
t = np.linspace(0, dur, int(sr * dur), endpoint=False)
f0 = 55.0          # lowest reference frequency
n = 7              # number of octave-spaced components
period = 4.0       # seconds to rise one octave
centre, width = 3.5, 1.8

out = np.zeros_like(t)
for k in range(n):
    octpos = k + t / period                 # this component rises one octave per period
    freq = f0 * 2.0 ** octpos
    phase = 2 * np.pi * np.cumsum(freq) / sr
    amp = np.exp(-0.5 * (((octpos % n) - centre) / width) ** 2)
    out += amp * np.sin(phase)

out *= 0.15 / np.max(np.abs(out))
print("An endlessly rising Shepard tone:")
display(Audio(out, rate=sr))
An endlessly rising Shepard tone:
Loading...
Source
import numpy as np
import matplotlib.pyplot as plt

# Generate missing fundamental example
print("--- Missing Fundamental Illusion ---")
f_fundamental = 100  # Hz (will be omitted)
sr_missing = 22050
duration_missing = 2
t_missing = np.linspace(0, duration_missing, int(sr_missing * duration_missing), endpoint=False)

# Create sound with harmonics but NO fundamental
missing_fundamental = np.zeros_like(t_missing)
for harmonic in range(2, 6):  # 2nd through 5th harmonics
    freq = f_fundamental * harmonic
    amplitude = 1.0 / harmonic  # Decrease amplitude for higher harmonics
    missing_fundamental += amplitude * np.sin(2 * np.pi * freq * t_missing)

missing_fundamental = missing_fundamental / np.max(np.abs(missing_fundamental)) * 0.7

# Compare with complete sound (with fundamental)
complete_sound = np.sin(2 * np.pi * f_fundamental * t_missing) * 0.3
for harmonic in range(2, 6):
    freq = f_fundamental * harmonic
    amplitude = 1.0 / harmonic
    complete_sound += amplitude * np.sin(2 * np.pi * freq * t_missing) * 0.3

complete_sound = complete_sound / np.max(np.abs(complete_sound)) * 0.7

# Visualize both
plt.figure(figsize=(12, 8))

# Missing fundamental spectrum
plt.subplot(2, 2, 1)
S_missing = librosa.amplitude_to_db(np.abs(librosa.stft(missing_fundamental, n_fft=2048)), ref=np.max)
librosa.display.specshow(S_missing, sr=sr_missing, x_axis='time', y_axis='log', cmap='viridis')
plt.title('Missing Fundamental: Harmonics Only (2nd-5th)')
plt.ylabel('Frequency (Hz)')
plt.ylim(50, 2000)

# Spectrum zoom on fundamental region
plt.subplot(2, 2, 2)
freqs = librosa.fft_frequencies(sr=sr_missing, n_fft=2048)
S_mag = np.abs(librosa.stft(missing_fundamental, n_fft=2048))
plt.semilogy(freqs[100:400], S_mag[100:400, 0])
plt.xlabel('Frequency (Hz)')
plt.ylabel('Magnitude')
plt.title('Frequency Spectrum (Missing Fundamental)')
plt.grid(True, alpha=0.3, which='both')
plt.axvline(x=f_fundamental, color='r', linestyle='--', label=f'Missing fundamental ({f_fundamental} Hz)')
plt.xlim(50, 600)
plt.legend()

# Complete sound spectrum
plt.subplot(2, 2, 3)
S_complete = librosa.amplitude_to_db(np.abs(librosa.stft(complete_sound, n_fft=2048)), ref=np.max)
librosa.display.specshow(S_complete, sr=sr_missing, x_axis='time', y_axis='log', cmap='viridis')
plt.title('Complete Sound: Fundamental + Harmonics')
plt.ylabel('Frequency (Hz)')
plt.ylim(50, 2000)

# Spectrum comparison
plt.subplot(2, 2, 4)
S_mag_complete = np.abs(librosa.stft(complete_sound, n_fft=2048))
plt.semilogy(freqs[100:400], S_mag_complete[100:400, 0])
plt.xlabel('Frequency (Hz)')
plt.ylabel('Magnitude')
plt.title('Frequency Spectrum (Complete)')
plt.grid(True, alpha=0.3, which='both')
plt.axvline(x=f_fundamental, color='r', linestyle='--', label=f'Fundamental ({f_fundamental} Hz)')
plt.xlim(50, 600)
plt.legend()

plt.tight_layout()
plt.show()

print(f"Both sounds should be perceived as having a pitch around {f_fundamental} Hz,")
print("even though the missing fundamental version lacks the {f_fundamental} Hz component!")
--- Missing Fundamental Illusion ---
<Figure size 1200x800 with 4 Axes>
Both sounds should be perceived as having a pitch around 100 Hz,
even though the missing fundamental version lacks the {f_fundamental} Hz component!

Chapter summary

Psychoacoustics links measurable sound to perceptual qualities: loudness, pitch, timbre, and space. This chapter covered thresholds, critical bands, masking, auditory scene analysis, illusions, and spatial cues, building a bridge from physics to the subjective experience of listening. It also looked at hearing health: safe listening levels, tinnitus, and how musicians’ earplugs protect without spoiling the music. A closer look at binaural beats showed how to weigh popular claims against mixed evidence.

Questions

  1. How does the ear transduce sound from outer ear to neural firing in the cochlea, and what is tonotopic organisation?
  2. What separates a temporary from a permanent threshold shift, and why do musicians’ earplugs with flat attenuation suit musical settings better than foam plugs?
  3. How do ITD and ILD contribute to localisation, and when might each cue dominate?
  4. How do continuity and streaming phenomena illustrate that perception is not a passive copy of the waveform?
  5. What do listeners actually hear in binaural beats, what is claimed about them commercially, and how well does the evidence support those claims?
References
  1. Moore, B. (1997). An Introduction to the Psychology of Hearing: Fourth Edition. BRILL. 10.1163/9789004658820
  2. Fastl, H., & Zwicker, E. (2007). Psychoacoustics. Springer Berlin Heidelberg. 10.1007/978-3-540-68888-4
  3. Cook, P. R. (Ed.). (1999). Music, Cognition, and Computerized Sound: An Introduction to Psychoacoustics. The MIT Press. 10.7551/mitpress/4808.001.0001
  4. Fletcher, H., & Munson, W. A. (1933). Loudness, Its Definition, Measurement and Calculation. The Journal of the Acoustical Society of America, 5(2), 82–108. 10.1121/1.1915637
  5. Bregman, A. S. (1990). Auditory Scene Analysis: The Perceptual Organization of Sound. The MIT Press. 10.7551/mitpress/1486.001.0001
  6. Shepard, R. N. (1964). Circularity in Judgments of Relative Pitch. The Journal of the Acoustical Society of America, 36(12), 2346–2353. 10.1121/1.1919362
  7. Garcia-Argibay, M., Santed, M. A., & Reales, J. M. (2019). Efficacy of Binaural Auditory Beats in Cognition, Anxiety, and Pain Perception: A Meta-Analysis. Psychological Research, 83(2), 357–372. 10.1007/s00426-018-1066-8