Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

4. Electroacoustics

How sound is captured, processed, digitised, and reproduced through technology

Electroacoustics is the branch of acoustics concerned with converting sound into electrical signals and back again. It covers the study and design of devices such as microphones, loudspeakers, and amplifiers, and it underpins nearly all modern music making, recording, and reproduction.

The central idea is transduction: converting energy from one form to another. A microphone turns acoustic pressure variations in the air into varying electrical voltages. A loudspeaker does the reverse, turning electrical signals back into pressure variations. These two transducers, together with the signal chain between them, sit behind almost everything we hear through technology.

Electroacoustics is a subdiscipline of acoustics, alongside the room and instrument acoustics we met in acoustics. The physics of sound propagation is the same, but electroacoustics adds electrical and electronic phenomena that shape how sound is captured, modified, transmitted, and reproduced.

This chapter is where the four levels of description cross from the physical signal into the digital representation. A microphone and a converter are the machinery of that crossing, and every choice in them, sample rate, bit depth, codec, decides what survives it.

A brief history

The field begins in the late nineteenth century; Katz (2022) gives a short account of how these developments reshaped music. Alexander Graham Bell’s telephone (1876) showed that speech could be turned into an electrical signal and sent down a wire. Thomas Edison’s phonograph (1877) showed that sound could be recorded and played back mechanically. The carbon microphone, the vacuum tube amplifier, and the moving-coil loudspeaker, all developed in the early twentieth century, made electrical sound reproduction practical.

The condenser (capacitor) microphone arrived in the 1920s. Magnetic recording, FM radio, and high-fidelity audio in the mid-twentieth century then transformed music production and broadcasting. Today the field keeps changing through digital signal processing, miniaturised microelectronics, and wireless transmission.

Why electroacoustics matters for music

Every recorded piece of music has passed through microphones, amplifiers, and loudspeakers, and the character of those devices shapes the timbre of the recording. Live performance leans on the same equipment. In concert halls and clubs, the choice and placement of microphones and speakers decides how a performance reaches the audience. Research has its own stake, since music-psychology experiments usually present stimuli through headphones or speakers, and you cannot interpret the results without knowing how those devices behave. And electronic instruments, from synthesisers to electric guitars, all involve electroacoustic transduction somewhere in the signal chain.

Digitised audio is also the usual input to computational descriptors of pitch and harmony (see machine listening), and it connects back to the music-theoretical representations in harmony and melody.

Transduction

A transducer is any device that converts energy from one form to another. In electroacoustics, we care mainly about devices that convert between acoustic energy (pressure waves in air) and electrical energy (varying voltages or currents).

We can describe the relationship between cause and effect in an electroacoustic system as follows:

The key transducers in this chain are the microphone (acoustic → electrical) and the loudspeaker (electrical → acoustic). The intermediate stages (preamplifier, signal processing, power amplifier) operate entirely in the electrical domain.

Transduction principles

Different electroacoustic transducers exploit different physical phenomena to convert energy:

  • Electromagnetic induction: a conductor moving in a magnetic field generates an electrical current (used in dynamic microphones and moving-coil loudspeakers).
  • Electrostatic / capacitive: changes in the distance between two charged plates alter the capacitance and produce a voltage (used in condenser microphones and electrostatic loudspeakers).
  • Piezoelectric: certain crystals and ceramics generate an electrical charge when mechanically deformed (used in contact microphones and some transducers).
  • Magnetostrictive and other effects: used in some specialist transducers.

The choice of principle is a set of trade-offs between frequency response, sensitivity, dynamic range, and cost, which is why the sections below treat each transducer type separately.

Microphones

A microphone is a transducer that converts acoustic pressure variations in the surrounding air into a corresponding electrical signal. The type, quality, and placement of a microphone strongly shape the character of any recording or live sound system. Rumsey & McCormick (2021) treats microphones, loudspeakers, and the studio signal chain in far more detail than we can here.

Dynamic (moving-coil) microphones

The dynamic microphone is the most widely used type in live sound. Its operating principle is electromagnetic induction:

  1. Sound waves cause a thin diaphragm to vibrate.
  2. A voice coil attached to the diaphragm moves within the field of a permanent magnet.
  3. This movement induces an electrical current in the coil, proportional to the velocity of the diaphragm.

Advantages: robust, tolerant of high sound pressure levels, requires no external power supply, relatively inexpensive.

Disadvantages: limited high-frequency response due to the inertia of the diaphragm and coil assembly; generally lower sensitivity than condenser microphones.

Typical examples include the Shure SM58 (vocal microphone) and Shure SM57 (instrument microphone), which are industry standards in live sound.

Condenser (capacitor) microphones

The condenser microphone uses an electrostatic principle:

  1. The diaphragm forms one plate of a capacitor, with a fixed backplate as the other.
  2. The capacitor is maintained at a fixed charge (either by a battery or by phantom power supplied through the microphone cable).
  3. Sound waves cause the diaphragm to move, varying the distance between the plates and hence the capacitance.
  4. This change in capacitance at constant charge produces a varying voltage, which is the output signal.

Advantages: very flat and extended frequency response, high sensitivity, low self-noise, available in small diaphragm (pencil) and large diaphragm varieties.

Disadvantages: requires power (phantom power at 48 V, or internal battery); more fragile than dynamic microphones; more expensive.

Condenser microphones are commonly chosen for studio recording of acoustic instruments, voices, and orchestras, because their flat, extended response and low self-noise capture fine detail. Large-diaphragm condensers (e.g., Neumann U87) are a staple of professional studios.

Ribbon microphones

The ribbon microphone uses a thin corrugated metallic ribbon suspended in a magnetic field:

  1. Sound waves cause the ribbon to vibrate within the magnetic field.
  2. The movement of the ribbon induces a voltage by electromagnetic induction, similar to a dynamic microphone, but the ribbon itself is both the diaphragm and the conductor.

Advantages: natural and smooth high-frequency roll-off, excellent transient response, often described as having a warm and vintage sound character. Ribbon microphones are inherently bidirectional (figure-8 polar pattern).

Disadvantages: fragile ribbon element; lower output level, requiring a high-gain preamplifier; sensitive to wind and plosive sounds; some classic ribbon microphones should not be connected to phantom power (though modern active ribbon microphones require it).

MEMS microphones

MEMS (Micro-Electro-Mechanical Systems) microphones are semiconductor-based transducers fabricated using integrated circuit manufacturing techniques. They are used in smartphones, laptops, hearing aids, and other consumer devices where small size matters most. Their performance has improved dramatically in recent years, and they are increasingly found in professional measurement applications as well.

Microphone polar patterns

A polar pattern (or directivity pattern) describes how sensitive a microphone is to sounds arriving from different directions. It is typically represented as a plot of sensitivity versus angle, in polar coordinates.

The main polar patterns are:

  • Omnidirectional: equal sensitivity in all directions. Captures the full acoustic environment, including room reflections.
  • Cardioid: most sensitive in front, least sensitive at the rear (heart-shaped pattern). The most common pattern for live sound and studio recording. Offers some rejection of unwanted sounds from behind.
  • Supercardioid / Hypercardioid: narrower front pickup than cardioid, with a small rear lobe. More directional; useful in noisy environments.
  • Figure-8 (bidirectional): picks up equally from front and rear, with strong rejection from the sides. Used in ribbon microphones and as a component in mid-side recording techniques.

In practice, polar patterns vary with frequency. Most directional microphones are more omnidirectional at low frequencies and become more directional as frequency increases. The plots below illustrate idealised patterns.

Source
import numpy as np
import matplotlib.pyplot as plt

theta = np.linspace(0, 2 * np.pi, 360)

# Standard polar pattern formulae (values clipped at 0 for plotting)
omni        = np.ones_like(theta)
cardioid    = np.clip(0.5 + 0.5 * np.cos(theta), 0, 1)
supercard   = np.clip(0.37 + 0.63 * np.cos(theta), 0, 1)
hypercard   = np.clip(0.25 + 0.75 * np.cos(theta), 0, 1)
figure8     = np.abs(np.cos(theta))

patterns = {
    "Omnidirectional": omni,
    "Cardioid":        cardioid,
    "Supercardioid":   supercard,
    "Hypercardioid":   hypercard,
    "Figure-8":        figure8,
}

fig, axes = plt.subplots(1, 5, subplot_kw=dict(polar=True), figsize=(14, 3))
colors = ['steelblue', 'darkorange', 'seagreen', 'firebrick', 'purple']

for ax, (name, pattern), color in zip(axes, patterns.items(), colors):
    ax.plot(theta, pattern, color=color, linewidth=2)
    ax.fill(theta, pattern, alpha=0.15, color=color)
    ax.set_theta_zero_location('N')  # 0 degrees at the top (front of microphone)
    ax.set_theta_direction(-1)       # Clockwise
    ax.set_ylim(0, 1)
    ax.set_yticks([0.5, 1.0])
    ax.set_yticklabels([])
    ax.set_xticks(np.radians([0, 90, 180, 270]))
    ax.set_xticklabels(['0°', '90°', '180°', '270°'], fontsize=7)
    ax.set_title(name, fontsize=9, pad=10)

plt.suptitle('Idealised Microphone Polar Patterns', fontsize=12, y=1.02)
plt.tight_layout()
plt.show()
<Figure size 1400x300 with 5 Axes>

Microphone frequency response

The frequency response of a microphone describes how its output sensitivity varies across the audio frequency range. An ideal microphone would have a perfectly flat frequency response, meaning it captures all frequencies equally. In practice, every microphone design imposes some deviation from flat response.

  • Dynamic microphones typically show a reduced sensitivity at high frequencies due to the inertia of the coil–diaphragm assembly, and may have a moderate presence peak (a boost around 3–8 kHz) to enhance clarity of voice or instruments.
  • Condenser microphones generally provide a flatter, more extended high-frequency response. Large-diaphragm condensers often have a subtle presence peak around 8–12 kHz.
  • Ribbon microphones exhibit a gentle high-frequency roll-off, which contributes to their characteristically warm sound.

At low frequencies, proximity effect is a common phenomenon in directional (cardioid, figure-8) microphones. The bass response increases when the sound source is very close to the microphone (within 30–60 cm). This effect is exploited by vocalists and broadcasters to achieve a richer, fuller sound, but can be a problem when recording at close range if unwanted bass boost is not desired.

Source
import numpy as np
import matplotlib.pyplot as plt

freq = np.logspace(np.log10(20), np.log10(20000), 1000)

def smooth_step(x, x0, width):
    """Logistic-style smooth transition centred on x0."""
    return 1 / (1 + np.exp(-(x - x0) / width))

# Simulate idealised frequency responses (in dB, relative to 1 kHz)

# Dynamic: slight presence peak ~5 kHz, high-frequency roll-off above 12 kHz
dynamic = (
    2.5 * np.exp(-((np.log10(freq) - np.log10(5000)) ** 2) / (2 * 0.15 ** 2))
    - 6 * smooth_step(np.log10(freq), np.log10(12000), 0.15)
)

# Condenser: very flat, small presence peak ~10 kHz, slight HF extension
condenser = (
    1.5 * np.exp(-((np.log10(freq) - np.log10(10000)) ** 2) / (2 * 0.2 ** 2))
    - 2 * smooth_step(np.log10(freq), np.log10(18000), 0.1)
)

# Ribbon: warm roll-off starting around 8 kHz
ribbon = (
    -0.5 * np.log10(freq / 1000)
    - 4 * smooth_step(np.log10(freq), np.log10(8000), 0.2)
)

# Low-frequency roll-off common to most mics (high-pass due to housing)
lf_rolloff = -10 * smooth_step(-np.log10(freq), -np.log10(80), 0.15)

dynamic   += lf_rolloff
condenser += lf_rolloff
ribbon    += lf_rolloff

fig, ax = plt.subplots(figsize=(10, 4))
ax.semilogx(freq, dynamic,   label='Dynamic (moving-coil)', color='steelblue',  linewidth=2)
ax.semilogx(freq, condenser, label='Condenser (capacitor)',  color='darkorange', linewidth=2)
ax.semilogx(freq, ribbon,    label='Ribbon',                 color='seagreen',   linewidth=2)
ax.axhline(0, color='grey', linestyle='--', linewidth=0.8, label='0 dB reference')
ax.set_xlim(20, 20000)
ax.set_ylim(-15, 8)
ax.set_xlabel('Frequency (Hz)', fontsize=11)
ax.set_ylabel('Relative level (dB)', fontsize=11)
ax.set_title('Idealised Frequency Responses of Microphone Types', fontsize=12)
ax.set_xticks([50, 100, 200, 500, 1000, 2000, 5000, 10000, 20000])
ax.set_xticklabels(['50', '100', '200', '500', '1k', '2k', '5k', '10k', '20k'])
ax.grid(True, which='both', alpha=0.3)
ax.legend(fontsize=10)
plt.tight_layout()
plt.show()
<Figure size 1000x400 with 1 Axes>

Microphone specifications

Manufacturers publish a specification sheet for every microphone. For most musical purposes four numbers do the work: how sensitive it is, how flat its frequency response is, how quiet it is when nothing is happening (self-noise), and how loud a source it can take before distorting (maximum SPL). The gap between those last two is the microphone’s usable dynamic range.

Loudspeakers

A loudspeaker (often simply called a speaker) converts electrical signals into acoustic pressure variations. It is the final electroacoustic transducer in most audio reproduction systems.

Moving-coil (dynamic) loudspeakers

The moving-coil loudspeaker is by far the most common loudspeaker type. Its operation is the electromagnetic inverse of the dynamic microphone:

  1. An alternating electrical current flows through a voice coil suspended in a permanent magnetic field.
  2. The interaction of the current and the magnetic field produces a force on the coil (Lorentz force), causing it to move back and forth.
  3. The coil is attached to a cone (diaphragm), whose movement creates pressure waves in the surrounding air.

Most practical loudspeaker systems use several drivers, each optimised for a different frequency range:

  • Woofer: handles low frequencies (bass), typically 20–500 Hz. Large diameter cone (15–40 cm) for efficient low-frequency radiation.
  • Midrange driver: covers middle frequencies (500 Hz – 5 kHz), where much of the energy of speech and music is concentrated.
  • Tweeter: handles high frequencies (5–20 kHz). Small diameter diaphragm (2–4 cm) for extended high-frequency response.
  • Subwoofer: dedicated low-frequency driver (20–200 Hz) for cinema, music production, and concert sound reinforcement.

A crossover network routes the appropriate frequencies to each driver, ensuring that each operates within its optimal range.

Electrostatic loudspeakers

Electrostatic loudspeakers use the same capacitive principle as condenser microphones, in reverse. A thin conductive diaphragm is stretched between two perforated stators. A high electrical field is maintained between the stators and the diaphragm; applying an audio signal to the stators causes the diaphragm to vibrate. Electrostatic loudspeakers are valued for their extremely low distortion and transient accuracy, but are expensive, large, and require high-voltage power supplies.

Headphones and earphones

Headphones are miniature loudspeakers worn directly on or in the ears. They are available in several configurations:

  • Over-ear (circumaural): pads surround the ears; typically offer good passive isolation and a more natural sound stage.
  • On-ear (supra-aural): pads rest on the ears; more compact but often less isolating.
  • In-ear monitors (IEMs): fit inside the ear canal; used extensively in professional monitoring on stage, as well as in consumer earbuds.
  • Bone-conduction headphones: transmit sound through vibrations of the skull bones rather than through the ear canal. Small transducers placed on the cheekbones (just in front of the ears) convert the electrical signal into mechanical vibration, which reaches the inner ear directly.

Most headphones use dynamic (moving-coil) drivers; high-end models may use planar magnetic or electrostatic elements. The acoustics of headphones differ fundamentally from loudspeakers in a room, since headphones bypass the effects of room acoustics and the pinnae, which can affect perceived spatial image. This makes them attractive for controlled listening experiments in music psychology research, but means that recordings mixed on headphones may translate differently to loudspeakers.

Loudspeaker frequency response and sensitivity

Like microphones, loudspeakers are characterised by their frequency response. Unlike microphones, achieving a flat frequency response across the full audio range from a single transducer is very difficult, which is why multi-driver systems with crossovers are standard.

The sensitivity of a loudspeaker is typically specified as the sound pressure level (in dB SPL) produced at a distance of 1 metre when driven with 1 watt (or 2.83 V into an 8 Ω load). A loudspeaker with higher sensitivity requires less amplifier power to achieve a given loudness.

The plot below shows idealised frequency responses for different loudspeaker configurations, illustrating how each driver covers a different portion of the audio spectrum.

Source
import numpy as np
import matplotlib.pyplot as plt

freq = np.logspace(np.log10(20), np.log10(20000), 1000)

def bandpass_response(f, f_low, f_high, slope_low=24, slope_high=24):
    """Simulate a bandpass driver response with gentle Butterworth-like slopes."""
    low_shelf  = 1 / np.sqrt(1 + (f_low / f) ** slope_low)
    high_shelf = 1 / np.sqrt(1 + (f / f_high) ** slope_high)
    return 20 * np.log10(low_shelf * high_shelf + 1e-10)

subwoofer  = bandpass_response(freq, 25,   200,  slope_low=24, slope_high=12)
woofer     = bandpass_response(freq, 50,   1000, slope_low=12, slope_high=12)
midrange   = bandpass_response(freq, 300,  5000, slope_low=12, slope_high=12)
tweeter    = bandpass_response(freq, 3000, 22000, slope_low=12, slope_high=24)

fig, ax = plt.subplots(figsize=(10, 4))
ax.semilogx(freq, subwoofer, label='Subwoofer (20–200 Hz)',     color='navy',       linewidth=2)
ax.semilogx(freq, woofer,    label='Woofer (50 Hz – 1 kHz)',    color='steelblue',  linewidth=2)
ax.semilogx(freq, midrange,  label='Midrange (300 Hz – 5 kHz)', color='darkorange', linewidth=2)
ax.semilogx(freq, tweeter,   label='Tweeter (3–20 kHz)',        color='firebrick',  linewidth=2)
ax.axhline(-3, color='grey', linestyle='--', linewidth=0.8, label='−3 dB point')
ax.set_xlim(20, 20000)
ax.set_ylim(-40, 5)
ax.set_xlabel('Frequency (Hz)', fontsize=11)
ax.set_ylabel('Relative level (dB)', fontsize=11)
ax.set_title('Idealised Frequency Responses of Loudspeaker Drivers', fontsize=12)
ax.set_xticks([50, 100, 200, 500, 1000, 2000, 5000, 10000, 20000])
ax.set_xticklabels(['50', '100', '200', '500', '1k', '2k', '5k', '10k', '20k'])
ax.grid(True, which='both', alpha=0.3)
ax.legend(fontsize=9)
plt.tight_layout()
plt.show()
<Figure size 1000x400 with 1 Axes>

Loudspeaker placement and room interaction

A loudspeaker does not operate in isolation, since it interacts strongly with the room in which it is placed. Reflections from walls, floor, and ceiling combine with the direct sound at the listener’s ears, modifying the perceived frequency response and spatial character.

  • Bass is position-dependent. Room modes (see acoustics) make low frequencies loud in some spots and almost absent in others. Moving a subwoofer, or your chair, can change the bass more than any equaliser will.
  • Walls add bass. A speaker close to a wall radiates more low end, because the reflection reinforces the direct sound; a corner adds more still. This is boundary loading, and it is why the same speaker sounds different once you move it.
  • Stereo wants symmetry. The usual starting point is an equilateral triangle: you at one corner, the two speakers at the others, tweeters at ear height and away from walls.

These effects mean that the same loudspeaker can sound quite different in different rooms, and that acoustic treatment of the room and careful speaker placement are as important as the choice of loudspeaker itself.

Amplifiers

An amplifier is a device that increases the power, voltage, or current of an electrical signal. In an audio signal chain, amplification is necessary because microphones produce very small voltages (millivolts) while loudspeakers require much larger voltages and currents to produce audible sound.

Signal levels

Audio signals travel at very different strengths at different points in the chain, and equipment expects a particular one at each input. From weakest to strongest: microphone level (needs a preamplifier before anything else can use it), instrument level (electric guitars and basses, usually via a DI box), line level (the standard for mixers, interfaces, and effects), and speaker level (the amplified signal driving a passive loudspeaker). Plugging a signal into an input expecting a different level is the single most common cause of a recording that is either buried in noise or badly distorted.

Preamplifiers

A preamplifier (preamp) amplifies a microphone-level signal to line level. A high-quality microphone preamplifier is essential for maintaining a low noise floor. If the preamp introduces noise, subsequent amplification will amplify that noise along with the signal. The key specification of a microphone preamplifier is its equivalent input noise (EIN), typically measured in dBu.

Modern audio interfaces for recording contain built-in microphone preamplifiers. Professional recording studios may use high-quality external preamps (e.g., API 312, Neve 1073) that contribute to the characteristic sound of recordings made in those facilities.

Power amplifiers

A power amplifier takes a line-level signal and amplifies it to the voltages and currents needed to drive a loudspeaker. Power amplifiers are characterised by their output power (in watts) and their efficiency. Modern Class D amplifiers (switching amplifiers) are efficient enough (>80%) to run on batteries without much heat, which is why portable speakers exist at all.

Phantom power

Most condenser microphones require an external power source to maintain the charge on the capacitor element. Phantom power (standardised at 48 V DC, denoted +48V) is supplied through the microphone cable by the preamplifier or audio interface. It is called ‘phantom’ because it does not interfere with the audio signal, which travels as a differential (balanced) pair on the same conductors. Dynamic and ribbon microphones generally do not require phantom power (and some older ribbon microphones can be damaged by it).

The signal chain

The signal chain is the complete path an audio signal follows from its acoustic source to the listener’s ears. Knowing the chain helps you troubleshoot problems and make sensible decisions about equipment and configuration.

A typical recording signal chain involves the following stages, here split into the capture half (from sound to computer) and the playback half (from computer to listener):

In live sound, the chain is similar but may include a digital mixing console and distribution to multiple loudspeakers covering different zones of the venue.

Balanced and unbalanced connections

Professional audio equipment uses balanced connections (XLR or TRS connectors) to carry signals over long cable runs without picking up interference. A balanced connection carries the audio signal twice: once as a normal polarity signal and once with the polarity inverted. Any electromagnetic interference picked up along the cable affects both conductors equally; at the receiving end, the differential input of the preamplifier or interface cancels the common-mode noise, leaving only the audio signal. This is known as common-mode rejection.

Consumer and instrument-level connections are often unbalanced (TS connectors, RCA/phono connectors), which are more susceptible to interference over long cable runs.

Gain staging

Gain staging is the practice of setting appropriate signal levels at each stage of the signal chain to maximise signal-to-noise ratio while avoiding clipping (distortion from exceeding the maximum signal level). Each stage should receive a signal that is neither so weak that noise becomes significant, nor so strong that it clips. In digital audio systems, clipping above 0 dBFS (decibels relative to full scale) causes severe digital distortion.

Signal processing

Signal processing is the field concerned with analysing, modifying, and synthesising signals such as sound, images, and scientific measurements. In electroacoustics and audio engineering, it is how we improve sound quality, pull out information, and adapt audio for specific uses.

Signal processing can happen in both the analogue and the digital domain. Its main uses in audio include:

  • Amplification: increasing the strength of electrical signals so they can drive loudspeakers or be recorded at usable levels.
  • Filtering: removing unwanted frequencies (such as noise or hum) or enhancing desired frequency ranges. Filters can be low-pass, high-pass, band-pass, or notch filters.
  • Equalisation (EQ): adjusting the balance between frequency components to shape the tonal quality of audio signals.
  • Dynamic range compression: reducing the difference between the loudest and quietest parts of a signal to make audio more consistent and prevent distortion.
  • Noise reduction: techniques such as gating, spectral subtraction, or adaptive filtering to minimise background noise.
  • Effects processing: adding reverberation, delay, chorus, distortion, or other effects to enhance or creatively alter the sound.
  • Modulation: changing aspects of the signal such as amplitude, frequency, or phase for transmission or synthesis.

You will find signal processing in microphones, mixing consoles, audio interfaces, hearing aids, mobile devices, loudspeaker management systems, and music production software.

Synthesising sound

So far the signal chain has started from an acoustic source, a voice or an instrument that a microphone captures. A synthesiser turns this around and creates the electrical signal from scratch, with no acoustic source at all Russ, 2009. Nothing vibrates until the loudspeaker at the end of the chain, which is where the sound first becomes acoustic. Synthesisers are everywhere in popular music, from the basslines of 1980s synth-pop to the pads and leads of current chart productions.

Oscillators

The classic recipe is subtractive synthesis, which starts with an oscillator: a circuit or algorithm that produces a repeating waveform at a chosen frequency. Four waveforms dominate, and each has its own spectrum and therefore its own timbre. The sine wave contains only its fundamental and sounds pure, like the deep sub-bass under much hip-hop. The triangle wave adds weak odd harmonics and sounds soft and flute-like. The square wave has strong odd harmonics and a hollow, reedy character familiar from chiptune game music. The sawtooth wave contains every harmonic and sounds bright and buzzy, which makes it the usual starting point for rich lead and string sounds.

Source
import numpy as np
import matplotlib.pyplot as plt
from scipy import signal

f0 = 220            # fundamental frequency (Hz)
sr = 44100
t = np.linspace(0, 3 / f0, int(sr * 3 / f0), endpoint=False)  # three cycles

waves = {
    'Sine':     np.sin(2 * np.pi * f0 * t),
    'Triangle': signal.sawtooth(2 * np.pi * f0 * t, width=0.5),
    'Square':   signal.square(2 * np.pi * f0 * t),
    'Sawtooth': signal.sawtooth(2 * np.pi * f0 * t),
}

# Relative harmonic amplitudes (idealised Fourier series, in dB re fundamental)
n = np.arange(1, 16)
spectra = {
    'Sine':     np.where(n == 1, 1.0, 0.0),
    'Triangle': np.where(n % 2 == 1, 1.0 / n**2, 0.0),
    'Square':   np.where(n % 2 == 1, 1.0 / n, 0.0),
    'Sawtooth': 1.0 / n,
}

fig, axes = plt.subplots(2, 4, figsize=(12, 4.5))
colors = ['steelblue', 'seagreen', 'firebrick', 'darkorange']

for col, (name, color) in enumerate(zip(waves, colors)):
    ax = axes[0, col]
    ax.plot(t * 1000, waves[name], color=color, linewidth=1.5)
    ax.set_title(name, fontsize=10)
    ax.set_ylim(-1.2, 1.2)
    ax.set_yticks([])
    ax.set_xlabel('Time (ms)', fontsize=8)
    ax.tick_params(labelsize=7)

    ax = axes[1, col]
    amps = spectra[name]
    present = amps > 0
    ax.vlines(n[present] * f0, -40, 20 * np.log10(amps[present]),
              color=color, linewidth=2)
    ax.set_ylim(-40, 3)
    ax.set_xlim(0, 15.5 * f0)
    ax.set_xlabel('Frequency (Hz)', fontsize=8)
    if col == 0:
        ax.set_ylabel('Level (dB)', fontsize=8)
    ax.tick_params(labelsize=7)

axes[0, 0].set_ylabel('Amplitude', fontsize=8)
plt.suptitle('The four classic oscillator waveforms and their harmonic spectra', fontsize=12)
plt.tight_layout()
plt.show()
<Figure size 1200x450 with 8 Axes>
Source
import numpy as np
from scipy import signal
from IPython.display import Audio, display

sr = 44100
f0 = 220
dur = 1.2
t = np.linspace(0, dur, int(sr * dur), endpoint=False)

waves = {
    'Sine':     np.sin(2 * np.pi * f0 * t),
    'Triangle': signal.sawtooth(2 * np.pi * f0 * t, width=0.5),
    'Square':   signal.square(2 * np.pi * f0 * t),
    'Sawtooth': signal.sawtooth(2 * np.pi * f0 * t),
}

fade = np.minimum(1, np.minimum(t / 0.02, (dur - t) / 0.02))  # avoid clicks

for name, x in waves.items():
    print(f"{name} wave at {f0} Hz:")
    display(Audio(0.3 * x * fade, rate=sr))
Sine wave at 220 Hz:
Loading...
Triangle wave at 220 Hz:
Loading...
Square wave at 220 Hz:
Loading...
Sawtooth wave at 220 Hz:
Loading...

Filters and envelopes

The oscillator’s raw output is then shaped by a filter, a circuit or algorithm that weakens some parts of the spectrum, which is why the method is called subtractive. The most common type is the low-pass filter, which lets frequencies below a chosen cutoff frequency pass and dampens everything above it. Closing the cutoff makes a buzzy sawtooth progressively darker and rounder, and sweeping it while a note sounds produces the classic filter sweep heard across dance music. Many filters also offer resonance, a boost of the frequencies just around the cutoff, which at high settings gives the squelchy character of acid house basslines.

An envelope then shapes how the loudness of each note unfolds over time. The standard version is the ADSR envelope, with four stages. Attack sets how quickly the sound reaches full level, decay how quickly it falls back, sustain the level held while the key stays down, and release how long the sound lingers after the key is lifted. A fast attack and short decay give a plucked, percussive note, while a slow attack and long release give a soft pad that swells and lingers. Envelopes can also steer the filter cutoff, so that each note opens brightly and then mellows.

Additive and FM synthesis

Additive synthesis works in the opposite direction: instead of carving away at a rich waveform, it builds timbre by adding sine-wave partials one by one. This is exactly the idea behind the harmonics demo in the acoustics chapter, where sine waves summed into a complex tone. The drawbars of a Hammond organ are an early mechanical version, with each drawbar adding one harmonic to the mix.

Frequency modulation (FM) synthesis instead uses one oscillator to wobble the frequency of another at audio rate, which spreads energy into dense clusters of partials at very little computational cost. The method powered the Yamaha DX7, whose glassy electric pianos and bell-like tones define countless 1980s pop ballads. Bells and metallic sounds suit FM particularly well because their partials do not follow the neat harmonic series of strings and pipes.

Sampling and granular synthesis

A sampler does not generate waveforms at all but plays back stored recordings, pitched up and down across the keyboard. Early hip-hop was built on sampled drum breaks looped from funk records, and modern film scores lean on sampled orchestras with thousands of recorded notes per instrument. Sampling blurs the border between recording and instrument: any captured sound can become playable material.

Granular synthesis pushes that idea further by chopping a recording into tiny grains, each lasting only tens of milliseconds, and re-scattering them in new patterns. Because the grains can be repeated and overlapped freely, a sound can be stretched in time without changing its pitch, or frozen into a shimmering, sustained cloud. Ambient and electronic producers use granular tools to turn a single vocal syllable into an evolving texture.

You can try all of these building blocks by ear in the book’s own Mini synthesiser app, which chains an oscillator, a lowpass filter, and an ADSR envelope under a one-octave keyboard, with live waveform and spectrum displays. Ableton Learning Synths offers a more guided browser-based tour of the same territory. The same blocks exist in physical form in modular synthesisers, and in the beginner-friendly littleBits kits used in this course, where oscillators, filters, and envelopes snap together with magnets so that every connection can be heard.

A closer look: does a better instrument need a simpler mapping?

On acoustic instruments the link between action and sound comes bundled with the physics. On electronic instruments it is a design decision called the mapping: which control changes which sound parameter. An obvious engineering instinct says one control per parameter, like the labelled knobs on a mixing desk.

  • The claim tested by Hunt et al. (2003) was that such one-to-one mappings are not necessarily best for musical instruments.
  • The evidence came from an experiment where participants performed simple musical tasks on interfaces controlling pitch, loudness, and timbre. One version mapped each control to one parameter; another cross-coupled them, so that a single movement affected several parameters at once, more like a violin bow.
  • The method combined measured task performance across several practice sessions with the participants’ own reports of engagement.
  • The limits are those of a small laboratory study: around a dozen participants, short sessions, simple sounds, and engagement judged partly from self-report.

The result was that the harder, cross-coupled mappings were more engaging and, given practice, supported better performance on the more demanding tasks. The finding became a touchstone in the New Interfaces for Musical Expression community: an instrument that is trivially easy to control can also be trivially boring, and a degree of resistance invites exploration and skill.

Digital audio

Digital audio extends the electroacoustic signal chain into storage, editing, transmission, and reproduction on computers and digital devices; Müller (2021) develops the signal-processing side with worked Python examples. The basic principles are simple, but they have a strong effect on fidelity, workflow, and file size.

Sound vs audio

It helps to keep two words apart: sound and audio. Sound is vibration travelling through a material medium; audio is the technological representation we use to capture, store, process, and reproduce those vibrations. Audio can exist in both analogue and digital forms, and it sits between transducers such as microphones and loudspeakers.

Digitisation

Digitisation converts physical sound into a digital audio signal using an analogue-to-digital converter (ADC). It happens in two steps: sampling and quantisation.

First, the continuous sound wave is measured at regular intervals, producing a stream of samples. The number of measurements per second is the sampling rate. Next, each sampled value is rounded to the nearest value that can be represented by a fixed number of bits, the bit depth. The result is a stream of numbers that digital systems can store, process, and transmit.

The plots below illustrate these two ideas: one shows how different sampling rates capture a continuous waveform, and the other shows how different bit depths quantise sampled values.

Sampling rate

The sampling rate is the number of samples per second taken from a continuous signal to create a discrete signal. According to the Nyquist–Shannon sampling theorem Shannon, 1949, the sampling rate must be at least twice the highest frequency present in the signal to reconstruct it accurately. This minimum limit is known as the Nyquist frequency.

The CD standard uses 44,100 Hz, which can represent frequencies up to 22,050 Hz. Professional recording often uses 48 kHz, 96 kHz, or higher. Higher sampling rates can reduce aliasing, where frequencies too high to be represented reappear as false lower tones (demonstrated below), and leave more headroom for some processing tasks, but they also increase file size and CPU load.

Bit depth

The audio bit depth is the amount of data used to represent each individual sample in a digital audio signal. Bit depth sets the precision of the stored amplitude values. Higher bit depths represent the original sound more accurately, giving lower quantisation noise and a greater dynamic range.

  • Low bit depth (for example 8-bit): few available amplitude levels, which can introduce audible distortion and noise.
  • 16-bit: the CD standard, allowing 65,536 possible amplitude values.
  • 24-bit: common in professional recording, offering over 16 million levels and much greater practical headroom.
  • 32-bit float: used in some modern recorders and production workflows, providing extremely large headroom and flexibility during recording and post-production.

Digital audio interfaces

A digital audio interface (often simply called an audio interface) connects microphones, instruments, and other audio sources to a computer for recording, and routes the computer’s audio output to loudspeakers or headphones for playback.

The audio interface contains:

  • Microphone preamplifiers: one or more high-quality preamps to bring microphone-level signals up to line level.
  • Analogue-to-digital converters (ADCs): convert the analogue input signal to a digital bitstream at the chosen sample rate and bit depth.
  • Digital-to-analogue converters (DACs): convert the digital output from the computer back to an analogue signal for monitoring.
  • Headphone amplifier: a low-impedance amplifier to drive headphones from the monitor output.
  • Digital audio connectivity: USB, Thunderbolt, FireWire, or Dante connections to the computer.

The converters in an audio interface apply the principles of sampling and quantisation described above. The quality of the ADCs and DACs, together with the chosen sample rate and bit depth, affects the transparency, noise floor, and working headroom of a recording system.

Demo: sampling and aliasing

When a signal contains frequencies above the Nyquist limit (half the sampling rate), they are aliased, folded down to lower frequencies that were never there. The demo uses a deliberately low sample rate so the effect is audible: a 7000 Hz tone sampled at 8000 Hz is heard as a 1000 Hz tone. The Sampling and quantisation app goes further: it degrades a looped riff live, from 48 kHz down to 2 kHz and from 16 bits down to 2, while a zoomed waveform shows the staircase forming.

Source
import numpy as np
import matplotlib.pyplot as plt
from IPython.display import Audio, display

sr = 8000          # low sample rate, so aliasing is easy to hear
dur = 1.5
t = np.linspace(0, dur, int(sr * dur), endpoint=False)
nyquist = sr / 2

for f in [1000, 3000, 7000]:
    x = 0.3 * np.sin(2 * np.pi * f * t)
    alias = abs(((f + nyquist) % sr) - nyquist)
    tag = "  <-- ALIAS" if f > nyquist else ""
    print(f"{f:>5} Hz at sr={sr} Hz (Nyquist {nyquist:.0f} Hz) -> heard as {alias:.0f} Hz{tag}")
    display(Audio(x, rate=sr))

tt = np.linspace(0, 0.005, 2000)
ts = np.arange(0, 0.005, 1 / sr)
fig, ax = plt.subplots(figsize=(10, 2.6))
ax.plot(tt * 1000, np.sin(2 * np.pi * 7000 * tt), color="0.7", lw=0.8, label="7000 Hz (true wave)")
ax.plot(ts * 1000, np.sin(2 * np.pi * 7000 * ts), "o-", ms=4, label=f"sampled at {sr} Hz (1000 Hz alias)")
ax.set_xlabel("Time (ms)")
ax.set_ylabel("Amplitude")
ax.set_title("Undersampling a 7000 Hz tone produces a 1000 Hz alias")
ax.legend(fontsize=8)
plt.tight_layout()
plt.show()
 1000 Hz at sr=8000 Hz (Nyquist 4000 Hz) -> heard as 1000 Hz
Loading...
 3000 Hz at sr=8000 Hz (Nyquist 4000 Hz) -> heard as 3000 Hz
Loading...
 7000 Hz at sr=8000 Hz (Nyquist 4000 Hz) -> heard as 1000 Hz  <-- ALIAS
Loading...
<Figure size 1000x260 with 1 Axes>

Demo: bit depth and quantisation noise

Sampling sets the time resolution; bit depth sets the amplitude resolution. Rounding each sample to a small number of levels adds audible quantisation noise. Listen as the same tone is reduced from 16 bits to 2 bits.

Source
import numpy as np
from IPython.display import Audio, display

sr = 22050
dur = 3.0
t = np.linspace(0, dur, int(sr * dur), endpoint=False)
x = 0.5 * np.sin(2 * np.pi * 220 * t)

def quantise(sig, bits):
    levels = 2 ** bits
    return np.round((sig * 0.5 + 0.5) * (levels - 1)) / (levels - 1) * 2 - 1

for bits in [16, 8, 4, 2]:
    print(f"{bits}-bit quantisation:")
    display(Audio(quantise(x, bits), rate=sr))
16-bit quantisation:
Loading...
8-bit quantisation:
Loading...
4-bit quantisation:
Loading...
2-bit quantisation:
Loading...

Real-time audio, buffers, and latency

In a digital audio workstation (DAW), a live performance setup, or a real-time synthesis patch, audio is processed in blocks (buffers) of samples rather than as one infinite stream. The buffer size, together with the sampling rate, sets how often the computer must finish a round of processing. A smaller buffer means lower latency, the delay between input and output, but it leaves less time per block and raises CPU load. A larger buffer is more forgiving with heavy plug-ins, but it increases latency and can make timing feel sluggish when you monitor yourself or play virtual instruments.

Round-trip latency comes from several places: analogue-to-digital and digital-to-analogue conversion, buffering in the driver and the application, and the digital signal processing in the chain. Interfaces and drivers often report separate input and output figures, but what matters to a performer is the total monitoring delay. To avoid hearing themselves late, many systems offer direct monitoring, an analogue path from input to headphones that bypasses the computer.

Source
import numpy as np
import matplotlib.pyplot as plt

frequency = 1
amplitude = 1
sampling_rates = [5, 10, 20, 50]

x = np.linspace(0, 2 * np.pi, 1000)
continuous_wave = amplitude * np.sin(2 * np.pi * frequency * x)

plt.figure(figsize=(12, 2))
plt.plot(x, continuous_wave, label='Continuous sine wave', color='blue', alpha=0.7)

for rate in sampling_rates:
    sampled_x = np.linspace(0, 2 * np.pi, rate, endpoint=False)
    sampled_wave = amplitude * np.sin(2 * np.pi * frequency * sampled_x)
    plt.scatter(sampled_x, sampled_wave, label=f'Sampled points ({rate} Hz)', zorder=5)

plt.title('Effect of Different Sampling Rates on a Sine Tone')
plt.xlabel('x')
plt.ylabel('Amplitude')
plt.axhline(0, color='black', linewidth=0.5, linestyle='--')
plt.legend()
plt.grid()
plt.show()
<Figure size 1200x200 with 1 Axes>
Source
import numpy as np
import matplotlib.pyplot as plt

# Illustrate quantisation at different bit depths
t = np.linspace(0, 2 * np.pi, 500)
signal = np.sin(t)

bit_depths = [2, 4, 8, 16]
fig, axes = plt.subplots(1, 4, figsize=(14, 3), sharey=True)

for ax, bits in zip(axes, bit_depths):
    levels = 2 ** bits
    quantised = np.round(signal * (levels / 2)) / (levels / 2)
    ax.plot(t, signal,    color='grey',       linewidth=1,   alpha=0.5, label='Original')
    ax.step(t, quantised, color='steelblue',  linewidth=1.2, where='mid', label=f'{bits}-bit')
    ax.set_title(f'{bits}-bit ({levels} levels)', fontsize=9)
    ax.set_xlim(0, 2 * np.pi)
    ax.set_ylim(-1.3, 1.3)
    ax.set_xticks([])
    ax.legend(fontsize=7, loc='upper right')
    ax.grid(alpha=0.3)

axes[0].set_ylabel('Amplitude', fontsize=10)
fig.suptitle('Quantisation at Different Bit Depths', fontsize=12)
plt.tight_layout()
plt.show()
<Figure size 1400x300 with 4 Axes>

Audio compression and file formats

Digital audio also has to be stored and transmitted efficiently. Audio containers are file formats that hold digital audio data together with metadata such as track information, album art, and artist details. Common containers include WAV and AIFF, both of which typically store uncompressed, high-quality audio.

Some containers, such as MKV, can hold uncompressed or compressed audio alongside video and metadata. Others, such as MP4, often store compressed audio, frequently together with video.

The container and the compression method are two separate things. Audio file compression reduces the file size of digital audio to make storage and transmission more efficient, and it is not the same as the dynamic range compression used in mixing and mastering. There are two main types:

  • Lossless compression: preserves all original audio data, allowing perfect reconstruction, as in FLAC and ALAC.
  • Lossy compression: removes some audio data, usually components that are less perceptible to human hearing, to achieve smaller file sizes, as in MP3 and AAC.
FormatCompression TypeTypical UseQuality
WAV, AIFFNoneRecording, editingExcellent
FLACLosslessArchiving, hi-fiExcellent
AACLossyStreamingGood

AAC generally gives better quality than MP3 at a similar file size, which is why it is commonly used for streaming. Open ecosystems also offer alternatives such as the Ogg container and related codecs.

Spatial audio

The ears locate a sound using three cues: the difference in arrival time between the ears, the difference in level, and the filtering the head and outer ear impose. How these cues work in perception is explained in the next chapter, psychoacoustics. Spatial audio is the engineering side of the same problem. Given loudspeakers or headphones, how do you deliver those cues convincingly? Three families of answer are in common use, and they differ in what the recording actually stores.

Channel-based audio

The oldest and still the most common approach stores one signal per loudspeaker. Stereo is the familiar case. A source panned between two speakers is not really located anywhere: sending the same signal to both at equal level produces a phantom image, which the listener hears as coming from a point between the speakers even though nothing is radiating from there. Adjusting the balance slides the image left or right. Surround formats such as 5.1 extend the idea with more speakers, adding rear channels for ambience and a dedicated low-frequency channel.

Channel-based audio is simple and robust, and its weakness follows directly from its definition. The mix is made for one specific loudspeaker layout, and the illusion is only reliable in the sweet spot, the small region roughly equidistant from the speakers. Move off-centre and the phantom images collapse towards whichever speaker you are nearest.

Binaural audio

If the spatial cues are things that happen at the two eardrums, then delivering them over headphones is a matter of reproducing the right signal at each ear. That is what binaural audio does, either by recording with microphones in the ears of a dummy head, or by filtering a mono source with a head-related transfer function (HRTF) for the desired direction, the measured filtering that the head and outer ears apply to sound arriving from that direction.

Binaural audio can be strikingly convincing, and it has two well-known problems. HRTFs are individual, since they depend on the shape of your head, torso, and outer ears, so a generic HRTF localises less well for most listeners, and front-back confusions are common. And a static binaural rendering turns with your head, which the real world does not do, so systems that track head orientation and rotate the scene to compensate sound considerably more stable.

Scene-based audio and ambisonics

The third approach stores neither speakers nor ears, but a description of the sound field itself, which is then decoded to whatever the playback system happens to be. Ambisonics is the main example: a full-sphere surround technique whose first-order recordings carry four channels, together known as B-format, that capture both the overall sound pressure at a point and which direction the sound is coming from. Higher orders add further channels, sharpening the spatial resolution at the cost of more of them.

The practical appeal is that recording and reproduction are decoupled. One ambisonic recording can be decoded to a pair of headphones, to a stereo pair, or to a dome of thirty speakers, and it can be rotated before decoding, which is what makes it the usual carrier format for virtual reality and 360-degree video. Because such a recording stores direction as well as level, it can also be analysed to reveal where sound comes from and how diffuse it is; machine listening puts this to use for describing everyday sound environments.

Object-based audio

A fourth approach, common in cinema, stores each source as a mono signal plus metadata saying where it should be, and leaves the renderer in the playback room to work out which speakers to use. This is object-based audio, and it is best understood as deferring the panning decision from the mixing stage to the listening stage.

Electroacoustic systems in practice

Studio recording

A professional recording studio provides a controlled acoustic environment with carefully treated rooms, high-quality microphones, preamplifiers, and monitoring loudspeakers. Common studio microphone setups include:

  • Close-miking: placing the microphone close to the instrument (within 30 cm). Minimises room reflections and maximises the direct-to-reverberant ratio. Subject to proximity effect in directional microphones.
  • Room miking: placing microphones further away (1–5 m) to capture the acoustic of the room as well as the instrument. Used for orchestral recording and when a natural, spacious sound is desired.
  • Stereo miking techniques: using a pair of microphones to capture a stereo image.

Live sound reinforcement

In live sound, the goal is to amplify performers so that the entire audience can hear clearly without the system causing feedback or colouring the sound. Key considerations include:

  • Feedback: occurs when the microphone picks up sound from the loudspeaker, which is re-amplified, creating a loop. Controlling feedback requires careful choice of polar patterns, microphone placement relative to speakers, and equalisation.
  • Monitor speakers (wedges/in-ear monitors): on-stage speakers facing the performers, allowing them to hear themselves and the rest of the ensemble.
  • Line arrays: the vertical columns of speakers hanging beside the stage at large concerts. Stacking elements this way narrows the vertical coverage, so the sound can be aimed at the audience rather than at the ceiling, which is what makes even coverage of an arena possible without drowning the stage in reflections.

Measurement microphones

Electroacoustic measurement uses specialised measurement microphones (typically small-diaphragm condensers with precisely flat, omnidirectional response) to characterise rooms, loudspeakers, and other acoustic systems. Room correction software (such as Dirac Live or Sonarworks) uses measurement microphone recordings to create correction filters that flatten the in-room frequency response of a loudspeaker system at the listening position.

Acoustic feedback and the Larsen effect

The Larsen effect, commonly known as audio feedback, is the howling or screeching sound produced when a microphone picks up the output of the loudspeaker it is feeding. This forms a closed acoustic-electrical loop in which any signal is continuously amplified. The frequency at which feedback occurs depends on the combined frequency response of the entire system (microphone, amplifier, loudspeaker, and room). Understanding electroacoustics helps engineers manage and prevent feedback through equalisation (notch filters), microphone placement, and careful gain management.

Source
import numpy as np
import matplotlib.pyplot as plt

# Illustrate the concept of signal-to-noise ratio and dynamic range
# for different microphone types (idealised)

mic_types = ['Dynamic\n(SM58)', 'Condenser\n(U87)', 'Ribbon\n(R84)', 'MEMS\n(smartphone)']
self_noise = [18, 12, 22, 32]    # dB SPL (A-weighted)
max_spl    = [150, 132, 130, 120] # dB SPL

dynamic_range = [m - n for m, n in zip(max_spl, self_noise)]

x = np.arange(len(mic_types))
width = 0.35

fig, ax = plt.subplots(figsize=(9, 4))
bars1 = ax.bar(x - width/2, max_spl,    width, label='Maximum SPL (dB)',   color='steelblue',  alpha=0.85)
bars2 = ax.bar(x + width/2, self_noise, width, label='Self-noise (dB SPL)', color='darkorange', alpha=0.85)

ax.set_xticks(x)
ax.set_xticklabels(mic_types, fontsize=10)
ax.set_ylabel('dB SPL', fontsize=11)
ax.set_title('Microphone Self-Noise, Maximum SPL, and Dynamic Range (Idealised)', fontsize=11)
ax.legend(fontsize=10)
ax.grid(axis='y', alpha=0.3)

# Annotate dynamic range
for i, (dr, ms, sn) in enumerate(zip(dynamic_range, max_spl, self_noise)):
    ax.annotate(
        f'DR: {dr} dB',
        xy=(i, (ms + sn) / 2),
        ha='center', va='center',
        fontsize=8, color='black',
        bbox=dict(boxstyle='round,pad=0.2', facecolor='white', edgecolor='grey', alpha=0.8)
    )

plt.tight_layout()
plt.show()
<Figure size 900x400 with 1 Axes>

Chapter summary

Electroacoustics follows sound through transduction, gain staging, analogue and digital processing, spatial audio, and reproduction; microphones, loudspeakers, interfaces, and formats connect acoustic reality to the signal chains you use in recording, performance, and analysis. Synthesisers extend the chain by generating signals from scratch, whether by subtracting from rich waveforms, adding partials, modulating frequencies, or replaying and reshaping recorded sound. The mapping case showed that the link between control and sound is a design decision with measurable consequences.

Questions

  1. What is transduction, and how does it appear at both capture (microphones) and reproduction (loudspeakers)?
  2. What are the main stages of a gain chain from mic to listener, and why does gain staging matter?
  3. What do the oscillator, filter, and envelope each contribute in subtractive synthesis, and how do additive, FM, sampling, and granular synthesis differ in how they construct timbre?
  4. How do sample rate, bit depth, and compression choices affect digital audio quality and workflow trade-offs?
  5. In the mapping experiment, why might a cross-coupled mapping outperform a one-to-one mapping, and what limits the strength of that conclusion?
References
  1. Katz, M. (2022). Music and Technology: A Very Short Introduction. Oxford University Press. 10.1093/actrade/9780199946983.001.0001
  2. Rumsey, F., & McCormick, T. (2021). Sound and Recording: Applications and Theory. Routledge. 10.4324/9781003092919
  3. Russ, M. (2009). Sound Synthesis and Sampling (3rd ed.). Focal Press. 10.4324/9780080926957
  4. Hunt, A., Wanderley, M. M., & Paradis, M. (2003). The Importance of Parameter Mapping in Electronic Instrument Design. Journal of New Music Research, 32(4), 429–440. 10.1076/jnmr.32.4.429.18853
  5. Müller, M. (2021). Fundamentals of Music Processing: Using Python and Jupyter Notebooks. Springer International Publishing. 10.1007/978-3-030-69808-9
  6. Shannon, C. E. (1949). Communication in the Presence of Noise. Proceedings of the IRE, 37(1), 10–21. 10.1109/jrproc.1949.232969