Part of Language AI Handbook
Explains how speech recognition systems transform raw audio waveforms into mel spectrograms using STFT, mel filterbanks.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Speech Representations: Mel Spectrograms
Speech exists as a continuous pressure wave traveling through air, a fundamentally different modality from the discrete symbols of text we have processed throughout this book. Unlike the clean, segmented tokens of written language, speech is a dense, high-dimensional signal where meaning is encoded in subtle variations of air pressure over time. Before transformers can process spoken language, we must bridge the gap between the analog acoustic world and the digital feature spaces where neural networks operate. This transformation involves capturing the spectral characteristics of sound while respecting the constraints of human hearing and the computational requirements of machine learning.
The raw waveform, a high-resolution sequence of air pressure measurements, contains all the information in speech but presents it in a form that is computationally unwieldy and perceptually redundant. A single second of CD-quality audio contains 44,100 samples, creating sequences thousands of times longer than typical text inputs. Feed this directly into a transformer, and the quadratic cost of self-attention becomes prohibitive: processing just one second would require computing attention over 44,100 positions, resulting in nearly two billion pairwise comparisons. More importantly, raw samples represent temporal evolution but obscure the frequency content that distinguishes phonemes. While the waveform shows how pressure changes from moment to moment, it does not directly reveal which frequencies are present, which makes it difficult for models to identify the vowel formants or consonant bursts that characterize speech sounds.
The spectral envelope, the distribution of energy across frequencies, determines whether we hear an "ah" or an "ee," while the fine temporal structure carries speaker-specific nuances and pitch information. These two aspects of speech are intertwined in the raw waveform but separable in the frequency domain, which is why almost every successful speech recognition system begins by transforming to a spectral representation.
This chapter develops the representations that make speech tractable for neural networks. We move from time-domain waveforms to frequency-domain spectrograms, then compress these into perceptually motivated mel spectrograms that align with human hearing. We examine the preprocessing pipeline that prepares these features for model training, including normalization techniques that handle the variability inherent in acoustic environments, recording equipment, and speaker characteristics. These foundations prepare us for the speech recognition architectures we will explore in subsequent chapters, where these representations serve as the input to encoder-decoder models like Whisper. Understanding this preprocessing pipeline is needed for implementing speech systems, diagnosing failures, optimizing for specific acoustic conditions, and adapting models to new languages or recording environments.
From Waveforms to Spectra
Speech production generates pressure fluctuations that propagate through air as longitudinal waves. When these pressure variations reach a microphone, the transducer converts the mechanical energy of air pressure into electrical signals through electromagnetic or piezoelectric effects. To process these signals digitally, we must sample the continuous voltage waveform at regular intervals, converting instantaneous amplitude measurements into discrete numerical values that computers can manipulate.
Sampling and the Nyquist Limit
The sampling theorem establishes the basic relationship between continuous and discrete representations of signals. To accurately capture a signal containing frequencies up to without aliasing (distortion from undersampling), we must sample at a rate . This Nyquist rate ensures that the highest frequency component is represented by at least two samples per cycle, letting perfect reconstruction of the original continuous signal from its discrete samples.
If we sample too slowly, high-frequency components fold back into lower frequencies, creating phantom tones that corrupt the signal. Imagine trying to capture a spinning wheel with a strobe light: if the flashes come too slowly, the wheel appears to spin backward or stand still. Similarly, undersampled audio creates false frequencies that were never present in the original sound. This aliasing is not recoverable after the fact; once the signal has been sampled below the Nyquist rate, the original information is irretrievably lost.
The following visualization demonstrates this phenomenon with a 5 Hz sine wave. When sampled at 20 Hz (above the Nyquist rate of 10 Hz), the signal is captured accurately. When sampled at 4 Hz (below Nyquist), the signal appears as a 1 Hz alias, illustrating how undersampling folds high frequencies into lower ones.

For speech, most phonetic information lies below 8 kHz. The basic frequency of human voices typically ranges from 80 Hz (deep male voices) to 250 Hz (high female voices), with important formant structures extending up to around 7-8 kHz. This spectral content suggests a minimum sampling rate of 16 kHz for general-purpose speech recognition. Telephone systems historically used 8 kHz (capturing up to 4 kHz), which preserves intelligibility but loses the high-frequency consonant information that distinguishes sounds like "s" from "f." Modern ASR systems typically work with 16 kHz or 22.05 kHz. This provides enough bandwidth for clear phonetic discrimination while maintaining manageable data volumes. Music and high-fidelity audio extend to 44.1 kHz or 48 kHz to capture the full range of human hearing up to 20 kHz.
The raw waveform represents amplitude at discrete time steps, but this representation is inefficient for phonetic analysis. The same phoneme spoken by different voices or in different contexts produces wildly different waveforms due to variations in pitch, speaking rate, and vocal tract anatomy, yet similar spectral patterns remain consistent. We need a representation that separates the slowly varying spectral envelope (carrying phonetic identity) from the rapid fine structure (carrying pitch and speaker characteristics). This separation allows recognition systems to identify phonemes regardless of whether they are spoken by a bass or soprano, while still preserving the acoustic cues that distinguish individual speakers.
Short-Time Fourier Transform
Speech is inherently non-stationary, meaning its statistical properties change continuously over time as speakers move their articulators and transition between sounds. However, over short intervals of 20-30 milliseconds, the vocal tract configuration remains approximately constant, and the signal behaves locally like a stationary process. This quasi-stationarity arises because human articulators, such as the tongue and lips, have mass and inertia that prevent instantaneous changes. A speaker transitioning from "ee" to "ah" does not flip a switch; instead, the tongue body moves over tens to hundreds of milliseconds, creating a gradual spectral trajectory that frames can track.
The Short-Time Fourier Transform (STFT) exploits this property by applying the Fourier transform to overlapping windows of the signal, analyzing each frame as if it were a stationary tone complex.
The STFT analyzes a signal's frequency content as it evolves over time by computing the Fourier transform of short overlapping windows. Each window produces a snapshot of the frequency spectrum at that moment, and stacking these snapshots produces the spectrogram, a two-dimensional time-frequency representation.
The STFT computes:
where:
- : the complex-valued STFT coefficient at time frame and angular frequency , representing both magnitude and phase
- : the discrete-time input signal at sample index
- : the window function centered at time , which tapers the signal to zero at frame boundaries to prevent spectral leakage
- : the imaginary unit (), representing the 90-degree phase shift in complex exponentials
- : the complex exponential basis function (cosine and sine components) for frequency analysis
- : the time frame index, showing the position of the analysis window
- : the angular frequency in radians per sample
- : the sample index within the summation window, ranging over the effective window length
The window function is centered at time and tapers the signal to zero at the edges, preventing spectral leakage. Without this tapering, the abrupt truncation of the signal at frame boundaries would create discontinuities, introducing high-frequency artifacts that spread across the spectrum and obscure the true frequency content. The window slides across the signal with a hop length typically set to 25% or 50% of the window size to ensure continuity between frames and prevent temporal aliasing.
The choice of window length creates an unavoidable tradeoff between time and frequency resolution, sometimes called the Heisenberg-Gabor limit in signal processing. A longer window improves frequency resolution (finer frequency bins) but reduces temporal resolution (each frame covers a longer time span, blurring rapid changes). A shorter window captures transient events precisely in time but cannot resolve closely spaced frequency components. For speech recognition, the standard 25ms window strikes a balance: long enough to resolve formant frequencies that are typically spaced 200-400 Hz apart, yet short enough to capture the rapid spectral changes of stop consonants and fricatives that evolve over 20-50ms.
Common window functions include:
- Hamming window:
- Hann window:
The Hamming window minimizes the maximum side lobe, reducing spectral leakage at the cost of slightly wider main lobes. The Hann window offers smoother spectral decay with lower side lobes but slightly worse frequency resolution. For speech, the Hamming window has been traditional due to its optimal side-lobe properties, though Hann and Blackman windows also see use in modern systems where reduced spectral leakage is prioritized.
@fig-window-functions compares these window functions in both time and frequency domains. The time domain shows the tapering applied to frame edges, while the frequency domain reveals the tradeoff between main lobe width (frequency resolution) and side lobe levels (spectral leakage).


The output of the STFT is a complex matrix representing magnitude and phase for each time-frequency bin. For speech recognition, we typically discard the phase and work with the power spectrum or magnitude spectrum, as human perception is relatively insensitive to phase relationships between frequencies. This phase insensitivity, sometimes called "phase deafness," arises from the way the cochlea processes sound: hair cells respond to the amplitude envelope of frequency bands rather than instantaneous phase. The squared magnitude of each complex coefficient gives us the power spectrum:
where:
- : the power spectral density (energy) at time frame and frequency bin
- : the real component of the complex STFT coefficient
- : the imaginary component of the complex STFT coefficient
- : the time frame index
- : the frequency bin index, corresponding to discrete frequencies from 0 to the Nyquist rate
Discarding phase simplifies the representation while retaining the phonetically relevant information for recognition. The tradeoff matters primarily when reconstructing audio from features (as in speech synthesis or enhancement), where phase recovery algorithms like Griffin-Lim must be applied.
Spectrograms
Visualizing the power spectrum across time produces the spectrogram, a two-dimensional heatmap showing how frequency content evolves throughout an utterance. The horizontal axis represents time, the vertical axis represents frequency, and color intensity represents energy (typically in decibels). This representation provides an intuitive view of speech structure, revealing horizontal bands of energy corresponding to formants, vertical striations showing periodic glottal pulses, and transient bursts marking consonant releases.
Standard spectrograms use a linear frequency scale, equally spacing Hz across the vertical axis. This allocation misaligns with human perception. We hear pitch logarithmically: the perceptual distance between 100 Hz and 200 Hz is much larger than between 1000 Hz and 1100 Hz, even though both differ by 100 Hz. Our frequency resolution decreases as frequency increases, following the organization of the basilar membrane in the cochlea. A linear spectrogram wastes resolution on high frequencies where we cannot discriminate fine differences, while undersampling the low frequencies where small changes matter perceptually. For instance, the distinction between the first formants of "bat" and "bet" occurs in the 500-700 Hz range, requiring fine resolution that linear spectrograms spread thin across the entire frequency axis.
A related concept is the distinction between "narrowband" and "broadband" spectrograms, which reflect different window length choices rather than frequency axis scaling. A narrowband spectrogram uses a long window (e.g., 100ms), creating fine frequency resolution that reveals individual harmonics as horizontal striations. A broadband spectrogram uses a short window (e.g., 5ms), sacrificing frequency resolution to reveal rapid temporal events like glottal pulses as vertical striations. Modern ASR systems typically use configurations closer to broadband, focusing on the spectral envelope rather than individual harmonics, since the harmonic spacing carries pitch information that is not reliably correlated with phonetic identity across speakers.
The Mel Scale and Filterbanks
The mel scale addresses the mismatch between linear frequency and human perception by warping the frequency axis to approximate human pitch perception. Developed through psychoacoustic experiments in the 1930s and 1940s, where listeners judged subjective pitch distances between pure tones, the mel scale spaces frequencies according to perceived equality rather than physical equality. On the mel scale, equal distances correspond to perceptually equal steps, compressing high frequencies and expanding low frequencies to match the non-linear frequency resolution of the human ear.
The name "mel" comes from "melody," which reflects its origins in musical pitch perception research. The psychoacoustic experiments involved subjects adjusting the frequency of one tone until it sounded exactly halfway between two reference tones in pitch, building up a perceptual scale of equal pitch steps. This differs from the physical frequency scale in important ways: the mel scale is nearly linear below 1000 Hz and approximately logarithmic above, following the pattern of the basilar membrane's mechanical frequency analysis.
Mel Scale Mathematics
The mel scale conversion from frequency (in Hz) to mel follows:
where:
- : the perceived pitch in mel units, spaced according to human perception
- : the physical frequency in Hertz (Hz)
- : the scaling constant that maps 1000 Hz to approximately 1000 mel, anchoring the scale at a convenient reference point
- : the base-10 logarithm that converts frequency ratios to perceived pitch differences, with approximately linear response below 1000 Hz and logarithmic response above
Alternatively, some implementations use the natural logarithm variant:
where:
- : the perceived pitch in mel units
- : the physical frequency in Hertz (Hz)
- : the scaling constant for the natural logarithm variant, mathematically equivalent to to maintain the same mapping
- : the natural logarithm (base ), giving the same perceptual warping as the base-10 variant
The two formulas are mathematically equivalent and produce the same frequency warping; they differ only in the numerical value of the mel units. The shape remains: approximately linear below 1000 Hz (where human hearing is roughly linear), transitioning to logarithmic above 1000 Hz. This reflects the structure of the cochlea, where the basal end (high frequencies) has logarithmic spacing of hair cells, while the apical end (low frequencies) has more linear organization. The constant 700 Hz represents the corner frequency where the transition occurs, determined empirically from psychoacoustic studies.
The inverse transformation from mel to Hz is also needed when constructing filterbanks:
This inverse allows us to convert equally spaced points in mel space back to their corresponding Hz values for constructing the actual filters.
@fig-mel-scale-mapping illustrates this nonlinear relationship, showing how the mel scale compresses high frequencies and expands low frequencies to match human perceptual resolution.

To create a mel spectrogram, we need filters that collect energy according to this perceptual scale rather than the linear bins of the FFT, aggregating fine frequency details where they matter most for phonetic discrimination.
Mel Filterbank Design
A mel filterbank consists of overlapping triangular filters spaced uniformly in the mel domain. The key insight is that "uniformly spaced in mel" means "logarithmically spaced in Hz," giving us dense coverage at low frequencies and sparse coverage at high frequencies, exactly matching the perceptual resolution of the ear.
The construction process involves five steps:
- Select frequency range: Typically 0 Hz to the Nyquist frequency (), or a narrower range like 0 to 8000 Hz for speech
- Convert to mel scale: Map the frequency bounds to mel values using the mel formula
- Create equally spaced points: Divide the mel range into points (where is the number of filters), creating boundaries for the triangular filters
- Convert back to Hz: Map these mel points back to the linear frequency scale to determine the center frequencies of each filter in physical frequency units
- Construct triangular filters: For each filter , create a triangle with its peak at mel point and zero crossings at mel points and
The unequal spacing in Hz means that low-frequency filters are narrow (covering perhaps 50-100 Hz each) while high-frequency filters are wide (covering 500-1000 Hz each). This is not a limitation but a feature: it matches the frequency resolution of the cochlea, which has more hair cells per octave at low frequencies than at high frequencies.
Mathematically, the -th filter applied to FFT bin (corresponding to frequency ) is defined as:
where:
- : the weight of the -th mel filter applied to FFT bin , ranging from 0 to 1
- : the center frequency of FFT bin in the linear frequency domain (Hz)
- : the center frequency of the -th mel filter in Hz (the peak of the triangle)
- and : the center frequencies of adjacent mel filters, defining the lower and upper bounds of the triangular passband
- The first case rejects frequencies below the filter's bandwidth
- The second case linearly interpolates from 0 to 1 as frequency increases from to
- The third case linearly decreases from 1 to 0 as frequency increases from to
- The fourth case rejects frequencies above the filter's bandwidth
The triangular shape has an elegant property: each pair of adjacent filters partitions their shared frequency band without gaps and without double-counting, since whenever filter 's weight increases from 0 to 1, filter 's weight correspondingly decreases from 1 to 0. This partition of unity property ensures that the total energy is conserved (no frequency component is lost or double-weighted) while giving smooth interpolation between filters.
The filters overlap, with each frequency bin adding to multiple mel bins. This smoothing reduces dimensionality while preserving perceptually relevant information, and provides a form of spectral averaging that makes the representation reliable to small frequency shifts. Typical ASR systems use 40, 80, or 128 mel filters, balancing resolution against computational cost. More filters provide finer frequency discrimination but increase the input dimensionality to the neural network. Whisper, the speech recognition model we will examine in subsequent chapters, uses 80 mel bins, a choice that provides good phonetic resolution while keeping the input compact.
Computing Mel Spectrogram Energies
To compute the mel spectrogram, we apply these filters to the power spectrum of each frame, effectively asking: "How much energy is present in each perceptual frequency band?"
where:
- : the mel spectrogram energy at mel filter and time frame , representing the total power in that mel band
- : the power spectrum at time frame and FFT bin , computed as the squared magnitude of the STFT
- : the weight of the -th mel filter for FFT bin , determining how much that frequency bin contributes to the mel band
- : the time frame index, corresponding to discrete time steps (typically every 10 ms)
- : the mel filter index, ranging from 0 to where is the total number of mel filters (typically 80)
- : the FFT bin index, ranging from 0 to (the number of unique frequency bins for a real signal)
- : the FFT size (number of points in the Fast Fourier Transform, typically 512 or 1024)
This summation aggregates energy across FFT bins according to the mel filter weights, creating a vector of values for each time frame. The result is a compressed representation where each mel bin represents the total energy in a perceptual frequency band rather than a single frequency component. The compression ratio is substantial: a 512-point FFT produces 257 unique frequency bins, but we condense these into 80 mel bins, reducing the frequency dimension by more than threefold while preserving the perceptually important structure.
Log Compression
Human loudness perception follows a logarithmic scale known as the Weber-Fechner law, which states that the just-noticeable difference in stimulus intensity is proportional to the current intensity level. A doubling of acoustic energy does not produce a doubling of perceived loudness; instead, we perceive loudness on a roughly logarithmic scale. This property has a practical consequence for speech processing: the raw energy values in mel spectrogram bins can span many orders of magnitude between a quiet fricative and a loud vowel. A softmax vowel might produce 1000 times the energy of a quiet "sh," meaning these two phonemes appear very different in linear scale but comparably different in log scale.
To match this perceptual non-linearity and to compress the dynamic range of the signal, we apply a logarithm to the mel energies:
where:
- : the log-compressed mel spectrogram value at mel filter and time
- : the mel spectrogram energy at mel filter and time (linear scale)
- : a small constant (typically or ) added for numerical stability
- : the logarithm function (natural log or base-10), compressing the dynamic range to match human loudness perception
The small constant prevents numerical instability when approaches zero, avoiding negative infinity values in logarithmic space. The result is often expressed in decibels relative to some reference, though for machine learning purposes, the natural log or base-10 log suffices without explicit dB conversion.
Log compression also aligns with the properties that make mel spectrograms easy for neural networks to process. After log compression, the numerical range of the features is typically -100 to 0 dB (or an equivalent log scale), which is much more tractable for gradient-based optimization than raw energies spanning many orders of magnitude. Gradient updates are more stable when input features have comparable magnitudes across all dimensions, and log compression contributes substantially to this stability.
The log-mel spectrogram has become the standard input representation for neural speech recognition systems, giving a good tradeoff between:
- Temporal resolution: Capturing rapid phonetic transitions and consonant bursts
- Frequency resolution: Respecting perceptual frequency spacing where low frequencies require fine discrimination
- Dimensionality: Typically 80 dimensions per frame versus 256+ for linear spectrograms, reducing computational load
- Invariance: Robust to minor pitch variations and phase shifts that do not affect phonetic identity
Audio Preprocessing Pipeline
Before computing mel spectrograms, several preprocessing steps condition the signal for analysis. These steps standardize the input, emphasize speech-relevant frequencies, and prepare the signal for spectral analysis. Each step addresses specific physical or perceptual properties of speech production and hearing, forming a principled pipeline that has evolved through decades of speech processing research.
Pre-emphasis
Speech signals have more energy at low frequencies (below 1 kHz) than at high frequencies, due to the physics of vocal fold vibration and vocal tract resonance. As the vocal folds open and close, they create a glottal pulse rich in harmonics, but the spectral envelope naturally rolls off at approximately -6 dB per octave due to the radiation characteristic of the lips and the physical properties of the glottal source. This spectral tilt causes high-frequency consonants (like fricatives) to be relatively weak compared to vowels. If we compute the mel spectrogram without pre-emphasis, the low-frequency vowel energy will dominate and high-frequency consonant information may be too faint for the network to learn from reliably.
Pre-emphasis applies a first-order high-pass filter to compensate:
where:
- : the pre-emphasized signal at sample
- : the input signal at sample
- : the pre-emphasis coefficient (typically 0.97), controlling the degree of high-frequency boost
- : the input signal at the previous sample (delay of one sample)
With , this filter boosts high frequencies by approximately 6 dB per octave, flattening the spectrum and compensating for the natural roll-off of speech. To understand why rather than some other value, consider the z-domain transfer function of this filter: . Its zero is at , very close to the unit circle at low frequencies, meaning the filter has weak attenuation at DC and increasingly strong attenuation as frequency decreases toward zero. The result is a gentle high-pass filter that amplifies high frequencies without introducing audible artifacts.
Pre-emphasis also compensates for the radiation characteristic of the lips, which attenuates low frequencies relative to high frequencies. By equalizing the spectrum, pre-emphasis ensures that high-frequency consonant information (necessary for distinguishing "s" from "sh" or "f" from "v") receives adequate representation in the subsequent spectral analysis.
The effect is reversible through de-emphasis, though in practice recognition systems work entirely in the pre-emphasized domain since the neural network learns to interpret the equalized spectral patterns directly.
Framing and Windowing
Speech analysis proceeds frame-by-frame to respect the quasi-stationarity assumption discussed earlier. Standard parameters include:
- Frame size: 25 milliseconds (400 samples at 16 kHz)
- Frame shift (hop length): 10 milliseconds (160 samples at 16 kHz)
- Overlap: 60% (15 ms of each frame overlaps with neighbors)
This overlap ensures temporal continuity and prevents information loss at frame boundaries. A phoneme transition falling between frame centers will still influence adjacent frames, maintaining acoustic context. The 25ms window provides sufficient duration to resolve frequency details (frequency resolution is inversely proportional to window length) while being short enough to capture rapid phonetic changes.
The relationship between window length and frequency resolution is quantitative. For a window of samples at sampling rate , the frequency resolution is . At 16 kHz with samples (25ms), this gives Hz per frequency bin. Formant frequencies in typical male speech are spaced roughly 1000-1200 Hz apart, meaning our 40 Hz resolution is more than sufficient to separate them. For higher-pitched voices with formants spaced closer together, the 25ms window remains adequate; the basic frequency (100-300 Hz for most speakers) determines the harmonic spacing, and the formant envelope is typically much smoother than individual harmonics.
After extracting each frame, we apply a window function (Hamming or Hann) to taper the signal to zero at the edges. Without windowing, the sharp discontinuities at frame boundaries would introduce spectral artifacts called sidelobes, spreading energy from strong frequencies into other frequency bins and obscuring the true spectral content.
FFT and Power Spectrum
Each windowed frame undergoes Fast Fourier Transform (typically with or points, zero-padded if necessary). Zero-padding interpolates the spectrum. This provides smoother frequency bins without adding actual frequency resolution (which is determined by the actual data length, not the FFT size). We compute the power spectrum by squaring the magnitude of the complex FFT output:
where:
- : the power spectral density at frequency bin
- : the FFT size (number of points in the Fast Fourier Transform)
- : the squared magnitude of the complex FFT output at bin
- : the frequency bin index, ranging from to
For to , the remaining bins are redundant for real signals due to conjugate symmetry. This yields unique frequency bins spanning 0 to Hz. The power spectrum represents the distribution of energy across frequencies for that brief moment in time.
The FFT is computationally efficient due to the Fast Fourier Transform algorithm, which reduces the complexity of computing the Discrete Fourier Transform from to . For typical frame sizes and batch processing, this makes spectrogram computation a small fraction of the total processing time compared to the neural network forward pass.
Complete Pipeline Summary
The full preprocessing pipeline turns raw audio into model-ready features through a sequence of operations that progressively extract and compress perceptually relevant information:
- Pre-emphasis: : Equalizes the spectrum by boosting high frequencies
- Framing: Extract overlapping 25ms frames every 10ms: Segments the signal into quasi-stationary chunks
- Windowing: Apply Hamming window to each frame: Tapers edges to prevent spectral leakage
- FFT: Compute 512-point FFT, extract magnitude: Transforms to frequency domain
- Power spectrum: Square magnitudes: Converts to energy representation
- Mel filtering: Apply 80 triangular mel filters: Aggregates linear frequencies into perceptual bands
- Log compression: Apply natural log with : Matches human loudness perception and compresses dynamic range
The output is a 2D tensor of shape where is the number of time frames (100 per second for 10ms shift) and is the number of mel bins (typically 80). This represents a compression from 16,000 samples per second to 8,000 values per second (100 frames 80 mel bins), a 50-fold reduction in dimensionality while preserving phonetically necessary information.
Feature Normalization
Raw mel spectrograms vary between speakers and utterances, especially when recording conditions change. Differences in microphone distance, room acoustics, and speaking volume create variations that obscure the linguistic content. Feature normalization stabilizes training and improves generalization by removing non-linguistic variation, letting the model to focus on phonetic patterns rather than recording artifacts.
To understand why normalization matters so much for speech, consider what happens without it. A speaker 30 cm from the microphone produces far more energy than one at 1 meter, shifting the entire spectrogram up or down uniformly. A phone conversation passes through bandwidth-limiting filters that attenuate both low and high frequencies. Background noise raises the energy floor in all frequency bins. These effects are consistent within a recording but vary dramatically between recordings, creating a distribution mismatch that degrades model performance on new recordings that differ in acoustic conditions from the training data.
Cepstral Mean and Variance Normalization
Cepstral Mean and Variance Normalization (CMVN) standardizes features to zero mean and unit variance, effectively removing channel characteristics and volume differences. For a sequence of feature vectors , we compute:
where:
- : the mean feature vector computed across all time frames, representing the average spectral characteristics
- : the variance vector (element-wise variance across frequency bins), measuring the spread of feature values
- : the normalized feature vector at time , with zero mean and unit variance (z-score normalization)
- : the original input feature vector (e.g., log-mel spectrogram frame) at time
- : the total number of time frames in the utterance or normalization window
- : a small constant added to the variance for numerical stability, preventing division by zero when features have near-zero variance
- The square root in the denominator converts variance to standard deviation for proper normalization scaling
The mean subtraction step removes additive channel effects. If a microphone consistently boosts a frequency band by a fixed amount, that offset will appear in the mean and be subtracted away. The variance normalization step addresses multiplicative effects: if one speaker consistently has higher-energy speech than another, dividing by the standard deviation brings all frequency bands to comparable scale.
CMVN can be applied per utterance (using statistics from the current speech segment) or globally (using statistics from the training corpus). Utterance-level CMVN handles channel and speaker variation effectively but requires the entire utterance before processing, introducing latency unsuitable for streaming recognition where decisions must be made with minimal delay. Global CMVN enables frame-by-frame processing but is less adaptive to specific recording conditions or speakers not represented in the training data.
For online recognition systems that process live audio, running estimates of mean and variance can be maintained and updated incrementally using exponential moving averages, approximating the full-sequence statistics without requiring future context. This approach effectively estimates the "current channel conditions" as audio streams in, adapting to changes in acoustic environment within a single session.
Delta and Delta-Delta Features
While static mel spectrograms capture spectral content at each frame, they do not explicitly represent how spectra change over time. Temporal dynamics are important for phonetic discrimination: vowels have formant transitions (moving resonances), and consonants have rapid spectral changes (bursts, friction noise). A recording of the vowel "ee" and a recording of "oo" differ primarily in their static spectral patterns. But the difference between the word "bad" and "bid" depends partly on the formant trajectory during the vowel, including the direction and rate of formant movement rather than only its static endpoint.
Delta features approximate the first derivative (velocity) of the feature trajectory, showing how quickly the spectrum is changing:
where:
- : the delta feature (first derivative approximation) at time , representing the rate of change of the spectral envelope
- : the static feature value (e.g., a single mel coefficient) at time
- : the static feature value frames ahead of time (future context)
- : the static feature value frames before time (past context)
- : the regression window size (typically 2), meaning the computation uses a 5-frame window ( to )
- The numerator computes a weighted sum of differences, giving more weight to frames further from the center (linear weighting by )
- The denominator normalizes by the sum of squared weights ( for ) to produce an unbiased estimate of the derivative
Delta-delta features (acceleration) are computed by applying the same formula to the delta features, capturing how the rate of change itself is changing. This second derivative is sensitive to the inflection points in spectral trajectories, useful for detecting the onset of formant transitions or the release of stop consonants.
Concatenating static, delta, and delta-delta features triples the feature dimension but provides explicit temporal context. Modern deep learning systems with large receptive fields (convolutional layers or long attention windows) often learn these temporal patterns implicitly from static features, making explicit delta features less common than in classical GMM-HMM systems. However, they remain useful for resource-constrained models or as auxiliary inputs that explicitly guide the model toward dynamic patterns.
Spectral Subtraction and Noise Robustness
In noisy environments, simple normalization may not suffice. Additive background noise raises the energy floor across frequencies, masking weak speech components and distorting spectral patterns. A classic example is the cocktail party problem: background conversations at comparable volume to the target speaker create spectral interference that appears across all frequency bands, making CMVN ineffective since the noise shifts the mean uniformly.
Spectral subtraction estimates the noise spectrum during non-speech segments (silence periods detected by voice activity detection) and subtracts it from the signal:
where:
- : the estimated clean signal power at frequency bin after noise removal
- : the observed noisy signal power at frequency bin , containing both speech and noise components
- : the estimated noise power at frequency bin , typically measured during non-speech segments (silence)
- : the over-subtraction factor (typically greater than 1, such as 1.5 to 3.0), which scales the noise estimate to account for noise variability and prevent musical noise artifacts
- : the frequency bin index across the spectrum
The over-subtraction factor accounts for the fact that noise varies over time; simply subtracting the average noise estimate leaves residual noise during high-noise frames. However, aggressive subtraction can introduce "musical noise," artifactual tones resulting from isolated spectral peaks that appear when most frequency bins are suppressed but a few retain significant energy. This musical noise is perceptually distracting and can confuse recognition systems. More sophisticated techniques like Wiener filtering or minimum mean-square error spectral estimation handle this more gracefully by adapting the subtraction amount per-bin based on local signal-to-noise estimates.
Vocal Tract Length Normalization (VTLN) addresses a different form of variability: anatomical differences between speakers. A child's shorter vocal tract produces formants at higher frequencies than an adult's longer tract. VTLN warps the frequency axis of the spectrogram by a speaker-specific factor , effectively normalizing for vocal tract length differences. Classical systems estimated this warp factor from the speech itself; modern deep learning systems often learn speaker normalization implicitly through multi-speaker training or explicit speaker embedding conditioning.
Deep learning has largely absorbed these preprocessing functions into learned representations that are reliable to such variations, but understanding the signal processing foundations helps you diagnose failures when models struggle with specific acoustic conditions, design targeted augmentation strategies, and adapt pretrained models to challenging environments.
Worked Example: Computing a Mel Spectrogram
Let's trace through a concrete example to solidify these concepts. Consider a 1-second audio clip sampled at 16,000 Hz containing the spoken word "hello." This word contains a variety of phonetic elements: the aspiration of the "h," the diphthong quality of the "e," the lateral approximant "l," and the rounded vowel "o."
With a frame size of 25ms (400 samples) and hop length of 10ms (160 samples), we extract frames starting at samples 0, 160, 320, up to 15,840. This yields 100 frames for the 1-second clip. The total frame count depends on boundary handling: with zero-padding at the end, we get exactly 100 frames; without padding, slightly fewer.
For each frame, the processing proceeds as follows. First, we apply the Hamming window: multiply the 400 samples by the window function, tapering the edges to zero. The center 200 samples receive weights close to 1.0, while the first and last samples receive weights close to 0. This smoothing prevents the FFT from seeing artificial discontinuities at frame boundaries.
Second, we compute the 512-point FFT (padded with 112 zeros at the end to reach 512 samples) to get smooth spectral interpolation. The zero-padding does not add information; it simply fills in additional points on the continuous spectral curve, making the discrete frequency bins more closely spaced without changing the basic frequency resolution. Third, we take the magnitude and square it to get the power spectrum (257 bins from 0 to 8 kHz), discarding the complex phase information.
Fourth, we apply 80 mel filters, each spanning a range of FFT bins. The low-frequency filters, covering 0-500 Hz, span very few FFT bins each (perhaps 3-5 bins) because the mel scale allocates many filters to this low-frequency region. The high-frequency filters, covering 4-8 kHz, span many FFT bins each (perhaps 20-40 bins) because the mel scale compresses this region into fewer filters. Fifth, for each filter, we compute the weighted sum of power spectrum values, creating one mel-band energy value. Sixth, we take the natural logarithm to compress the dynamic range.
The result is a matrix representing the time-frequency evolution of the utterance. The first few rows show low energy across all frequencies (silence before speech begins), followed by the broad-spectrum noise of the "h" aspirate. The "e" vowel appears as concentrated energy bands at specific frequencies (the first formant near 400 Hz and the second formant near 2000 Hz for this vowel). The lateral "l" sounds show a specific resonance pattern, and the rounded "o" shifts the formant pattern to lower frequencies (first formant near 500 Hz, second formant near 1000 Hz for a rounded back vowel). The final frames return to near-silence as the utterance ends.
This spatial structure in the mel spectrogram, where phonemes appear as distinctive horizontal patterns that shift vertically over time, is what transformer encoders learn to read. The self-attention mechanism can relate any two time positions, discovering that the formant pattern at frame 30 during the "e" corresponds to a particular phoneme while the shifted pattern at frame 70 during the "o" corresponds to another, even if the patterns are visually similar but physically at different frequency positions.
Code Implementation
Let's implement the complete speech preprocessing pipeline using NumPy, SciPy, and Matplotlib. We will create a speech-like sample, compute mel spectrograms, and visualize the representations.
import numpy as np
# Set random seed for reproducibility
np.random.seed(42)
# Generate a synthetic speech-like signal
# In practice, you would load audio from disk and resample it to this rate.
sr = 16000 # Sample rate
duration = 2.0 # seconds
t = np.linspace(0, duration, int(sr * duration))
# Create a synthetic signal with formant-like structure
# Fundamental frequency (pitch) around 150 Hz with harmonics
f0 = 150
signal = np.zeros_like(t)
for harmonic in range(1, 6):
# Add harmonics with decreasing amplitude (simulating glottal source)
signal += (0.5**harmonic) * np.sin(2 * np.pi * f0 * harmonic * t)
# Add formant structure (resonances) by filtering
# Simulate two formants at 500 Hz and 1500 Hz
from scipy.signal import butter, lfilter
def add_formant(signal, sr, freq, bandwidth=100):
"""Add a formant resonance using a bandpass filter"""
nyquist = sr / 2
low = (freq - bandwidth / 2) / nyquist
high = (freq + bandwidth / 2) / nyquist
b, a = butter(2, [low, high], btype="band")
return lfilter(b, a, signal)
signal = add_formant(signal, sr, 500, 200)
signal = add_formant(signal, sr, 1500, 300)
signal = signal * 0.5 # Normalize amplitude
# Add some noise for realism
signal += 0.01 * np.random.randn(len(signal))We have created a synthetic speech signal with harmonic structure and formant resonances. In real applications, you would load recorded speech from an audio file and resample it to the target rate, such as 16 kHz for many ASR systems. Now let's examine the raw waveform and compute its spectrographic representations.

Sample rate: 16000 Hz Duration: 2.0 seconds Total samples: 32000
The waveform shows the oscillating pressure variations characteristic of voiced speech. The dense oscillations represent the high-frequency components, while the amplitude envelope varies slowly according to the syllabic structure. Next, we compute the Short-Time Fourier Transform to examine the frequency content over time.
import numpy as np
# Pre-emphasis
pre_emphasis = 0.97
emphasized_signal = np.append(
signal[0], signal[1:] - pre_emphasis * signal[:-1]
)
# STFT parameters
n_fft = 512
hop_length = 160 # 10ms at 16kHz
win_length = 400 # 25ms at 16kHz
# Compute STFT using scipy
from scipy import signal as sp_signal
# Create Hamming window
hamming_window = 0.54 - 0.46 * np.cos(
2 * np.pi * np.arange(win_length) / (win_length - 1)
)
# Compute STFT
frequencies, times, stft_matrix = sp_signal.stft(
emphasized_signal,
fs=16000,
nperseg=win_length,
noverlap=win_length - hop_length,
nfft=n_fft,
window=hamming_window,
boundary="zeros",
)
magnitude = np.abs(stft_matrix)
power_spectrogram = magnitude**2
# Convert to dB scale for visualization
db_spectrogram = 20 * np.log10(magnitude + 1e-10)
db_spectrogram = db_spectrogram - np.max(
db_spectrogram
) # Normalize to max = 0 dB@fig-preemphasis-effect demonstrates the spectral effect of pre-emphasis on our synthetic speech signal. The filter boosts high-frequency components, flattening the spectral envelope and compensating for the natural roll-off of the glottal source.


Spectrogram shape: (257, 201) Frequency bins: 257 Time frames: 201
The linear spectrogram shows frequency on the vertical axis with equal spacing per Hz. Notice how the harmonic structure appears as horizontal bands, but the low-frequency region (below 2 kHz) appears compressed while high frequencies are stretched. The fine vertical striations indicate the periodic glottal pulses. Now let's compute the mel spectrogram with perceptual frequency scaling.
import numpy as np
# Mel spectrogram parameters
n_mels = 80
f_min = 0
f_max = 8000 # 8 kHz
def hz_to_mel(freq_hz):
return 2595 * np.log10(1 + freq_hz / 700)
def mel_to_hz(mel):
return 700 * (10 ** (mel / 2595) - 1)
def create_mel_filterbank(sr, n_fft, n_mels, fmin, fmax):
mel_points = np.linspace(hz_to_mel(fmin), hz_to_mel(fmax), n_mels + 2)
hz_points = mel_to_hz(mel_points)
freqs = np.linspace(0, sr / 2, 1 + n_fft // 2)
filters = np.zeros((n_mels, len(freqs)))
for i in range(n_mels):
left, center, right = hz_points[i : i + 3]
left_slope = (freqs - left) / (center - left)
right_slope = (right - freqs) / (right - center)
filters[i] = np.maximum(0, np.minimum(left_slope, right_slope))
return filters, freqs
mel_filters, mel_filter_freqs = create_mel_filterbank(
sr=sr, n_fft=n_fft, n_mels=n_mels, fmin=f_min, fmax=f_max
)
mel_spectrogram = mel_filters @ power_spectrogram
# Convert to log scale (dB)
log_mel_spectrogram = 10 * np.log10(
np.maximum(mel_spectrogram, 1e-10) / np.max(mel_spectrogram)
)@fig-pipeline-stages visualizes the transformation of a single speech frame through the preprocessing stages: from the raw waveform through windowing, Fourier analysis, and mel filtering.






The comparison reveals how the mel scale redistributes frequency resolution. In the linear spectrogram, the region below 1 kHz occupies a small portion of the vertical space, yet this is where important phonetic information resides, including the first two formants that distinguish most vowels. The mel spectrogram expands this region, giving the model more parameters to distinguish low-frequency content while compressing the high-frequency region where fine discrimination is less perceptually important.
Now let's implement feature normalization and examine the effect of CMVN.
# Compute statistics for normalization
# Mean and variance across time frames (axis=1)
mean = np.mean(log_mel_spectrogram, axis=1, keepdims=True)
std = np.std(log_mel_spectrogram, axis=1, keepdims=True)
# Apply CMVN
normalized_mel = (log_mel_spectrogram - mean) / (std + 1e-8)
# Alternative: Global normalization (using full spectrogram statistics)
global_mean = np.mean(log_mel_spectrogram)
global_std = np.std(log_mel_spectrogram)
global_normalized = (log_mel_spectrogram - global_mean) / (global_std + 1e-8)

Original range: [-55.49, 0.00] Normalized range: [-5.60, 3.04] Normalized mean: -0.0000, std: 1.0000
Normalization centers the features and scales them to comparable ranges across frequency bands. This prevents high-energy low-frequency bands from dominating the learning process and makes the model more reliable to variations in recording volume and microphone characteristics. The normalized spectrogram shows enhanced contrast in mid-frequency regions that might have been overshadowed by strong low-frequency energy in the original.
Finally, let's examine the mel filterbank itself to understand how frequencies are aggregated.
# Frequency axis for plotting
freqs = mel_filter_freqs
Filterbank shape: (80, 257) Number of filters: 80 FFT bins: 257 First filter center: ~31 Hz Last filter center: ~7719 Hz
The filterbank visualization demonstrates the non-linear spacing. The leftmost filters are narrow and steep. This provides fine resolution in the low-frequency range where formant information is necessary. The rightmost filters are broad and gentle, aggregating energy across wide swaths of high frequencies where precise discrimination is less perceptually important. This allocation reflects the frequency resolution of the human ear, optimizing the representation for speech perception.
Key Parameters
The key parameters for the speech preprocessing pipeline are summarized in Table preprocessing-params.
| Parameter | Typical Value | Description |
|---|---|---|
| sr | 16000 Hz | Sample rate capturing frequencies up to 8 kHz |
| pre_emphasis | 0.97 | High-pass filter coefficient boosting high frequencies |
| n_fft | 512 | FFT points determining frequency resolution (bin width = sr/n_fft) |
| hop_length | 160 samples (10 ms) | Frame shift controlling temporal step between frames |
| win_length | 400 samples (25 ms) | Analysis window length affecting frequency resolution |
| n_mels | 80 | Number of mel frequency bins (40-128 typical range) |
| f_min | 0 Hz | Lower frequency bound for mel filterbank |
| f_max | 8000 Hz | Upper frequency bound for mel filterbank |
Higher sample rates preserve more high-frequency information but increase processing requirements linearly. Values closer to 1.0 for pre-emphasis provide a stronger high-frequency boost. Larger n_fft values provide finer frequency bins but reduce temporal localization. More mel bins provide finer frequency resolution but increase feature dimensionality. These parameters interact: increasing n_fft without increasing win_length amounts to zero-padding, which smooths the spectrum but does not add frequency resolution. Truly improving frequency resolution requires a longer analysis window, which in turn reduces temporal resolution. These constraints are basic and cannot be engineered away; the goal is choosing parameters that best serve the phonetic discrimination task given available compute.
Limitations and Design Tradeoffs
While mel spectrograms dominate neural speech recognition, they embody specific assumptions that may not suit all applications. Understanding these limitations helps practitioners choose appropriate representations and recognize when alternative approaches might be necessary.
Time-Frequency Resolution Tradeoff
The STFT assumes stationarity within each frame, fixing a tradeoff between time and frequency resolution determined by the window length. A 25ms window provides sufficient frequency resolution to resolve formants (typically requiring 40-50 Hz resolution) but blurs rapid temporal events. This Heisenberg-Gabor limit is basic to Fourier analysis: you cannot simultaneously achieve arbitrarily fine time and frequency resolution.
Transients like stop consonant releases (plosives) or fricative onsets have microsecond-scale structure that 25ms frames smear, potentially obscuring necessary cues for distinguishing "p" from "t" or "s" from "sh." Alternative representations like wavelets or constant-Q turns adapt resolution to frequency. This provides better time resolution for high frequencies where transients matter most. However, mel spectrograms remain standard due to computational efficiency and proven effectiveness across diverse speech tasks. Modern architectures like Whisper have demonstrated that the 25ms standard window, despite this limitation, is enough for state-of-the-art recognition across many languages and acoustic conditions, suggesting that the network learns to compensate for frame-level smearing through its contextual attention mechanism.
Phase Information Loss
Mel spectrograms discard phase, keeping only magnitude information. For human listening, phase is relatively unimportant for speech intelligibility, a property called "phase deafness." However, for high-fidelity synthesis or specific discriminative tasks like speaker verification or music analysis, phase contains important information about temporal fine structure and glottal pulse timing.
Griffin-Lim algorithms can iteratively estimate phase from magnitude for signal reconstruction, but the process is approximate and computationally expensive, typically requiring 50-100 iterations to converge to an acceptable reconstruction. End-to-end models that learn directly from waveforms (WaveNet, EnCodec, SoundStream) or use complex spectrograms (keeping real and imaginary components) can use phase information at the cost of increased dimensionality and data requirements. For recognition tasks, discarding phase is generally the right choice; for synthesis tasks, phase becomes necessary.
Fixed Mel Scale
The mel scale approximates average human hearing for typical listeners. Individual variation in cochlear anatomy, hearing impairment, or age-related hearing loss changes frequency perception, potentially making the standard mel scale suboptimal for specific populations. The mel scale also optimizes for perceptual quality rather than phonetic discrimination. Some phonetic contrasts occur in frequency regions where the mel scale provides limited resolution.
Learned filterbanks, where neural networks optimize frequency analysis jointly with downstream tasks, sometimes discover non-mel frequency allocations that improve recognition accuracy. The SincNet architecture replaces hand-designed mel filters with learnable sinc functions initialized to approximate mel filters but allowed to adapt during training. Experimental results suggest that learned filters often converge to mel-like distributions but with subtle differences in bandwidth and spacing that improve task-specific performance.
Environmental Robustness
The mel spectrogram is sensitive to additive noise, reverberation (multipath reflections), and channel distortion (microphone frequency response). While log compression provides some robustness to multiplicative noise (which in log space becomes an additive offset removed by CMVN), additive noise raises the floor of the spectrogram, obscuring weak speech components. Normalization helps, but convolutional noise (channel effects) and non-stationary noise (overlapping speakers) distort the spectral patterns in ways that simple mean subtraction cannot fix.
Robust ASR systems often front mel spectrograms with neural enhancement networks trained specifically to suppress noise before feature extraction. Alternatively, multi-condition training exposes the model to diverse acoustic environments (different noise types, room impulse responses, microphone characteristics) during training, learning noise-invariant representations through data augmentation rather than explicit front-end processing. SpecAugment, a widely used augmentation technique, randomly masks frequency bands and time steps in the spectrogram during training, forcing the model to be reliable to missing or corrupted features and giving significant improvements in noisy conditions.
Alternatives to Mel Spectrograms
Raw waveform models process audio samples directly using learned convolutional frontends, eliminating hand-designed features entirely. Models like WaveNet and modern conformer variants demonstrate that neural networks can learn mel-like representations from data, potentially discovering features better suited to the specific task than human-designed mel scales. Self-supervised models like wav2vec 2.0 and HuBERT learn representations directly from waveforms without any hand-designed features, achieving state-of-the-art performance with limited labeled data.
However, mel spectrograms remain prevalent because they provide an informative, compressed representation that reduces computational requirements, speeds convergence, and provides an interpretable intermediate representation for debugging and analysis. For most applications, the inductive bias of the mel scale is beneficial rather than restrictive. This provides a strong prior that guides learning efficiently. Whisper's success using standard 80-bin log-mel spectrograms on massively diverse multilingual data confirms that the mel representation, properly combined with a large encoder, is enough for near-human recognition performance.
As we move toward the speech recognition architectures in upcoming chapters, particularly the Whisper architecture we will explore next, these mel spectrograms act as the basic input representation. The encoder turns these time-frequency patterns into latent representations, mapping acoustic similarity to linguistic meaning. The quality of this initial transformation fundamentally constrains the performance of the downstream recognition system. This makes careful preprocessing design needed for successful speech applications.
Summary
Speech representations bridge the continuous acoustic world and discrete neural processing. The journey from air pressure waves to model input involves several key transformations that progressively extract and compress phonetically relevant information:
- Sampling converts analog signals to digital at rates satisfying the Nyquist criterion (typically 16 kHz for speech), preserving all frequency content up to 8 kHz without aliasing.
- Pre-emphasis compensates for spectral tilt and boosts high-frequency content. This keeps consonant information is not overshadowed by dominant low-frequency vowel energy.
- Framing and windowing segment the signal into quasi-stationary chunks (25ms frames, 10ms shift) tapered to prevent spectral leakage, respecting the time-varying nature of speech while letting spectral analysis.
- STFT turns time-domain frames into frequency-domain representations, revealing the spectral envelope that carries phonetic identity.
- Mel filterbanks aggregate linear frequency bins into perceptually spaced channels (typically 80 filters) following the logarithmic sensitivity of human hearing, allocating computational resources where they matter most for perception.
- Log compression matches the non-linear perception of loudness and stabilizes numerical ranges, preventing high-energy regions from dominating the learning process.
- Normalization (CMVN) removes channel and speaker variability, standardizing features to zero mean and unit variance so the model focuses on linguistic content rather than recording conditions.
The resulting mel spectrogram compresses a 16,000-dimensional second of audio into a matrix (for 10ms frames and 80 mel bins), preserving phonetically relevant information while letting efficient processing by transformer encoders. This represents a 50-fold reduction in data rate from raw samples, yet retains the needed cues for phonetic discrimination. The representation balances the temporal resolution needed for consonant discrimination with the frequency resolution needed for vowel identification, all within a perceptual framework that aligns with human auditory processing.
Understanding these preprocessing steps is needed for debugging speech recognition systems (recognizing when normalization fails or when frame sizes are inappropriate), designing augmentation strategies (knowing which spectral regions to mask or shift), and adapting models to new acoustic environments (adjusting for different noise profiles or channel characteristics). The choices made in feature extraction, window sizes, and normalization strategies fundamentally constrain what the neural network can learn, making this pipeline a necessary foundation for all subsequent speech processing.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about speech representations and preprocessing pipelines.
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!