TL;DR — This paper investigates whether gender differences in /s/ sibilant acoustics are driven by physical vocal tract length (VTL) or social performance, using causal mediation analysis on a new multilingual database of 12 languages and 1,386 speakers. The results show that while VTL reliably predicts /s/ peak frequency across nearly all languages, the relative contributions of anatomy and social performance vary drastically—revealing hidden social effects or cancelling effects in specific languages like Mandarin.
Key contributions
- Assembled a new multilingual phonetic database combining GlobalPhone, NCHLT, and LibriSpeech, spanning 12 languages, 1,386 speakers, 3.8 million vowel tokens, and 128,949 /s/ tokens.
- Applied causal mediation analysis to phonetics to mathematically disentangle the direct effect of gender (social performance) from the indirect effect mediated by vocal tract length (anatomy).
- Demonstrated that estimated VTL () predicts /s/ peak frequency consistently across 11 out of 12 languages (with Arabic as the sole exception).
- Uncovered hidden social effects in languages like Mandarin and Arabic, where direct (social) and indirect (physiological) effects run in opposite directions and cancel out total effects.
Problem
Sibilant fricatives like /s/ often exhibit acoustic differences between male and female speakers, but it is historically difficult to determine from acoustic data alone whether this variation stems from physiological differences (such as vocal tract length and front cavity size) or from the active performance of social identity. Prior approaches either ignored VTL or assumed a direct mapping, missing cases where speakers override anatomy for social reasons (e.g., working-class Glaswegian girls sounding more masculine) or where social performance masks physiological realities. Without proper causal modeling, researchers risk misattributing anatomical differences to social performance or overlooking true sociophonetic variation.
Method
The study analyzes read speech sampled at 16 kHz from a curated set of ASR corpora. Vowel formants (F1–F3) were extracted at 33% duration for all vowels longer than 50 ms using PolyglotDB and a basic LPC refinement procedure with acoustic prototypes mapped from Buckeye and Quebec French. Speaker-level vocal tract length proxy (average formant spacing) was computed using the inverse relationship. For sibilants, multitaper spectra (8 tapers, time-bandwidth 4) were computed over a 20 ms window at the midpoint of word-initial pre-vocalic /s/ tokens longer than 50 ms, with amplitude normalized against local silence and speech. The highest spectral peak above 1 kHz was measured as a proxy for front cavity length.
To separate physiological and social pathways, the authors implemented causal mediation analysis using linear mixed-effects regression models via lme4 and the mediation package in R. The full outcome model regressed /s/ peak frequency on , GENDER (binary-coded and mean-centered), and their interaction, with maximal by-language random effects. The full mediator model regressed on GENDER. The direct effect of gender is captured by the GENDER coefficient in the outcome model (controlling for VTL), while the indirect effect is the product of the GENDER coefficient in the mediator model and the coefficient in the outcome model. By-language models without random effects were also fitted individually to evaluate cross-linguistic heterogeneity.
Experimental setup
The final dataset contains 1,386 speakers across 12 languages (Czech, Polish, Russian, Afrikaans, English, Swedish, French, Spanish, Arabic, Japanese, Korean, Mandarin), encompassing 3,806,272 vowel tokens and 128,949 /s/ tokens. Force alignment was performed using the Montreal Forced Aligner (MFA) with custom acoustic models and pronunciation dictionaries (supplemented by Epitran for Japanese and G2P for Arabic). Statistical significance for mediation pathways was assessed using quasi-Bayesian confidence intervals.
Results
The full outcome model yielded a conditional of 0.444 and a marginal of 0.247. There is a strong, statistically significant effect of ( Hz, ), where shorter VTL () predicts higher /s/ peak frequency. The direct effect of GENDER is smaller and marginal ( Hz, ), while the indirect effect mediated by is highly significant ( Hz, ). Across individual languages, Afrikaans, Czech, and English show both significant direct and indirect effects. Mandarin exhibits significant direct and indirect effects of similar magnitude but opposite signs, which cancel out to yield a non-significant total effect. Arabic is the only language showing no significant relationship between and peak frequency.
| System / Language Condition | Direct Effect (Social) | Indirect Effect (VTL-Mediated) | Total Effect |
|---|---|---|---|
| Full Model (Average) | Hz () | Hz () | Hz () |
| English | Significant | Significant | Significant |
| Mandarin | Significant (Positive) | Significant (Negative) | Non-significant (Cancel out) |
| Japanese | Non-significant | Non-significant | Significant |
| Arabic | Non-significant | Non-significant | Borderline significant |
Limitations
The dataset relies exclusively on read speech, which is typically less conducive to social identity performance than spontaneous conversational speech. Gender and sex are conflated in the metadata, meaning sex-based morphological differences in vocal tract shape rather than pure social performance could bleed into the direct effect. The 16 kHz sampling rate limits spectral resolution above 8 kHz, potentially capping peak measurements for female speakers with extremely short vocal tracts. Additionally, using overall peak frequency is a noisy index for front cavity resonance compared to narrow-band targeted searches.
Why read this
Phoneticians, sociolinguists, and speech researchers should read this paper to understand how to rigorously separate anatomical traits from social performance in acoustic data without requiring physical body measurements. It provides a blueprint for applying causal mediation analysis across multilingual corpora, demonstrating that standard acoustic gender differences cannot be taken at face value.
Code
Applications
Improving sociolinguistic speech analysis, building more robust and culturally aware speaker normalization algorithms, and designing unbiased speech recognition systems that distinguish physiological anatomy from stylistic speech variation.
Institutions
McGill University
Funding / 經費: Social Sciences and Humanities Research Council, Fonds de recherche du Quebec - Societe et culture, Canada Research Chairs, Natural Sciences and Engineering Research Council of Canada
Related
- Vocal Tract Disparity and Potential Implications for Speaker Recognition — same problem · relatedness 2.0/3
- A Corpus-Based Study of Creaky Voice Production in English and Mandarin — same problem · relatedness 1.8/3
- Tongue-Shape Strategies for Standard Mandarin Retroflex Sibilants: A Preliminary Ultrasound and Unsupervised Clustering Study — same problem · relatedness 1.8/3
- Towards Language-Agnostic Speech Inversion — complementary · relatedness 1.8/3
- Speaker or Language? Explaining Variance in Charismatic Prosody Across Luxembourgish and French — relatedness 1.7/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-2975