TL;DR — This study evaluates whether wav2vec2.0 models exhibit human-like perceptual compensation for tonal context in Mandarin Chinese, finding no evidence of compensation in purely self-supervised representations and only weak, non-human-like shifts in fine-tuned models.
Key contributions
- Conducts a computational pseudo-replication of a classic Mandarin Chinese psycholinguistic tone perception experiment using ~13,700 resynthesized tonal continua (~192,000 stimuli).
- Compares internal representations between purely self-supervised pre-trained (PT) and Mandarin ASR fine-tuned (FT) wav2vec2.0 models using both layer-wise embedding similarities and linear probing classifiers.
- Demonstrates that purely pre-trained wav2vec2.0 embeddings completely lack sensitivity to preceding tonal context.
- Shows that while supervised fine-tuning and linear probes introduce some context sensitivity, they fail to replicate human perceptual baselines on isolated test syllables, exhibiting strong frequency-like biases.
Problem
Prior work has claimed that self-supervised learning (SSL) models implicitly acquire phonological structure and context-dependent phonetic adaptation (perceptual compensation) without explicit symbolic supervision, mirroring human behavior. However, these claims have primarily focused on segmental contrasts (like English [r/l]). This paper investigates whether such purely unsupervised contextual encoding extends to suprasegmental features like lexical tone in Mandarin Chinese, where acoustic realizations of fundamental frequency are jointly influenced by numerous extraneous physiological and prosodic factors.
Method
The study analyzes two wav2vec2.0 model checkpoints: a pre-trained model (1,000 hours of untranscribed Mandarin) and a fine-tuned Mandarin ASR model (178 hours of transcribed speech), both featuring 7 CNN feature extractor layers and 12 Transformer layers. Stimuli were generated by extracting disyllables from the AISHELL-3 test split and applying Parselmouth duration and F0 manipulations to create 14-step T4-T3 target continua preceded by context tones (T1, T2, T4) or presented in isolation.
Two analysis methods were used across layer 0 (CNN output) and layers 1-12 (Transformer outputs): embedding similarities and probing classifiers. Embedding similarities calculated the relative distance of stimulus vectors to step-1 (T4) and step-14 (T3) reference endpoints, modeled via generalized additive mixed models (GAMMs) with beta-distributed response variables. Probing classifiers consisted of binary logistic regression neural networks trained with cross-entropy loss and the Adam optimizer (lr 1e-3) on 100 T3 and 100 T4 utterances from 36 speakers for 5 epochs, then tested on the manipulated continuum stimuli.
Experimental setup
Evaluated on ~13,700 Mandarin tonal continua (~192,000 stimuli) derived from the AISHELL-3 corpus test split, with probes trained on data from 36 speakers and validated on 4 speakers. Compared purely pre-trained (PT) wav2vec2.0 against ASR fine-tuned (FT) wav2vec2.0. Metrics include cosine embedding similarities modeled via GAMMs and probing classification accuracy/responses modeled via Bernoulli GAMMs.
Results
Embedding similarities for the purely pre-trained model showed zero evidence of contextual compensation across any layer, contrasting with previous claims about segmental SSL representations. The fine-tuned model's embeddings exhibited slight context sensitivity where T1 differed from T2/T4, but these shifts were small, unanchored to no-context baselines, and qualitatively distinct from human psycholinguistic patterns. Linear probes on fine-tuned embeddings (e.g., layer 8) showed qualitative sensitivity to context and continuum steps, but completely failed to recover the sigmoidal response curve for no-context isolated syllables, suffering from an extreme T4 response bias.
| System / Condition | Layer 0 (CNN) Accuracy | Layer 4 Accuracy | Layer 8 Accuracy | Layer 12 Accuracy |
|---|---|---|---|---|
| PT Probes (Validation) | ~82% | ~90% | ~95% | ~99% |
| FT Probes (Validation) | ~82% | ~90% | ~95% | ~99% |
Limitations
The investigation is restricted to a single architecture (wav2vec2.0) and a single suprasegmental domain (Mandarin lexical tones). Probes tested on isolated syllables suffered from severe boundary and frequency biases (e.g., favoring T4), potentially exacerbated by mismatch between utterance-level training contexts and single-syllable test evaluations.
Why read this
Speech and ML researchers studying SSL model interpretability should read this to understand the limitations of unsupervised pre-training in capturing complex suprasegmental phonological phenomena, challenging the assumption that human-like phonological abstraction emerges automatically from self-supervision alone.
Code
Applications
Improving the phonological fidelity and perceptual alignment of speech representation models, speech recognition, and diagnostic evaluation of self-supervised speech architectures.
Institutions
LMU Munich
Related
- Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations — same problem · relatedness 2.2/3
- Using Phonological-Level Wav2Vec2 for Mandarin Automatic Mispronunciation Detection and Diagnosis — shared technique · relatedness 1.9/3
- Probing the Layer-wise Geometry of Chinese Dialect Representations in Wav2Vec 2.0 — shared technique · relatedness 1.9/3
- DEBATE: A Dataset for Disentangling Textual Ambiguity in Mandarin Through Speech — same problem · relatedness 1.9/3
- English Vowel Perceptual Training under Multitalker Babble: A Comparison of Humans and Large Language Models — shared technique · relatedness 1.9/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-2409