All papers
Phonetics & linguisticsFull-paper digest

Perceptual compensation for tonal context in self-supervised speech models

James Kirby, Ioana Krehan, Michele Gubian

Code & resourcesgithub.com/kehanlu/mandarin-wav2vec2

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.0 KB · Ready to paste

Preview copied content

TL;DR — This study evaluates whether wav2vec2.0 models exhibit human-like perceptual compensation for tonal context in Mandarin Chinese, finding no evidence of compensation in purely self-supervised representations and only weak, non-human-like shifts in fine-tuned models.

Key contributions

  • Conducts a computational pseudo-replication of a classic Mandarin Chinese psycholinguistic tone perception experiment using ~13,700 resynthesized tonal continua (~192,000 stimuli).
  • Compares internal representations between purely self-supervised pre-trained (PT) and Mandarin ASR fine-tuned (FT) wav2vec2.0 models using both layer-wise embedding similarities and linear probing classifiers.
  • Demonstrates that purely pre-trained wav2vec2.0 embeddings completely lack sensitivity to preceding tonal context.
  • Shows that while supervised fine-tuning and linear probes introduce some context sensitivity, they fail to replicate human perceptual baselines on isolated test syllables, exhibiting strong frequency-like biases.

Problem

Prior work has claimed that self-supervised learning (SSL) models implicitly acquire phonological structure and context-dependent phonetic adaptation (perceptual compensation) without explicit symbolic supervision, mirroring human behavior. However, these claims have primarily focused on segmental contrasts (like English [r/l]). This paper investigates whether such purely unsupervised contextual encoding extends to suprasegmental features like lexical tone in Mandarin Chinese, where acoustic realizations of fundamental frequency are jointly influenced by numerous extraneous physiological and prosodic factors.

Method

The study analyzes two wav2vec2.0 model checkpoints: a pre-trained model (1,000 hours of untranscribed Mandarin) and a fine-tuned Mandarin ASR model (178 hours of transcribed speech), both featuring 7 CNN feature extractor layers and 12 Transformer layers. Stimuli were generated by extracting disyllables from the AISHELL-3 test split and applying Parselmouth duration and F0 manipulations to create 14-step T4-T3 target continua preceded by context tones (T1, T2, T4) or presented in isolation.

Two analysis methods were used across layer 0 (CNN output) and layers 1-12 (Transformer outputs): embedding similarities and probing classifiers. Embedding similarities calculated the relative distance of stimulus vectors to step-1 (T4) and step-14 (T3) reference endpoints, modeled via generalized additive mixed models (GAMMs) with beta-distributed response variables. Probing classifiers consisted of binary logistic regression neural networks trained with cross-entropy loss and the Adam optimizer (lr 1e-3) on 100 T3 and 100 T4 utterances from 36 speakers for 5 epochs, then tested on the manipulated continuum stimuli.

Experimental setup

Evaluated on ~13,700 Mandarin tonal continua (~192,000 stimuli) derived from the AISHELL-3 corpus test split, with probes trained on data from 36 speakers and validated on 4 speakers. Compared purely pre-trained (PT) wav2vec2.0 against ASR fine-tuned (FT) wav2vec2.0. Metrics include cosine embedding similarities modeled via GAMMs and probing classification accuracy/responses modeled via Bernoulli GAMMs.

Results

Embedding similarities for the purely pre-trained model showed zero evidence of contextual compensation across any layer, contrasting with previous claims about segmental SSL representations. The fine-tuned model's embeddings exhibited slight context sensitivity where T1 differed from T2/T4, but these shifts were small, unanchored to no-context baselines, and qualitatively distinct from human psycholinguistic patterns. Linear probes on fine-tuned embeddings (e.g., layer 8) showed qualitative sensitivity to context and continuum steps, but completely failed to recover the sigmoidal response curve for no-context isolated syllables, suffering from an extreme T4 response bias.

System / ConditionLayer 0 (CNN) AccuracyLayer 4 AccuracyLayer 8 AccuracyLayer 12 Accuracy
PT Probes (Validation)~82%~90%~95%~99%
FT Probes (Validation)~82%~90%~95%~99%

Limitations

The investigation is restricted to a single architecture (wav2vec2.0) and a single suprasegmental domain (Mandarin lexical tones). Probes tested on isolated syllables suffered from severe boundary and frequency biases (e.g., favoring T4), potentially exacerbated by mismatch between utterance-level training contexts and single-syllable test evaluations.

Why read this

Speech and ML researchers studying SSL model interpretability should read this to understand the limitations of unsupervised pre-training in capturing complex suprasegmental phonological phenomena, challenging the assumption that human-like phonological abstraction emerges automatically from self-supervision alone.

Code

Applications

Improving the phonological fidelity and perceptual alignment of speech representation models, speech recognition, and diagnostic evaluation of self-supervised speech architectures.

Institutions

LMU Munich

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-2409