---
id: kirby26_interspeech
title: Perceptual compensation for tonal context in self-supervised speech models
authors:
  - James Kirby
  - Ioana Krehan
  - Michele Gubian
year: 2026
doi: 10.21437/Interspeech.2026-2409
isca_url: https://www.isca-archive.org/interspeech_2026/kirby26_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/kirby26_interspeech.pdf
session: Tones
topics:
  - self-supervised
  - evaluation
  - phonetics
category: phonetics-linguistics
labels:
  - self-supervised
institutions:
  - LMU Munich
code:
  url: https://github.com/kehanlu/mandarin-wav2vec2
  stars: 44
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: kirby26_interspeech
  category: phonetics-linguistics
  labels:
    - self-supervised
  institutions:
    - LMU Munich
  code: https://github.com/kehanlu/mandarin-wav2vec2
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-2409
  pdf: https://www.isca-archive.org/interspeech_2026/kirby26_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/kirby26_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/kirby26_interspeech/markdown.md
---

# Perceptual compensation for tonal context in self-supervised speech models

*James Kirby, Ioana Krehan, Michele Gubian*

[PDF](https://www.isca-archive.org/interspeech_2026/kirby26_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/kirby26_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-2409)

**Category:** `phonetics-linguistics` · **Labels:** `self-supervised`

**TL;DR** — This study evaluates whether wav2vec2.0 models exhibit human-like perceptual compensation for tonal context in Mandarin Chinese, finding no evidence of compensation in purely self-supervised representations and only weak, non-human-like shifts in fine-tuned models.

## Key contributions

- Conducts a computational pseudo-replication of a classic Mandarin Chinese psycholinguistic tone perception experiment using ~13,700 resynthesized tonal continua (~192,000 stimuli).
- Compares internal representations between purely self-supervised pre-trained (PT) and Mandarin ASR fine-tuned (FT) wav2vec2.0 models using both layer-wise embedding similarities and linear probing classifiers.
- Demonstrates that purely pre-trained wav2vec2.0 embeddings completely lack sensitivity to preceding tonal context.
- Shows that while supervised fine-tuning and linear probes introduce some context sensitivity, they fail to replicate human perceptual baselines on isolated test syllables, exhibiting strong frequency-like biases.

## Problem

Prior work has claimed that self-supervised learning (SSL) models implicitly acquire phonological structure and context-dependent phonetic adaptation (perceptual compensation) without explicit symbolic supervision, mirroring human behavior. However, these claims have primarily focused on segmental contrasts (like English [r/l]). This paper investigates whether such purely unsupervised contextual encoding extends to suprasegmental features like lexical tone in Mandarin Chinese, where acoustic realizations of fundamental frequency are jointly influenced by numerous extraneous physiological and prosodic factors.

## Method

The study analyzes two wav2vec2.0 model checkpoints: a pre-trained model (1,000 hours of untranscribed Mandarin) and a fine-tuned Mandarin ASR model (178 hours of transcribed speech), both featuring 7 CNN feature extractor layers and 12 Transformer layers. Stimuli were generated by extracting disyllables from the AISHELL-3 test split and applying Parselmouth duration and F0 manipulations to create 14-step T4-T3 target continua preceded by context tones (T1, T2, T4) or presented in isolation. 

Two analysis methods were used across layer 0 (CNN output) and layers 1-12 (Transformer outputs): embedding similarities and probing classifiers. Embedding similarities calculated the relative distance of stimulus vectors to step-1 (T4) and step-14 (T3) reference endpoints, modeled via generalized additive mixed models (GAMMs) with beta-distributed response variables. Probing classifiers consisted of binary logistic regression neural networks trained with cross-entropy loss and the Adam optimizer (lr 1e-3) on 100 T3 and 100 T4 utterances from 36 speakers for 5 epochs, then tested on the manipulated continuum stimuli.

## Experimental setup

Evaluated on ~13,700 Mandarin tonal continua (~192,000 stimuli) derived from the AISHELL-3 corpus test split, with probes trained on data from 36 speakers and validated on 4 speakers. Compared purely pre-trained (PT) wav2vec2.0 against ASR fine-tuned (FT) wav2vec2.0. Metrics include cosine embedding similarities modeled via GAMMs and probing classification accuracy/responses modeled via Bernoulli GAMMs.

## Results

Embedding similarities for the purely pre-trained model showed zero evidence of contextual compensation across any layer, contrasting with previous claims about segmental SSL representations. The fine-tuned model's embeddings exhibited slight context sensitivity where T1 differed from T2/T4, but these shifts were small, unanchored to no-context baselines, and qualitatively distinct from human psycholinguistic patterns. Linear probes on fine-tuned embeddings (e.g., layer 8) showed qualitative sensitivity to context and continuum steps, but completely failed to recover the sigmoidal response curve for no-context isolated syllables, suffering from an extreme T4 response bias.

| System / Condition | Layer 0 (CNN) Accuracy | Layer 4 Accuracy | Layer 8 Accuracy | Layer 12 Accuracy |
|---|---|---|---|---|
| PT Probes (Validation) | ~82% | ~90% | ~95% | ~99% |
| FT Probes (Validation) | ~82% | ~90% | ~95% | ~99% |

## Limitations

The investigation is restricted to a single architecture (wav2vec2.0) and a single suprasegmental domain (Mandarin lexical tones). Probes tested on isolated syllables suffered from severe boundary and frequency biases (e.g., favoring T4), potentially exacerbated by mismatch between utterance-level training contexts and single-syllable test evaluations.

## Why read this

Speech and ML researchers studying SSL model interpretability should read this to understand the limitations of unsupervised pre-training in capturing complex suprasegmental phonological phenomena, challenging the assumption that human-like phonological abstraction emerges automatically from self-supervision alone.

## Code

- https://github.com/kehanlu/mandarin-wav2vec2

## Applications

Improving the phonological fidelity and perceptual alignment of speech representation models, speech recognition, and diagnostic evaluation of self-supervised speech architectures.

## Institutions / 機構

LMU Munich

## Related

- [Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations](sun26_interspeech.md) — same problem · relatedness 2.2/3
- [Using Phonological-Level Wav2Vec2 for Mandarin Automatic Mispronunciation Detection and Diagnosis](chen26g_interspeech.md) — shared technique · relatedness 1.9/3
- [Probing the Layer-wise Geometry of Chinese Dialect Representations in Wav2Vec 2.0](peng26c_interspeech.md) — shared technique · relatedness 1.9/3
- [DEBATE: A Dataset for Disentangling Textual Ambiguity in Mandarin Through Speech](guo26e_interspeech.md) — same problem · relatedness 1.9/3
- [English Vowel Perceptual Training under Multitalker Babble: A Comparison of Humans and Large Language Models](dong26b_interspeech.md) — shared technique · relatedness 1.9/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
