All papers
Speech codingFull-paper digest

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

Laurin Wagner

Code & resourcesgithub.com/nyrahealth/PINT

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.0 KB · Ready to paste

Preview copied content

TL;DR — PINT (Parallel INvariant Tokenization) fine-tunes a HuBERT encoder using parallel-utterance alignment losses and aggressive augmentations to eliminate non-linguistic nuisance variation, achieving a 98.7% relative reduction in speaker probe accuracy, a 42% lower ABX error rate, and a 27–30% lower language model perplexity versus baselines.

Key contributions

  • Formalizes semantic speech tokenization through a dual criterion: sufficient phonetic content capture and strict nuisance invariance to collapse conditional entropy.
  • Proposes the PINT training framework which combines sequence-level soft-DTW, word-level contrastive loss, and autoregressive phoneme distillation across parallel data.
  • Introduces a discrete Stage B optimization using student-teacher CTC alignment to guarantee sequence-level consistency in discrete code spaces.
  • Demonstrates massive down-stream gains: 98.7% lower speaker probe accuracy, 42% lower across-speaker ABX error, and 27-30% lower LM test perplexity.

Problem

Discrete speech tokens derived from standard self-supervised learning (SSL) models like HuBERT and WavLM suffer from a severe nuisance-leakage problem, where speaker identity, prosody, and channel conditions remain heavily recoverable. This variance causes token sequences to fluctuate drastically for the same underlying phonetic content, inflating conditional entropy and hurting compressibility and autoregressive predictability. Prior semantic tokenizers and neural codecs inherit this underlying feature brittleness rather than fixing it at the encoder root, forcing downstream generative and editing models to waste capacity undoing entanglement.

Method

PINT builds upon a HuBERT-base encoder and processes audio in two stages. Stage A enforces continuous invariance and content capture via a multi-task loss combining sequence-level soft Dynamic Time Warping (sDTW, λsdtw=0.5\lambda_{\text{sdtw}}=0.5) over parallel utterance pairs, a word-level contrastive loss (λword=2\lambda_{\text{word}}=2, duration-weighted cosine similarity with Jaccard-filtered negative pairs), and a teacher-forced phoneme cross-entropy loss (λce=10\lambda_{\text{ce}}=10) from a two-layer Transformer decoder.

Stage B transitions to discrete representations by attaching a linear projection layer mapping frames to a vocabulary of K=200K=200 codes plus a blank symbol. Using a student-teacher EMA setup, an anchor utterance provides argmax token IDs that are deduplicated to form a stable reference sequence for each parallel group. The student is trained via Connectionist Temporal Classification (CTC) alongside a marginal entropy regularization loss (LmargL_{\text{marg}}) and an orthogonality loss (LorthL_{\text{orth}}) to prevent codebook collapse and encourage feature separability, weighted at λid=0.6\lambda_{\text{id}}=0.6 and λm,o=0.3\lambda_{\text{m,o}}=0.3.

The training recipe leverages diverse parallel datasets (ARCTIC, CHAINS, CSTR-VCTK, EnDialects, ESD, SynSpeech, TIMIT, and noise corpora) totaling over 500 hours of natural and synthetic parallel groups, alongside large non-parallel corpora augmented on-the-fly with additive noise, reverberation, and pitch/speed perturbations. At inference, PINT functions as a direct drop-in feature encoder or tokenizer yielding highly deterministic, deduplication-friendly sequences.

Experimental setup

Evaluations utilize LibriSpeech (train-clean-360 for ASR BLSTM training, dev-clean/test-clean for ABX and CER/WER), CSTR-VCTK (held-out speakers for speaker probes and DTW invariance), RAVDESS (emotion probes), and LibriLight (6,000h clean subset for training 85M-parameter decoder-only transformer LMs). Baselines are HuBERT-base layer 9 and WavLM-base layer 12 using 200 k-means centroids. Models are compared using CER, WER, ABX error rates, TDNN probe accuracies, sequence DTW/edit distances, noise entropy/RMS SD, and autoregressive perplexity.

Results

PINT achieves an outstanding reduction in speaker probe accuracy down to 1.2% (compared to 93.1% for HuBERT and 78.9% for WavLM) and lowers cross-speaker ABX error from 0.042 (WavLM) down to 0.040. In terms of downstream language modeling on LibriLight, an identical 85M-parameter transformer LM trained on PINT tokens reaches a test perplexity of 1.95, which is a 27–30% improvement over HuBERT (2.78) and WavLM (2.67), while converging 23× faster. Furthermore, PINT RLE compression achieves 152 bits/s (a 2.6× reduction from raw tokens), closely approaching pure text BPE at 116 bits/s.

Ablations demonstrate that omitting real parallel data (synth-only) degrades ASR/ABX performance, while dropping word-level contrastive loss or using CTC/AR-only decoders leads to severe code collapse or high speaker leakage (e.g., dec-AR retains 20.1% speaker accuracy).

SystemCER ct/disWER ct/disABX acr.Spk Probe (%)LM Perplexity
HuBERT-base4.33 / 7.5510.99 / 21.370.06693.12.78
WavLM-base4.03 / 6.4011.53 / 18.420.05978.92.67
PINT (ours)3.84 / 4.659.79 / 12.130.0421.21.95

Limitations

While PINT successfully removes speaker and noise variance, some emotional variation still leaks through (RAVDESS emotion probe accuracy is 32.1%). The approach relies heavily on parallel data availability, and although synthetic data augmentation via text-to-speech engine Kokoro alleviates this, pure synthetic training degrades absolute acoustic-phonetic performance. The evaluation is currently restricted to English speech and requires forced alignments during training preprocessing.

Why read this

Speech and ML engineers building autoregressive speech generators, audio codecs, or disentangled voice conversion models should read this paper to learn how upstream encoder invariance can eliminate nuisance leakage, drastically boost sequence compressibility, and lower language model perplexity.

Code

Applications

Neural audio codecs, autoregressive speech generation, expressive text-to-speech, voice and accent conversion, and speech language models.

Institutions

nyra health

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-2817