All papers
Speech codingFull-paper digest

ContextCodec: Content-Focused Context Guidance for Ultra-Low Bitrate Speech Coding

Chengbin Liang, Wenqi Guo, Hao Cao, Zhijin Qin

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

14.1 KB · Ready to paste

Preview copied content

TL;DR — ContextCodec is a content-first, ultra-low-bitrate neural speech codec that uses a dual-branch encoder with CLIP-style phoneme alignment and stage-wise context injection, achieving strong intelligibility and perceptual quality down to 500 bps.

Key contributions

  • Proposes a content-first dual-branch encoder that explicitly decouples acoustic details from linguistic context to resolve the zero-sum bit allocation problem at ultra-low bitrates.
  • Introduces a CLIP-style frame-level contrastive alignment loss against forced-aligned phonemes to maximize linguistic information while minimizing paralinguistic speaker and dialect leakage.
  • Designs a context-guided attention decoder featuring acoustic pre-conditioning, time-varying residual modulation, and stage-wise context feature injection via lightweight gates.
  • Integrates a lightweight autoregressive (AR) latent refinement module operating over interleaved phase sequences to progressively predict and remove means/scales before quantization.

Problem

As speech coding is pushed to ultra-low bitrates below 1000 bps, it becomes a zero-sum bit allocation problem where bits spent on fine acoustic details starve the core linguistic message, degrading intelligibility. Traditional neural acoustic codecs prioritize timbre and waveform fidelity, while hybrid codecs incorporating self-supervised learning (SSL) features often suffer from paralinguistic leakage (speaker identity, emotion) and let semantic priors attenuate across successive decoding stages. This creates an urgent need for a codec that explicitly prioritizes and actively reinforces linguistic content throughout the decoding process.

Method

ContextCodec builds upon a GAN-based quantized autoencoder framework using finite scalar quantization (FSQ) and DAC-style base blocks, mapping an input waveform to shared features that are split into an acoustic stream (da = 512) and a context stream (dc = 512) via a dual-branch encoder. The context head utilizes 2 Transformer layers with 4 attention heads. Prior to decoding, the concatenated latents y are partitioned into P = 4 interleaved phase sequences and refined using a reusable predictor module gp that estimates per-time means and scales from already restored phases using depthwise-separable 1D convolutions with Snake activations, forming normalized residuals that are independently quantized.

The CLIP-style phoneme alignment loss uses Montreal Forced Aligner (MFA) frame-level phoneme IDs mapped via an embedding table, computing a masked symmetric InfoNCE contrastive loss with temperature tau = 0.07 between l2-normalized quantized context features and phoneme embeddings.

The context-guided attention decoder takes the restored latents, splits them into acoustic and context streams, and passes the acoustic features through a context feature enhancer (global channel-wise gate and local time-varying residual pathway). At each upsampling stage (using transposed convolutions and residual 1D conv blocks with hop size h = 640 for a 16 kHz model), the context stream is linearly interpolated to match temporal resolution, projected via pointwise convolution, and applied via sigmoid gating and residual fusion to actively guide waveform reconstruction before a final tanh output layer.

Experimental setup

Trained on LibriTTS and AISHELL-3 datasets with audio segmented into 3-second clips at 16 kHz sample rate. Evaluated on 6,000 VCTK utterances and 300 utterances per language across 10 Common Voice 21.0 languages, comparing against baselines like EnCodec, DAC, SNAC, Secousticodec, SemantiCodec, FACodec, SpeechTokenizer, X-Codec, and Mimi. Metrics include PESQ, STOI, SI-SDR, Word Error Rate (WER) using Whisper-Turbo, and subjective pairwise preference listening tests. Implemented with AdamW optimizer (lr 2e-4, betas 0.8, 0.99), trained for 1M steps on a single NVIDIA RTX 4090 GPU with batch size 8, using loss weights lambda_m=15, lambda_adv=1, lambda_fm=2, lambda_clip=3.

Results

At 1000 bps on the multilingual set, ContextCodec achieves a PESQ of 2.140, STOI of 0.866, SI-SDR of 2.110 dB, and WER of 28.31%, outperforming Mimi (PESQ 2.028, STOI 0.852, SI-SDR 1.614 dB, WER 33.60%) and X-Codec. On VCTK at 1000 bps, it yields a PESQ of 2.476, STOI of 0.880, SI-SDR of 3.614 dB, and a low WER of 2.25%. Pushed down to 500 bps, ContextCodec achieves a VCTK PESQ of 2.120 and STOI of 0.846, outperforming Mimi at 550 bps (PESQ 1.685) and SemantiCodec at 625 bps (PESQ 1.910). Subjectively, ContextCodec at 500 bps is preferred over Opus 6K (97.92% to 0.0%) and SemantiCodec (52.92% to 40.83%).

Ablations demonstrate that replacing CLIP phoneme alignment with Wav2Vec 2.0 SSL distillation increases WER from 5.56% to 7.91%, removing stage-wise context injection degrades WER to 8.20%, and eliminating AR latent refinement drops PESQ from 2.048 down to 1.887. Attribute probing on TIMIT shows ContextCodec's Phoneme-CLIP objective achieves higher phone accuracy (88.7% vs 70.2%) and lower speaker leakage (51.8% vs 91.0%) compared to SSL distillation.

ModelTypeBitrate (bps)Multilingual PESQ_↑Multilingual STOI_↑Multilingual WER_↓VCTK PESQ_↑VCTK STOI_↑VCTK WER_↓
Mimi [22]Hybrid11002.0280.85233.60%2.2560.8404.57%
X-Codec [18]Hybrid10001.8460.81231.06%2.2570.8223.29%
ContextCodec (ours)Hybrid10002.1400.86628.31%2.4760.8802.25%
Mimi [22]Hybrid5501.5530.79060.35%1.6850.77210.22%
SemantiCodec [29]Hybrid6251.6600.78852.69%1.9100.82011.42%
ContextCodec (ours)Hybrid5001.7580.81252.11%2.1200.8465.85%

Limitations

Phonological inventories differ across languages, meaning rare or unseen phonemes in the training data may be underrepresented and suffer in cross-lingual generalization. The current evaluation focuses on batch processing, and a fully streaming, stateful implementation requires further engineering validation to minimize algorithmic delay.

Why read this

Researchers and engineers working on speech coding, generative audio tokenization, or low-bandwidth communication should read this paper to see how explicit linguistic supervision via CLIP-style alignment and stage-wise context injection can overcome the intelligibility bottlenecks of ultra-low-bitrate neural codecs.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Satellite communications, bandwidth-constrained IoT voice links, emergency radio systems, and tokenization backbones for ultra-low-bitrate speech language models.

Institutions

Tsinghua University

Funding / 經費: National Key Research and Development Program of China, National Natural Science Foundation of China, Beijing Natural Science Foundation

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-3355