All papers
Speech recognitionFull-paper digest

ArtNet: A JEPA-Like Articulatory Predictive Framework for Robust Zero-Shot Phoneme Recognition

Zeqian Hu, Fuliang Weng, Shu Shang, Yaqian Zhou

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.2 KB · Ready to paste

Preview copied content

TL;DR — ArtNet introduces a JEPA-like articulatory predictive framework with a variational information bottleneck to suppress language-specific variations, achieving a 20.56% relative phoneme error rate reduction in zero-shot cross-lingual transfer.

Key contributions

  • Formulates zero-shot cross-lingual phoneme recognition as a non-generative, structured articulatory prediction task inspired by JEPA.
  • Integrates a lightweight articulatory predictor with a variational information bottleneck (VIB) to strip language-specific variations from SSL representations.
  • Proposes vector-space inventory alignment (VSIA) as a zero-shot inference strategy using angular proximity in a continuous articulatory space.
  • Demonstrates mitigation of the substitution error bottleneck in zero-shot transfer, improving both in-vocabulary and out-of-vocabulary generalization.

Problem

Direct acoustic-to-symbol mapping in standard SSL-based phoneme recognizers is brittle because models overfit to source-language acoustic fluctuations and prosodic patterns. Cross-lingual performance collapses primarily due to massive substitution errors—contrarily affecting both in-vocabulary (61.6%) and out-of-vocabulary phonemes—when transferring to unseen languages. Existing auxiliary linguistic methods fail to enforce intrinsic frontend robustness or adequately disentangle universal phonetic properties from language-specific traits.

Method

The framework leverages mHuBERT-147 (95M parameters, 12 layers, 768 hidden dimensions) as an SSL context encoder, initialized via a CTC objective and frozen during ArtNet training. Orthographic transcripts are converted to IPA using Epitran and mapped via the Panphon database into a static 24-dimensional trinary articulatory matrix, transformed into numeric values {-1, 0, 1}. Frame-level pseudo-labels generate ground-truth articulatory vectors.

To eliminate noise, hidden states are fed into a VIB encoder predicting multivariate Gaussian parameters (mu and sigma) for a 128-dimensional stochastic latent variable z_t. An articulatory predictor (AP) then minimizes mean squared error (MSE) against the ground-truth articulatory vector, regularized by KL divergence against a standard normal prior with trade-off parameter beta = 0.001.

During inference, a segment pooling strategy groups consecutive non-blank frames, and vector-space inventory alignment (VSIA) performs a nearest-neighbor angular proximity search (cosine similarity) against the target language phoneme inventory to output final symbols.

Experimental setup

The model is trained solely on the 100-hour LibriSpeech train-clean-100 corpus (English) and evaluated zero-shot on the test sets of 7 Multilingual LibriSpeech languages (German, Dutch, French, Spanish, Italian, Portuguese, Polish). Evaluations use Phoneme Error Rate (PER) and Phoneme Feature Error Rate (PFER). Optimization uses the Adam optimizer with a learning rate warming up to 1e-3, utilizing LoRA for the initial phoneme recognizer and training TDNN, MLP, or LSTM configurations for the AP and VIB modules.

Results

The complete ArtNet system paired with VSIA achieves an average PER of 45.54% across the 7 unseen languages, representing a 20.56% relative reduction over the baseline (57.33%), alongside a 7.01% relative improvement in PFER (down to 12.73% from 13.69%). In Spanish, absolute PER drops by approximately 28.26 percentage points when combining baseline mapping with VSIA/ArtNet. Ablations of the AP backbone show that a local context Time Delay Neural Network (TDNN) achieves the lowest average PER (54.94% unaligned) compared to a global context LSTM (56.51%) and context-free MLP (55.33%), proving that global source prosody introduces harmful language-specific biases.

SystemDutchFrenchGermanItalianPolishPortugueseSpanishAvg PER
Baseline59.6759.3352.6354.4355.0861.3858.7657.33
ArtNet58.6456.8451.6351.7350.2459.9455.5754.94
Baseline+tr2tgt56.1256.5450.4846.0741.0158.0234.2748.93
ArtNet+VSIA55.4053.7550.0439.9635.1853.9330.5045.54

Limitations

The evaluation is restricted to Indo-European languages within the Multilingual LibriSpeech dataset, leaving tonal languages and non-alphabetic writing systems unexplored. The dependency on Epitran and Panphon limits language coverage to those with well-mapped IPA phonological definitions. Furthermore, the framework relies heavily on a pre-existing source language pairing and does not investigate fully unsupervised phoneme discovery without a G2P tool.

Why read this

Speech researchers and engineers working on low-resource or zero-shot ASR should read this to understand how JOINT-Embedding Predictive Architecture (JEPA) principles and articulatory bottleneck representations can replace fragile direct acoustic-to-symbol decoders.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Low-resource zero-shot automatic speech recognition and cross-lingual phonetic transfer for unwritten or under-resourced languages.

Institutions

Fudan University, Logos & Dialogos

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-304