TL;DR — This paper evaluates self-supervised learning (SSL) feature vectors combined with a two-stage Pre-net/Post-net classifier for multilingual and cross-lingual lexical stress detection in Arabic and English, achieving near-monolingual accuracy (~97% English, ~90% Arabic) under joint training. It demonstrates that cross-lingual transfer is asymmetric and benefits significantly from multilingual base representations (XLS-R) and word-level temporal modeling.
Key contributions
- Evaluates lexical stress detection in typologically distinct languages (Modern Standard Arabic's quantity-sensitive weight system vs. English's contrastive/morphological lexical stress) using modern SSL features.
- Proposes a two-stage classification framework (Pre-net DNN for syllable-level features and Post-net TDNN for word-context dependencies) with weight initialization transfer.
- Demonstrates that multilingual joint training retains near-monolingual accuracy across both languages while handling class imbalances (3:7 stressed/unstressed for Arabic, 5:6 for English).
- Analyzes cross-lingual transfer asymmetry, showing that English-to-Arabic transfer is harder (68.12% average) than Arabic-to-English (76.09% average), and that multilingual encoders like XLS-R mitigate this gap.
Problem
Lexical stress detection is critical for computer-aided pronunciation learning (CAPL) to evaluate prosodic realization, but prior work has focused almost exclusively on monolingual setups using handcrafted prosodic features or early neural architectures like GMMs, DBNs, and simple DNNs. Transferring these models across typologically distant languages—such as Modern Standard Arabic (MSA), which relies on strict syllable-weight and vowel-length rules, and English, which uses lexically specified, morphological contrastive stress—is severely challenging. Without proper temporal context modeling and cross-lingual feature spaces, performance drops steeply when models are applied outside their training language.
Method
The framework utilizes forced alignment and pre-trained SSL encoders (WavLM, HuBERT, and XLS-R) to extract 1024-dimensional frame-level representations from 16 kHz audio. Syllable embeddings are generated by averaging frame vectors over time stamps obtained via phonetic alignments and a phonetiser (Buckwalter for MSA, CMUdict for English). Primary and secondary stress labels are collapsed into a binary stressed/unstressed class.
The classification architecture comprises two sequential stages. First, a Pre-net Deep Neural Network (DNN) takes syllable-level embeddings and predicts syllable-wise stress using fully connected hidden layers with ReLU activations and a binary cross-entropy loss. Second, a Post-net Time-Delay Neural Network (TDNN) models inter-syllable dependencies within words. Syllable sequences are padded and masked to a fixed maximum word length of x = 9 syllables. The TDNN layers are initialized with the Pre-net's trained weights to accelerate convergence, sharing identical layer dimensions. Models are optimized using early stopping with a patience of 10 epochs on validation loss.
Key architectural design choices—such as utilizing 1024-dimensional SSL features to encapsulate global prosodic relationships (duration, pitch, intensity) and using 3 to 4 hidden layers—were selected to balance capacity and prevent overfitting. The Post-net was specifically introduced to capture word-level contextual dependencies, which are vital for distinguishing stress patterns across languages with contrasting phonetic rules.
Experimental setup
Experiments use the Arabic (82,601 utterances, ~1.2M syllables) and English (~80,000 utterances, ~1.31M syllables) subsets of Common Voice 12, split into 70/15/15 partitions for training, validation, and test by filename. Baselines compare monolingual training, joint multilingual training, and cross-lingual transfer across three SSL feature extractors (WavLM trained on 94k hours, HuBERT-Large on 60k hours of Libri-Light, and XLS-R on 436k hours across 128 languages). Metrics include classification accuracy and F1-score implemented via scikit-learn in TensorFlow/Keras.
Results
Joint multilingual training preserves near-monolingual performance, yielding average accuracies of 87.87% for Arabic and 95.24% for English, with WavLM Post-net reaching 89.77% on Arabic and HuBERT Post-net hitting 97.08% on English (close to monolingual benchmarks of 90.56% and 97.49% respectively). In contrast, cross-lingual transfer exhibits performance degradation and asymmetry: English-to-Arabic transfer averages 68.12% accuracy (besting at 79.17% with XLS-R Post-net), while Arabic-to-English averages 76.09% (besting at 86.69% with XLS-R Post-net).
Ablations on network depth show that configurations with 3 to 4 hidden layers offer the most stable accuracy across conditions. The Post-net TDNN consistently yields steady gains over the Pre-net DNN, especially in restoring performance in challenging cross-lingual setups by compensating for structural differences between quantity-sensitive and morphologically driven stress.
| System / Condition | Encoder | Arabic Accuracy (%) | English Accuracy (%) |
|---|---|---|---|
| Monolingual Benchmark | WavLM / HuBERT | 90.56 | 97.49 |
| Multilingual (Post-net) | WavLM (Ar) / HuBERT (En) | 89.77 | 97.08 |
| Cross-Lingual (En -> Ar) | XLS-R Post-net | 79.17 | - |
| Cross-Lingual (Ar -> En) | XLS-R Post-net | - | 86.69 |
Limitations
Speaker-disjoint data splits could not be explicitly verified due to a lack of speaker metadata in the Common Voice subsets. The study is restricted to Modern Standard Arabic (ignoring dialectal diversity) and an unverified American-leaning English subset, leaving open how accent variation affects performance. Furthermore, cross-lingual transfer remains bounded by double-digit drops compared to in-language models, and the massive multilingual pre-training scale required (e.g., XLS-R's 436k hours) limits on-device adaptability.
Why read this
Speech researchers and engineers working on multilingual prosody or computer-aided pronunciation training (CAPL) should read this paper to understand how large multilingual SSL embeddings and lightweight word-context TDNNs can overcome typological mismatches in lexical stress.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Computer-Aided Pronunciation Learning (CAPL) systems, automated spoken language assessment tools, and prosody-aware multilingual speech analytics.
Institutions
University of New South Wales
Related
- Boundaryless Speech-to-Syllable Representations with Hierarchical CNN for Linguistically Inspired Automatic Stress Detection — same problem · relatedness 2.8/3
- Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations — complementary · relatedness 2.0/3
- A Novel Sentence Stress Detection Framework Leveraging Auxiliary Word-Stress Modeling and Loss Optimization — shared technique · relatedness 2.0/3
- Evaluating and Preserving Lexical Stress in English-to-Chinese Speech-to-Speech Translation — complementary · relatedness 1.9/3
- Multilingual Phonological Feature Recognition with Self-Supervised Speech Models — shared technique · relatedness 1.9/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-2914