All papers
Applications & otherFull-paper digest

A Novel Sentence Stress Detection Framework Leveraging Auxiliary Word-Stress Modeling and Loss Optimization

Tien-Hong Lo, Fong-Chun Tsai, Ting-An Hung, Yu-Hsuan Hsieh, Yao-Ting Sung, Berlin Chen

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.8 KB · Ready to paste

Preview copied content

TL;DR — We propose STRAW, an automatic pronunciation assessment framework built on a frozen Whisper backbone that unifies sentence stress detection (SSD) and word stress detection (WSD) via a novel word-span stress regularizer (WSR), achieving an F1 score of 0.934 on the TinyStress-15K benchmark.

Key contributions

  • A unified framework for prosodic stress detection featuring task-specific heads for token-level Sentence Stress Detection (SSD) and phone-level Word Stress Detection (WSD).
  • A word-span stress regularizer (WSR) that enforces a single dominant stress position within subword-tokenized stressed word spans, mitigating underconstrained probability allocations.
  • Empirical demonstration that adding auxiliary WSD and WSR improves SSD F1 from 0.909 (WhiStress baseline) to 0.934 on the TinyStress-15K corpus.
  • Comprehensive part-of-speech error analysis showing reduced false negative rates on functional categories like particles and determiners.

Problem

Prior automatic pronunciation assessment (APA) systems treat sentence stress detection (SSD) and word stress detection (WSD) as completely independent tasks, failing to exploit their shared reliance on acoustic-prosodic cues such as pitch, duration, and intensity. Furthermore, alignment-based methods heavily rely on fragile forced alignments, while alignment-free token-level models suffer from spatial ambiguity where subword tokenization spreads prominence probabilities across multiple tokens within a single word span. Addressing these gaps is crucial for robust computer-assisted language learning (CALL) feedback.

Method

The framework utilizes a frozen whisper-small backbone, extracting contextualized representations from the 9th layer of both encoder and decoder blocks. The input speech log-mel spectrogram is processed by the encoder to yield acoustic states e1:Te_{1:T}, while the reference transcript provides tokens t1:Mt_{1:M} and a G2P-derived phone sequence p1:Np_{1:N} with CMU-style stress markers.

The SSD module uses Whisper decoder states as queries and encoder states as keys/values in a Transformer decoder block, followed by a fully connected neural network (FCNN) to predict binary token-level stress posteriors y^mssdŷ_m^{ssd} optimized via cross-entropy (LSSDL_{SSD}). Simultaneously, the WSD module computes phone-aware acoustic representations via cross-attention (phone embeddings as queries, encoder states as keys/values) passed through an FCNN to predict binary phone-level stress posteriors y^nwsdŷ_n^{wsd} optimized via cross-entropy (LWSDL_{WSD}).

To resolve multi-token ambiguity within a stressed word span wmw_m spanning tokens [bm,em][b_m, e_m], a word-span stress regularizer (LWSRL_{WSR}) is introduced. It penalizes deviations of the cumulative token stress probability from 1 and sharpens the distribution by pushing the maximum token probability y^issd⋆ŷ_i^{ssd⋆} toward 1 while suppressing other tokens in that span. The multi-task objective combines all three losses weighted by hyperparameters α\alpha, β\beta, and λ\lambda set to 1.0.

Experimental setup

Experiments use the TinyStress-15K synthetic benchmark corpus comprising 13,500 training, 1,500 validation, and 1,000 test utterances. Baselines include GT alignment, Montreal Forced Aligner (MFA) with BLSTM, and the alignment-free WhiStress model. Evaluation metrics are precision, recall, and F1-score with respect to the stressed class, evaluated at the word level for SSD and phone level for WSD. Models are trained on an NVIDIA 3090 GPU using AdamW (batch size 16, initial learning rate 1e-4, 20 epochs), choosing the checkpoint with the highest validation SSD F1.

Results

On the TinyStress-15K test set, the proposed STRAW framework achieves an SSD F1 of 0.934 (0.945 precision, 0.924 recall), outperforming the GT alignment baseline (0.858 F1), MFA (0.815 F1), and the alignment-free WhiStress baseline (0.909 F1). Ablation experiments demonstrate that removing WSD drops F1 to 0.922, removing WSR drops F1 to 0.929, and removing both drops F1 to 0.915. For WSD, STRAW achieves an F1 of 0.920, which is largely unaffected by the removal of WSR (0.921 F1). Error analysis across part-of-speech categories reveals that STRAW reduces false negative rates on functional categories like particles (PART) and determiners (DET), though high false negative rates persist on rare categories such as subordinating conjunctions (SCONJ) due to data scarcity.

ModelPrecisionRecallF1
GT alignment [14]0.8620.8530.858
MFA [14]0.7760.8590.815
WhiStress [14]0.9120.9060.909
STRAW (Complete)0.9450.9240.934
- WSD0.9380.9060.922
- WSR0.9420.9170.929

Limitations

The framework simplifies lexical stress by modeling only a single primary position per word without capturing full syllabic structures or secondary stress. The frozen backbone and separate task-specific heads prevent direct hidden-state interaction or knowledge transfer between the SSD and WSD branches. Evaluation is restricted to synthetic speech due to a lack of large-scale fine-grained human stress annotations.

Why read this

Speech and ML researchers building automatic pronunciation assessment or prosody models will learn how to design alignment-free auxiliary regularizers that fix subword tokenization ambiguities in transformer-based architectures.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Computer-assisted language learning (CALL) systems, automated pronunciation scoring, and prosody feedback tools for second-language learners.

Institutions

National Taiwan Normal University

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1494