TL;DR — This paper presents a weakly supervised framework that uses acoustic cues and positive-unlabeled (PU) learning to predict continuous prosodic boundary strengths without requiring manual ToBI labels, achieving a Prec@1% of 0.956 against punctuation proxies on Japanese speech.
Key contributions
- A weakly supervised approach for identifying high-confidence prosodic boundary anchors from raw acoustic cues without manual label dependencies.
- A hierarchical anchor design (B1 and B2) that explicitly controls the precision-coverage trade-off using data-driven quantile thresholds.
- The application of non-negative positive-unlabeled (nnPU) learning to infer graded, continuous prosodic boundary strength scores over all candidate junctures.
- Comprehensive validation combining anchor statistics, acoustic cue-score consistency analyses, transcript punctuation enrichment, and a blinded perceptual audit (Fleiss' kappa of 0.657).
Problem
Traditional prosodic boundary detection relies heavily on manual annotations like ToBI labels, which require extensive expert knowledge, are labor-intensive, and are completely absent in most speech corpora. Existing learning-based models typically assume fully labeled datasets or resort to rigid, dataset-specific heuristic rules combining acoustic and linguistic features. Furthermore, defining reliable negative "no-boundary" examples is fundamentally difficult because non-boundary cases are highly ambiguous and context-dependent. This paper addresses this gap by showing how reliable positive anchors can be automatically derived from acoustic evidence to enable principled weak supervision.
Method
The framework operates on 164,323 word junctures extracted from a Japanese speech corpus, focusing on a pitch-valid subset of 132,031 instances where robust F0 is available. Stage 1 constructs a conservative positive anchor set (B1) using long pauses where pause duration . Stage 2 refines B1 using data-driven quantile thresholds on pitch reset magnitude (, yielding 5.58 semitones) and optional energy reduction (, yielding ), producing a looser positive set (, 1,161 items) and a stricter subset (, 250 items). A voicing gate requires at least 15 voiced frames (ratio ) in both pre- and post-boundary 0.3-second windows.
The system frames boundary detection as a strength estimation problem using the non-negative PU (nnPU) risk estimator with logistic loss. The classifier is a lightweight 1-layer MLP with 32 hidden units, ReLU activation, and a sigmoid output mapping to a continuous boundary strength score. Input features consist of five standardized cue variables: pause duration (), pitch reset times voicing (), energy change (), and pre/post voicing ratios. The model is optimized using AdamW (learning rate , weight decay ) for up to 4,000 steps with early stopping (patience 8), treating as the positive set and the remaining candidates as unlabeled, with class prior .
Experimental setup
Experiments are conducted on a large-scale Japanese speech corpus derived from the JVS corpus containing studio-quality read speech from 100 native speakers, yielding 164,323 total word junctures. The dataset is split speaker-disjointly into 80% training and 20% development sets. Evaluation metrics include class prior sensitivity (), transcript punctuation proxy enrichment (Precision@1% = 0.956, Prec@5% = 0.926, AUC = 0.740), and a blinded perceptual audit involving 60 high-score and 60 low-score samples judged by three raters.
Results
The nnPU model achieves strong alignment with external proxies, yielding a Prec@1% of 0.956, a Prec@5% of 0.926, and an AUC of 0.740 against transcript punctuation marks. In a blinded perceptual audit, high-confidence anchors secured a perceived break rate of 0.467 compared to just 0.033 for low-score unlabeled candidates, with substantial inter-rater reliability (Fleiss' ). Sensitivity analysis over class priors demonstrates stable ranking capability with Spearman correlations between and .
| System / Condition | Count | Punctuation Rate | Perceived Break Rate |
|---|---|---|---|
| Unlabeled (U) / Low-Score | 116,833 | 0.124 | 0.033 |
| B1 (Pause ms) | 12,482 | 0.989 | - |
| B2_base (Voice + ) | 898 | 0.972 | - |
| B2_strict (Voice + + RMS) | 243 | 0.975 | 0.467 |
Limitations
The study is currently restricted to read speech in a single language (Japanese), leaving cross-linguistic generalization and spontaneous speech adaptation unverified. The framework relies heavily on reliable automatic word alignments and stable F0 extraction, which can degrade in noisy acoustic environments or dense conversational dialogue. Additionally, the approach currently omits lexical, syntactic, and higher-level semantic contexts that often govern natural prosodic phrasing.
Why read this
Speech researchers and engineers working on prosody, text-to-speech, or punctuation restoration will find this paper a practical blueprint for bypassing manual ToBI annotation bottlenecks using positive-unlabeled learning.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Improving prosodic phrasing in text-to-speech (TTS) synthesis, automatic punctuation restoration, and acoustic modeling for spoken language understanding.
Institutions
East China Normal University
Related
- Conflict-Aware Pseudo-Labeling via Acoustic Signals for Multi-Task Speech Emotion Recognition — shared technique · relatedness 1.9/3
- Progressive Weak Supervision for Speech Emotion Recognition — shared technique · relatedness 1.9/3
- Prosodic Boundary-Aware Streaming Generation for LLM-Based TTS with Streaming Text Input — complementary · relatedness 1.9/3
- LLM-Based Multi-Reference Evaluation for Efficient and Robust Assessment of Phrase Break Annotations — same problem · relatedness 1.8/3
- P-SED : Asymmetric Prototype Metric Learning for Weakly Supervised Speech Emotion Diarization — shared technique · relatedness 1.7/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-209