TL;DR — This paper proposes a hierarchical conditional continuous normalizing flow architecture that structurally disentangles high-level speaker attributes from low-level voice qualities to edit creaky voice without altering speaker identity. It achieves improved speaker similarity in subjective tests compared to flat baseline flows.
Key contributions
- A hierarchical conditional continuous normalizing flow (CCNF) that decomposes speaker representation transformation into a multi-stage sequence governed by feature type.
- Integration of Domain Adversarial Training (DAT) with gradient reversal in the intermediate latent space to enforce invariance to high-level attributes like mean pitch and gender.
- A stage-wise conditioning scheme separating high-level speaker descriptors (mean pitch, gender) from low-level paralinguistic voice qualities (breathiness, roughness, creak).
- Comprehensive subjective evaluations by 11 voice quality experts showing significantly better preservation of speaker similarity (SMOS) during creak amplification.
Problem
Standard generative voice editing models struggle with selective control because they learn natural inter-speaker correlations from training data—such as the strong negative correlation between a speaker's mean pitch and creak probability. When a model attempts to modify low-level perceptual voice qualities (PVQs) like creak, it inadvertently drags along correlated high-level traits like pitch and gender, severely damaging perceived speaker identity. While data augmentation strategies can artificially decorrelate specific feature pairs, they are rigid and task-specific rather than providing a general architectural solution.
Method
The model utilizes a two-stage Conditional Continuous Normalizing Flow (CCNF) where the conditioning vector is partitioned into . The transformation is formulated as a sequence of two Ordinary Differential Equation (ODE) initial value problems connected at intermediate time step . The first stage maps the speaker representation to an intermediate latent space conditioned on high-level attributes (mean pitch and gender). The second stage maps down to the base normal distribution conditioned on low-level voice qualities (breathiness, roughness, creak).
To ensure the intermediate representation is truly invariant to high-level traits, Domain Adversarial Training (DAT) is applied using auxiliary predictors (a feed-through regression network for mean pitch with MSE loss and a classification network for binary gender with cross-entropy loss). Gradients from these predictors are reversed using a Gradient Reversal Layer (GRL) and propagated back to the first flow network , forcing it to strip out high-level attribute information.
The YourTTS architecture serves as the speech synthesis backend, operating on extracted durations and text features. The networks use CCNF blocks with hidden dimensions of 128 for and 64 for . The auxiliary predictors use a single hidden layer of dimension 64. Training optimizes the joint log-likelihood plus adversarial losses with regularization weights , a batch size of 200, and an initial learning rate of . Inference is performed by applying the forward flow path to standard speaker representations and re-sampling through the inverse path using modified target conditioning values .
Experimental setup
Experiments are conducted on the LibriTTS-R dataset using YourTTS. Comparisons are made against three systems: base-flow (standard single-stage CCNF), base-extd. (base flow augmented with pitch conditioning and adversarial learning in the final space), and data-mod.-flow (training data modification via pitch-creak augmentation). Metrics include Equal Error Rate (EER) for speaker verification, absolute mean pitch deviation (), gender classification accuracy, Mean Opinion Score (MOS) for naturalness, Speaker Similarity MOS (SMOS), and perceived creak probability on a 0-100 scale rated by 11 trained voice quality experts.
Results
The hierarchical flow achieves the most stable performance across manipulation strengths , exhibiting minimal mean pitch deviation and near-constant gender classification accuracy and low EER, whereas the base model suffers severe pitch drift and identity degradation. In subjective evaluations, both models successfully amplify perceived creak (increasing scores from ~27 to ~61-77), but the hierarchical model incurs a significantly lower drop in Speaker Similarity MOS (SMOS drops to 3.0 vs 1.6 for the base model under amplification, ).
While the hierarchical model successfully preserves speaker identity better during strong modifications, neither model shows statistically significant improvements over the base model regarding naturalness (MOS) changes, and creak suppression differences were less salient due to lower baseline creak levels.
| Model | Condition | MOS | SMOS | Perceived Creak | Speaker Sim. Drop |
|---|---|---|---|---|---|
| base [5] | Unmanipulated | — | |||
| base [5] | Amplified | Severe | |||
| hierarch | Unmanipulated | — | |||
| hierarch | Amplified | Moderate |
Limitations
The evaluation is restricted to a single paralinguistic voice quality (creak) and English data from LibriTTS-R, leaving multi-attribute interactions across diverse languages untested. The approach relies heavily on external estimators (CreaPy, pitch trackers, and auxiliary regressors) whose errors can propagate into the adversarial training loop. Furthermore, listening tests were limited to 11 expert annotators, and absolute creak suppression was difficult to measure conclusively due to floor effects in naturally non-creaky samples.
Why read this
Read this if you work on fine-grained control or disentanglement in speech synthesis and want to see how to combine continuous normalizing flows with adversarial objectives to cleanly decouple correlated acoustic dimensions without manual data hacking.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Training data augmentation for speech therapy educational tools, fine-grained expressive text-to-speech synthesis, and precise voice style editing.
Institutions
Paderborn University, Bielefeld University
Related
- RIVET: Robust Idempotent Voice Attribute Editing — same problem · relatedness 2.3/3
- VoiceQualityGUI: A Tool for Word-Level Voice Quality Modifications — same problem · relatedness 2.1/3
- CFLOW-VC: An unsupervised cycle training strategy based on normalizing flows for Voice Conversion — shared technique · relatedness 2.0/3
- FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech — same problem · relatedness 2.0/3
- SRF-SVB: Style-Consistent Singing Voice Beautifying via Rectified Flow — shared technique · relatedness 2.0/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-1341