All papers
Enhancement & separationFull-paper digest

VeRe-Flow: Guiding Flow Matching toward Clean Speech via Velocity Contrastive Regularization and Representation Alignment for Noise-Robust Bandwidth Expansion

Sujin Koo, Sangyoon Kim, Ji Sub Um, Hoirin Kim

Code & resourcesvere-flow.github.io/VeRe-Flow-Demo

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.5 KB · Ready to paste

Preview copied content

TL;DR — VeRe-Flow is a clean-guided flow matching framework for noise-robust bandwidth expansion that uses velocity contrastive regularization and representation alignment to suppress noise and restore high frequencies, achieving state-of-the-art LSD (1.10) and DNSMOS OVRL (3.12).

Key contributions

  • Introduces velocity contrastive regularization (VeCoR) to provide two-sided supervision in velocity space, attracting predicted velocity toward clean trajectories while repelling noisy ones.
  • Integrates a representation alignment objective (REPA) that encourages intermediate model features to match clean self-supervised learning (SSL) representations.
  • Combines convolutional residual blocks and noise-robust SSL conditioning (XEUS) within a unified flow-based noise-robust bandwidth expansion (NR-BWE) architecture.
  • Demonstrates superior performance on Valentini-Botinhao, outperforming generative and non-generative baselines across LSD, all DNSMOS metrics, and subjective MOS.

Problem

Traditional bandwidth expansion (BWE) models assume clean inputs and degrade severely when exposed to background noise, whereas standard speech enhancement methods remove noise but fail to reconstruct missing high-frequency components. Prior joint approaches struggle with the fundamental trade-off between accurate high-frequency spectral recovery and effective noise suppression. Furthermore, standard flow matching relies on one-sided supervision, which leads to ambiguous velocity estimation under noisy conditions and causes generative trajectories to drift away from the clean speech manifold.

Method

The model parameterizes a conditional flow matching velocity field using a sandwich architecture containing a convolutional pre-stage, a central transformer stage, and a convolutional post-stage (each convolutional stage built from 4 DiC-style Conv ResBlocks using GroupNorm, activations, kernel size 3, and mid-block scale-and-shift time conditioning). The network receives a noisy low-resolution mel-spectrogram concatenated with frame-wise noise-robust SSL features from frozen XEUS (extracted every 20 ms and projected into the input space). Unlike standard flow matching which starts from a data-dependent prior, VeRe-Flow employs a Gaussian prior (x0∼N(0,I)x_0 \sim \mathcal{N}(0, I)) mapped to high-resolution clean mel targets (x1=xHRcleanx_1 = x_{\text{HR}}^{\text{clean}}).

The training objective combines a velocity contrastive regularization loss (LVeCoR\mathcal{L}_{\text{VeCoR}}) and a representation alignment loss (Lalign\mathcal{L}_{\text{align}}). VeCoR pulls the predicted velocity toward the clean velocity vector while explicitly repelling it from the noisy velocity vector (xHRnoisyx_{\text{HR}}^{\text{noisy}}) using a repulsion weight λVeCoR=0.05\lambda_{\text{VeCoR}} = 0.05. Representation alignment (adapted from REPA) minimizes cosine similarity losses between intermediate hidden states and clean XEUS SSL features with a weight λalign=0.25\lambda_{\text{align}} = 0.25. At inference, the model solves the learned ordinary differential equation using an Euler solver with an extremely low number of function evaluations (NFE = 2), and wave reconstruction is performed using a pre-trained BigVGAN vocoder operating at 16 kHz with 80 mel bins.

Experimental setup

Evaluated on the merged 84-speaker Valentini-Botinhao parallel clean-noisy corpus (combining the 28- and 56-speaker sets) with 20 unseen noise conditions in the test set. Inputs are simulated via Chebyshev Type-I low-pass filtering and downsampled to 8 kHz (reconstructed outputs evaluated at 16 kHz). Compared against non-generative baselines (UEE, MTL-MBE, EP-WUN, I-DTLN+, SDNet, Liu et al.) and generative baselines (NU-Wave2 and FLowHigh retrained under NR-BWE conditions). Metrics include Log-Spectral Distance (LSD), DNSMOS (SIG, BAK, OVRL), and 5-point Mean Opinion Score (MOS) evaluated via Amazon Mechanical Turk. Implementation uses the Adam optimizer for 400k iterations, batch size 16, learning rate 3×10−43 \times 10^{-4} with a cosine annealing schedule, and random SNR sampling between 5 dB and 20 dB during training.

Results

VeRe-Flow achieves the lowest LSD (1.10) and highest DNSMOS OVRL (3.12) among all compared methods, alongside a top generative MOS of 4.14. Against retrained generative baseline FLowHigh (NFE=2, Euler solver), VeRe-Flow improves LSD from 1.12 to 1.10, SIG from 3.40 to 3.43, BAK from 3.91 to 3.97, and OVRL from 3.07 to 3.12. Diffusion baseline NU-Wave2 achieves an LSD of 1.35 and OVRL of 2.98 at NFE=48. Ablations show that XEUS SSL features outperform WavLM (LSD 1.15) and Wav2Vec 2.0 (LSD 1.47), and component-wise additions verify that XEUS primarily drives LSD reduction while REPA and VeCoR specifically boost DNSMOS SIG and BAK scores.

MethodNFELSD ↓SIG ↑BAK ↑OVRL ↑MOS ↑
GT–0.003.514.043.224.30
EP-WUN [3]11.233.502.942.86–
SDNet [5]11.163.293.322.92–
NU-Wave2† [9]481.353.293.932.983.76
FLowHigh† [7]21.123.403.913.074.03
Proposed VeRe-Flow21.103.433.973.124.14

Limitations

The evaluation is restricted to English speech corpora and a fixed 8 kHz input to 16 kHz output bandwidth expansion setting. The method relies on a frozen external SSL model (XEUS) for feature extraction during both training and inference, adding computational overhead and dependency on pre-trained semantic representations. Additionally, performance has not been tested on extremely low signal-to-noise ratios below 5 dB or highly non-stationary real-world acoustic environments outside the Valentini-Botinhao dataset distribution.

Why read this

Speech and generative modeling researchers should read this paper to learn how to inject two-sided velocity guidance and representation alignment into continuous flow matching frameworks to prevent trajectory drift in noise-corrupted generative tasks.

Code

Applications

Noise-robust telephony, hearing aids, and legacy audio archiving where low-resolution, noisy speech must be restored to wideband studio quality.

Institutions

MAGO, KAIST

Funding / 經費: Institute of Information & Communications Technology Planning & Evaluation

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-712