TL;DR — The paper introduces directivity-regularized FastMNMF for DOA-informed speech separation, incorporating soft von Mises angular weights into full-rank spatial covariance matrix diagonalization to achieve robust separation. It yields a striking 4.4 dB SDR improvement over standard FastMNMF for single-source separation in reverberant environments.
Key contributions
- Proposes a directivity-regularized objective for FastMNMF that shapes beam patterns across a continuous angular domain rather than enforcing rigid point-wise constraints.
- Formulates angle-dependent weights using von Mises probability density functions to robustly handle DOA estimation errors, head movement, and speech directivity spread.
- Derives a closed-form vector coordinate descent (VCD) and multiplicative update optimization recipe for the regularized model.
- Demonstrates substantial performance and convergence speed gains over classical BSS (IVA, ILRMA, FastMNMF) and hard-constrained variants in reverberant multi-speaker environments.
Problem
Blind source separation (BSS) methods like standard MNMF and FastMNMF suffer from high degrees of freedom in their full-rank spatial covariance matrices (SCMs), making iterative optimization prone to poor local optima in unknown acoustic environments. While visual sensors on wearable devices can now provide reliable direction of arrival (DOA) estimates, conventional BSS frameworks cannot natively exploit this prior metadata. Existing geometrically constrained approaches rely on hard point-wise steering vector constraints, which fail catastrophically under DOA estimation errors, speaker movement, and reverberation.
Method
The proposed method builds on FastMNMF by integrating a directivity regularization term into the negative log-likelihood objective function. The spatial model restricts source SCMs via a joint-diagonalization matrix , where row vectors act as virtual spatial separation filters. The regularization term encourages a distortionless response toward target speaker DOAs while penalizing responses toward null directions .
Unlike rigid point-constraint models, the angular weights are modeled using von Mises probability density functions centered on the target and null angles: , where controls directional concentration (fixed empirically to ) and determines the overall regularization weight. This smoothly accommodates continuous spatial spread and DOA uncertainty. As the concentration parameter , the von Mises distribution collapses to conventional hard point constraints.
All non-spatial parameters (source PSD bases , activations , and non-negative spatial vectors ) are optimized using standard FastMNMF multiplicative updates, while the joint diagonalization matrix is updated via closed-form vector coordinate descent (VCD). The system uses 5 microphones in an arc array (5 cm inter-mic spacing), STFT features at 16 kHz, and NMF bases.
Experimental setup
Simulated acoustic mixtures in an m room using an arc-shaped 5-microphone array () with an RT60 of 0.3 s and 20 dB SNR environmental interfering noise. Evaluated across sources using CMU ARCTIC corpus speech data (10 speakers). Baselines include IVA, ILRMA, FastMNMF, SR-FastMNMF, and their point-constraint regularized variants. Evaluated via source-to-distortion ratio (SDR), PESQ, and STOI over 10 independent trials.
Results
For source, the proposed DR-FastMNMF with von Mises weighting achieves a headline SDR of dB, PESQ of , and STOI of , outperforming standard FastMNMF ( dB SDR) and the point-constraint variant ( dB SDR). For sources, the proposed method achieves dB SDR compared to dB for FastMNMF and dB for the point-constraint variant, proving the necessity of soft probabilistic weighting over hard constraints. Convergence analysis shows the von Mises variant hits ~6 dB SDR within 10 iterations and converges around 10.5 dB.
In congested conditions ( and ), the proposed method loses its top-1 advantage to SR-FastMNMF with point weighting, indicating that distortionless-response assumptions struggle when source overlap becomes extreme.
| System | Weight | N=1 SDR (dB) | N=2 SDR (dB) | N=3 SDR (dB) | N=4 SDR (dB) |
|---|---|---|---|---|---|
| FastMNMF [6] | - | 13.8 (4.0) | 4.9 (3.4) | 2.6 (2.0) | -0.6 (1.6) |
| SR-FastMNMF [21] | - | 10.1 (1.8) | 5.1 (3.5) | 3.0 (2.4) | -0.5 (1.6) |
| SR-FastMNMF | Point | 9.6 (2.1) | 7.6 (1.1) | 5.9 (0.9) | 3.4 (0.7) |
| DR-FastMNMF (Proposed) | von Mises | 18.2 (2.4) | 9.2 (2.5) | 4.1 (0.6) | -0.1 (0.6) |
Limitations
Evaluated exclusively on simulated room impulse responses with a fixed RT60 of 0.3 s and clean oracle DOAs rather than estimated tracker outputs. Performance degrades in dense acoustic scenes () where spatial congestion limits the effectiveness of distortionless-response beam patterns. Lacks validation on real hardware recordings, mobile form factors, and dynamic head movements.
Why read this
Speech signal processing researchers and audio engineers working on wearable front-ends will learn how to inject probabilistic spatial priors into full-rank matrix factorization frameworks. It offers a mathematically rigorous alternative to hard geometrical constraints for multi-microphone source enhancement.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Smart glasses, wearable audio enhancement, hearing aids, and distant-speech recognition front-ends operating in noisy multi-talker environments.
Institutions
Kyoto University, RIKEN, National Institute of Advanced Industrial Science and Technology
Related
- SPOT-TSE: Spatial Point-Guided Target Speech Extraction — same problem · relatedness 2.4/3
- Geometrically Constrained Decentralized Independent Vector Analysis for Distributed Microphone Arrays — same problem · relatedness 2.2/3
- MeCo: One-Step MeanFlow-based Corrector for Multi-Channel Speech Separation — same problem · relatedness 2.2/3
- HRTF-guided Binaural Target Speaker Extraction with Real-World Validation — same problem · relatedness 2.1/3
- TF-MoE: Time-Frequency Mixture-of-Experts for Efficient Speech Separation — same problem · relatedness 2.1/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-2139