TL;DR — SpkGuideDOA enhances multi-speaker direct-path inter-channel phase difference (DP-IPD) localization by injecting implicit speaker-wise representations—generated via an auxiliary VAD-supervised guidance branch—into the spatial cue estimator through residual modulation, achieving superior directional accuracy and lower miss rates with minimal compute overhead.
Key contributions
- Proposes a decoupled guidance architecture where an auxiliary Guidance Generator (GG) extracts speaker-specific representations from multichannel log-magnitudes and modulates the Spatial Cue Estimator (SCE) via residual gating at a pooled time-frequency resolution.
- Introduces Joint-Permutation-Invariant Training (Joint-PIT) to enforce consistent speaker assignment across the localization and guidance streams without letting auxiliary gradients degrade the primary SCE backbone.
- Replaces computationally heavy mask-based supervision with a lightweight, pooled binary cross-entropy VAD loss confined strictly to the guidance branch, preventing destructive multi-task gradient interference.
- Demonstrates robust performance improvements across both highly challenging simulated close-speaker scenarios and real-world LOCATA dataset evaluations.
Problem
Localizing multiple moving speakers in dynamic acoustic environments under severe spatial-spectral overlap causes conventional direct-path inter-channel phase difference (DP-IPD) frameworks to suffer from nearly identical spatial cues and ambiguous track assignments when speaker trajectories intersect. While existing deep spatial-state models like IPDNet, IPDNet2, and TF-Mamba offer efficient tracking, they rely exclusively on spatial representations without an explicit mechanism to enforce speaker-specific differentiation. Decoupled separation-and-association tracking strategies exist, but they fail when intermediate DOA estimation breaks down under high spatial ambiguity and struggle with real-time constraints due to long-term mask estimation.
Method
The framework features a dual-stream architecture consisting of a Spatial Cue Estimator (SCE) and a Guidance Generator (GG). The SCE processes concatenated real and imaginary multi-channel STFTs () through point-wise convolutions, grouped 1D frequency convolutions (1D F-GConv), cross-band, and narrow-band modules to predict DP-IPD targets. Simultaneously, the GG processes log-magnitude features () through a multi-resolution band-wise feature extractor ( sub-bands), mobile inverted bottleneck convolution (MBConv1) blocks, a segmented-and-pooled simple softmax-free attention (SP-SimA) mechanism with sequence length , a T-Mamba layer, and a classifier to produce per-speaker, frequency-dependent guidance features at a compressed pooled resolution (Formula could not be rendered; check the original paper. Source: T'_ \times F'_).
Guidance is injected into the SCE features via element-wise sigmoid gating () combined with a residual connection to let the SCE recover from imperfect guidance. The training pipeline optimizes a combined objective using Joint-PIT: a mean squared error (MSE) loss on DP-IPD estimation for localization and a binary cross-entropy (BCE) VAD loss supervised via an auxiliary decoder (). Crucially, the VAD loss gradients are detached from the SCE parameters, ensuring the guidance module acts purely as an efficient localization enhancer rather than a multi-task bottleneck.
Experimental setup
Evaluated on 5-second simulated LibriSpeech mixtures generated via gpuRIR using a 12-channel 3D array across randomized room sizes ( to m), reverberation s, and SNR from to dB, alongside the real-world LOCATA dataset (Tasks 3-6). Compared against IPDNet, TF-Mamba, and IPDNet2 under identical training regimens. Metrics include Miss Detection Rate (MDR, %), False Alarm Rate (FAR, %), and Mean Absolute Error (MAE, ) with a tolerance threshold. Notable hyperparameters: 16 kHz sampling rate, 512-point STFT window with 256-sample hop, temporal pooling , frequency pooling , hidden dimensions , batch size 4, 15 epochs, and Adam optimizer with initial learning rate decaying by per epoch.
Results
On simulated mixtures, SpkGuideDOA achieves an MDR of 4.0%, FAR of 15.8%, and MAE of 5.7, substantially outperforming IPDNet (11.2% MDR, 23.8% FAR, 9.0 MAE), TF-Mamba (5.2% MDR, 18.6% FAR, 6.2 MAE), and capacity-matched IPDNet2 (12.6% MDR, 24.4% FAR, 10.9 MAE). On real-world LOCATA data, it records an MDR of 9.2%, FAR of 8.9%, and MAE of 5.5, leading all baselines. Ablations show that removing the Guidance Generator (w/o GG) spikes simulated MDR to 14.1% and MAE to 11.8, removing the auxiliary VAD loss degrades MDR to 9.3%, and decoupling permutations (w/o Joint-PIT) increases MDR to 6.4% and MAE to 7.3. The model maintains a compact footprint of 2.8M parameters and 1.1 GFLOPs/s.
| System | Sim. MDR (%) | Sim. FAR (%) | Sim. MAE () | Real MDR (%) | Real FAR (%) | Real MAE () |
|---|---|---|---|---|---|---|
| IPDNet [9] | 11.2 | 23.8 | 9.0 | 14.9 | 11.1 | 9.0 |
| TF-Mamba [13] | 5.2 | 18.6 | 6.2 | 9.7 | 13.1 | 6.0 |
| IPDNet2 [10] | 12.6 | 24.4 | 10.9 | 16.9 | 16.3 | 9.7 |
| SpkGuideDOA (Ours) | 4.0 | 15.8 | 5.7 | 9.2 | 8.9 | 5.5 |
Limitations
Evaluated exclusively up to a maximum of two concurrent active speakers (), leaving higher-density multi-speaker scaling untested. The framework relies on simulated room impulse responses and limited real-world LOCATA evaluations, which may not fully capture extreme acoustic anomalies or non-stationary noise distributions found in real-world deployments. Auxiliary guidance assumes a fixed number of tracking slots determined prior to inference.
Why read this
Speech and ML researchers focusing on multi-source spatial localization should read this to see how injecting decoupled, auxiliary VAD-supervised representations via residual gating can resolve spatial overlap ambiguities without ballooning computational complexity.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Real-time multi-speaker tracking, smart conferencing systems, acoustic surveillance, and robust front-end spatial preprocessing for conversational speech recognition.
Institutions
Hanyang University
Funding / 經費: National Research Foundation of Korea
Related
- BiEAR: A Human Auditory-Inspired Adaptive Binaural Front-end for Multi-Speaker Localisation and Distance Estimation — same problem · relatedness 2.3/3
- G2C-NET: A Grid-to-Continuous Neural Network for Sound Source Localization in Distributed Microphone Arrays — same problem · relatedness 2.1/3
- End-Fire Degradation-Robust DOA Estimation for Compact Linear Microphone Arrays — same problem · relatedness 2.0/3
- Position-Aware Target Speaker Extraction for Long-Form Multi-Party Conversations: A Diarization-Free Framework for ASR — complementary · relatedness 2.0/3
- Ego-Noise-Aware Spatial Filtering for Reliable UAV Audition in Extreme Low-SNR Conditions — same problem · relatedness 2.0/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-3131