TL;DR — This paper demonstrates that end-fire localization degradation in compact linear microphone arrays stems from an ill-conditioned TDOA-to-azimuth mapping, and proposes two training-free, reliability-aware methods—W-SRP-PHAT and GCC-WLS—that reduce end-fire root-mean-square error from 4.83° to 3.18° at a 1 m source distance.
Key contributions
- Derived a theoretical Fisher information and CRLB analysis proving that end-fire degradation arises from severe error amplification caused by the ill-conditioned inverse mapping of time-difference-of-arrival (TDOA) to azimuth.
- Proposed a CRLB-driven inverse-variance weighting strategy (W-SRP-PHAT) embedded with a coherence-based confidence factor to optimally suppress variance-driven peak flattening in end-fire regions.
- Introduced a TDOA refinement and cosine-domain weighted least squares fusion scheme (GCC-WLS) that avoids nonlinearly amplified angular errors by fusing estimates linearly before inversion.
- Collected and released a new real-world reverberant dataset using a compact 4-mic uniform linear array to benchmark end-fire DOA estimation performance.
Problem
Direction-of-arrival (DOA) estimation using compact uniform linear arrays (ULAs) consistently suffers from degraded accuracy, severe bias, and broadened spatial spectrum main lobes when sources approach the end-fire axis ( or ). Prior literature treated this primarily as an empirical observation rather than a fundamental geometric constraint. Conventional spatial processors like SRP-PHAT, MVDR, and MUSIC fail in these regions because compact apertures combined with indoor noise and reverberation cause small TDOA perturbations to map into massive angular errors.
Method
The paper investigates a compact ULA setup and models propagation delays under a far-field plane-wave assumption. Through sensitivity analysis, the authors show that the derivative of azimuth with respect to TDOA approaches infinity as , rendering the inversion highly ill-conditioned. To counter this, two complementary strategies are designed within a reliability-aware phase transform framework without modifying array geometry.
First, the weighting-enhanced SRP approach (W-SRP-PHAT) constructs wideband SRP scores by incorporating a CRLB-inspired inverse-variance weighting structure , where is the inter-microphone distance. This is modulated by a coherence-based confidence factor derived from auto- and cross-spectra to suppress spurious local maxima and emphasize informative high-resolution frequency components and baselines. The optimal hyperparameters are empirically and theoretically set to and .
Second, the GCC-WLS method avoids direct angle averaging by performing least-squares fusion in the linear cosine domain () prior to applying the inverse cosine mapping. TDOA values are initially extracted via GCC-PHAT cross-correlation over a physically constrained delay search space with 16 oversampling and parabolic interpolation. By keeping perturbations additive and linear during aggregation, the method avoids repeated error amplification, applying the ill-conditioned arccos operation only once after fusion.
Experimental setup
Experiments are conducted in a m room using a 4-microphone compact ULA with an inter-element spacing of 3.5 cm (total aperture 10.5 cm), positioned 1.5 m high and 0.2 m from a wall. A loudspeaker broadcasts clean speech across azimuths – in increments at distances of 1 m and 2 m, generating roughly 100–200 one-second speech segments per condition under 10–20 dB SNR and typical home reverberation. Baselines include standard training-free zero-shot methods SRP-PHAT and SRP-MVDR, evaluated using overall RMSE (), end-fire RMSE ( for – and –), Cauchy soft-accuracy (S-ACC@), and degraded span width.
Results
At a 1 m source distance, the proposed GCC-WLS achieves an overall RMSE of and an end-fire RMSE () of , substantially outperforming baseline SRP-PHAT ( overall, end-fire) and SRP-MVDR ( overall, end-fire). W-SRP-PHAT performs comparably with an overall RMSE of and end-fire RMSE of , while both proposed methods achieve an end-fire soft-accuracy (S-ACC) of 0.82 (compared to 0.65 for SRP-PHAT and 0.51 for SRP-MVDR) and entirely eliminate the degraded span down to (from in baselines).
At a more challenging 2 m distance, W-SRP-PHAT leads with an overall RMSE of and an end-fire RMSE of (vs. SRP-PHAT's and ), maintaining a degraded span whereas baseline degradation spans widen to –. The methods do not win in terms of added computational complexity over basic frame-level argmax search due to oversampling and cosine-domain fusion overhead, though they remain entirely training-free.
| System/Condition | () | () | S-ACC | Deg. Span () |
|---|---|---|---|---|
| 1m: SRP-MVDR | 5.31 | 9.01 | 0.51 | 30 |
| 1m: SRP-PHAT | 3.31 | 4.83 | 0.65 | 30 |
| 1m: W-SRP-PHAT | 2.45 | 3.18 | 0.82 | 0 |
| 1m: GCC-WLS | 2.43 | 3.14 | 0.82 | 0 |
| 2m: SRP-PHAT | 5.38 | 7.71 | 0.47 | 40 |
| 2m: W-SRP-PHAT | 3.71 | 4.46 | 0.70 | 0 |
Limitations
The evaluation is restricted to a single 4-element linear array configuration with a fixed 3.5 cm spacing in a single indoor reverberant environment. The analysis assumes a single far-field sound source, leaving multi-source scenarios and near-field spherical wavefront models unaddressed. Furthermore, performance is only tested on horizontal azimuths from to , excluding extreme axial end-fire angles ( and ) where division-by-zero singularities occur in the mathematical formulation.
Why read this
Audio engineers and researchers working on resource-constrained edge hardware will learn how to stabilize traditional, zero-training spatial processors without resorting to heavy neural networks. It provides rigorous theoretical framing and practical signal-processing fixes for the long-standing problem of end-fire variance in compact linear microphone arrays.
Code
Applications
Smart televisions, voice-controlled home appliances, and low-complexity edge devices requiring robust spatial audio capture and direction-of-arrival tracking.
Institutions
Wuhan University, Harbin Engineering University, Guangdong Murora Intelligent Lighting Co., Ltd
Related
- SpkGuideDOA: Speaker-wise Representation Guidance for Multiple Moving Speaker Localization — same problem · relatedness 2.0/3
- What Do Neural Networks Learn for TDOA Estimation? A Cross-Architecture Probing Study — same problem · relatedness 2.0/3
- Ego-Noise-Aware Spatial Filtering for Reliable UAV Audition in Extreme Low-SNR Conditions — same problem · relatedness 1.9/3
- Position-Aware Target Speaker Extraction for Long-Form Multi-Party Conversations: A Diarization-Free Framework for ASR — complementary · relatedness 1.9/3
- Ada-Mic: Orientation-Adaptive and Robust Close-to-Mic Speech Detection on Smartphone Using Generalized Cross-Correlation Features — shared technique · relatedness 1.9/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-2156