All papers
Enhancement & separationFull-paper digest

Optimal Source Placement for TDoA-based Geometry Calibration of Distributed Microphone Arrays

Xu Wang, Qintuya Si, Qingying Zhao, De Hu

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.4 KB · Ready to paste

Preview copied content

TL;DR — This paper presents an optimal source placement strategy using a mobile robot to improve the time-difference-of-arrival (TDoA) based geometry calibration of distributed microphone arrays, significantly reducing self-localization mean squared error compared to random placement. By minimizing the Cramer-Rao lower bound (CRLB) of microphone positions and capture time offsets (CTOs), the approach achieves superior geometric conditioning.

Key contributions

  • Derives the Fisher Information Matrix (FIM) and Cramér-Rao Lower Bound (CRLB) for TDoA-based geometry calibration explicitly accounting for capture time offsets (CTOs).
  • Proposes a one-stage optimal source placement method that jointly optimizes all robot calibration source positions via Adam combined with simultaneous perturbation stochastic approximation (SPSA) and projection constraints.
  • Develops a computationally efficient multi-stage placement solution that iteratively optimizes source positions to balance runtime and localization accuracy.
  • Provides extensive simulation validation under both synthetic Gaussian noise and realistic acoustic room impulse responses (RIRs) generated via the image source method.

Problem

Distributed microphone arrays (DMAs) require precise geometric knowledge of sensor locations to perform high-resolution tasks like beamforming and sound source localization. Traditional self-localization methods often rely on randomly placed calibration sources, which lead to suboptimal relative geometries and inflate the Cramér-Rao lower bound (CRLB). Furthermore, unknown capture time offsets (CTOs) between asynchronous nodes bias time-of-arrival measurements, complicating calibration. Addressing this requires controllable source trajectories—such as those executed by a mobile robot—to strategically place calibration nodes and minimize calibration error.

Method

The system models a 2D or 3D distributed microphone array with MM nodes and a mobile robot emitting acoustic signals from NN positions. The TDoA measurement model incorporates microphone positions rir_i, source positions sns_n, sound speed cc, and node capture time offsets ηi\eta_i. The authors construct the Fisher Information Matrix (FIM) and derive the CRLB of the unknown parameter vector x=[r1T,…,rMT,δT]Tx = [r_1^T, \dots, r_M^T, \delta^T]^T (where δi=c⋅ηi\delta_i = c \cdot \eta_i), proving that the CRLB is independent of absolute CTOs but highly dependent on the source-sensor relative geometry.

To find optimal source coordinates θ\theta, a non-convex cost function minimizing the trace of the CRLB inverse is established under box constraints defining a restricted room boundary. Because exact gradients are intractable, the method employs the Adam optimizer coupled with simultaneous perturbation stochastic approximation (SPSA) for gradient estimation, alongside a projection operator to enforce room boundary feasibility. Initial microphone position estimates obtained from preliminary random sources initialize the loop, with experiments demonstrating robustness against initial position errors.

To mitigate the linear scaling of computational complexity with respect to the source count NN, a multi-stage solution is introduced. The first stage jointly optimizes a small subset of KK sources (K=4K=4 for M>2M>2), while subsequent stages iteratively optimize remaining source positions one by one. This decouples the 2N2N-dimensional optimization problem into smaller sub-problems, significantly cutting runtime while preserving calibration performance.

Experimental setup

Simulations were conducted in MATLAB (R2022b) on an Intel i7-12700KF CPU with 32GB RAM. Rooms were sized dynamically between 6m×6m6\text{m} \times 6\text{m} and 10m×10m10\text{m} \times 10\text{m} (or 8×108\times10 for RIR tests), with sound speed c=343m/sc = 343\text{m}/\text{s} and random CTOs drawn from [0,1000] μs[0, 1000]\,\mu\text{s}. Baselines included random source placement. Metrics evaluated include Cramér-Rao Lower Bound and Mean Squared Error for microphone positions (MSEr\text{MSE}_r) and distance offsets (MSEδ\text{MSE}_\delta) over L=100L=100 Monte Carlo trials.

Results

The proposed one-stage optimal placement method consistently outperforms random placement across varying Gaussian noise levels (σd\sigma_d), yielding substantially lower MSEr\text{MSE}_r and MSEδ\text{MSE}_\delta. In acoustic simulations with GCC-PHAT estimated TDoAs (RT60=0.3sRT_{60}=0.3\text{s}, SNR=15dB\text{SNR}=15\text{dB}, M=N=10M=N=10), the multi-stage method with K=6K=6 cuts the position MSE from 0.240m20.240\text{m}^2 (random) down to 0.0645m20.0645\text{m}^2 while running in 0.8670.867 seconds compared to 1.591.59 seconds for the full one-stage method. Ablations over varying KK in the multi-stage algorithm show that increasing KK improves localization accuracy at the expense of runtime. The method experiences graceful degradation under harsher reverberation (RT60RT_{60} up to 0.4s0.4\text{s}) and lower SNR (down to 0dB0\text{dB}), maintaining an advantage over random setups except in extreme noise conditions where TDoA outliers dominate.

Placement MethodsRunning Times (s)MSEr\text{MSE}_r (m2\text{m}^2)MSEδ\text{MSE}_\delta (m2\text{m}^2)
Random PlacementN/A0.24000.1141
Multi-Stage (K=4K=4)0.7430.11280.0785
Multi-Stage (K=6K=6)0.8670.06450.0373
Multi-Stage (K=8K=8)1.2700.06340.0353
One-Stage Method1.5900.04890.0294

Limitations

The current framework assumes ideal line-of-sight propagation and lacks explicit modeling for TDoA outliers caused by severe multipath reflections or background noise clipping. Evaluations are strictly simulation-based (both synthetic Gaussian and image-source RIR models) without real-world hardware deployment validation. Furthermore, the optimization relies on rough initial microphone position estimates, which could degrade if initial errors exceed the evaluated tolerance bounds.

Why read this

Speech and ML engineers building distributed acoustic sensor networks or smart-home multi-device audio systems should read this paper to learn how active, trajectory-optimized calibration sources can drastically reduce geometric self-localization errors.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Distributed microphone array geometry calibration, wireless acoustic sensor networks, smart speaker spatial configuration.

Institutions

Inner Mongolia University

Funding / 經費: National Natural Science Foundation of China, Natural Science Foundation of Inner Mongolia Autonomous Region

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1691