All papers
Enhancement & separationFull-paper digest

Improving Multichannel Speech Enhancement through Accurate Room-Acoustic Simulations

Georg Götz, Alessia Milo, Steinar Guðjónsson, Daniel Gert Nielsen, Jesper Pedersen, Finnur Pind

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.3 KB · Ready to paste

Preview copied content

TL;DR — This paper investigates how the fidelity of room-acoustic simulation data augmentation impacts multichannel speech enhancement, showing that training SpatialNet on high-fidelity hybrid wave-geometrical simulations yields up to a 38% relative reduction in median word error rate on measured data compared to standard uninformed geometrical acoustics.

Key contributions

  • Evaluates the impact of room-acoustic simulation fidelity (wave-based vs. geometrical acoustics) on downstream multichannel speech enhancement using real-world measured evaluation data.
  • Introduces LibriCSS-EM6, a 6-channel Eigenmike evaluation dataset comprising ~5,000 utterances across six speaker overlap conditions (0S to OV40) derived from real-world RIRs.
  • Demonstrates that high-fidelity hybrid simulation augmentation reduces median word error rate by up to 38% overall relative to uninformed image-source methods, without modifying model architecture.
  • Isolates the effect of dataset design by comparing uninformed vs. informed geometrical acoustics alongside hybrid wave-based modeling.

Problem

Deep learning models for multichannel speech enhancement rely heavily on data augmentation via simulated room impulse responses (RIRs). However, most state-of-the-art pipelines use simplistic synthetic RIRs generated by shoebox-geometry image-source methods (ISM) with frequency-independent absorption coefficients, ignoring low-frequency room modes, complex boundary conditions, diffraction, and detailed scattering objects. Prior studies have not rigorously explored how simulation fidelity—specifically incorporating wave-based physics and realistic receiver modeling—affects multichannel speech enhancement with rigid arrays in real-world conditions, leaving a simulation-to-real performance gap.

Method

The study trains the SpatialNet-small architecture at a 16 kHz sampling rate using end-to-end multichannel STFT processing, incorporating narrow-band blocks for temporal filtering and cross-band blocks for cross-frequency correlations. SpatialNet is microphone-array-dependent and configured here for a 6-channel subset of the Eigenmike array (channels 1, 19, 11, 27, 21, and 9) simulating a rigid spherical scatterer. Three distinct training datasets (4,801 scenes total, up to 3 overlapping speakers) are compared: (1) ISM-U (uninformed dataset using gpuRIR with randomly sampled shoebox dimensions x in [3, 33]m, y in [3, 26]m, z in [2.5, 4.7]m, and T20 in [0.2, 1.6]s); (2) ISM-M (informed dataset matching the exact room dimensions, source-receiver positions, and T20 targets of the hybrid dataset, modeled as an open array); and (3) Hybrid (generated via Treble SDK combining a low/mid-frequency wave solver up to 1-2 kHz with a geometrical acoustics ray-radiosity/ISM solver up to 12 kHz across 133 living rooms, 103 classrooms, and 88 restaurants, with frequency-dependent complex surface impedances and full-wave 16th-order Ambisonics device-related transfer functions for the Eigenmike).

Models are trained for 30 epochs using direct-path speech targets, with network weights averaged over the last 10 epochs. Evaluation uses the newly introduced LibriCSS-EM6 test set, where clean LibriSpeech utterances are convolved with real-world measured RIRs from the Motus and Arni6DoF datasets, mixed with diffuse noise at 0-20 dB SNR, and transcribed via a 3-layer BLSTM acoustic model and Kaldi 4-gram language model pipeline.

Experimental setup

Evaluated on LibriCSS-EM6 (~5,000 utterances across 60 sessions covering 6 overlap conditions: 0S, 0L, OV10, OV20, OV30, and OV40) augmented with diffuse noise (0-20 dB SNR) and real measured Eigenmike RIRs from Motus and Arni6DoF datasets. Compared against SpatialNet trained on uninformed ISM (ISM-U) and informed ISM (ISM-M). The primary metric is Word Error Rate (WER) with bootstrapped 95% confidence intervals based on paired utterance-level differences. SpatialNet-small is trained for 30 epochs with checkpoint weight averaging over the final 10 epochs.

Results

Training on the high-fidelity hybrid dataset achieves a substantial, statistically significant median WER reduction across all overlap conditions compared to the baseline simulators. Overall, the Hybrid model achieves a 30.0% relative median WER reduction (absolute drop of 4.29%) compared to ISM-U, and a 16.3% relative reduction (absolute drop of 1.93%) compared to the informed ISM-M baseline. Under high speaker overlap conditions (OV40), the relative median WER reduction reaches up to 38.3% compared to ISM-U and 23.5% compared to ISM-M. The informed ISM-M model consistently outperforms the uninformed ISM-U model, confirming that matching room volumes, materials, and source-receiver distributions provides intermediate gains.

Reference SystemCondition: 0S Rel. Imp. (%)Condition: OV20 Rel. Imp. (%)Condition: OV40 Rel. Imp. (%)Overall Rel. Imp. (%)
ISM-U26.733.338.330.0
ISM-M16.715.623.516.3

Limitations

The study is restricted to a single network architecture (SpatialNet-small) and a specific 6-channel rigid spherical array geometry (Eigenmike subset), meaning array-dependent generalization to arbitrary consumer geometries remains unverified. The evaluation relies exclusively on simulated mixtures built from measured RIRs (Motus and Arni6DoF) rather than fully in-the-wild recorded multi-talker directional meeting data. Furthermore, high-fidelity wave-based simulation is computationally heavier than standard geometric methods, though it is used here strictly as an offline data augmentation pre-computation step.

Why read this

Researchers and audio engineers working on multichannel speech enhancement and data augmentation will learn how physical simulation fidelity (wave-based solvers, frequency-dependent materials, and accurate scatterer DRTFs) directly dictates neural network generalization to real-world acoustic environments without changing model code.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Smart speakers, conference room hardware, automotive voice assistants, and hearing aids operating in reverberant environments with background noise and overlapping speakers.

Institutions

Treble Technologies

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-2512