TL;DR — This paper investigates how the fidelity of room-acoustic simulation data augmentation impacts multichannel speech enhancement, showing that training SpatialNet on high-fidelity hybrid wave-geometrical simulations yields up to a 38% relative reduction in median word error rate on measured data compared to standard uninformed geometrical acoustics.
Key contributions
- Evaluates the impact of room-acoustic simulation fidelity (wave-based vs. geometrical acoustics) on downstream multichannel speech enhancement using real-world measured evaluation data.
- Introduces LibriCSS-EM6, a 6-channel Eigenmike evaluation dataset comprising ~5,000 utterances across six speaker overlap conditions (0S to OV40) derived from real-world RIRs.
- Demonstrates that high-fidelity hybrid simulation augmentation reduces median word error rate by up to 38% overall relative to uninformed image-source methods, without modifying model architecture.
- Isolates the effect of dataset design by comparing uninformed vs. informed geometrical acoustics alongside hybrid wave-based modeling.
Problem
Deep learning models for multichannel speech enhancement rely heavily on data augmentation via simulated room impulse responses (RIRs). However, most state-of-the-art pipelines use simplistic synthetic RIRs generated by shoebox-geometry image-source methods (ISM) with frequency-independent absorption coefficients, ignoring low-frequency room modes, complex boundary conditions, diffraction, and detailed scattering objects. Prior studies have not rigorously explored how simulation fidelity—specifically incorporating wave-based physics and realistic receiver modeling—affects multichannel speech enhancement with rigid arrays in real-world conditions, leaving a simulation-to-real performance gap.
Method
The study trains the SpatialNet-small architecture at a 16 kHz sampling rate using end-to-end multichannel STFT processing, incorporating narrow-band blocks for temporal filtering and cross-band blocks for cross-frequency correlations. SpatialNet is microphone-array-dependent and configured here for a 6-channel subset of the Eigenmike array (channels 1, 19, 11, 27, 21, and 9) simulating a rigid spherical scatterer. Three distinct training datasets (4,801 scenes total, up to 3 overlapping speakers) are compared: (1) ISM-U (uninformed dataset using gpuRIR with randomly sampled shoebox dimensions x in [3, 33]m, y in [3, 26]m, z in [2.5, 4.7]m, and T20 in [0.2, 1.6]s); (2) ISM-M (informed dataset matching the exact room dimensions, source-receiver positions, and T20 targets of the hybrid dataset, modeled as an open array); and (3) Hybrid (generated via Treble SDK combining a low/mid-frequency wave solver up to 1-2 kHz with a geometrical acoustics ray-radiosity/ISM solver up to 12 kHz across 133 living rooms, 103 classrooms, and 88 restaurants, with frequency-dependent complex surface impedances and full-wave 16th-order Ambisonics device-related transfer functions for the Eigenmike).
Models are trained for 30 epochs using direct-path speech targets, with network weights averaged over the last 10 epochs. Evaluation uses the newly introduced LibriCSS-EM6 test set, where clean LibriSpeech utterances are convolved with real-world measured RIRs from the Motus and Arni6DoF datasets, mixed with diffuse noise at 0-20 dB SNR, and transcribed via a 3-layer BLSTM acoustic model and Kaldi 4-gram language model pipeline.
Experimental setup
Evaluated on LibriCSS-EM6 (~5,000 utterances across 60 sessions covering 6 overlap conditions: 0S, 0L, OV10, OV20, OV30, and OV40) augmented with diffuse noise (0-20 dB SNR) and real measured Eigenmike RIRs from Motus and Arni6DoF datasets. Compared against SpatialNet trained on uninformed ISM (ISM-U) and informed ISM (ISM-M). The primary metric is Word Error Rate (WER) with bootstrapped 95% confidence intervals based on paired utterance-level differences. SpatialNet-small is trained for 30 epochs with checkpoint weight averaging over the final 10 epochs.
Results
Training on the high-fidelity hybrid dataset achieves a substantial, statistically significant median WER reduction across all overlap conditions compared to the baseline simulators. Overall, the Hybrid model achieves a 30.0% relative median WER reduction (absolute drop of 4.29%) compared to ISM-U, and a 16.3% relative reduction (absolute drop of 1.93%) compared to the informed ISM-M baseline. Under high speaker overlap conditions (OV40), the relative median WER reduction reaches up to 38.3% compared to ISM-U and 23.5% compared to ISM-M. The informed ISM-M model consistently outperforms the uninformed ISM-U model, confirming that matching room volumes, materials, and source-receiver distributions provides intermediate gains.
| Reference System | Condition: 0S Rel. Imp. (%) | Condition: OV20 Rel. Imp. (%) | Condition: OV40 Rel. Imp. (%) | Overall Rel. Imp. (%) |
|---|---|---|---|---|
| ISM-U | 26.7 | 33.3 | 38.3 | 30.0 |
| ISM-M | 16.7 | 15.6 | 23.5 | 16.3 |
Limitations
The study is restricted to a single network architecture (SpatialNet-small) and a specific 6-channel rigid spherical array geometry (Eigenmike subset), meaning array-dependent generalization to arbitrary consumer geometries remains unverified. The evaluation relies exclusively on simulated mixtures built from measured RIRs (Motus and Arni6DoF) rather than fully in-the-wild recorded multi-talker directional meeting data. Furthermore, high-fidelity wave-based simulation is computationally heavier than standard geometric methods, though it is used here strictly as an offline data augmentation pre-computation step.
Why read this
Researchers and audio engineers working on multichannel speech enhancement and data augmentation will learn how physical simulation fidelity (wave-based solvers, frequency-dependent materials, and accurate scatterer DRTFs) directly dictates neural network generalization to real-world acoustic environments without changing model code.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Smart speakers, conference room hardware, automotive voice assistants, and hearing aids operating in reverberant environments with background noise and overlapping speakers.
Institutions
Treble Technologies
Related
- Explicit Context-Driven Neural Acoustic Modeling for High-Fidelity RIR Generation — complementary · relatedness 2.3/3
- Towards Array-Invariant Speech Enhancement via Geometry-Aware Dynamic Convolution — same problem · relatedness 2.2/3
- Bridging the Distribution Gap in Real-World Far-Field Speech Enhancement via Lightweight Latent Representation Alignment — same problem · relatedness 2.1/3
- Scalable Audio Scene Generation with the Treble SDK — complementary · relatedness 2.1/3
- A Novel Transfer Learning Approach for Room Impulse Response Estimation and Speech Dereverberation Across Geometrically Diverse and Data-Scarce Environments — same problem · relatedness 2.0/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-2512