---
id: gotz26_interspeech
title: Improving Multichannel Speech Enhancement through Accurate Room-Acoustic
  Simulations
authors:
  - Georg Götz
  - Alessia Milo
  - Steinar Guðjónsson
  - Daniel Gert Nielsen
  - Jesper Pedersen
  - Finnur Pind
year: 2026
doi: 10.21437/Interspeech.2026-2512
isca_url: https://www.isca-archive.org/interspeech_2026/gotz26_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/gotz26_interspeech.pdf
session: Challenges in Speech Data Collection, Curation, and Annotation
topics:
  - speech-enhancement
  - self-supervised
  - evaluation
category: enhancement-separation
labels:
  - robustness-noise
institutions:
  - Treble Technologies
code:
  url: ""
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: gotz26_interspeech
  category: enhancement-separation
  labels:
    - robustness-noise
  institutions:
    - Treble Technologies
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-2512
  pdf: https://www.isca-archive.org/interspeech_2026/gotz26_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/gotz26_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/gotz26_interspeech/markdown.md
---

# Improving Multichannel Speech Enhancement through Accurate Room-Acoustic Simulations

*Georg Götz, Alessia Milo, Steinar Guðjónsson, Daniel Gert Nielsen, Jesper Pedersen, Finnur Pind*

[PDF](https://www.isca-archive.org/interspeech_2026/gotz26_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/gotz26_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-2512)

**Category:** `enhancement-separation` · **Labels:** `robustness-noise`

**TL;DR** — This paper investigates how the fidelity of room-acoustic simulation data augmentation impacts multichannel speech enhancement, showing that training SpatialNet on high-fidelity hybrid wave-geometrical simulations yields up to a 38% relative reduction in median word error rate on measured data compared to standard uninformed geometrical acoustics.

## Key contributions

- Evaluates the impact of room-acoustic simulation fidelity (wave-based vs. geometrical acoustics) on downstream multichannel speech enhancement using real-world measured evaluation data.
- Introduces LibriCSS-EM6, a 6-channel Eigenmike evaluation dataset comprising ~5,000 utterances across six speaker overlap conditions (0S to OV40) derived from real-world RIRs.
- Demonstrates that high-fidelity hybrid simulation augmentation reduces median word error rate by up to 38% overall relative to uninformed image-source methods, without modifying model architecture.
- Isolates the effect of dataset design by comparing uninformed vs. informed geometrical acoustics alongside hybrid wave-based modeling.

## Problem

Deep learning models for multichannel speech enhancement rely heavily on data augmentation via simulated room impulse responses (RIRs). However, most state-of-the-art pipelines use simplistic synthetic RIRs generated by shoebox-geometry image-source methods (ISM) with frequency-independent absorption coefficients, ignoring low-frequency room modes, complex boundary conditions, diffraction, and detailed scattering objects. Prior studies have not rigorously explored how simulation fidelity—specifically incorporating wave-based physics and realistic receiver modeling—affects multichannel speech enhancement with rigid arrays in real-world conditions, leaving a simulation-to-real performance gap.

## Method

The study trains the SpatialNet-small architecture at a 16 kHz sampling rate using end-to-end multichannel STFT processing, incorporating narrow-band blocks for temporal filtering and cross-band blocks for cross-frequency correlations. SpatialNet is microphone-array-dependent and configured here for a 6-channel subset of the Eigenmike array (channels 1, 19, 11, 27, 21, and 9) simulating a rigid spherical scatterer. Three distinct training datasets (4,801 scenes total, up to 3 overlapping speakers) are compared: (1) ISM-U (uninformed dataset using gpuRIR with randomly sampled shoebox dimensions x in [3, 33]m, y in [3, 26]m, z in [2.5, 4.7]m, and T20 in [0.2, 1.6]s); (2) ISM-M (informed dataset matching the exact room dimensions, source-receiver positions, and T20 targets of the hybrid dataset, modeled as an open array); and (3) Hybrid (generated via Treble SDK combining a low/mid-frequency wave solver up to 1-2 kHz with a geometrical acoustics ray-radiosity/ISM solver up to 12 kHz across 133 living rooms, 103 classrooms, and 88 restaurants, with frequency-dependent complex surface impedances and full-wave 16th-order Ambisonics device-related transfer functions for the Eigenmike).

Models are trained for 30 epochs using direct-path speech targets, with network weights averaged over the last 10 epochs. Evaluation uses the newly introduced LibriCSS-EM6 test set, where clean LibriSpeech utterances are convolved with real-world measured RIRs from the Motus and Arni6DoF datasets, mixed with diffuse noise at 0-20 dB SNR, and transcribed via a 3-layer BLSTM acoustic model and Kaldi 4-gram language model pipeline.

## Experimental setup

Evaluated on LibriCSS-EM6 (~5,000 utterances across 60 sessions covering 6 overlap conditions: 0S, 0L, OV10, OV20, OV30, and OV40) augmented with diffuse noise (0-20 dB SNR) and real measured Eigenmike RIRs from Motus and Arni6DoF datasets. Compared against SpatialNet trained on uninformed ISM (ISM-U) and informed ISM (ISM-M). The primary metric is Word Error Rate (WER) with bootstrapped 95% confidence intervals based on paired utterance-level differences. SpatialNet-small is trained for 30 epochs with checkpoint weight averaging over the final 10 epochs.

## Results

Training on the high-fidelity hybrid dataset achieves a substantial, statistically significant median WER reduction across all overlap conditions compared to the baseline simulators. Overall, the Hybrid model achieves a 30.0% relative median WER reduction (absolute drop of 4.29%) compared to ISM-U, and a 16.3% relative reduction (absolute drop of 1.93%) compared to the informed ISM-M baseline. Under high speaker overlap conditions (OV40), the relative median WER reduction reaches up to 38.3% compared to ISM-U and 23.5% compared to ISM-M. The informed ISM-M model consistently outperforms the uninformed ISM-U model, confirming that matching room volumes, materials, and source-receiver distributions provides intermediate gains.

| Reference System | Condition: 0S Rel. Imp. (%) | Condition: OV20 Rel. Imp. (%) | Condition: OV40 Rel. Imp. (%) | Overall Rel. Imp. (%) |
| :--- | :--- | :--- | :--- | :--- |
| ISM-U | 26.7 | 33.3 | 38.3 | 30.0 |
| ISM-M | 16.7 | 15.6 | 23.5 | 16.3 |

## Limitations

The study is restricted to a single network architecture (SpatialNet-small) and a specific 6-channel rigid spherical array geometry (Eigenmike subset), meaning array-dependent generalization to arbitrary consumer geometries remains unverified. The evaluation relies exclusively on simulated mixtures built from measured RIRs (Motus and Arni6DoF) rather than fully in-the-wild recorded multi-talker directional meeting data. Furthermore, high-fidelity wave-based simulation is computationally heavier than standard geometric methods, though it is used here strictly as an offline data augmentation pre-computation step.

## Why read this

Researchers and audio engineers working on multichannel speech enhancement and data augmentation will learn how physical simulation fidelity (wave-based solvers, frequency-dependent materials, and accurate scatterer DRTFs) directly dictates neural network generalization to real-world acoustic environments without changing model code.

## Code

None released (as of this page's `updated` date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

## Applications

Smart speakers, conference room hardware, automotive voice assistants, and hearing aids operating in reverberant environments with background noise and overlapping speakers.

## Institutions / 機構

Treble Technologies

## Related

- [Explicit Context-Driven Neural Acoustic Modeling for High-Fidelity RIR Generation](si26_interspeech.md) — complementary · relatedness 2.3/3
- [Towards Array-Invariant Speech Enhancement via Geometry-Aware Dynamic Convolution](liu26d_interspeech.md) — same problem · relatedness 2.2/3
- [Bridging the Distribution Gap in Real-World Far-Field Speech Enhancement via Lightweight Latent Representation Alignment](liu26s_interspeech.md) — same problem · relatedness 2.1/3
- [Scalable Audio Scene Generation with the Treble SDK](gotz26b_interspeech.md) — complementary · relatedness 2.1/3
- [A Novel Transfer Learning Approach for Room Impulse Response Estimation and Speech Dereverberation Across Geometrically Diverse and Data-Scarce Environments](pasha26_interspeech.md) — same problem · relatedness 2.0/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
