---
id: pratiwi26_interspeech
title: "Adaptation to Room Acoustics in Understanding Vocoded Speech: A
  Comparison Between Listeners With Varying Immersion Age"
authors:
  - Epri Pratiwi
  - C. T. Justine Hui
  - Yusuke Hioka
year: 2026
doi: 10.21437/Interspeech.2026-113
isca_url: https://www.isca-archive.org/interspeech_2026/pratiwi26_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/pratiwi26_interspeech.pdf
session: Assistive Technologies 1
topics:
  - speech-enhancement
  - paralinguistics
  - evaluation
category: health-clinical
labels:
  - robustness-noise
institutions:
  - University of Auckland
code:
  url: ""
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: pratiwi26_interspeech
  category: health-clinical
  labels:
    - robustness-noise
  institutions:
    - University of Auckland
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-113
  pdf: https://www.isca-archive.org/interspeech_2026/pratiwi26_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/pratiwi26_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/pratiwi26_interspeech/markdown.md
---

# Adaptation to Room Acoustics in Understanding Vocoded Speech: A Comparison Between Listeners With Varying Immersion Age

*Epri Pratiwi, C. T. Justine Hui, Yusuke Hioka*

[PDF](https://www.isca-archive.org/interspeech_2026/pratiwi26_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/pratiwi26_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-113)

**Category:** `health-clinical` · **Labels:** `robustness-noise`

**TL;DR** — This study investigates how listeners with varying language immersion ages adapt to room acoustics under cochlear-implant (CI) simulated vocoded speech, measuring both speech intelligibility and listening effort via pupillometry. Results show that while immersion age affects overall intelligibility of degraded speech, room acoustic adaptation rates and reductions in listening effort are comparable across all immersion groups.

## Key contributions

- Evaluated room acoustics adaptation in non-native and late-immersed listeners under simulated CI vocoded speech.
- Combined behavioral word-recognition tasks with physiological pupillometric measures (peak pupil dilation, PPD) to track listening effort dynamically across successive sentences.
- Implemented a 31-channel spherical loudspeaker array spatial sound reproduction system to preserve listener head-related transfer functions (HRTFs), avoiding the limitations of direct audio input.
- Demonstrated that while overall baseline performance depends on language immersion age, trial-by-trial room adaptation trajectories remain uniform across early-immersed, late-immersed, and non-immersed cohorts.

## Problem

Cochlear implant users struggle significantly in everyday reverberant environments because reverberation degrades speech temporal envelopes. Although normal-hearing native listeners can adapt to consistent room acoustics over blocked presentations, it remains unknown how linguistic background and immersion age influence this adaptation under CI processing. Existing studies largely ignore non-native listeners and rarely track the cognitive listening effort required alongside raw intelligibility scores.

## Method

Speech stimuli were drawn from the simplified University of Canterbury Auditory–Visual Matrix Sentence Test (UCAMST-P), consisting of 3-word pseudo-sentences ('Quantity + Adjective + Object'). Vocoded speech simulating CI processing was generated by decomposing signals into 8 mel-spaced frequency bands (400–7000 Hz) via FIR band-pass filters, extracting envelopes using the Hilbert transform, and modulating white noise carriers. Room acoustics were simulated using room impulse responses (RIRs) measured via a 32-channel spherical microphone array, decoded into 31-channel RIRs, and reproduced via a 31-channel Genelec loudspeaker array in an anechoic chamber.

Participants were seated centrally with a Tobii Pro Spark eye tracker monitoring pupils at 60 Hz. The experiment employed a blocked presentation design: 100 total trials comprising 10 consecutive sentences per acoustic environment across 5 conditions (anechoic, seminar room at 2m and 5m, chapel at 2m and 5m) for both non-vocoded and vocoded speech. Pupil traces were cleaned by blink interpolation (80 ms pre- to 150 ms post-blink), moving-average smoothed (117 ms window), baseline-corrected over a 2-s pre-stimulus interval, and normalized via 0.025 and 0.075 quantiles. Peak pupil dilation (PPD) served as the index of listening effort, while word-by-word correctness measured intelligibility. Statistical evaluations utilized linear mixed-effects models (lme4) with likelihood ratio tests and Tukey-corrected post-hoc comparisons.

## Experimental setup

51 adult participants divided into three immersion groups based on New Zealand English (NZE) exposure: early-immersed (n=17, age 28.0±5.7), late-immersed (n=14, age 39.0±6.7), and non-immersed (n=20, age 31.5±4.8). Acoustical spaces tested included anechoic, a seminar room (Volume 400 m³, RT 0.7s), and a chapel (Volume 1200 m³, RT 1.8s) at 2m and 5m source-listener distances. Metrics included proportion correct (intelligibility) and peak pupil dilation (PPD).

## Results

Speech intelligibility exhibited a significant three-way interaction between environment, speech type, and sentence order (chi-squared(36) = 146.57, p < 0.001), alongside an interaction between speech type and immersion group (chi-squared(2) = 43.48, p < 0.001). Under non-vocoded speech, intelligibility remained near ceiling across all environments, whereas vocoded speech showed strong dependence on acoustic context and sentence order without significant group-by-order interaction effects. PPD increased under vocoded speech and challenging room conditions, but progressively decreased across successive sentence blocks, indicating systematic reduction in cognitive listening effort as familiarity increased.

| Condition / Group | Intelligibility (Proportion Correct - Vocoded) | Peak Pupil Dilation (PPD - Vocoded) |
|---|---|---|
| Early Immersed (Seminar 2m) | ~0.65 - 0.85 | ~0.35 - 0.60 |
| Late Immersed (Seminar 2m) | ~0.50 - 0.75 | ~0.40 - 0.65 |
| Non-Immersed (Seminar 2m) | ~0.45 - 0.70 | ~0.45 - 0.70 |
| Early Immersed (Chapel 5m) | ~0.40 - 0.65 | ~0.45 - 0.68 |
| Non-Immersed (Chapel 5m) | ~0.30 - 0.55 | ~0.50 - 0.72 |

## Limitations

The study uses simulated CI vocoded speech on normal-hearing individuals rather than actual CI patients, bypassing patient-specific neural variability and device adaptation quirks. The vocabulary is restricted to a closed 18-word matrix structure (UCAMST-P), limiting generalization to natural, open-set conversational dialogue. Furthermore, the participant pool is limited to three specific NZE immersion brackets, and visual luminance was strictly controlled, which might not reflect dynamic real-world visual environments.

## Why read this

Read this paper if you design speech enhancement or CI signal processing algorithms and need to understand how human cognitive load and acoustic adaptation operate across diverse linguistic populations. It provides rare physiological evidence that adaptation rates to reverberant environments are robust to language background, even when absolute intelligibility baselines vary.

## Code

None released (as of this page's `updated` date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

## Applications

Cochlear implant speech processor design, adaptive hearing aids, and acoustic training software for second-language learners in noisy or reverberant classrooms.

## Institutions / 機構

University of Auckland

## Related

- [Relative Importance of Formants to the Intelligibility of Vocoded Speech in Cochlear Implant Simulation](cai26_interspeech.md) — same problem · relatedness 2.0/3
- [Deep learning-based predictions of perceived listening effort and intelligibility across enhanced, synthetic, natural, and binaural speech](hoffner26_interspeech.md) — same problem · relatedness 1.9/3
- [Bayesian Model-Based Assessment of Spatial and Source Priors in Sagittal-Plane Sound Localization](chen26fa_interspeech.md) — same problem · relatedness 1.8/3
- [Improving Cross-Dataset Speech Intelligibility Prediction for Hearing-Impaired Listeners with Few-Shot Adaptation](lin26h_interspeech.md) — same problem · relatedness 1.8/3
- [Steps toward a wearable-informed model of real-world listening effort and fatigue among adults with hearing loss](meng26e_interspeech.md) — same problem · relatedness 1.7/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
