All papers
Enhancement & separationFull-paper digest

Dual-Geometry Manifolds for Few-shot RIR Prediction

Swapnil Bhosale, Gordon Wichern, Yoshiki Masuyama, Moitreya Chatterjee, Christoph Boeddeker, Julius Richter, Xiatian Zhu, Jonathan Le Roux

Code & resourcesgithub.com/merlresearch/janus-rir

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.9 KB · Ready to paste

Preview copied content

TL;DR — Janus-RIR is a temporally gated, dual-geometry aggregation model that matches latent acoustic geometry to physical room evolution, reducing late reverberation (T60T_{60}) error by 26% over Euclidean baselines.

Key contributions

  • Empirically proves using Gromov δ\delta-hyperbolicity that the acoustic manifold's geometry evolves dynamically from a hierarchical tree structure during early reflections to a flat Euclidean diffuse tail.
  • Proposes Janus-RIR, featuring parallel hyperbolic (Poincaré ball) and Euclidean aggregation branches to properly model early specular paths versus late statistical reverberation.
  • Introduces a context-aware dynamic signature gate conditioned on temporal embeddings, target queries, and reference-derived acoustic context to handle room-specific mixing times.
  • Achieves state-of-the-art performance on the AcousticRooms dataset, improving speech clarity (C50C_{50}) and reducing error across EDT, C50C_{50}, and T60T_{60} metrics.

Problem

Current cross-scene few-shot room impulse response (RIR) models (such as xRIR and Few-shot RIR) rely on static, zero-curvature Euclidean latent spaces for all temporal phases. This architectural choice conflicts with the physical reality of room acoustics: early RIR portions are discrete, specular reflection paths forming an exponentially growing tree structure, while late portions form a statistically uniform diffuse tail. Forcing Euclidean geometry onto hierarchical trees causes geometric smearing, destructively interfering with early acoustic transients and degrading speech clarity metrics like C50C_{50}.

Method

Janus-RIR aligns reference RIRs using Time-of-Flight delays relative to a target, converts them to log-magnitude spectrograms via a ResNet-18 backbone, and merges them with coordinate and temporal embeddings. The architecture splits features into two parallel branches: a Hyperbolic branch and a Euclidean branch.

The Hyperbolic branch maps Euclidean features to the Poincaré ball using the exponential map at the origin. It computes hyperbolic attention logits using scaled negative hyperbolic distances and weighted Einstein midpoints (Möbius gyromidpoints) to preserve the hierarchical geometry of early reflections. An entropy penalty mask with a time-decaying weight prevents the hyperbolic expert from over-averaging early branches. The Euclidean branch uses standard dot-product similarity and weighted sums to enable smooth statistical averaging suited for late reverberation.

A context-aware dynamic signature gate (GϕG_\phi) predicts a mixture coefficient α(t)\alpha(t) conditioned on the temporal embedding, target query, and an acoustic context vector obtained by max-pooling reference features. The outputs are combined via a product fusion layer, projecting back to RF\mathbb{R}^F. The network is trained end-to-end using L1L_1 spectrogram reconstruction, energy decay losses, and entropy regularization. During inference, time-domain RIR waveforms are reconstructed from predicted magnitude spectrograms using the Griffin-Lim algorithm.

Experimental setup

Evaluated on the AcousticRooms dataset containing 260 rooms across 10 categories, utilizing K∈{1,4,8}K \in \{1, 4, 8\} reference shots. Compared against Few-shot RIR (Euclidean cross-attention) and xRIR (visual-acoustic Euclidean baseline). Evaluated using Early Decay Time (EDT in seconds), Clarity (C50C_{50} in dB), and Reverberation Time (T60T_{60} percentage error). Includes both full-featured variants (with ViT-encoded depth maps) and 'NoVision' variants.

Results

At K=8K=8, the Janus-RIR (NoVision) model achieves a T60T_{60} error of 7.75%, representing a 26% relative error reduction over the Euclidean xRIR baseline (10.53%). This is accompanied by reductions in Early Decay Time error (0.055s to 0.045s) and Clarity error (1.457 dB to 1.126 dB). In spatial extrapolation difficulty tiers, Janus-RIR maintains tight error bounds for targets farther than 2.1m from a reference, whereas Euclidean baselines degrade severely.

Ablation studies show that forcing a purely hyperbolic space (α=1\alpha=1) degrades T60T_{60} error back to 10.52%, and replacing the hyperbolic branch with a second Euclidean branch worsens C50C_{50} error from 1.127 dB to 1.328 dB, proving that Euclidean space inherently smears early specular reflections. Removing acoustic context from the gating network degrades T60T_{60} error from 7.75% to 9.46%.

System / ConditionEDT (s) ↓\downarrowC50C_{50} (dB) ↓\downarrowT60T_{60} (%) ↓\downarrow
Few-shot RIR (K=8K=8) [1]0.1874.47021.15
xRIR (K=8K=8) [4]0.0551.45010.53
xRIR (NoVision, K=8K=8) [4]0.0501.49011.93
Janus-RIR (K=8K=8)0.0471.2007.97
Janus-RIR (NoVision, K=8K=8)0.0451.1207.75

Limitations

The evaluation is restricted to simulated environments within the AcousticRooms dataset, which may not capture all real-world acoustic anomalies or complex scattering phenomena. Inference relies on Griffin-Lim phase reconstruction from magnitude spectrograms, which can introduce phase artifacts. The approach currently relies on sparse spatial reference points and assumes adequate microphone distribution across the target room.

Why read this

Researchers working on spatial audio, few-shot acoustic modeling, or non-Euclidean representation learning will learn how to build mixed-curvature latent spaces that respect physical dynamics rather than enforcing static Euclidean assumptions.

Code

Applications

Spatial audio rendering, augmented reality acoustic simulation, binaural synthesis, and cross-scene room impulse response prediction.

Institutions

Mitsubishi Electric Research Laboratories, University of Surrey

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-2630