---
id: silverman26_interspeech
title: Learning Self-Supervised Spatial Representations via Soft Acoustic
  Contrastive Alignment
authors:
  - Yotam Silverman
  - Bracha Laufer-Goldshtein
year: 2026
doi: 10.21437/Interspeech.2026-641
isca_url: https://www.isca-archive.org/interspeech_2026/silverman26_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/silverman26_interspeech.pdf
session: Spatial Audio 4
topics:
  - self-supervised
  - spatial-audio
  - speech-enhancement
category: enhancement-separation
labels:
  - self-supervised
institutions:
  - Tel Aviv University
funding:
  - Israel Science Foundation
  - Israeli Ministry of Innovation, Science and Technology
code:
  url: https://github.com/Yotamsil/SAC_Alignment
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: silverman26_interspeech
  category: enhancement-separation
  labels:
    - self-supervised
  institutions:
    - Tel Aviv University
  code: https://github.com/Yotamsil/SAC_Alignment
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-641
  pdf: https://www.isca-archive.org/interspeech_2026/silverman26_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/silverman26_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/silverman26_interspeech/markdown.md
---

# Learning Self-Supervised Spatial Representations via Soft Acoustic Contrastive Alignment

*Yotam Silverman, Bracha Laufer-Goldshtein*

[PDF](https://www.isca-archive.org/interspeech_2026/silverman26_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/silverman26_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-641)

**Category:** `enhancement-separation` · **Labels:** `self-supervised`

**TL;DR** — This paper introduces the Soft Acoustic Contrastive (SAC) loss, a self-supervised objective that aligns the latent space of binaural audio encoders with analytical acoustic priors to improve room parameter estimation. Combining SAC with cross-channel signal reconstruction reduces mean absolute error across five spatial tasks compared to reconstruction-only baselines.

## Key contributions

- Proposes a geometry-aware soft contrastive loss (LSAC) for binaural self-supervised learning using continuous similarity targets derived from physical acoustic features.
- Eliminates the need for manual spatial metadata, ground-truth labels, or isolated source signals during pre-training.
- Integrates seamlessly with existing multi-channel masked reconstruction tasks (CCSR), structuring the latent manifold for downstream regression.
- Demonstrates consistent performance gains across five distinct spatial parameter estimation tasks under both linear evaluation and full fine-tuning protocols.

## Problem

Current supervised spatial estimation methods rely on large-scale labeled datasets or true room impulse response recordings, which are expensive and difficult to scale. Existing self-supervised learning approaches focus on single-channel semantic features or general reconstruction, treating acoustic reflections and channel differences as nuisance factors rather than extracting critical spatial information. Prior multi-channel methods like Cross-Channel Signal Reconstruction (CCSR) implicitly learn physical dependencies but fail to explicitly structure the latent space around physical room properties, limiting their generalization on spatial tasks.

## Method

The framework adopts a dual-stream Multi-Channel Conformer (MC-Conformer) architecture containing a spectral encoder (1 Conformer block, Dspec=512) and a spatial encoder (3 Conformer blocks, Dspat=256), totaling approximately 17.6M parameters. During pre-training, input binaural STFT spectrograms undergo a 50% channel-wise masking strategy to force the spatial encoder to capture inter-channel differences for a Mean Squared Error (MSE) signal reconstruction loss. In parallel, a projection head maps spatial embeddings to a 128-dimensional latent space optimized via the Soft Acoustic Contrastive (SAC) loss. The SAC loss uses a Gaussian kernel with temperature scaling (tau = 0.15, sigma = 1.0) to softly align latent similarities with 5 analytical acoustic features extracted via classical signal processing: TDOA (via GCC-PHAT), GCC-PHAT peak magnitude, and frequency-band spectral coherence across Low (<500 Hz), Mid (500-2500 Hz), and High (>2500 Hz) bands.

Pre-training runs for 30 epochs with a batch size of 128 using the Adam optimizer and a cosine-decay learning rate schedule starting at 1e-3. The loss balances reconstruction and SAC alignment equally (lambda = 1). For downstream evaluation, the decoder, spectral encoder, and projection head are discarded, and a linear prediction head is attached to the frozen or fine-tuned spatial encoder. Downstream training uses a batch size of 8, Adam optimizer, and learning rate plateaus patience of 10 epochs, with final models constructed via ensembling the best checkpoint and its 4 preceding epochs.

## Experimental setup

The pre-training dataset consists of 50,000 unlabeled two-channel audio samples generated by convolving Wall Street Journal clean speech with simulated RIRs (room sizes 3x3x2.5m to 15x10x6m, T60 in [0.2, 1.0] s, microphone spacing 3-20 cm, and diffuse white noise at 15-30 dB SNR). Downstream evaluation uses non-overlapping splits comprising 8, 20, and 20 rooms with 100, 50, and 200 utterances respectively across 4 independent trials. Baselines include CCSR, an unmasked Autoencoder variant with LSAC, standalone LSAC, and a 'No Pretrain' model. Evaluation metrics are Mean Absolute Error (MAE) across five tasks: TDOA, T60, Direct-to-Reverberant Ratio (DRR), Clarity Index (C50), and Mean Absorption Coefficient (ABS).

## Results

Under linear evaluation, combining CCSR with LSAC yields an MAE of 1.06 samples for TDOA and 0.113 seconds for T60, outperforming the CCSR baseline (1.88 samples and 0.174 s). For full fine-tuning, CCSR+LSAC achieves 0.288 samples for TDOA and 0.075 s for T60, compared to CCSR's 0.404 samples and 0.083 s. Ablations confirm that combining masked reconstruction with the soft acoustic contrastive loss consistently outperforms unmasked autoencoder variants or standalone losses. However, the method does not consistently outperform supervised or un-pretrained baselines on simpler spectral metrics like the Mean Absorption Coefficient (ABS), where the 'No Pretrain' model achieves competitive error rates (0.063 vs 0.067 ABS).

| Protocol | Model | TDOA (sample) | T60 (s) | DRR (dB) | C50 (dB) | ABS |
|---|---|---|---|---|---|---|
| Linear Eval | CCSR (Baseline) | 1.88 | 0.174 | 2.67 | 1.63 | 0.075 |
| Linear Eval | CCSR + LSAC (Ours) | 1.06 | 0.113 | 1.86 | 0.998 | 0.067 |
| Fine-Tune | CCSR (Baseline) | 0.404 | 0.083 | 2.18 | 0.662 | 0.079 |
| Fine-Tune | LSAC (Ours) | 0.314 | 0.079 | 1.70 | 0.611 | 0.064 |
| Fine-Tune | CCSR + LSAC (Ours) | 0.288 | 0.075 | 1.72 | 0.580 | 0.066 |

## Limitations

The evaluation relies entirely on simulated room impulse responses, lacking validation on real-world acoustic recordings with complex reflections and hardware imperfections. The effectiveness of the SAC loss is strictly bound to the reliability of upstream analytical feature extractors like GCC-PHAT, which degrade in severely adverse environments with very low signal-to-noise ratios. The scope is limited to binaural two-channel audio configurations and standard room acoustic parameter estimation tasks.

## Why read this

Researchers and engineers working on spatial audio representation learning or blind room acoustic parameter estimation should read this paper to learn how to incorporate continuous, physics-guided soft contrastive alignment into multi-channel self-supervised pipelines.

## Code

- https://github.com/Yotamsil/SAC_Alignment

## Applications

Autonomous robotics, immersive AR/VR spatial audio rendering, and advanced hearing aid algorithms requiring blind room acoustic parameter estimation.

## Institutions / 機構

Tel Aviv University

**Funding / 經費:** Israel Science Foundation, Israeli Ministry of Innovation, Science and Technology

## Related

- [Probing Spatial Structure in Pretrained Audio Representations](chen26ba_interspeech.md) — same problem · relatedness 2.3/3
- [Blind Room Impulse Response Identification via Reverberant Speech Spectrum Reconstruction](wang26c_interspeech.md) — same problem · relatedness 2.1/3
- [Dual-Geometry Manifolds for Few-shot RIR Prediction](bhosale26b_interspeech.md) — same problem · relatedness 2.1/3
- [Quantifying the Uncertainty of Blindly Estimated Room Embeddings Using a Dispersion-Calibrated Score](xiang26_interspeech.md) — same problem · relatedness 2.1/3
- [Beyond Cross-Reconstruction: Probing-Based Disentanglement Evaluation for Acoustic Teleportation Codecs](grundhuber26_interspeech.md) — same problem · relatedness 2.0/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
