All papers
Applications & otherFull-paper digest

HRTF Personalization via Sim-to-Real Neural Field

Yoshiki Masuyama, Gordon Wichern, Christoph Boeddeker, Julius Richter, Takahiro Edo, Swapnil Bhosale, Jonathan Le Roux

Code & resourcesgithub.com/merlresearch/s2rnf

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

11.9 KB · Ready to paste

Preview copied content

TL;DR — The paper introduces Sim-to-Real Neural Field (S2RNF), a measurement-free HRTF personalization method that maps physics-simulated head-related transfer functions to realistic ones, outperforming baseline simulations and prior neural refinement models.

Key contributions

  • Proposes S2RNF, incorporating domain-specific parameters alongside subject-specific parameters into a neural field to bridge the domain gap between simulated and real HRTFs.
  • Eliminates the need for specialized anechoic chamber measurements by inferring subject-specific latent parameters directly from simulated HRTFs computed on 3D head meshes.
  • Leverages the Implicit Gradient Origin Network (IGON) framework for one-step gradient-based encoding, removing the need for a separate encoder network.
  • Demonstrates superior spectral distortion metrics compared to previous sim-to-real and anthropometry-based baselines on both the SONICOM and HUTUBS datasets.

Problem

Individual head-related transfer functions (HRTFs) are essential for high-fidelity immersive binaural audio, but acquiring them requires tedious physical measurements inside specialized anechoic chambers. While alternative measurement-free approaches attempt to simulate HRTFs from 3D meshes or predict them from anthropometric features using deep learning, simulations suffer from severe high-frequency artifacts. Prior deep learning models also require extra encoder architectures or manual feature curation, highlighting the need for robust sim-to-real transfer mechanisms.

Method

S2RNF models the log-magnitude response of HRTFs using a neural field architecture that takes random Fourier features (RFF) of the sound source direction as input. The network comprises fully connected layers with BitFit-style conditioning, injecting subject-specific biases (z_i) and domain-specific biases (w_j, where j=0 for simulated and j=1 for measured) into the hidden core blocks.

During inference, subject-specific parameters z_i are extracted directly from simulated HRTFs via one-step gradient descent (IGON) using domain parameters w_0. Realistic HRTFs are then synthesized by switching the domain parameter to w_1. The model is trained end-to-end by jointly minimizing a loss on predicted realistic HRTFs (L_1) and simulated HRTFs (L_0) with a balancing weight of lambda = 1, ensuring the computational graph flows cleanly through the gradient-based parameter extraction step.

Experimental setup

Evaluated on an extended SONICOM dataset (200 pairs of measured/simulated HRTFs at 48 kHz across 793 directions, split into 160/15/25 for train/val/test) and the HUTUBS dataset (85 training, 5 validation, 6 test subjects, 440 directions). Compared against raw Mesh2HRTF simulations, non-personalized average baselines, SPCA-DNN, BEM-DNN, and Proto. DNN. Metrics include Mean Absolute Error (MAE), Root Mean Squared Error (RMSE) / Log-Spectral Distortion (LSD), and Polar RMSE (PolRMSE). Implemented with 4 hidden layers of 512 units using Swish activations, optimized via RAdam at a constant learning rate of 0.0005 for 1500 epochs.

Results

On the SONICOM dataset, S2RNF achieves an RMSE of 4.90 dB and MAE of 3.60 dB, outperforming the non-personalized average baseline (RMSE 5.05 dB, MAE 3.78 dB) and raw Mesh2HRTF simulation (RMSE 10.53 dB, MAE 7.50 dB). Ablation studies confirm that removing the simulation loss L_0 degrades RMSE to 4.98 dB. On the HUTUBS dataset, S2RNF achieves an MAE of 3.66 dB and RMSE of 5.02 dB, beating BEM-DNN (4.80 dB MAE) and performing competitively with Proto. DNN (3.69 dB MAE) while requiring no dedicated encoder network.

SystemMAE [dB]RMSE [dB]PolRMSE [degree]
Mesh2HRTF (SONICOM)7.5010.5342.27
Avg. HRTF (SONICOM)3.785.0541.63
S2RNF (SONICOM)3.604.9040.89
BEM-DNN (HUTUBS)4.80--
Proto. DNN (HUTUBS)3.695.1032.90
S2RNF (HUTUBS)3.665.0233.22

Limitations

The evaluation relies on pre-computed official simulated HRTFs derived from high-fidelity 3D head meshes rather than noisy meshes scanned in everyday environments. The study focuses purely on log-magnitude responses and minimum-phase reconstructions, deferring individual interaural time difference (ITD) personalization to future work.

Why read this

Speech and audio researchers building accessible immersive spatial audio pipelines will find this a clean blueprint for bridging the gap between physics-based acoustic simulators and real-world neural representations without requiring manual anthropometric data or physical lab sessions.

Code

Applications

Virtual reality, augmented reality, immersive gaming, and personal audio listening devices requiring customized binaural sound rendering.

Institutions

Mitsubishi Electric Research Laboratories, University of Surrey

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1391