TL;DR — The paper introduces YT-SPEECH, an 8.9-hour dataset of speech-dominant 360-degree video and First-Order Ambisonics (FOA) audio, alongside a Localizer-Renderer framework that achieves superior spatial accuracy and speech quality via confidence-gated audio-visual priors.
Key contributions
- YT-SPEECH: The first speech-oriented 360-degree video-FOA dataset (8.9 hours, 197 source videos, 24 kHz) curated via multi-stage filtering for multi-channel audio and on-screen speaker activity.
- A Localizer-Renderer framework that utilizes an audio-visual segmentation (AVS) backbone to produce fine-grained spatial heatmaps for direction-consistent FOA reconstruction.
- A confidence-based gating mechanism (combining peak concentration and entropy) and Feature-wise Linear Modulation (FiLM) to dynamically adapt conditioning strength under acoustically ambiguous conditions.
- Comprehensive validation demonstrating consistent improvements in reconstruction fidelity, spatial angular error (∆ang), and perceptual speech quality (PESQ, MOS-P).
Problem
Recovering high-quality spatial First-Order Ambisonics (FOA) audio from unconstrained 360-degree video and omnidirectional microphone recordings remains challenging due to a lack of paired real-world datasets and suboptimal prior methods. Existing spatial audio datasets are predominantly binaural or rely on synthetic/simulated labels rather than real panoramic video cues. Meanwhile, prior explicit spatial reconstruction approaches (like SpatialAudioGen) utilize self-supervised label-free separation that introduces acoustic artifacts, and end-to-end models prioritize semantic consistency over precise spatial localization.
Method
The proposed architecture features a two-stage Localizer-Renderer framework operating in the complex short-time Fourier transform (STFT) domain. Given an equirectangular projection (ERP) video frame and an omnidirectional FOA channel , a fine-tuned Audio-Visual Segmentation (AVS) backbone acts as the Localizer to output a dense 7x14 normalized spatial heatmap , using circular padding along ERP dimensions to preserve horizontal wrap-around continuity.
The Renderer consists of a complex-domain U-Net that takes the omnidirectional spectrum and predicts complex-valued masks for missing directional channels (). To handle noisy or diffuse visual frames, a frame-level confidence gate is derived from the spatial prior's peak concentration and entropy. This scalar gate computes modulation parameters via Feature-wise Linear Modulation (FiLM) applied across the U-Net decoder blocks to balance visual guidance and audio evidence adaptively.
The training recipe involves pretraining the Renderer on Sphere360 data using AdamW (learning rate ), followed by joint fine-tuning on YT-SPEECH (Localizer lr , Renderer lr ). Data augmentation includes random horizontal yaw rotations () and horizontal flipping applied with probabilities of 0.8 and 0.2, respectively, with corresponding FOA channel rotations. The composite loss function combines multi-resolution STFT loss (), magnitude loss (), and waveform loss, weighted by confidence scores.
Experimental setup
Evaluations utilize the newly curated YT-SPEECH dataset (8.9 hours total, 5-second clips at 24 kHz) split into train, validation, and test sets. The method is compared against ablation variants (NoVideo-Renderer, VidEnc-Renderer, FrozenLoc-Renderer, Ours-NoPT), analytic baselines (Localizer-AmbiEnc, Localizer-Pyroom), and SpatialAudioGen (SAG) on benchmark datasets including YT-ALL, YT-MUSIC, YT-CLEAN, and YT-SPEECH. Metrics include reconstruction errors (, , , ), DOA spatial errors (, , ), ear-wise PESQ (with ), and subjective MOS scores (MOS-Q for audio quality, MOS-P for spatial accuracy) evaluated by 9 human listeners.
Results
On the YT-SPEECH test set, the full proposed model achieves the lowest reconstruction error (), minimal angular error (), and top-tier speech quality with a PESQ of 3.42 and MOS-P of 3.16. Compared to analytic baselines like Localizer-AmbiEnc, the learned Renderer offers substantial gains in spatial accuracy (reducing from to ) and subjective spatial preference. In compatibility comparisons against SpatialAudioGen (SAG) across YT-ALL, YT-MUSIC, YT-CLEAN, and YT-SPEECH, the proposed model consistently outperforms SAG on complex STFT and envelope (ENV) distance metrics, showing the most pronounced advantages on visually clean speech (YT-SPEECH STFT: 0.85 vs SAG's 1.14; YT-CLEAN STFT: 0.91 vs SAG's 1.58). The framework does not win as decisively on acoustically diverse open-domain datasets like YT-ALL where speech is less dominant, and struggles in outdoor or highly noisy backgrounds containing overlapping sound sources which trigger spurious spatial activations.
| Model | () | PESQ | MOS-Q | MOS-P | |
|---|---|---|---|---|---|
| NoVideo-Renderer | 1.59 | 0.73 | 2.50 | 3.24 | 2.28 |
| VidEnc-Renderer | 1.95 | 0.66 | 1.78 | 3.03 | 3.02 |
| Localizer-AmbiEnc | 1.55 | 0.83 | 3.35 | 3.68 | 2.97 |
| Ours (NoPT) | 1.71 | 0.92 | 1.48 | 2.94 | 2.90 |
| Ours (Full) | 1.15 | 0.64 | 3.42 | 3.63 | 3.16 |
Limitations
The YT-SPEECH dataset is relatively small at 8.9 hours, constraining the diversity of learned acoustic interactions. The system exhibits reduced stability and produces spurious heatmap activations when handling overlapping multiple sound sources, complex reverberation, or heavy background noise where visual grounding becomes ambiguous.
Why read this
Speech and audio researchers working on immersive 360-degree video, ambisonic spatialization, or audio-visual multi-modal learning should read this to see how audio-visual segmentation priors can be effectively coupled with complex-domain U-Nets via confidence gating.
Code
Applications
Immersive telepresence, virtual reality (VR) video players, panoramic streaming platforms, and interactive 360-degree media applications.
Institutions
University of Surrey
Funding / 經費: Bang & Olufsen A/S, AURIC
Related
- Sweep-RSE: Streaming Region-of-Interest Speech Extraction in Multi-Talker Scenarios via Explicit Spatial Sweeping — same problem · relatedness 2.1/3
- Online Audiovisual Speaker Separation Using Efficient Visual Knowledge Distillation — same problem · relatedness 2.0/3
- Multi-View Based Audio Visual Target Speaker Extraction — same problem · relatedness 2.0/3
- SPOT-TSE: Spatial Point-Guided Target Speech Extraction — same problem · relatedness 2.0/3
- Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement — same problem · relatedness 2.0/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-2577