All papers
Audio understandingFull-paper digest

CoSTALA: Compositional Spatio-Temporal Audio-Language Alignment via Multi-Grain Hierarchical Contrastive Learning

Peiwei Ren, Jinbo Hu, Fang Kang, Shan Liang, Yin Cao

Code & resourcesgithub.com/Cell778/CoSTALA26.git

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.8 KB · Ready to paste

Preview copied content

TL;DR — CoSTALA is an audio-language alignment training paradigm designed to transition from coarse global representations to fine-grained spatio-temporal reasoning for multi-event spatial audio. It achieves an absolute text-to-audio Recall@1 of 8.10% (vs. 5.66% for T-CLAP) on a newly constructed spatial audio-text benchmark.

Key contributions

  • Proposes a multi-grain hierarchical contrastive learning framework that replaces single global vectors with a combination of hard, soft, and global embeddings.
  • Introduces a 3-way spatio-temporal loss that jointly penalizes chronological (temporal) errors and spatial localization errors.
  • Applies a local alignment loss and a feature consistency loss (with stop-gradient) to prevent temporal entanglement and semantic drift in long-duration audio streams.
  • Constructs a 375-hour multi-event spatial dataset derived from Clotho via LLM-assisted (Qwen3-8B) spatial caption rewriting and simulated First-Order Ambisonics (FOA) SRIRs.

Problem

Traditional audio language models (ALMs) like CLIP, CLAP, LAION-CLAP, and SALM rely heavily on static, global alignment which compresses dynamic sound streams into holistic vectors. This causes 'context drift' and information bottlenecks, blinding models to multi-event audio sequences where distinct sound occurrences unfold across different spatial coordinates and time steps. Prior temporal approaches like T-CLAP handle chronological sequences but are limited to two-channel audio and lack spatial awareness, failing to parse multi-event, multi-directional scenes.

Method

CoSTALA integrates a RoBERTa-based text encoder with a hierarchical audio backbone driven by HTSAT, augmented by a dedicated Transformer temporal encoder equipped with Rotated Position Embedding (RoPE) and a dual-branch design separating acoustic semantics from spatial localization. The architecture processes audio via three pathways: Hard Embeddings (EhardciE_{hard}^{c_i}) from isolated spatial events to preserve semantic purity, Soft Embeddings (EsoftciE_{soft}^{c_i}) from chunked representations of concatenated sequences (C1∥C2C_1 \parallel C_2), and Global Embeddings (EglobalE_{global}) for macroscopic features. The Soft embeddings feed the temporal encoder to extract chronological dependencies (EtempE_{temp}), which are then combined with global features via a learnable residual scalar α\alpha to yield spatio-temporal representations (EstE_{st}).

The training objective is a weighted linear combination of four losses: (1) Contrastive Learning Loss (LclL_{cl}) combining semantic and macroscopic pairs against batch-mined hard negatives; (2) A 3-way Spatio-Temporal Loss (LstL_{st}) utilizing a learnable temperature parameter τ\tau to explicitly force text and audio representations to distinguish temporal reversals from spatial swapping; (3) Local Alignment Loss (LlocalL_{local}) using InfoNCE on isolated hard chunks and directional spatial captions; and (4) Feature Consistency Loss (LconsistL_{consist}) applying an MSE constraint between hard anchors (with stop-gradients) and soft chunk representations to prevent representational collapse.

Experimental setup

The model is trained on a 375-hour synthetic spatial dataset containing 30,000 training samples and 9,000 evaluation samples sampled at 24 kHz (64-dimensional log-mel spectrograms and intensity vectors from First-Order Ambisonics FOA using a 1,024-point Hanning window). Evaluated against SALM and T-CLAP using bi-directional spatial retrieval metrics (Recall@1, Recall@5, Recall@10). Trained for 15 epochs using the AdamW optimizer with a peak learning rate of 10−410^{-4}, a 3-epoch linear warm-up, and a cosine annealing schedule.

Results

CoSTALA achieves an 8.10% Text-to-Audio R@1 (compared to 5.66% for T-CLAP and 4.77% for SALM) and an 8.16% Audio-to-Text R@1 (compared to 5.92% for T-CLAP and 4.57% for SALM). Ablation studies show that removing the 3-way spatio-temporal loss drops Text-to-Audio R@1 down to 7.16%, while removing local alignment or feature consistency individually degrades performance. A zero-shot test using only semantic audio representations (EsemE_{sem}) without spatial components collapses performance, underscoring the necessity of the spatial pathway.

System / Loss ConfigT2A R@1T2A R@5T2A R@10A2T R@1A2T R@5A2T R@10
SALM4.7714.3721.274.5714.2821.23
T-CLAP5.6616.8424.885.9216.7124.71
CoSTALA (LclL_{cl})5.8417.6826.076.5818.4926.01
CoSTALA (Lcl+LstL_{cl}+L_{st})7.1617.4524.526.9617.7823.84
CoSTALA (Full)8.1019.8627.688.1620.4927.08

Limitations

The dataset is entirely synthetic, created by convolving monophonic Clotho audio with simulated Spatial Room Impulse Responses (SRIRs), which may limit zero-shot generalization to complex real-world acoustic environments and reverberation profiles. The spatial resolution is restricted to eight azimuth directions at 45-degree intervals, bypassing continuous 360-degree elevation and distance modeling.

Why read this

Speech and ML researchers working on spatial audio understanding and multi-modal contrastive learning should read this paper to learn how to combine fine-grained local anchoring with global multi-event temporal loss objectives.

Code

Applications

Complex acoustic scene monitoring, spatially aware conversational agents, and multi-event augmented reality sound indexing.

Institutions

Xi'an Jiaotong-Liverpool University, Xiaomi, University of Oulu, Chinese Academy of Sciences

Funding / 經費: Xi'an Jiaotong-Liverpool University

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1110