All papers
Audio understandingFull-paper digest

Enhancing Temporal Prediction Consistency for Short-Duration Acoustic Scene Classification via Semantic Adversarial Training

Yiqiang Cai, Yizhou Tan, Peihong Zhang, Yuxuan Liu, Shengchen Li, Xi Shao

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.9 KB · Ready to paste

Preview copied content

TL;DR — This paper identifies temporal prediction inconsistency as the root cause of performance collapse in short-duration acoustic scene classification (ASC) and proposes a Semantic Adversarial Training (SAT) framework, improving 1-second segment classification accuracy to 63.5% on TAU20.

Key contributions

  • Identified and formulated temporal prediction inconsistency as the core failure mode when restricting ASC models to short observation windows.
  • Proposed the Local-Global Prediction Discrepancy (LGPD) metric, which uses Kullback-Leibler divergence between full-duration and short-duration predictions to quantify model sensitivity to context loss.
  • Introduced the Semantic Adversarial Training (SAT) framework, leveraging a gradient reversal layer and an auxiliary event discriminator to suppress duration-dependent local event biases.
  • Demonstrated consistent improvements over standard multi-task learning, knowledge distillation, and joint learning baselines, particularly on high-discrepancy unstable subsets.

Problem

Standard acoustic scene classification (ASC) models suffer from severe performance degradation when processing short observation windows (e.g., 1 second) instead of full-duration recordings (e.g., 10 seconds). Prior deep learning approaches relied on heavier architectures, external data pre-training, or transfer learning, but failed to address why models struggle specifically with short inputs. The underlying issue is that long-duration models exploit transient foreground events as discriminative shortcuts, whereas short windows introduce high semantic variability and temporal instability, causing standard models to yield highly inconsistent, overfitted predictions.

Method

The SAT framework is built around a standard convolutional feature extractor F(⋅∣θf)F(\cdot|\theta_f) and a linear scene classifier H(⋅∣θh)H(\cdot|\theta_h) optimized with cross-entropy loss on short segments xx. To counter event-shortcut exploitation, an auxiliary event discriminator D(⋅∣θd)D(\cdot|\theta_d) with an output dimension of 527 estimates multi-label event probabilities via a multi-label binary cross-entropy loss, using pseudo event labels generated by a frozen BEATs model pre-trained on AudioSet.

A Gradient Reversal Layer (GRL) with a negative scalar −αadv-\alpha_{\text{adv}} is inserted between the feature extractor and the event discriminator. This creates a min-max adversarial game where the scene classifier minimizes classification error while the feature extractor maximizes the event discriminator's loss (scaled by λadv=5\lambda_{\text{adv}} = 5), forcing latent representations zz to discard local event variances while retaining global scene semantics. A random subset of 50% of each batch is utilized for adversarial optimization, the discriminator learning rate is scaled down by 0.01, the feature extractor undergoes 5 update steps per iteration, and the auxiliary event branch is completely discarded at inference time.

Training uses the TF-SepNet architecture with 150 epochs, batch size 512, Adam optimizer, and an exponential warm-up for 14 epochs followed by a linear decay to 0.5% of peak. Augmentations include Mixup, Freq-MixStyle, and DIR-Aug. Audio is resampled to 32 kHz and converted to log-Mel spectrograms using a 4096-point FFT, 96 ms window, 16 ms hop size, and 512 Mel bins.

Experimental setup

Evaluated on the DCASE Task 1 TAU Urban Acoustic Scenes 2020 Mobile Development Dataset (TAU20), which contains ~64 hours across 10 classes, using the official 7:3 train-test split with 10-second clips sliced into non-overlapping 1-second segments. Compared against baselines including TF-SepNet, Multi-Task Learning (MTL), Knowledge Distillation (KD), and Joint Learning (JTL). Metrics include classification accuracy (%) and the proposed LGPD metric to stratify test samples into stable and unstable subsets.

Results

On the 1-second TAU20 test set, SAT achieves an overall accuracy of 63.5%, outperforming the TF-SepNet baseline (59.6%), MTL (60.3%), KD (61.1%), and JTL (61.8%). On the hardest subset (unstable top-50% LGPD), SAT achieves 50.9% accuracy compared to the baseline's 43.1%, reducing the stability performance gap from 33.0 down to 25.2. In quantile-level evaluations, SAT yields a +12.0% absolute accuracy improvement over the baseline in the highest-discrepancy quantile (46.7% vs 34.7%), while incurring a negligible accuracy drop on the simplest stable samples (94.6% vs 95.8%).

MethodStableUnstableGap (↓)Overall
Baseline76.143.133.059.6
+ MTL [14]74.546.128.460.3
+ KD [23]75.147.227.961.1
+ JTL [13]75.448.327.161.8
+ SAT (Ours)76.250.925.263.5

Limitations

The framework relies on off-the-shelf pre-trained sound event taggers (like BEATs), which introduces potential label noise and external dependency. The evaluation is restricted to a single standard dataset (TAU20) and a single short-duration target window (1 second), leaving open how well the approach scales to diverse acoustic environments or varying sub-second granularities.

Why read this

Researchers and engineers working on low-latency environmental audio sensing or short-window acoustic scene classification will find a principled metric (LGPD) and adversarial formulation to prevent short-duration performance collapse.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Responsive edge-AI environmental sensing systems, real-time context-aware wearable devices, and smart-city surveillance.

Institutions

Xi'an Jiaotong-Liverpool University, Nanjing University of Posts and Telecommunications

Funding / 經費: Jiangsu Provincial Major Science and Technology Project

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1955