TL;DR — This paper identifies temporal prediction inconsistency as the root cause of performance collapse in short-duration acoustic scene classification (ASC) and proposes a Semantic Adversarial Training (SAT) framework, improving 1-second segment classification accuracy to 63.5% on TAU20.
Key contributions
- Identified and formulated temporal prediction inconsistency as the core failure mode when restricting ASC models to short observation windows.
- Proposed the Local-Global Prediction Discrepancy (LGPD) metric, which uses Kullback-Leibler divergence between full-duration and short-duration predictions to quantify model sensitivity to context loss.
- Introduced the Semantic Adversarial Training (SAT) framework, leveraging a gradient reversal layer and an auxiliary event discriminator to suppress duration-dependent local event biases.
- Demonstrated consistent improvements over standard multi-task learning, knowledge distillation, and joint learning baselines, particularly on high-discrepancy unstable subsets.
Problem
Standard acoustic scene classification (ASC) models suffer from severe performance degradation when processing short observation windows (e.g., 1 second) instead of full-duration recordings (e.g., 10 seconds). Prior deep learning approaches relied on heavier architectures, external data pre-training, or transfer learning, but failed to address why models struggle specifically with short inputs. The underlying issue is that long-duration models exploit transient foreground events as discriminative shortcuts, whereas short windows introduce high semantic variability and temporal instability, causing standard models to yield highly inconsistent, overfitted predictions.
Method
The SAT framework is built around a standard convolutional feature extractor and a linear scene classifier optimized with cross-entropy loss on short segments . To counter event-shortcut exploitation, an auxiliary event discriminator with an output dimension of 527 estimates multi-label event probabilities via a multi-label binary cross-entropy loss, using pseudo event labels generated by a frozen BEATs model pre-trained on AudioSet.
A Gradient Reversal Layer (GRL) with a negative scalar is inserted between the feature extractor and the event discriminator. This creates a min-max adversarial game where the scene classifier minimizes classification error while the feature extractor maximizes the event discriminator's loss (scaled by ), forcing latent representations to discard local event variances while retaining global scene semantics. A random subset of 50% of each batch is utilized for adversarial optimization, the discriminator learning rate is scaled down by 0.01, the feature extractor undergoes 5 update steps per iteration, and the auxiliary event branch is completely discarded at inference time.
Training uses the TF-SepNet architecture with 150 epochs, batch size 512, Adam optimizer, and an exponential warm-up for 14 epochs followed by a linear decay to 0.5% of peak. Augmentations include Mixup, Freq-MixStyle, and DIR-Aug. Audio is resampled to 32 kHz and converted to log-Mel spectrograms using a 4096-point FFT, 96 ms window, 16 ms hop size, and 512 Mel bins.
Experimental setup
Evaluated on the DCASE Task 1 TAU Urban Acoustic Scenes 2020 Mobile Development Dataset (TAU20), which contains ~64 hours across 10 classes, using the official 7:3 train-test split with 10-second clips sliced into non-overlapping 1-second segments. Compared against baselines including TF-SepNet, Multi-Task Learning (MTL), Knowledge Distillation (KD), and Joint Learning (JTL). Metrics include classification accuracy (%) and the proposed LGPD metric to stratify test samples into stable and unstable subsets.
Results
On the 1-second TAU20 test set, SAT achieves an overall accuracy of 63.5%, outperforming the TF-SepNet baseline (59.6%), MTL (60.3%), KD (61.1%), and JTL (61.8%). On the hardest subset (unstable top-50% LGPD), SAT achieves 50.9% accuracy compared to the baseline's 43.1%, reducing the stability performance gap from 33.0 down to 25.2. In quantile-level evaluations, SAT yields a +12.0% absolute accuracy improvement over the baseline in the highest-discrepancy quantile (46.7% vs 34.7%), while incurring a negligible accuracy drop on the simplest stable samples (94.6% vs 95.8%).
| Method | Stable | Unstable | Gap (↓) | Overall |
|---|---|---|---|---|
| Baseline | 76.1 | 43.1 | 33.0 | 59.6 |
| + MTL [14] | 74.5 | 46.1 | 28.4 | 60.3 |
| + KD [23] | 75.1 | 47.2 | 27.9 | 61.1 |
| + JTL [13] | 75.4 | 48.3 | 27.1 | 61.8 |
| + SAT (Ours) | 76.2 | 50.9 | 25.2 | 63.5 |
Limitations
The framework relies on off-the-shelf pre-trained sound event taggers (like BEATs), which introduces potential label noise and external dependency. The evaluation is restricted to a single standard dataset (TAU20) and a single short-duration target window (1 second), leaving open how well the approach scales to diverse acoustic environments or varying sub-second granularities.
Why read this
Researchers and engineers working on low-latency environmental audio sensing or short-window acoustic scene classification will find a principled metric (LGPD) and adversarial formulation to prevent short-duration performance collapse.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Responsive edge-AI environmental sensing systems, real-time context-aware wearable devices, and smart-city surveillance.
Institutions
Xi'an Jiaotong-Liverpool University, Nanjing University of Posts and Telecommunications
Funding / 經費: Jiangsu Provincial Major Science and Technology Project
Related
- Branch-wise Complementary Attention for Acoustic Scene Classification — same problem · relatedness 2.7/3
- Towards Event-Robust Acoustic Scene Classification — same problem · relatedness 2.0/3
- Consistency-Regularized Dual-Branch Network with Performance-Aware Mean Teacher for Sound Event Detection — shared technique · relatedness 1.8/3
- Robust Multi-Source-Free Domain Adaptation via Posterior Adjustment and Label Agreement — same problem · relatedness 1.8/3
- A Closer Look at Failure Modes in Temporal Understanding of Large Audio-Language Models — same problem · relatedness 1.7/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-1955