---
id: cai26b_interspeech
title: Enhancing Temporal Prediction Consistency for Short-Duration Acoustic
  Scene Classification via Semantic Adversarial Training
authors:
  - Yiqiang Cai
  - Yizhou Tan
  - Peihong Zhang
  - Yuxuan Liu
  - Shengchen Li
  - Xi Shao
year: 2026
doi: 10.21437/Interspeech.2026-1955
isca_url: https://www.isca-archive.org/interspeech_2026/cai26b_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/cai26b_interspeech.pdf
session: Acoustic Event Detection 2
topics:
  - acoustic-scene-classification
  - self-supervised
  - evaluation
category: audio-understanding
institutions:
  - Xi'an Jiaotong-Liverpool University
  - Nanjing University of Posts and Telecommunications
funding:
  - Jiangsu Provincial Major Science and Technology Project
code:
  url: ""
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: cai26b_interspeech
  category: audio-understanding
  institutions:
    - Xi'an Jiaotong-Liverpool University
    - Nanjing University of Posts and Telecommunications
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-1955
  pdf: https://www.isca-archive.org/interspeech_2026/cai26b_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/cai26b_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/cai26b_interspeech/markdown.md
---

# Enhancing Temporal Prediction Consistency for Short-Duration Acoustic Scene Classification via Semantic Adversarial Training

*Yiqiang Cai, Yizhou Tan, Peihong Zhang, Yuxuan Liu, Shengchen Li, Xi Shao*

[PDF](https://www.isca-archive.org/interspeech_2026/cai26b_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/cai26b_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-1955)

**Category:** `audio-understanding`

**TL;DR** — This paper identifies temporal prediction inconsistency as the root cause of performance collapse in short-duration acoustic scene classification (ASC) and proposes a Semantic Adversarial Training (SAT) framework, improving 1-second segment classification accuracy to 63.5% on TAU20.

## Key contributions

- Identified and formulated temporal prediction inconsistency as the core failure mode when restricting ASC models to short observation windows.
- Proposed the Local-Global Prediction Discrepancy (LGPD) metric, which uses Kullback-Leibler divergence between full-duration and short-duration predictions to quantify model sensitivity to context loss.
- Introduced the Semantic Adversarial Training (SAT) framework, leveraging a gradient reversal layer and an auxiliary event discriminator to suppress duration-dependent local event biases.
- Demonstrated consistent improvements over standard multi-task learning, knowledge distillation, and joint learning baselines, particularly on high-discrepancy unstable subsets.

## Problem

Standard acoustic scene classification (ASC) models suffer from severe performance degradation when processing short observation windows (e.g., 1 second) instead of full-duration recordings (e.g., 10 seconds). Prior deep learning approaches relied on heavier architectures, external data pre-training, or transfer learning, but failed to address why models struggle specifically with short inputs. The underlying issue is that long-duration models exploit transient foreground events as discriminative shortcuts, whereas short windows introduce high semantic variability and temporal instability, causing standard models to yield highly inconsistent, overfitted predictions.

## Method

The SAT framework is built around a standard convolutional feature extractor $F(\cdot|\theta_f)$ and a linear scene classifier $H(\cdot|\theta_h)$ optimized with cross-entropy loss on short segments $x$. To counter event-shortcut exploitation, an auxiliary event discriminator $D(\cdot|\theta_d)$ with an output dimension of 527 estimates multi-label event probabilities via a multi-label binary cross-entropy loss, using pseudo event labels generated by a frozen BEATs model pre-trained on AudioSet.

A Gradient Reversal Layer (GRL) with a negative scalar $-\alpha_{\text{adv}}$ is inserted between the feature extractor and the event discriminator. This creates a min-max adversarial game where the scene classifier minimizes classification error while the feature extractor maximizes the event discriminator's loss (scaled by $\lambda_{\text{adv}} = 5$), forcing latent representations $z$ to discard local event variances while retaining global scene semantics. A random subset of 50% of each batch is utilized for adversarial optimization, the discriminator learning rate is scaled down by 0.01, the feature extractor undergoes 5 update steps per iteration, and the auxiliary event branch is completely discarded at inference time.

Training uses the TF-SepNet architecture with 150 epochs, batch size 512, Adam optimizer, and an exponential warm-up for 14 epochs followed by a linear decay to 0.5% of peak. Augmentations include Mixup, Freq-MixStyle, and DIR-Aug. Audio is resampled to 32 kHz and converted to log-Mel spectrograms using a 4096-point FFT, 96 ms window, 16 ms hop size, and 512 Mel bins.

## Experimental setup

Evaluated on the DCASE Task 1 TAU Urban Acoustic Scenes 2020 Mobile Development Dataset (TAU20), which contains ~64 hours across 10 classes, using the official 7:3 train-test split with 10-second clips sliced into non-overlapping 1-second segments. Compared against baselines including TF-SepNet, Multi-Task Learning (MTL), Knowledge Distillation (KD), and Joint Learning (JTL). Metrics include classification accuracy (%) and the proposed LGPD metric to stratify test samples into stable and unstable subsets.

## Results

On the 1-second TAU20 test set, SAT achieves an overall accuracy of 63.5%, outperforming the TF-SepNet baseline (59.6%), MTL (60.3%), KD (61.1%), and JTL (61.8%). On the hardest subset (unstable top-50% LGPD), SAT achieves 50.9% accuracy compared to the baseline's 43.1%, reducing the stability performance gap from 33.0 down to 25.2. In quantile-level evaluations, SAT yields a +12.0% absolute accuracy improvement over the baseline in the highest-discrepancy quantile (46.7% vs 34.7%), while incurring a negligible accuracy drop on the simplest stable samples (94.6% vs 95.8%).

| Method | Stable | Unstable | Gap (↓) | Overall |
|---|---|---|---|---|
| Baseline | 76.1 | 43.1 | 33.0 | 59.6 |
| + MTL [14] | 74.5 | 46.1 | 28.4 | 60.3 |
| + KD [23] | 75.1 | 47.2 | 27.9 | 61.1 |
| + JTL [13] | 75.4 | 48.3 | 27.1 | 61.8 |
| + SAT (Ours) | 76.2 | 50.9 | 25.2 | 63.5 |

## Limitations

The framework relies on off-the-shelf pre-trained sound event taggers (like BEATs), which introduces potential label noise and external dependency. The evaluation is restricted to a single standard dataset (TAU20) and a single short-duration target window (1 second), leaving open how well the approach scales to diverse acoustic environments or varying sub-second granularities.

## Why read this

Researchers and engineers working on low-latency environmental audio sensing or short-window acoustic scene classification will find a principled metric (LGPD) and adversarial formulation to prevent short-duration performance collapse.

## Code

None released (as of this page's `updated` date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

## Applications

Responsive edge-AI environmental sensing systems, real-time context-aware wearable devices, and smart-city surveillance.

## Institutions / 機構

Xi'an Jiaotong-Liverpool University, Nanjing University of Posts and Telecommunications

**Funding / 經費:** Jiangsu Provincial Major Science and Technology Project

## Related

- [Branch-wise Complementary Attention for Acoustic Scene Classification](han26b_interspeech.md) — same problem · relatedness 2.7/3
- [Towards Event-Robust Acoustic Scene Classification](cai26d_interspeech.md) — same problem · relatedness 2.0/3
- [Consistency-Regularized Dual-Branch Network with Performance-Aware Mean Teacher for Sound Event Detection](dai26_interspeech.md) — shared technique · relatedness 1.8/3
- [Robust Multi-Source-Free Domain Adaptation via Posterior Adjustment and Label Agreement](yoon26_interspeech.md) — same problem · relatedness 1.8/3
- [A Closer Look at Failure Modes in Temporal Understanding of Large Audio-Language Models](kulkarni26_interspeech.md) — same problem · relatedness 1.7/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
