TL;DR — This paper introduces Cross-Domain Few-Shot Class-Incremental Audio Classification (CD-FCAC), a task where a model must continually learn new classes from an unseen target domain using very few samples without forgetting base classes from a source domain. The authors propose an adversarial contrastive learning strategy that achieves state-of-the-art average accuracy across six cross-domain dataset pairs.
Key contributions
- Formulates the CD-FCAC problem, recognizing that domain shift and class increment frequently coexist in practical audio applications.
- Proposes an adversarial contrastive training strategy combining a spectral disruptor, adversarial training, and supervised contrastive loss to extract domain-invariant and discriminative embeddings.
- Implements a memory-preserving incremental update framework where the encoder is frozen and class mean vectors are retained alongside a dynamically updated classifier.
- Demonstrates superior average accuracy compared to nine baseline systems (including DFSL, CEC, PAN, AMFO adaptations, and PCR) across six cross-domain dataset combinations.
Problem
Traditional Few-Shot Class-Incremental Audio Classification (FCAC) methods—such as DFSL, CEC, PAN, and AMFO—assume that samples across both base and incremental sessions originate from the exact same data distribution. In real-world deployments like smart homes or autonomous agents, however, new classes arrive from entirely different acoustic environments or recording devices, introducing a severe domain shift. Tackling this CD-FCAC setting requires overcoming catastrophic forgetting of old classes, bridging domain gaps between source and target data, and avoiding few-shot overfitting on sparse new-class samples.
Method
The model consists of a ResNet-18 encoder and a multi-layer perceptron (MLP) classifier operating on 128-dimensional log Mel-spectrograms to output 512-dimensional embeddings. During the base training session (session 0), an adversarial contrastive training strategy is executed in two alternating steps per episode: first, adversarial samples are generated by perturbing the source log Mel-spec using a spectral disruptor (comprising two frozen random convolutional layers) while maximizing cross-entropy loss penalized by an enforced semantic distance constraint; second, the model parameters are updated by minimizing a joint loss comprising cross-entropy and supervised contrastive loss (SCL) with scaling coefficients alpha=1 and beta=0.2. This encourages intra-class compactness and inter-class separability in the embedding space.
In subsequent incremental sessions (sessions 1 through M-1), the encoder is completely frozen to preserve base knowledge, and only the classifier is updated using incoming support samples alongside saved mean vectors of old classes' embeddings. The incremental objective minimizes a combined cross-entropy loss over new classes and old-class mean vectors (weighted by lambda=0.6). During inference, classification relies on cosine distance between test sample embeddings and the classifier's weight vectors.
Experimental setup
Experiments utilize three public audio datasets: LS-100, NSynth-100, and FSC-89, combined into six cross-domain pairs (FS->NS, FS->LS, NS->FS, NS->LS, LS->FS, and LS->NS) where source and target datasets differ in acoustic domain and class sets. The system is evaluated using Average Accuracy (AA) across all sessions under a standard 5-way 5-shot setting (N=5, K=5). The base session is trained for 200 epochs (50 episodes/epoch) and incremental sessions for 100 epochs, using learning rates of 0.1.
Results
The proposed method achieves top-tier average accuracy across all six cross-domain evaluation pairs, outperforming standard FCAC baselines and cross-domain adaptations. Specifically, it yields AA scores of 46.89% on FS->NS, 41.67% on FS->LS, 85.17% on NS->FS, 79.09% on NS->LS, 80.05% on LS->FS, and 79.78% on LS->NS, beating the strongest competitors (such as AMFO+AFA and PCR) by notable margins. Ablation studies on the NS->LS pair confirm that combining adversarial training and supervised contrastive loss yields the highest AA (79.09%), outperforming models trained with neither (75.46%) or only one of the components. Extended evaluations show that AA scales positively with shot count K, while setting N=5 optimal balances session length against inter-class confusion.
| Systems / Conditions | FS->NS (%) | FS->LS (%) | NS->FS (%) | NS->LS (%) | LS->FS (%) | LS->NS (%) |
|---|---|---|---|---|---|---|
| DFSL | 37.21 | 37.54 | 82.58 | 76.62 | 77.83 | 75.30 |
| CEC | 40.75 | 39.29 | 83.75 | 77.31 | 78.57 | 76.86 |
| PAN | 40.88 | 39.35 | 83.83 | 77.50 | 78.71 | 77.34 |
| AMFO+AFA | 44.75 | 41.52 | 84.46 | 78.50 | 79.47 | 78.68 |
| PCR | 38.70 | 36.98 | 75.61 | 78.22 | 78.79 | 78.80 |
| Ours | 46.89 | 41.67 | 85.17 | 79.09 | 80.05 | 79.78 |
Limitations
The evaluation is restricted to static cross-domain pairs derived from three datasets (LS-100, NSynth-100, FSC-89), leaving multi-domain sequences or unconstrained open-world acoustic streams untested. The frozen encoder strategy, while effective at preventing catastrophic forgetting of old classes, limits the capacity of the model to adapt its feature extractor to entirely novel acoustic characteristics encountered late in incremental sessions.
Why read this
Speech and machine learning researchers working on lifelong or continual audio learning will find this paper valuable for its clean formulation of cross-domain few-shot class-incremental audio classification and its adversarial contrastive training recipe.
Code
Applications
Smart home voice assistants and autonomous mobile robots incrementally learning to recognize new sound events, musical instruments, and voice commands across varying acoustic environments.
Institutions
South China University of Technology
Funding / 經費: National Natural Science Foundation of China, China-Croatia Science and Technology Cooperation Committee
Related
- Few-shot Class-variable Incremental Audio Classification via Prototype Adaptation and Pseudo Class-variable Training — same problem · relatedness 2.2/3
- Scaling few-shot spoken word classification with generative meta-continual learning — same problem · relatedness 2.0/3
- Audio-Language Prompt Learning for Few-Shot Audio Classification — same problem · relatedness 2.0/3
- Hybrid Continual Learning for Low-Resource Australian Aboriginal Language Identification — same problem · relatedness 2.0/3
- Robust Language Identification Using Semi-positive Contrastive Learning — shared technique · relatedness 2.0/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-1250