All papers
Audio understandingFull-paper digest

Robust Multi-Source-Free Domain Adaptation via Posterior Adjustment and Label Agreement

Hoyoung Yoon, U Kang

Code & resourcesgithub.com/snudatalab/FASOLA

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.2 KB · Ready to paste

Preview copied content

TL;DR — FASOLA is a robust multi-source-free domain adaptation framework for acoustic scene classification that corrects device-induced prediction biases and weights models via label agreement, improving average accuracy to 54.43% on DCASE 2020 Task 1A.

Key contributions

  • Proposes FASOLA, an MSFDA framework combining posterior adjustment and label agreement for privacy-compliant acoustic scene classification across heterogeneous devices.
  • Introduces a posterior adjustment module that estimates target class priors without labels and recalibrates source model logits to eliminate device-specific prediction biases.
  • Develops a label agreement module that replaces unreliable confidence measures with prediction consistency to dynamically prioritize trustworthy source models.
  • Demonstrates state-of-the-art performance across six unseen target devices (S1-S6) on DCASE 2020 Task 1A, outperforming prior unsupervised aggregation baselines.

Problem

Real-world acoustic scene classification suffers from severe performance degradation due to device heterogeneity, such as varying microphone frequency responses. While multi-source-free domain adaptation (MSFDA) protects privacy by avoiding raw source data sharing, aggregating diverse pre-trained models on unlabeled target domains is hindered by persistent prediction biases and unknown acoustic similarities. Prior methods rely heavily on confidence or entropy measures that fail under distribution misalignment, often misguiding adaptation by reinforcing biased predictions or overemphasizing miscalibrated models.

Method

FASOLA takes a set of KK frozen source models and an unlabeled target dataset D={xj}j=1∣D∣D = \{\mathbf{x}_j\}_{j=1}^{|D|} to predict target labels without accessing source training data. The pipeline consists of three sequential steps per target sample: logit calibration via posterior adjustment, source reliability weighting via label agreement, and weighted probability aggregation.

The posterior adjustment module addresses class-frequency prediction bias by re-centering each source model's outputs. It computes the empirical marginal prediction distribution p^k\hat{\mathbf{p}}^k for the kk-th model over the unlabeled target set and estimates the global target class prior q^\hat{\mathbf{q}} iteratively using a momentum update rate α\alpha. The calibrated logit z~jk\tilde{\mathbf{z}}_j^k is obtained by subtracting the log source prediction bias and adding the log estimated target prior, scaled by a correction strength hyperparameter τ\tau: z~jk=zjk−τ(log⁡p^k−log⁡q^)\tilde{\mathbf{z}}_j^k = \mathbf{z}_j^k - \tau (\log \hat{\mathbf{p}}^k - \log \hat{\mathbf{q}}).

The label agreement module computes source importance weights without relying on uncalibrated confidence scores. For each model, it measures prediction consistency against the majority vote of the other source models using an indicator function over argmax predictions. The resulting agreement scores are normalized using a sharpness control parameter γ\gamma to yield adaptive ensemble weights wkw_k. Finally, the predictions are combined via weighted aggregation of the adjusted softmax probabilities to yield the final classification.

Experimental setup

Evaluated on the DCASE 2020 Task 1A dataset, which features 10 acoustic scene classes across real (A, B, C) and simulated (S1-S6) recording devices using a Leave-One-Domain-Out protocol where each simulated device serves as an unseen target in turn. The backbone architecture is CP-ResNet with models pre-trained independently and frozen, and AdaBN pre-processing applied to align batch normalization statistics. Compared against Oracle, Uniform ensemble, DECISION, CAiDA, DATE, and Bi-ATEN.

Results

FASOLA achieves an average accuracy of 54.43% and an average F1-score of 54.04% across all six unseen target devices (S1-S6), outperforming the Uniform ensemble (50.68% Acc), DECISION (51.88%), CAiDA (51.57%), and Bi-ATEN (50.39%). On Target S6 specifically, FASOLA reaches 50.74% accuracy and 50.53% F1-score. Ablation studies confirm that combining posterior adjustment and label agreement outperforms using either component in isolation under both raw and AdaBN pre-processed settings. Furthermore, FASOLA achieves the highest correlation with true target accuracy among evaluated source weighting techniques, with a Pearson correlation r=0.9365r = 0.9365 and Spearman rank correlation ρ=0.9286\rho = 0.9286.

MethodTarget S1 AccTarget S2 AccTarget S3 AccTarget S4 AccTarget S5 AccTarget S6 AccAverage Acc
Oracle38.8940.3748.0649.8155.0043.1545.88
Uniform42.6942.6956.9455.5659.1747.0450.68
DECISION [9]43.5244.0757.3157.8760.8347.6951.88
CAiDA [30]43.1544.6359.8156.6758.2446.9451.57
Bi-ATEN [32]42.5943.3356.1155.1958.0647.0450.39
FASOLA (proposed)46.2048.8959.8160.0060.9350.7454.43

Limitations

The framework relies on offline global target statistics across the entire target dataset to estimate priors, making it unsuitable for real-time streaming or single-sample edge deployment without modification. Evaluation is restricted to acoustic scene classification on the DCASE 2020 dataset, leaving broader speech domains like ASR or speaker verification unverified.

Why read this

Researchers working on source-free domain adaptation and multi-model ensembling will learn how explicit posterior alignment and consensus-based weighting overcome the failure modes of confidence-based selection under severe domain shifts.

Code

Applications

Deploying robust acoustic scene classification models on edge devices with heterogeneous microphones without accessing raw user audio data.

Institutions

Seoul National University

Funding / 經費: Institute of Information & Communications Technology Planning & Evaluation, Korea government (MSIT), XVoice: Multi-Modal Voice Meta Learning, AI Star Fellowship Support Program, Global AI Frontier Lab, Artificial Intelligence Graduate School Program

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-370