---
id: ding26d_interspeech
title: "SGAD: A State-Guided Adaptive Decision Framework for Robust EEG-Based
  Auditory Attention Switch Decoding"
authors:
  - Yuting Ding
  - Xuefei Wang
  - Ximin Chen
  - Chunlin Li
  - Fei Chen
year: 2026
doi: 10.21437/Interspeech.2026-1815
isca_url: https://www.isca-archive.org/interspeech_2026/ding26d_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/ding26d_interspeech.pdf
session: Speech Production and Perception 1
topics:
  - spoken-language-understanding
  - paralinguistics
  - evaluation
category: applications-other
labels:
  - streaming-real-time
institutions:
  - Southern University of Science and Technology
  - Capital Medical University
funding:
  - National Key Research and Development Program of China
  - National Natural Science Foundation of China
code:
  url: ""
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: ding26d_interspeech
  category: applications-other
  labels:
    - streaming-real-time
  institutions:
    - Southern University of Science and Technology
    - Capital Medical University
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-1815
  pdf: https://www.isca-archive.org/interspeech_2026/ding26d_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/ding26d_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/ding26d_interspeech/markdown.md
---

# SGAD: A State-Guided Adaptive Decision Framework for Robust EEG-Based Auditory Attention Switch Decoding

*Yuting Ding, Xuefei Wang, Ximin Chen, Chunlin Li, Fei Chen*

[PDF](https://www.isca-archive.org/interspeech_2026/ding26d_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/ding26d_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-1815)

**Category:** `applications-other` · **Labels:** `streaming-real-time`

**TL;DR** — The paper introduces the State-Guided Adaptive Decision (SGAD) framework to resolve the stability-latency trade-off in EEG-based auditory attention switch decoding, boosting switch-F1 score by over 13 percentage points over exponential moving average baselines. It also establishes six hierarchical evaluation protocols to rigorously test generalization across audio content, speaker identity, and subject boundaries.

## Key contributions

- Proposes the SGAD decision-level, encoder-agnostic framework featuring a causal state detector (CSD) and adaptive gating mechanism (AGM) to dynamically adjust temporal smoothing.
- Eliminates the traditional fixed trade-off between steady-state stability and switch-detection latency by lowering system inertia only during estimated attentional transitions.
- Establishes six hierarchical evaluation protocols combining trial, audio, speaker, and subject constraints (from LOTO to the most stringent LOSSO) to expose data-partition biases.
- Demonstrates consistent improvements in decoding accuracy (up to 85.6%) and switch-F1 across both CNNT and FCTNet backbones.

## Problem

Early auditory attention decoding (AAD) studies relied on static attention paradigms or window-wise decoding (WD), which generates independent predictions per short segment and suffers from severe temporal instability due to EEG non-stationarity. While traditional post-processing methods like persistence-based decision (PBD) and exponential moving average (EBD) apply fixed temporal smoothing to suppress fluctuations, they treat stable and transitional states uniformly, leading to high latency and sluggish adaptation during attention switches. Furthermore, prior evaluations often overlook confounding factors such as audio content and speaker identity overlap, allowing models to exploit non-attentional cues and overinflate performance.

## Method

The pipeline starts by mapping multi-band EEG segments ($N \times C \times T$) and two competing speech streams into 64-dimensional embeddings via frozen encoders (CNNT or FCTNet for EEG; a pretrained wav2vec 2.0 with a 1x1 conv projection for speech). Raw correlation scores $r_t^{(i)}$ yield window-wise margins $\Delta r_t = r_t^{(2)} - r_t^{(1)}$.

To overcome static smoothing, the CSD models causal temporal dependencies using a sliding-window KV cache of length $W=15$ over multi-head self-attention ($d_k=16$). It outputs an attention-switch probability $\hat{s}_t$. The AGM then maps $\hat{s}_t$ through a monotonically decreasing function with learnable parameters $a$ and $b$ (clamped between $g_{\min}=0.05$ and $g_{\max}=0.95$) to produce a dynamic smoothing factor $g_t$. Higher switch probability reduces $g_t$ to minimize inertia and speed up adaptation, whereas lower probability increases $g_t$ to enforce steady-state smoothing via recursive updating: $\hat{\Delta r}_t = g_t \hat{\Delta r}_{t-1} + (1 - g_t) \Delta r_t$.

SGAD is trained with frozen encoders using a multi-task loss: $L_{\text{SGAD}} = L_{\text{dec}} + \lambda_1 L_{\text{state}} + \lambda_2 L_{\text{smooth}}$, where $L_{\text{dec}}$ is BCE on the filtered margin, $L_{\text{state}}$ uses triangular soft labels ($R=5\text{ s}$) to supervise CSD transition states ($\lambda_1=0.5$), and $L_{\text{smooth}}$ regularizes gate fluctuations ($\lambda_2=0.05$).

## Experimental setup

Evaluated on the MS-AASD dataset consisting of 13 hours of 64-channel EEG recordings from 13 normal-hearing adults performing self-initiated attention switching across 60 trials per participant (3.37 switches/trial). Evaluated via 6 protocols (LOTO, LOAO, LOSpO, LOSO, LOSAO, LOSSO) measuring decoding accuracy (Acc), switch-F1 score (Sw-F1), and switch detection latency (SDL). Implemented in PyTorch on NVIDIA V100 GPUs using AdamW (weight decay $1 \times 10^{-4}$), with encoder pretraining for 50 epochs and SGAD trained with a 2-window burn-in and early stopping.

## Results

Across all protocols and backbones, SGAD improves mean accuracy to 79.8% and Sw-F1 to 67.3% while containing response latency to a competitive 1.18 s (outperforming EBD's 1.38 s SDL and 53.7% Sw-F1). Using the FCTNet backbone under the easiest LOTO protocol, SGAD achieves 85.6% Acc and 72.5% Sw-F1, whereas under the most stringent LOSSO protocol, performance drops to 76.7% Acc and 62.3% Sw-F1, illustrating the impact of rigorous cross-condition constraints. Ablation studies under LOSSO confirm that removing the CSD severely harms switch tracking, dropping FCTNet's Sw-F1 from 62.3% down to 48.5%.

| System / Condition | Acc (%) | Sw-F1 (%) | SDL (s) |
|---|---|---|---|
| CNNT + WD (Mean) | 77.4 | 29.0 | 0.89 |
| CNNT + PBD (Mean) | 78.6 | 45.6 | 1.23 |
| CNNT + EBD (Mean) | 79.0 | 53.7 | 1.38 |
| CNNT + SGAD (Mean) | 79.8 | 67.3 | 1.18 |
| FCTNet + SGAD (LOTO) | 85.6 | 72.5 | 1.12 |
| FCTNet + SGAD (LOSSO) | 76.7 | 62.3 | 1.42 |

## Limitations

The evaluation relies on a single dataset (MS-AASD) with 13 normal-hearing participants, leaving real-world hearing-impaired cohort generalization and cross-device electrode variability unverified. The framework assumes clean binary attention paradigms without testing multi-speaker cocktail party scenes beyond two competing streams. Furthermore, compute overhead introduced by the causal KV cache and hyperparameter sensitivity require careful tuning for real-time, ultra-low-power hearing aid deployment.

## Why read this

Read this paper if you work on EEG-based auditory attention decoding or real-time neural interfaces and want to master state-dependent temporal decision-making over window-wise predictions. It provides a blueprint for decoupling stability from latency while establishing a rigorous hierarchy of cross-condition evaluation protocols to expose data partitioning biases.

## Code

None released (as of this page's `updated` date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

## Applications

Neuro-steered hearing aids, brain-computer interfaces (BCIs), and cognitive load monitoring systems.

## Institutions / 機構

Southern University of Science and Technology, Capital Medical University

**Funding / 經費:** National Key Research and Development Program of China, National Natural Science Foundation of China

## Related

- [DisenEEG-Net: Disentangling EEG features via sufficient information bottleneck and adversarial learning for cross-subject auditory attention detection](kausar26_interspeech.md) — same problem · relatedness 2.2/3
- [SAGE: Switch-Aware EEG-Guided Soft Gating for Target Speaker Extraction with In-Trial Switching](wang26p_interspeech.md) — same problem · relatedness 2.0/3
- [MOV-AAD: A Large-Scale Multimodal Dataset for Auditory Attention Decoding During Moving Conversations](he26g_interspeech.md) — same problem · relatedness 2.0/3
- [NeuroMultiSpEx: Neuro-Guided Target Speaker Extraction for Multi-Speaker Scenarios](silva26_interspeech.md) — same problem · relatedness 1.7/3
- [Breaking Shortcut Learning for Cross-Trial EEG-Guided Target Speech Extraction via Two-Stage Training](shin26_interspeech.md) — same problem · relatedness 1.6/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
