All papers
Enhancement & separationFull-paper digest

Breaking Shortcut Learning for Cross-Trial EEG-Guided Target Speech Extraction via Two-Stage Training

Wonchul Shin, Inyong Choi, Kyogu Lee

Code & resourcesgithub.com/argaaw/TRUST-TSE

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.8 KB · Ready to paste

Preview copied content

TL;DR — TRUST-TSE is a two-stage training framework for EEG-guided target speech extraction that breaks shortcut learning caused by trial-specific EEG structures, raising cross-trial selection accuracy from 37.56% (NeuroHeed) to 62.27% on the KUL dataset.

Key contributions

  • Identifies and systematically analyzes cross-trial generalization failure in end-to-end EEG-guided target speech extraction due to trial-specific EEG shortcut learning.
  • Proposes Stage 1 cross-modal contrastive pretraining with attended-speaker negative sampling to suppress trial identity cues and encourage fine-grained EEG-speech alignment.
  • Introduces a confidence-weighted SI-SDR extraction objective (CWS) using EEG-source similarity margins to couple frozen representations and handle ambiguous segments.
  • Demonstrates consistent cross-trial performance gains on both the KUL and DTU public datasets compared to end-to-end baselines like NeuroHeed and M3ANet.

Problem

Recent end-to-end EEG-guided target speech extraction models report high within-trial performance, often claiming SI-SDR above 10 dB and attended-source accuracy near 90%. However, continuous EEG exhibits slow temporal drifts and high autocorrelation over seconds to minutes. This allows end-to-end models to exploit trial-specific nuisance patterns as shortcuts to identify the target speaker rather than decoding actual attention. Consequently, these models suffer severe performance degradation when evaluated under strict cross-trial protocols where training and test segments originate from completely different recording sessions.

Method

TRUST-TSE decouples representation learning from waveform extraction via a two-stage training framework. In Stage 1, an EEG encoder (F_theta) and a mel-spectrogram audio encoder (H_psi) are pretrained via a contrastive loss (L_NCE) with temperature tau = 0.07. To prevent trial-identity shortcuts, the authors introduce attended-speaker negative sampling, where negative audio segments are drawn from non-aligned blocks spoken by the same attended speaker within the same trial. Both encoders use lightweight convolutional structures. L2-normalized embeddings at time step t are computed across a shared D-dimensional space, with audio embeddings linearly interpolated to match the temporal resolution of the EEG sequence.

In Stage 2, the pretrained EEG encoder is frozen to prevent drift toward shortcut cues, and a time-domain Dual-Path RNN (DPRNN) speech extractor (G_phi) maps a mixed waveform and the frozen EEG embedding to target speech. Conditioning is applied by linearly interpolating the EEG sequence and using channel-wise concatenation. To mitigate unstable gradients from ambiguous EEG guidance, Stage 2 optimizes a confidence-weighted SI-SDR objective (L_CWS). The similarity margin delta = s(z, a_att) - s(z, a_ign) is converted via a tanh weighting function with sharpness parameter kappa = 5. Segments where the embedding aligns more closely with the ignored source (negative weights) retain their magnitude to train the extractor to remain responsive without dropping data.

Experimental setup

Evaluated on two public auditory attention datasets: KUL (16 subjects, 64-channel EEG, 8 trials per subject of ~6 mins, Dutch speech) and DTU (18 subjects, 64-channel EEG, 60 trials per subject of 50 s, Danish speech). Evaluated using strict 4-fold cross-trial protocols (trial-level splits avoiding data leakage). Compared against end-to-end baselines NeuroHeed and M3ANet. Metrics include target selection accuracy (%), SI-SDRatt-All, SI-SDRatt-Correct, and SI-SDRign-Wrong. Stage 1 is trained for up to 50 epochs using AdamW (lr 5e-4, batch size 64). Stage 2 is trained for up to 100 epochs using AdamW (lr 3e-4, batch size 8).

Results

Under strict cross-trial evaluation with 5 s windows on the KUL dataset, TRUST-TSE achieves a selection accuracy of 62.27%, significantly outperforming M3ANet (48.42%) and NeuroHeed (37.56%), while reaching an SI-SDRatt-Correct of 15.23 dB. On the DTU dataset, TRUST-TSE reaches 70.40% accuracy and an SI-SDRatt-Correct of 19.21 dB, compared to NeuroHeed's 55.79% and M3ANet's 50.23%. Ablation studies confirm that attended-speaker negative sampling yields the most robust pretraining selection accuracy (65.62% on KUL) compared to in-batch (64.05%) or ignored-speaker sampling (39.63%). Furthermore, the confidence-weighted positive-and-negative tanh objective outperforms non-weighted (39.07%) and positive-only variants (42.73%) on KUL.

SystemAccuracy (%)SI-SDRatt-All (dB)SI-SDRatt-Correct (dB)
NeuroHeed (KUL)37.56-12.099.52
M3ANet (KUL)48.42-0.193.62
TRUST-TSE (KUL, Ours)62.270.2615.23
NeuroHeed (DTU)55.790.0810.40
M3ANet (DTU)50.23-0.574.83
TRUST-TSE (DTU, Ours)70.404.8519.21

Limitations

The evaluation is restricted to modest public EEG datasets (KUL and DTU) featuring limited speaker and language diversity. While the framework suppresses trial-specific shortcuts, broader multi-session deployment data, cross-device electrode shifts, and unseen-subject domain shifts remain challenging real-world bottlenecks.

Why read this

Speech and ML researchers building neuro-steered hearing aids should read this paper to understand how within-trial evaluation protocols artificially inflate speech extraction performance, and how a two-stage contrastive framework resolves this reliability gap.

Code

Applications

Neuro-steered hearing aids, brain-computer interfaces, and selective auditory attention prosthetics.

Institutions

Seoul National University, University of Iowa

Funding / 經費: National Research Foundation of Korea, Institute of Information & Communications Technology Planning & Evaluation

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-595