All papers
Enhancement & separationFull-paper digest

Through-Wall Radar Speech Acquisition via Cascaded Attention Fusion

Ruotong Ding, Zhi-Wei Tan, V.G. Reju, Andy W. H. Khong

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.9 KB · Ready to paste

Preview copied content

TL;DR — The Cascaded Attention Fusion Transformer (CAF-Former) tackles through-wall radar speech recovery by combining progressive spectral expansion with a cascaded attention hierarchy, achieving superior intelligibility and perceptual scores over existing models under severe band limitation.

Key contributions

  • Proposes a sensing-aware progressive bandwidth extension framework that gradually expands reliable low-frequency radar representations to full-band spectrum across 10 layers.
  • Introduces Temporal Multi-Query Attention (TMQA), where parallel query projections share keys and values to capture diverse temporal patterns while preventing head fragmentation in low-SNR sub-bands.
  • Develops a Frequency Attention Fusion (FAF) module to model cross-frequency interactions and leverage informative low-frequency cues for high-frequency harmonic recovery.
  • Demonstrates consistent improvements over state-of-the-art speech enhancement baselines (RANet, Wave-Voice Net, DPTNet, TF-Locoformer, EBENet) on both simulated and real through-wall radar recordings.

Problem

Speech acquisition via through-barrier radio-frequency sensing suffers from severe clutter, noise, and high-frequency attenuation above 1 kHz due to barrier penetration and low signal-to-noise ratios. Conventional deep learning architectures like CNN-based RANet and Wave-Voice Net struggle with long-range temporal-spectral dependencies, while standard Transformer variants and Conformer models utilize multi-head self-attention (MHSA) that suffers from head fragmentation toward noise-dominated high-frequency sub-bands. These limitations prevent effective restoration of harmonic structures and intelligibility cues, necessitating a dedicated sensing-aware architecture.

Method

The CAF-Former processes single-channel FMCW radar signals transformed via STFT, retaining only the lowest 50 reliable frequency bins (0 to 765.6 Hz) as input. The model stacks K=10K=10 progressive transformer layers that gradually expand the spectral dimension, where each layer utilizes positional encoding and residual connections. Each layer consists of a cascaded attention mechanism: first, Temporal Multi-Query Attention (TMQA) employs Nq=8N_q = 8 independent query branches attending to a shared set of keys (dk=64d_k = 64) and values to extract diverse temporal maps without head fragmentation. Second, a Frequency Attention Fusion (FAF) module reorganizes these attention maps into a frequency-centric matrix, applying scaled dot-product attention along the frequency axis to model inter-frequency dependencies and guide high-frequency recovery.

Training is performed using the log-spectral amplitude distance loss function on noisy input phase combined with magnitude estimation. The network is optimized using the Adam optimizer with a learning rate of 1×10−41 \times 10^{-4} for up to 50 epochs. The input comprises 4-second utterances sampled at 8 kHz, processed using 512-sample frames with a 75% overlap and a Hamming window, retaining 257 frequency bins.

Experimental setup

Evaluated on approximately 312 hours of simulated training data (derived from LibriSpeech convolved with measured radar impulse responses and uniform SNR levels from -5 to 15 dB) and 6 hours of validation data, alongside real-world recordings through a 15 cm concrete wall using a 5.31 GHz FMCW radar platform. Compared against RANet, Wave-Voice Net, DPTNet, TF-Locoformer, and EBENet. Metrics include PESQ, STOI, DNSMOS, and CS-MFCC. The model contains 2.726M parameters.

Results

CAF-Former achieves top performance on intelligibility and perceptual quality metrics across conditions. On recorded test data, it attains an STOI of 0.617, DNSMOS of 2.423, and CS-MFCC of 0.591, outperforming EBENet (STOI 0.598, DNSMOS 2.384, CS-MFCC 0.585) and TF-Locoformer (STOI 0.612, DNSMOS 2.351, CS-MFCC 0.583). Ablations confirm that combining TMQA and FAF yields the highest scores (PESQ 1.782, ESTOI 0.617), whereas standard MHSA collapses to a PESQ of 1.451. While TF-Locoformer and EBENet achieve slightly higher raw PESQ scores in certain isolated configurations, CAF-Former delivers more balanced superiority across overall intelligibility and naturalness metrics.

ModelSimulated PESQSimulated STOIRecorded PESQRecorded STOIRecorded DNSMOS
Noisy input1.0640.5351.0590.4621.753
RANet [11]1.5730.6181.4620.5562.231
Wave-Voice Net [12]1.6300.6071.4820.5482.213
TF-Locoformer [17]1.9020.6951.8210.6122.351
EBENet [18]1.9100.6871.7980.5982.384
CAF-Former (proposed)1.8520.6991.7820.6172.423

Limitations

The evaluation is restricted to controlled through-wall scenarios using a single 15 cm concrete wall and a fixed 5.31 GHz FMCW radar geometry, meaning generalization to varied wall thicknesses, materials, and dynamic human subjects remains unproven. The system relies on magnitude spectrum estimation while retaining noisy input phase, which caps phase-correction performance. Dataset scale is bounded to simulated LibriSpeech mixtures without extensive real-world human speaker evaluations in the primary benchmark.

Why read this

Speech and ML engineers working on non-acoustic sensing or challenging bandwidth extension tasks should read this to see how cascaded temporal multi-query attention and frequency-domain fusion can effectively stabilize high-frequency reconstruction under extremely low SNR conditions.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Non-acoustic speech enhancement, through-wall surveillance, search and rescue communications, and disaster recovery audio processing.

Institutions

Nanyang Technological University, Cochin University of Science and Technology

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1034