All papers
Speech codingFull-paper digest

CAPS: A Cascaded Reconstruction Model to Power Saving in Hearables Using Sub-Nyquist Sampling with Bandwidth Extension

Tarikul Islam, Sajid F. Dipto, Luke B. Baja-Ricketts, David C. Vergano, Anomadarshi Barua

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.0 KB · Ready to paste

Preview copied content

TL;DR — CAPS is a cascaded reconstruction framework that jointly performs sub-Nyquist sampling (4 kHz) and low bit resolution (8-bit) in hearables to achieve a 3.3x reduction in hardware power consumption, while utilizing a mobile-hosted network to restore high-resolution wideband audio in 55.11 ms.

Key contributions

  • Proposes a multimodal sub-Nyquist sampling scheme for hearables combining air conduction microphones (ACMs) and bone conduction microphones (BCMs) to achieve a 3.31x ADC power reduction (dropping from 24 kHz/12-bit to 4 kHz/8-bit).
  • Introduces a cascaded architecture featuring a Spectral Enhancement Network (SEN) with a 2D-adapted Mamba bottleneck, an upsampling network, and an Amplitude-Phase Enhancement Network (APEN) for cross-modal fusion.
  • Engineers custom anti-wrapping instantaneous phase and group delay loss functions alongside a light multi-period time-domain loss to maintain perceptual fidelity.
  • Demonstrates real-time streaming capability on mobile hardware (Google Pixel 7 and Samsung Galaxy S21) with an inference time of 55.11 ms and a memory footprint of 11.04 MB.

Problem

Hearable devices require high sampling frequencies and bit resolutions to capture wideband audio, which drains their limited battery capacity. Existing speech enhancement (SE) and bandwidth extension (BWE) frameworks (such as ATS-UNet, TFiLM, AFiLM, TRAMBA, AERO, EBEN, HiFi++, NVSR, and NU-Wave) either target single-modal audio or fail to operate under reduced bit-resolution and sub-Nyquist ADC settings. Consequently, there is a lack of efficient multimodal SE methods capable of reconstructing wideband, high-fidelity signals from low-power, narrowband hearable inputs.

Method

CAPS processes low-resolution (4 kHz), noisy 2D T-F spectrograms from ACMs through a Spectral Enhancement Network (SEN). The SEN utilizes a 2D convolutional U-Net with 5 residual encoder/decoder layers and a Mamba sequence bottleneck—created by flattening spatial dimensions into a sequence for linear-time complexity and reshaping back—to model inter-phoneme dependencies. The spectrum is then converted to a 1D waveform via an Upsampling Network (UN) inspired by HiFi-GAN v2, employing four transposed convolution stages (8x, 8x, 2x, 2x) combined with dilated residual blocks (dilation rates 1, 3, 9) to achieve a 256x resolution increase.

Next, the Amplitude-Phase Enhancement Network (APEN) concatenates the 1D waveform with raw, less-noisy 1D BCM vibration signals. The APEN processes amplitude and phase streams in parallel using mutual coupling through large-kernel depth-wise 1D convolutions (kernel size 7x1) and point-wise linear layers, interleaved with Layer Normalization and GELU activations. The model is trained using a multi-period loss (periods 5 and 7 computed via MAE on reshaped 2D tensors), anti-wrapping instantaneous phase and group delay losses, and a 3-scale waveform MAE loss (1x, 2x, 4x downsampling).

The complete PyTorch model is converted via ONNX to TensorFlow Lite (TFLite) and deployed onto mobile platforms utilizing GPU/TPU delegates, running at an inference latency of 1.36 ms on a desktop GPU and 55.11 ms on a Google Pixel 7.

Experimental setup

Evaluated on a collected multimodal dataset of 20 speakers (45 minutes per speaker at 22 kHz, 12/10/8-bit) using VCTK-derived text, high-pass filtered at 5 Hz, mixed with non-speech noises and LibriSpeech speech noise at -7 to 5 dB. Also evaluated on the MagnaTagATune music dataset. Compared against six baselines (TFiLM, VibVoice, AERO, EBEN, HiFi++, SEANet). Metrics include LSD, VISQOL, NISQA-MOS, SI-SDR, PESQ, and STOI, evaluated on hardware including an Intel Silver 4310 desktop with an RTX 4090, Google Pixel 7 (Tensor G2, Mali-G710 MP7), and Samsung Galaxy S21.

Results

CAPS achieves superior or competitive performance against heavier SOTA models while operating at a fraction of the footprint. On 4-16 kHz VCTK noisy evaluations, CAPS achieves an LSD of 0.87, VISQOL of 4.15, NISQA-MOS of 4.13, SI-SDR of 16.99 dB, PESQ of 2.99, and STOI of 0.90, outperforming HiFi++ (LSD 0.89, PESQ 2.85) while using 25x fewer parameters (2.85M vs 72.2M) and a 4.63x faster desktop inference time (1.36 ms vs 6.3 ms). When replacing the Mamba bottleneck with Transformers, training time per epoch surges from 212s to 369s with higher parameter count (2.98M) for negligible metric gains. Removing APEN severely degrades performance (SI-SDR drops from 16.99 to 7.12 dB).

SystemLSD ↓VISQOL ↑NISQA ↑SI-SDR ↑PESQ ↑STOI ↑
Unprocessed2.781.841.278.541.110.79
TFiLM [15]4.851.683.7310.282.030.81
AERO [9]0.974.164.0317.032.930.89
HiFi++ [17]0.894.184.1117.482.850.90
SEANet [2]1.393.893.7814.312.430.89
CAPS (Ours)0.874.154.1316.992.990.90

Limitations

The framework does not incorporate audio stream encryption between the hearable device and the mobile platform. Furthermore, standard audio codecs were omitted during the evaluation of transmission power, performance, and efficiency, and language coverage is constrained by the English-centric VCTK and LibriSpeech data sourcing.

Why read this

Researchers and embedded engineers building resource-constrained hearables or low-latency mobile streaming audio pipelines should read this paper to learn how to combine sub-Nyquist sensor sampling with Mamba-backed neural bandwidth extension and cross-modal fusion.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Real-time voice communication, hearable audio enhancement, and low-power wearable audio streaming.

Institutions

George Mason University

Funding / 經費: Office of Naval Research

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-506