All papers
Enhancement & separationFull-paper digest

CLEAR: Clinical LLM Embedding and Attention-based Reconstruction

Eashita Wazed, Hieyong Jeong, Choonsung Shin, Shima Okada

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.7 KB · Ready to paste

Preview copied content

TL;DR — CLEAR is a three-stage heart sound restoration framework combining a pretrained Whisper encoder, a Masked Graph Attention Network, and a Latent GAN to denoise clinical auscultation recordings, achieving a peak PESQ of 4.64 and SI-SDR of 64.35 dB under 10 dB SNR environmental noise.

Key contributions

  • Introduces CLEAR, a modular framework using pretrained Whisper embeddings to capture long-range physiological periodicity (S1, S2, and murmurs) in heart sounds.
  • Proposes a Masked Graph Attention Network (GAT) with a bounded hypothesis space that guarantees a tighter generalization error bound and avoids unconstrained direct regression.
  • Applies a Latent GAN operating in a low-dimensional embedding space to heal multiplicative masking holes and resolve exponential Wasserstein distance penalties found in high-dimensional waveform GANs.
  • Demonstrates exceptional restoration performance on the BUET Multi-Disease Heart Sound Dataset mixed with 10 dB SNR hospital background noise.

Problem

Traditional heart sound enhancement relies on simple bandpass filters and wavelet transforms, which fail under complex, non-stationary clinical noise or when noise overlaps with cardiac frequency bands. Modern data-driven deep learning models using CNNs or RNNs struggle because their local receptive fields lack high-level semantic understanding of the global cardiac cycle, leading to musical noise and obscured murmurs. This failure in noise suppression severely obstructs both physician auscultation and downstream computer-aided diagnostic systems, causing critical misdiagnoses.

Method

The architecture operates in three sequential stages. First, raw audio is passed through a pretrained Whisper encoder to extract high-dimensional semantic embeddings. Unlike local CNNs with exponential mutual information decay, Whisper's self-attention maintains a stable lower bound of non-zero mutual information globally across the entire cardiac cycle sequence.

Second, these embeddings form a semantic graph fed into a Masked Graph Attention Network (GAT), which computes attention coefficients to predict a bounded noise-suppression mask within [0, 1]. Applying this mask via element-wise multiplication isolates pure heart sounds while leveraging statistical learning theory to bound Rademacher complexity and generalization error.

Third, to repair holes created by multiplicative masking, a Latent GAN operates entirely within a low-dimensional dense embedding space. The generator heals the masked representations, regularized by a Cosine Similarity semantic loss combined with an adversarial loss from a 1-Lipschitz continuous discriminator. Finally, a lightweight acoustic decoder maps the restored embeddings back to the time-domain waveform.

Experimental setup

Evaluated using the BUET Multi-Disease Heart Sound Dataset containing normal and pathological recordings (S1, S2, murmurs). Input mixtures were constructed by synthesizing environmental noise (conversations, medical equipment, reverberations) at a harsh 10 dB SNR. Baselines compared in ablations include CNN encoders, direct GAT regression, CNN separation modules, and Diffusion models. Metrics include PESQ, SI-SDR, Log-Spectral Distance (LSD), CBAK, and COVL, evaluated via a dual-input reference protocol.

Results

Under peak conditions with 10 dB SNR environmental noise, the noisy input degraded to a PESQ of 1.03 and SI-SDR of -29.37 dB, whereas the proposed CLEAR framework restored it to a peak PESQ of 4.64 and SI-SDR of 64.35 dB. Dataset-averaged evaluations show the full Whisper + GAT + GAN architecture achieves an average PESQ of 2.99, LSD of 1.78, CBAK of 2.67, and COVL of 3.38.

Ablations demonstrate that replacing Whisper with a CNN encoder causes severe degradation, substituting GAT with a CNN worsens spectral distortion (LSD rising to 3.55), and replacing the Latent GAN with a Diffusion model leads to overall metric drops.

System / ConditionPESQSI-SDR (dB)LSDCBAKCOVL
Noisy Input1.03-29.37---
CNN + GAT + GAN (Avg)1.04-1.932.132.42
Whisper + CNN + GAN (Avg)1.03-3.552.142.58
Whisper + GAT + Diff (Avg)1.03-3.862.172.47
Whisper + GAT + GAN (Proposed, Avg)3.00-1.792.673.39
Enhanced Output (Proposed, Peak)4.6464.36---

Limitations

The evaluation relies on artificially synthesized noise mixtures at a fixed 10 dB SNR rather than unconstrained in-the-wild recordings captured directly from diverse stethoscope hardware. The framework's generalization across various unseen clinical pathologies, extremely low-resource hardware deployment constraints, and real-time streaming latency are not thoroughly quantified.

Why read this

Speech and audio engineers working on bio-acoustic enhancement or representation learning will learn how to adapt pretrained speech LLMs and latent GANs for non-speech physiological signals while avoiding high-dimensional manifold collapse.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Smart electronic stethoscopes, telemedicine diagnostic platforms, and robust frontend denoising for automated computer-aided cardiac disease classification systems.

Institutions

Chonnam National University, Ritsumeikan University

Funding / 經費: National Research Foundation of Korea, Institute of Information & Communications Technology Planning & Evaluation, Korea Creative Content Agency

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-883