---
id: wazed26_interspeech
title: "CLEAR: Clinical LLM Embedding and Attention-based Reconstruction"
authors:
  - Eashita Wazed
  - Hieyong Jeong
  - Choonsung Shin
  - Shima Okada
year: 2026
doi: 10.21437/Interspeech.2026-883
isca_url: https://www.isca-archive.org/interspeech_2026/wazed26_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/wazed26_interspeech.pdf
session: Acoustic Event Detection 3
topics:
  - speech-enhancement
  - source-separation
  - self-supervised
category: enhancement-separation
labels:
  - self-supervised
  - generative-model
  - robustness-noise
institutions:
  - Chonnam National University
  - Ritsumeikan University
funding:
  - National Research Foundation of Korea
  - Institute of Information & Communications Technology Planning & Evaluation
  - Korea Creative Content Agency
code:
  url: ""
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: wazed26_interspeech
  category: enhancement-separation
  labels:
    - self-supervised
    - generative-model
    - robustness-noise
  institutions:
    - Chonnam National University
    - Ritsumeikan University
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-883
  pdf: https://www.isca-archive.org/interspeech_2026/wazed26_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/wazed26_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/wazed26_interspeech/markdown.md
---

# CLEAR: Clinical LLM Embedding and Attention-based Reconstruction

*Eashita Wazed, Hieyong Jeong, Choonsung Shin, Shima Okada*

[PDF](https://www.isca-archive.org/interspeech_2026/wazed26_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/wazed26_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-883)

**Category:** `enhancement-separation` · **Labels:** `self-supervised`, `generative-model`, `robustness-noise`

**TL;DR** — CLEAR is a three-stage heart sound restoration framework combining a pretrained Whisper encoder, a Masked Graph Attention Network, and a Latent GAN to denoise clinical auscultation recordings, achieving a peak PESQ of 4.64 and SI-SDR of 64.35 dB under 10 dB SNR environmental noise.

## Key contributions

- Introduces CLEAR, a modular framework using pretrained Whisper embeddings to capture long-range physiological periodicity (S1, S2, and murmurs) in heart sounds.
- Proposes a Masked Graph Attention Network (GAT) with a bounded hypothesis space that guarantees a tighter generalization error bound and avoids unconstrained direct regression.
- Applies a Latent GAN operating in a low-dimensional embedding space to heal multiplicative masking holes and resolve exponential Wasserstein distance penalties found in high-dimensional waveform GANs.
- Demonstrates exceptional restoration performance on the BUET Multi-Disease Heart Sound Dataset mixed with 10 dB SNR hospital background noise.

## Problem

Traditional heart sound enhancement relies on simple bandpass filters and wavelet transforms, which fail under complex, non-stationary clinical noise or when noise overlaps with cardiac frequency bands. Modern data-driven deep learning models using CNNs or RNNs struggle because their local receptive fields lack high-level semantic understanding of the global cardiac cycle, leading to musical noise and obscured murmurs. This failure in noise suppression severely obstructs both physician auscultation and downstream computer-aided diagnostic systems, causing critical misdiagnoses.

## Method

The architecture operates in three sequential stages. First, raw audio is passed through a pretrained Whisper encoder to extract high-dimensional semantic embeddings. Unlike local CNNs with exponential mutual information decay, Whisper's self-attention maintains a stable lower bound of non-zero mutual information globally across the entire cardiac cycle sequence.

Second, these embeddings form a semantic graph fed into a Masked Graph Attention Network (GAT), which computes attention coefficients to predict a bounded noise-suppression mask within [0, 1]. Applying this mask via element-wise multiplication isolates pure heart sounds while leveraging statistical learning theory to bound Rademacher complexity and generalization error.

Third, to repair holes created by multiplicative masking, a Latent GAN operates entirely within a low-dimensional dense embedding space. The generator heals the masked representations, regularized by a Cosine Similarity semantic loss combined with an adversarial loss from a 1-Lipschitz continuous discriminator. Finally, a lightweight acoustic decoder maps the restored embeddings back to the time-domain waveform.

## Experimental setup

Evaluated using the BUET Multi-Disease Heart Sound Dataset containing normal and pathological recordings (S1, S2, murmurs). Input mixtures were constructed by synthesizing environmental noise (conversations, medical equipment, reverberations) at a harsh 10 dB SNR. Baselines compared in ablations include CNN encoders, direct GAT regression, CNN separation modules, and Diffusion models. Metrics include PESQ, SI-SDR, Log-Spectral Distance (LSD), CBAK, and COVL, evaluated via a dual-input reference protocol.

## Results

Under peak conditions with 10 dB SNR environmental noise, the noisy input degraded to a PESQ of 1.03 and SI-SDR of -29.37 dB, whereas the proposed CLEAR framework restored it to a peak PESQ of 4.64 and SI-SDR of 64.35 dB. Dataset-averaged evaluations show the full Whisper + GAT + GAN architecture achieves an average PESQ of 2.99, LSD of 1.78, CBAK of 2.67, and COVL of 3.38.

Ablations demonstrate that replacing Whisper with a CNN encoder causes severe degradation, substituting GAT with a CNN worsens spectral distortion (LSD rising to 3.55), and replacing the Latent GAN with a Diffusion model leads to overall metric drops.

| System / Condition | PESQ | SI-SDR (dB) | LSD | CBAK | COVL |
|---|---|---|---|---|---|
| Noisy Input | 1.03 | -29.37 | - | - | - |
| CNN + GAT + GAN (Avg) | 1.04 | - | 1.93 | 2.13 | 2.42 |
| Whisper + CNN + GAN (Avg) | 1.03 | - | 3.55 | 2.14 | 2.58 |
| Whisper + GAT + Diff (Avg) | 1.03 | - | 3.86 | 2.17 | 2.47 |
| **Whisper + GAT + GAN (Proposed, Avg)** | **3.00** | - | **1.79** | **2.67** | **3.39** |
| **Enhanced Output (Proposed, Peak)** | **4.64** | **64.36** | - | - | - |

## Limitations

The evaluation relies on artificially synthesized noise mixtures at a fixed 10 dB SNR rather than unconstrained in-the-wild recordings captured directly from diverse stethoscope hardware. The framework's generalization across various unseen clinical pathologies, extremely low-resource hardware deployment constraints, and real-time streaming latency are not thoroughly quantified.

## Why read this

Speech and audio engineers working on bio-acoustic enhancement or representation learning will learn how to adapt pretrained speech LLMs and latent GANs for non-speech physiological signals while avoiding high-dimensional manifold collapse.

## Code

None released (as of this page's `updated` date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

## Applications

Smart electronic stethoscopes, telemedicine diagnostic platforms, and robust frontend denoising for automated computer-aided cardiac disease classification systems.

## Institutions / 機構

Chonnam National University, Ritsumeikan University

**Funding / 經費:** National Research Foundation of Korea, Institute of Information & Communications Technology Planning & Evaluation, Korea Creative Content Agency

## Related

- [UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement](yan26_interspeech.md) — shared technique · relatedness 2.1/3
- [QC-GAN: A Parameter-Efficient Quaternion Conformer GAN for High-Fidelity Speech Enhancement](yamauchi26_interspeech.md) — shared technique · relatedness 2.1/3
- [Seed-Enh: Generative Speech Enhancement in Decoupled Semantic and Timbre Spaces](shang26_interspeech.md) — same problem · relatedness 2.1/3
- [UFL-GAN: A Multi-Discriminator GAN for Unsupervised Speech Enhancement](bejugam26_interspeech.md) — shared technique · relatedness 2.0/3
- [Bridging Self-Supervised Learning and Speech Enhancement: A Wav2Vec2-Conditioned Framework](ojha26_interspeech.md) — shared technique · relatedness 2.0/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
