---
id: jeon26b_interspeech
title: Ego-Noise-Aware Spatial Filtering for Reliable UAV Audition in Extreme
  Low-SNR Conditions
authors:
  - Chanhong Jeon
  - Jeongmin Lee
  - Hyungjoo Seo
  - Kyuhong Shim
  - Taewook Kang
year: 2026
doi: 10.21437/Interspeech.2026-1530
isca_url: https://www.isca-archive.org/interspeech_2026/jeon26b_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/jeon26b_interspeech.pdf
session: Multi-Channel Processing and Specialized Acquisition (UAV, Radar, Hearables)
topics:
  - speech-enhancement
  - spatial-filtering
  - adaptive-beamforming
category: enhancement-separation
labels:
  - robustness-noise
institutions:
  - Sungkyunkwan University
  - University of Illinois Urbana-Champaign
funding:
  - National Research Foundation
  - IITP
  - KIAT
code:
  url: ""
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: jeon26b_interspeech
  category: enhancement-separation
  labels:
    - robustness-noise
  institutions:
    - Sungkyunkwan University
    - University of Illinois Urbana-Champaign
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-1530
  pdf: https://www.isca-archive.org/interspeech_2026/jeon26b_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/jeon26b_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/jeon26b_interspeech/markdown.md
---

# Ego-Noise-Aware Spatial Filtering for Reliable UAV Audition in Extreme Low-SNR Conditions

*Chanhong Jeon, Jeongmin Lee, Hyungjoo Seo, Kyuhong Shim, Taewook Kang*

[PDF](https://www.isca-archive.org/interspeech_2026/jeon26b_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/jeon26b_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-1530)

**Category:** `enhancement-separation` · **Labels:** `robustness-noise`

**TL;DR** — This paper presents a lightweight, model-free speech enhancement framework tailored for UAV audition systems operating under extreme ego-noise, achieving 98.12% DoA accuracy at -25 dB SNR. It uses an ego-noise-aware spatial filtering pipeline that jointly optimizes DoA estimation, crossover-based hybrid noise covariance matrix estimation, and a variance-modulated post-filter.

## Key contributions

- Phase-consistency-guided time-frequency bin selection for robust direction-of-arrival (DoA) estimation that mitigates false alarms from temporally correlated drone noise.
- A crossover-based hybrid noise covariance matrix (NCM) estimation strategy that combines closeness gains at low frequencies and directional gains at high frequencies.
- A null-reference post-filter utilizing variance modulation to suppress residual stationary ego-noise without distorting the target speech.
- An ultra-lightweight signal processing architecture requiring only 0.021 GMAC/s without relying on pre-trained neural networks.

## Problem

Unmanned aerial vehicle (UAV) audition systems struggle under severe ego-noise characterized by strong temporal correlation, frequency-varying composition, and amplitude stationarity. Traditional spatial filtering methods fail because these noise properties destabilize direction-of-arrival (DoA) estimation, corrupt noise covariance matrix (NCM) estimation, and leave persistent residual noise after beamforming. Prior deep learning solutions demand heavy computation exceeding several GMAC/s, restricting real-time on-device deployment.

## Method

The system processes multichannel audio captured by a circular microphone array (radius 0.1 m, 4 mics) via the STFT domain (512 sample window, 75% overlap). First, a phase-consistency-guided temporal similarity metric computes phase delays between microphone pairs, and a dynamic threshold based on Median Absolute Deviation (MAD with k=2.5) selects speech-informative TF bins. Spatial likelihood is then aggregated over these selected bins to determine the global DoA via statistical mode. This shared DoA reference drives subsequent spatial filtering stages.

Next, a frequency-adaptive hybrid masking strategy is used to estimate the signal covariance matrix (SCM) and NCM for an MVDR beamformer. A crossover frequency nu is defined at 3 kHz (separating regimes based on the Speech Intelligibility Index). Below nu, a closeness gain (assuming a 10-degree DoA standard deviation) handles lower frequencies where directional resolution is poor; above nu, a directional gain combining maximum directivity factor (mDF) and delay-and-sum (DS) beamformers takes over to overcome high-frequency ego-noise dominance.

Finally, a null-reference post-filter suppresses stationary residual ego-noise from the MVDR output. Using a null direction steering vector corresponding to the minimum speech likelihood over selected bins, a null-space projection estimates residual noise. A variance-modulated primary Wiener gain scales suppression depending on signal versus noise variance, driving the gain toward unity when speech is present to prevent signal distortion.

## Experimental setup

Evaluated on 160 TIMIT speech sentences synthesized with real Syma Z4W drone ego-noise via Pyroomacoustics under both anechoic and reverberant (T60 = 0.3 s) conditions. Target DoAs were set to {-45, 0, 45} degrees across five input SNR levels (-25, -20, -15, -10, -5 dB), yielding 480 samples per SNR. Baselines include MPDR, direction-gain MVDR, closeness-gain MVDR, and MMSE beamformers. Metrics include SI-SDR, PESQ, ESTOI, DNSMOS P.835 (SIG, BAK, OVRL), DoA accuracy, and Mean Absolute Error (MAE).

## Results

At an input SNR of -25 dB under anechoic conditions, the proposed method achieves an SI-SDR of -2.446 dB (outperforming the best baseline MPDR variant at -10.150 dB), a PESQ of 1.449, an ESTOI of 0.249, and a DNSMOS OVRL of 2.245. Under reverberant conditions at -25 dB SNR, it reaches an SI-SDR of -5.134 dB and a PESQ of 1.379. DoA estimation accuracy reaches 98.12% in anechoic and 94.79% in reverberant conditions at -25 dB SNR. Ablations show that adding the proposed post-filter consistently improves PESQ and background noise (BAK) scores without harming speech intelligibility.

| System (-25dB Anechoic) | SI-SDR (dB) | PESQ | ESTOI | DNSMOS OVRL |
|---|---|---|---|---|
| Noisy Mixture | -24.826 | 1.110 | 0.138 | 1.196 |
| MPDR | -12.957 | 1.214 | 0.233 | 1.582 |
| MVDR (Directional) | -11.023 | 1.230 | 0.243 | 1.684 |
| MVDR (Closeness) | -10.150 | 1.252 | 0.242 | 1.774 |
| MMSE | -11.823 | 1.233 | 0.213 | 1.686 |
| Proposed Full System | -2.446 | 1.449 | 0.249 | 2.245 |

## Limitations

The evaluation is restricted to a simulated setup using static hovering drone recordings and synthetic room impulse responses. The framework assumes a stationary UAV state and does not account for acoustic dynamics or aerodynamic turbulence shifts during active flight or non-hovering movement maneuvers. Generalization to variable multi-drone environments or uncalibrated array geometries remains unproven.

## Why read this

Researchers and engineers designing real-time, resource-constrained audio front-ends for robotic platforms will learn how to tightly couple DoA estimation and hybrid spatial covariance masking using domain-specific noise properties rather than heavy neural networks.

## Code

None released (as of this page's `updated` date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

## Applications

On-board speech recognition, emergency communication, and acoustic human-robot interaction for unmanned aerial vehicles in noisy outdoor environments.

## Institutions / 機構

Sungkyunkwan University, University of Illinois Urbana-Champaign

**Funding / 經費:** National Research Foundation, IITP, KIAT

## Related

- [DroFiT: A Lightweight Band-Fused Frequency Attention Toward Real-Time UAV Speech Enhancement](lee26n_interspeech.md) — same problem · relatedness 2.5/3
- [Joint Learning of Covariance Estimation and White Noise Gain for Robust MVDR Beamforming](deng26d_interspeech.md) — same problem · relatedness 2.0/3
- [Fast Multichannel Nonnegative Matrix Factorization with Directivity Regularization for DOA-Informed Speech Separation](ono26_interspeech.md) — same problem · relatedness 2.0/3
- [SpkGuideDOA: Speaker-wise Representation Guidance for Multiple Moving Speaker Localization](choi26g_interspeech.md) — same problem · relatedness 2.0/3
- [Towards Array-Invariant Speech Enhancement via Geometry-Aware Dynamic Convolution](liu26d_interspeech.md) — same problem · relatedness 2.0/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
