TL;DR — Speech2See is an end-to-end framework for speech-driven open-world object detection that bypasses intermediate text generation by coupling HuBERT speech features with Grounding DINO visual representations via a progressive pre-training and fine-tuning paradigm, achieving 56.2 AP on closed-set COCO-val and 42.7 AP on zero-shot COCO.
Key contributions
- Proposes a fully end-to-end architecture for direct 'listening-to-localization' object detection without requiring an intermediate Automatic Speech Recognition (ASR) text-to-detector pipeline.
- Introduces the Query-Guided Semantic Aggregation (QSA) module, which employs learnable queries via cross-attention to resolve temporal redundancy in continuous audio sequences.
- Incorporates a Mixture-of-LoRA-Experts (MoLE) module into the cross-modal decoder during fine-tuning to handle paralinguistic variations and speaker diversity without full model retraining.
- Demonstrates state-of-the-art performance across COCO and LVIS benchmarks, outperforming two-stage speech-vision baselines by +17.0 AP on closed-set COCO.
Problem
Prior audio grounding systems, such as YOSS, rely on decoupled two-stage pipelines that combine separately trained audio, text (CLIP), and vision models, which introduces error propagation and convergence bottlenecks. Furthermore, cascaded ASR-plus-text-detector pipelines suffer from accumulation of transcription errors and high inference latency. Directly learning open-set detection from raw speech is further compounded by extreme data scarcity for paired audio-image data and the continuous, entangled nature of acoustic signals compared to discrete text tokens.
Method
Speech2See takes multi-scale visual features via a Swin Transformer and raw speech embeddings via a HuBERT-Base encoder. The raw HuBERT feature sequence is processed through the Query-Guided Semantic Aggregation (QSA) module, where learnable queries act as a semantic adapter to condense temporal speech representations into compact semantic tokens via cross-attention. These tokens are fed into a Grounding DINO-inspired vision-language detector with a feature enhancer and speech-guided query selection.
During the pre-training phase (10 epochs on AdamW with lr and batch size 64), the Swin-T and HuBERT backbones and main detector are frozen, optimizing only the QSA and projection layers. During the fine-tuning phase (2 epochs), the decoder weights are frozen while a Mixture-of-LoRA-Experts (MoLE) module with rank and experts is inserted into the feed-forward network layers. A Top-1 router assigns tokens to experts to capture diverse acoustic patterns, backed by an auxiliary load balancing loss () to prevent expert underutilization.
Experimental setup
Experiments are conducted on COCO 2017 (80 categories, 123K images), Objects365 (365 categories, 609K images), Flickr30k (159K audio files), and LVIS (1,000 categories, 164K images). Audio annotations are synthesized from object category text using edge-TTS across 10 distinct speakers. Models use Swin-T and HuBERT-Base backbones (197.8M total parameters) and are trained on 8 NVIDIA A800 GPUs. Baselines include YOSS (base/large), text-driven Grounding DINO, and a cascaded Whisper-plus-Grounding-DINO system. Evaluation metrics include standard COCO AP, , , and LVIS category splits.
Results
On closed-set COCO-val, Speech2See with MoLE achieves 56.2 AP, 71.3 , and 60.7 , outperforming the two-stage YOSS-large baseline (39.2 AP) by 17.0 AP points. In zero-shot COCO evaluations trained on Objects365, Flickr30k, and GQA, the model reaches 42.7 AP (55.7 , 46.8 ), surpassing the closed-set supervised performance of YOSS-large. On the zero-shot LVIS benchmark, it achieves 19.9 AP overall, outperforming YOSS-large (16.3 AP). Ablation studies show that replacing the QSA module with a standard MLP drops zero-shot AP catastrophically from 42.7 down to 27.2 AP, confirming QSA's vital role in parsing temporal speech redundancy. Increasing MoLE experts from to raises AP from 40.7 to 42.7, whereas yields no further gains.
| System / Condition | Pre-Training Data | MoLE | AP | ||
|---|---|---|---|---|---|
| YOSS-large (+GQA) | Flickr, COCO | No | 39.2 | 53.3 | 42.6 |
| Ours (Closed-set) | COCO | No | 54.1 | 70.2 | 59.6 |
| Ours (Closed-set) | COCO | Yes | 56.2 | 71.3 | 60.7 |
| Ours (Zero-shot) | Obj365, Flickr, GQA | No | 40.4 | 52.7 | 44.2 |
| Ours (Zero-shot) | Obj365, Flickr, GQA | Yes | 42.7 | 55.7 | 46.8 |
Limitations
The work relies entirely on synthesized speech data generated via edge-TTS rather than real-world human-spoken audio grounding datasets, which limits claims regarding robustness in authentic acoustic environments. Performance still lags behind text-driven counterparts (e.g., Grounding DINO at 25.6 LVIS AP vs. 19.9 for Speech2See) due to continuous acoustic variability and paralinguistic entanglement.
Why read this
Researchers and engineers building multimodal speech-vision systems or looking for alternatives to cascaded ASR-plus-text-detector pipelines will find this a blueprint for bypassing text bottlenecks via parameter-efficient cross-modal adaptation.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Robotic manipulation, voice-guided navigation, and multimodal human-computer interaction where users query visual scenes directly via speech commands.
Institutions
Tianjin University, PipeChina Institute of Science and Technology
Related
- EchoLoc: Audio-Aware Object Grounding via Joint Heatmap and Box-Level Localization — same problem · relatedness 2.1/3
- Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation — shared technique · relatedness 1.6/3
- Speech Codec Probing from Semantic and Phonetic Perspectives — complementary · relatedness 1.6/3
- A Multimodal Semi-Supervised Framework for Automatic Construction of a Cross-Lingual Taigi Speech-Chinese Subtitle Corpus — shared technique · relatedness 1.6/3
- Discrete vs. Continuous: A Comprehensive Study of Unified Audio Understanding in LALMs — shared technique · relatedness 1.6/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-2183