TL;DR — Step-Audio-R1 introduces Modality-Grounded Reasoning Distillation (MGRD) to solve the inverted reasoning scaling problem in audio LLMs, achieving an average score of 83.6% on speech-to-text benchmarks that outperforms Gemini 2.5 Pro.
Key contributions
- Identifies textual surrogate reasoning as the root cause of performance degradation during extended audio model deliberation.
- Proposes Modality-Grounded Reasoning Distillation (MGRD), an iterative framework combining self-distillation and multimodal RL to anchor reasoning in acoustic features.
- Implements a composite reward strategy incorporating a think-format penalty to prevent systematic reasoning collapse and maintain 2,300-2,800 token thought chains.
- Develops a DPO-based self-cognition calibration pipeline that reduces audio-perception denial errors from 6.76% down to 0.02%.
Problem
Audio language models consistently suffer from inverted scaling behavior where performance degrades as test-time reasoning length increases, unlike text and vision domains. Prior work attributes this to audio being inherently resistant to reasoning or uses language-model judges that only enforce consistency. The authors uncover that models engage in textual surrogate reasoning—deliberating over transcript-like abstractions rather than low-level acoustic properties like pitch contours, timbre, or rhythmic structures. This stems from initializing CoT capabilities via text-derived supervised fine-tuning, which creates a modality mismatch.
Method
Step-Audio-R1 is built on a Qwen2 audio encoder, a downsampling audio adaptor that compresses frame rate to 12.5 Hz, and a Qwen2.5-32B LLM decoder. The framework processes latent audio features directly through the LLM decoder to produce explicit reasoning chains inside <think> tags followed by the final answer.
The training pipeline begins with a Cold-Start phase using 5M samples (1B text and 4B audio tokens) for joint SFT and reinforcement learning with verified rewards (RLVR) with 10% distilled audio CoT. Following this, the MGRD framework executes iterative cycles of self-distillation where the model generates candidate reasoning responses on perceptual audio data. These are filtered for acoustic grounding, logic, and correctness to curate .
The subsequent multimodal reinforcement learning uses PPO with a zero KL penalty coefficient, a clipping range of 0.2, a discount factor of 1.0, and a maximum sequence length of 10,240 tokens. The reward function combines a 0.8 weight for answer accuracy and a 0.2 weight for the presence of the reasoning format. A difficulty-based filtering strategy selects samples where pass@8 falls within [3, 6], discarding trivial or unsolvable tasks to maintain stable reward convergence.
Experimental setup
Evaluated on speech-to-text benchmarks including Big Bench Audio (BBA), Spoken MQA (SMQA), MMSU, MMAU, and Wild Speech (WS), alongside the Big Bench Audio speech-to-speech benchmark. Baselines include Step-Audio 2, Gemini 2.5 Pro, Gemini 3 Pro, GPT-4o mini Realtime, and Gemini 2.5 Flash. The architecture utilizes a 32-billion parameter Qwen2.5 LLM decoder with PPO rollout sampling of 16 candidates per prompt.
Results
Step-Audio-R1 achieves an average score of 83.6% across S2T benchmarks, outperforming Gemini 2.5 Pro (81.5%) and approaching Gemini 3 Pro (85.1%). On individual tasks, it achieves 98.7% on Big Bench Audio and 95.2% on Spoken MQA. In the speech-to-speech setting, Step-Audio-R1 Realtime reaches 96.1% reasoning accuracy with a first-packet latency of 0.92 seconds. Ablations show that removing format rewards causes reasoning length to collapse from 3,000 tokens down to below 1,500 tokens, whereas format-rewarded models maintain stable lengths of 2,300–2,800 tokens and improve MMAU accuracy from 76.5 to 77.7.
| Model | Avg. | BBA | SMQA | MMSU | MMAU | WS |
|---|---|---|---|---|---|---|
| Step-Audio 2 | 68.3 | 59.1 | 88.8 | 64.3 | 78.0 | 51.1 |
| Gemini 2.5 Pro | 81.5 | 96.1 | 94.8 | 79.3 | 77.4 | 60.0 |
| Gemini 3 Pro | 85.1 | 92.1 | 95.3 | 82.9 | 78.9 | 76.4 |
| Step-Audio-R1 | 83.6 | 98.7 | 95.2 | 75.9 | 77.7 | 70.6 |
Limitations
The approach relies heavily on a multi-stage iterative pipeline that requires careful difficulty-based filtering and custom reward shaping, which can be computationally expensive. The evaluation focuses primarily on English-centric or standard multilingual speech/audio perception benchmarks, leaving broader zero-shot dialect or extreme low-resource acoustic adaptation unverified. Furthermore, first-packet latency sits at 0.92 seconds for the realtime variant, which is slightly higher than some non-reasoning low-latency counterparts.
Why read this
Speech and ML researchers working on audio LLMs or test-time compute scaling should read this to understand why traditional CoT fails for audio and how modality-grounded reinforcement learning can successfully unlock deliberation scaling.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Real-time conversational speech assistants, complex acoustic scene analysis, and multi-modal dialogue systems requiring deep audio reasoning.
Institutions
Shanghai Jiao Tong University, University of New South Wales, Nanyang Technological University, StepFun
Related
- Audio-DeepThinker: Progressive Reasoning-Aware Reinforcement Learning for High-Quality Chain-of-Thought Emergence in Audio Language Models — same problem · relatedness 2.6/3
- Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models — same problem · relatedness 2.5/3
- ALARM: Audio–Language Alignment for Reasoning Models — same problem · relatedness 2.4/3
- Nudging Hidden States: Training-Free Model Steering for Chain-of-Thought Reasoning in Large Audio-Language Models — same problem · relatedness 2.2/3
- Enhancing Audio Reasoning via Semantic Summary Prediction — same problem · relatedness 2.2/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-256