All papers
Speech codingFull-paper digest

OmniCodec: Low Frame Rate Universal Audio Codec with Semantic–Acoustic Disentanglement

Jingbin Hu, Haoyu Zhang, Dake Guo, Qirui Zhan, Wenhao Li, Huakang Chen, Guobin Ma, Hanke Xie, Chengyou Wang, Pengyuan Xie, Chuan Xie, Qiang Zhang, Lei Xie

Code & resourcesgithub.com/ASLP-lab/OmniCodec

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.0 KB · Ready to paste

Preview copied content

TL;DR — OmniCodec is a universal neural audio codec operating at low frame rates (12.5 Hz / 6.25 Hz) that achieves semantic-acoustic disentanglement across speech, music, and general sound by leveraging a pre-trained understanding model's audio encoder and a self-guidance training strategy.

Key contributions

  • Replaces unsupervised models (e.g., WavLM) with the supervised Qwen3-Omni-AuT-Encoder to provide robust, cross-domain semantic representations for the codec's semantic branch.
  • Proposes a hierarchical multi-codebook design with explicit semantic-acoustic decoupling via adapter layers (subtracting quantified semantic hidden features from acoustic latents and recombining them).
  • Introduces a self-guidance loss to handle quantization error by forcing the decoder to produce consistent outputs for both continuous pre-quantized latents and discrete tokens, boosting codebook utilization.
  • Achieves pure causality and low frame rates (12.5 Hz and 6.25 Hz), supporting fully streaming and real-time fast inference across speech, music, and general sound domains.

Problem

Existing neural audio codecs prioritize high-fidelity reconstruction at extreme frame rates and bitrates, leading to multi-codebook structures that lack structural semantic alignment and are poorly suited for Large Language Model (LLM) generation tasks. Conversely, single-codebook or semantic-distilled codecs often struggle with low-frame-rate reconstruction quality across diverse, multi-domain audio (speech, music, and general sound). Prior attempts either lack unified cross-domain handling, rely on high frame rates incompatible with LLM scaling, or fail to decouple semantic properties from fine acoustic details.

Method

OmniCodec features a dual-branch architecture consisting of a semantic branch and an acoustic branch. The acoustic encoder and decoder use SEANet with streaming convolutions, combined with a causal Transformer operating with a purely causal receptive field. The semantic branch processes audio via the pre-trained Qwen3-Omni-AuT-Encoder (which compresses 16 kHz audio down to 12.5 Hz), using vector quantization (VQ) with an exponential moving average (EMA) codebook update, followed by an adapter linear layer. The acoustic branch uses residual vector quantization (RVQ) with 8 to 32 stages. Disentanglement is achieved by subtracting the quantified semantic hidden features from the acoustic hidden features and subsequently adding the quantized acoustic hidden features back.

The training objective combines multi-scale mel reconstruction loss (weighted 15.0), semantic representation reconstruction loss, VQ commitment loss, self-guidance loss (weighted 0.1), adversarial losses (using multi-scale STFT, MPD, MSD, MRD, and a WavLM-based discriminator), and feature matching loss. The self-guidance loss enforces similarity between decoder outputs driven by continuous pre-quantized latents and those driven by discrete tokens, thereby maximizing codebook utilization.

Experimental setup

Trained on ~160,000 hours of data: 95K hours of Emilia and LibriTTS for speech, ~60K hours of in-house data for music, and an 800-hour filtered AudioSet subset for general sound. Evaluated on LibriSpeech test-clean (speech), GTZAN testset (music), and an AudioSet eval subset (general sound). Baselines include WavTokenizer, UniCodec, AUV, X-codec, and Mimi. Metrics include PESQ (WB/NB), STOI, Mel distance, MCD, N-MOS, S-MOS, Audiobox Aesthetics, and LLM perplexity (PPL0, PPL mean). Implemented across 4 A100 GPUs with a global batch size of 24, AdamW optimizer (peak LR 1e-4, 2.5K warmup, 500K cosine decay steps), totaling ~134M parameters.

Results

OmniCodec-32L achieves superior performance on LibriSpeech test-clean with a PESQ-WB of 3.02, STOI of 0.96, Mel distance of 0.75, MCD of 2.54, and N-MOS of 3.63, outperforming Mimi-16L and single-codebook models like UniCodec across multiple domains at competitive bitrates. Even at an extremely low frame rate of 6.25 Hz (OmniCodec-F-32L), it outperforms the 75 Hz UniCodec model on STOI, Mel distance, and MCD. Ablation studies demonstrate that removing the self-guidance loss drops codebook utilization from 0.982 to 0.974, while dropping the semantic branch severely degrades downstream LLM perplexity from ~10.02 to 18.44.

ModelBitrate (bps)Frame Rate (Hz)PESQ-WB (Speech)STOI (Speech)Mel dis. (Speech)MCD (Speech)
WavTokenizer4804801.880.870.974.79
UniCodec1050752.650.920.793.46
X-codec-8L400050x82.960.930.793.35
Mimi-16L220012.5x162.880.941.123.93
OmniCodec-16L220012.5x162.760.940.813.05
OmniCodec-32L440012.5x323.020.960.752.54

Limitations

While OmniCodec excels in music and general sound semantic representation, its perplexity (PPL) in the speech domain lags behind models specifically distilled using WavLM due to differences in semantic architecture and training data proportions. The study relies heavily on an in-house music dataset, and fine-tuning specifically for speech (OmniCodec-8L-FT) trades away music and sound modeling capacity.

Why read this

Speech and ML researchers building audio LLMs or streaming generative models should read this to see how supervised understanding-model encoders can replace self-supervised distillations for universal, low-bitrate semantic-acoustic decoupling.

Code

Applications

Streaming speech-to-speech translation, universal audio generation via language models, real-time interactive voice assistants, and multi-domain audio compression.

Institutions

Northwestern Polytechnical University, Shanghai Lingguang Zhaxian Technology

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-494