All papers
Speech codingFull-paper digest

Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio

Phillip Long, Zachary Novack, Chris Donahue

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.7 KB · Ready to paste

Preview copied content

TL;DR — The paper introduces Trilobyte, a hierarchical byte-level tokenization scheme that reduces autoregressive language model vocabulary scaling from exponential to constant, enabling the first tractable neural lossless compression of full-fidelity 24-bit audio.

Key contributions

  • Proposed Trilobyte, a hierarchical byte-level tokenization scheme that reduces vocabulary scaling from O(2^b) to O(1) in bit depth.
  • Provided the first comprehensive benchmarking of autoregressive language model-based lossless compression on full-fidelity 16-bit and 24-bit audio across diverse domains (speech, music, bioacoustics).
  • Introduced a multi-bit-rate masking strategy allowing a single generalist model to losslessly compress arbitrary bit depths (8-bit, 16-bit, 24-bit) without retraining.
  • Demonstrated that learned autoregressive compression consistently outperforms FLAC at 8-bit and modestly at 16-bit, while trailing FLAC by 9% at 24-bit due to inherent least-significant-bit noise.

Problem

Prior autoregressive language model lossless compression approaches were strictly limited to 8-bit audio at 16 kHz because sample-level tokenization creates exponentially scaling vocabularies—65K tokens for 16-bit and 16.7M tokens for 24-bit—which require tens of billions of parameters just for output projections. This makes them computationally intractable for professional or CD-quality audio (16-bit/24-bit at 44.1 kHz+). Without addressing this vocabulary explosion, it remained an open question whether neural compression could scale to practical fidelities and compete with traditional codecs like FLAC.

Method

The authors utilize decoder-only Transformer architectures (GPT-2 style, 90M parameters for 8/16-bit and larger for 24-bit) combined with arithmetic coding to perform lossless compression via next-token prediction, where negative log-likelihood directly dictates the expected bits per token. Standard sample-level tokenization treats every PCM sample as a discrete token, yielding vocabularies of size 2^b that quickly exceed memory limits. To solve this, Trilobyte decomposes each b-bit sample into B = ceil(b/8) bytes and interleaves them (MSB to LSB), maintaining a constant vocabulary size of |V| = 256 regardless of bit depth while implicitly learning byte-position distributions autoregressively.

For stereo audio, left and right channels are concatenated in random order rather than sample-level interleaved to allow the model to leverage cross-channel correlations across longer spans. Training is performed using cross-entropy loss over 300K steps. For multi-bit-rate transfer experiments, a single model is trained by randomly dropping out lower-significance bytes using a learned null token (p = 0.1), enabling arbitrary bit-depth compression at inference time.

Experimental setup

Evaluated across multiple domains: music (MusDB18, Beethoven, YouTube Mix, Commercial 16/24-bit up to 192 kHz), speech (LibriSpeech, LJSpeech, SC09, VCTK), bioacoustics (Birdvox), and sound effects (Epidemic Sound). Evaluated at native bit depths (8-bit, 16-bit, 24-bit) and sample rates ranging from 16 kHz to 48 kHz. Baselines include FLAC (compression level 8), sample-level tokenization, and in-context compression using pre-trained Llama-2-7B. Models were trained for a fixed 300K steps with a standard 90M parameter backbone.

Results

At 8-bit, Trilobyte matches standard tokenization and substantially outperforms FLAC, achieving 370%, 163%, and 119% improvements on Beethoven, YouTube Mix, and SC09 respectively. At 16-bit, Trilobyte yields modest but consistent gains over FLAC (e.g., 31% on MusDB18 Mono, 21% on LibriSpeech, 29% on Epidemic Sound). At 24-bit commercial music, Trilobyte achieves a 1.48x compression rate, trailing FLAC (1.63x) by 9%, which the authors attribute to FLAC's Rice coding efficiently compressing imperceptible least-significant-bit noise. The pre-trained Llama-2-7B in-context baseline underperforms both FLAC and trained models across nearly all settings.

System / ConditionSC09 (8-bit)LibriSpeech (16-bit)MusDB18 Mixes (16-bit)Commercial 24-bitAverage Gain vs FLAC
FLAC (Level 8)0.95x1.74x1.87x1.63xBaseline
In-context (Llama-2-7B)1.80x1.64x1.27x1.07xWorse
Standard (Sample-level)2.08x2.06x1.85xIntractableStrong at low bit depth
Trilobyte (Proposed)2.08x2.11x2.08x1.48x+18% at 16-bit
Trilobyte (Transfer)2.88x2.10x1.98x1.47xUniversal multi-bit model

Limitations

ML-based compression models are orders of magnitude slower than FLAC, meaning their modest compression wins do not currently justify their high computational and inference costs for real-world deployment. At 24-bit, the model trails FLAC by 9%, likely struggling with high-resolution least-significant-bit noise patterns. The evaluation is primarily constrained to datasets with controlled resampling (mostly 44.1 kHz/48 kHz) and does not explore extreme multi-channel setups beyond stereo.

Why read this

Speech and ML researchers working on autoregressive waveform modeling or neural compression should read this paper to understand how byte-level tokenization overcomes vocabulary bottlenecks at scale, and to realistically assess the limits of neural lossless compression against traditional algorithms like FLAC.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Lossless audio archiving, on-device audio storage optimization, and general-purpose byte-stream neural compression.

Institutions

University of California, San Diego, Carnegie Mellon University

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1748