---
id: gourav26_interspeech
title: "DsNA(Digital sigNature for Audios): A Unique Method to Fingerprint Audio
  Files Generated by Text to Speech"
authors:
  - Vishal Gourav
  - Phanindra Mankale
year: 2026
isca_url: https://www.isca-archive.org/interspeech_2026/gourav26_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/gourav26_interspeech.pdf
session: Speech Synthesis, Voice Conversion and Audio Generation
topics:
  - tts
  - evaluation
  - audio-deepfake
category: deepfake-security
institutions:
  - Oracle
code:
  url: ""
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: gourav26_interspeech
  category: deepfake-security
  institutions:
    - Oracle
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://www.isca-archive.org/interspeech_2026/gourav26_interspeech.html
  pdf: https://www.isca-archive.org/interspeech_2026/gourav26_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/gourav26_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/gourav26_interspeech/markdown.md
---

# DsNA(Digital sigNature for Audios): A Unique Method to Fingerprint Audio Files Generated by Text to Speech

*Vishal Gourav, Phanindra Mankale*

[PDF](https://www.isca-archive.org/interspeech_2026/gourav26_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/gourav26_interspeech.html)

**Category:** `deepfake-security`

**TL;DR** — DsNA (Digital sigNature for Audios) is a lightweight post-generation framework that embeds verifiable cryptographic provenance directly into text-to-speech audio by interleaving signature shards with audio chunks, achieving identical signal-to-noise ratios (45.2 dB) compared to original outputs.

## Key contributions

- Introduces DsNA, a post-generation audio fingerprinting framework that operates independently of underlying text-to-speech architectures.
- Proposes a structured embedding mechanism using RSA-inspired cryptographic hashing, partitioning digital signatures into shards interleaved with segmented audio chunks.
- Eliminates reliance on external databases or metadata lookup tables by making the audio artifact self-contained and intrinsically verifiable.
- Demonstrates preservation of audio fidelity, maintaining an unchanged signal-to-noise ratio of 45.2 dB alongside stable perceptual and transcription metrics.

## Problem

As neural text-to-speech systems such as WaveNet and Tacotron achieve high realism, ensuring audio provenance, authenticity, and preventing misuse become critical challenges. Traditional multimedia fingerprinting methods focus heavily on content identification or copyright protection rather than embedding verifiable provenance information directly into generated outputs. Furthermore, external matching systems or lookup tables are fragile and prone to metadata stripping or synchronization loss. Addressing this requires an intrinsic, self-contained embedding scheme that preserves acoustic quality without modifying internal neural TTS pipelines.

## Method

The DsNA pipeline operates as a post-generation wrapper on synthetic speech waveforms paired with associated metadata. In the injection stage, an audio chunker divides the waveform into two equal chunks. Simultaneously, a fingerprint generator creates a compact digital signature derived from the audio and metadata using principles inspired by RSA-based cryptographic authentication. A shard generator splits this signature into three distinct shards, which are then interleaved with the audio chunks to produce the final fingerprinted output.

To authenticate an audio file, the reverse DsNA pipeline extracts the interleaved shards, reassembles the cryptographic signature, and validates it against the embedded metadata and audio content. This design avoids modifications to the input text processing, acoustic model, and neural vocoder stages of standard TTS pipelines. By operating strictly at the waveform level post-generation, the framework guarantees compatibility with diverse acoustic architectures while keeping computational overhead minimal.

## Experimental setup

The framework was evaluated using a set of 250 audio samples generated from a baseline TTS model. Evaluation metrics include Signal-to-Noise Ratio (SNR) for acoustic fidelity, Mean Opinion Score (MOS) for perceptual quality, and Character Error Rate (CER) and Word Error Rate (WER) to assess transcription consistency through an evaluation pipeline. The prototype demonstrates a baseline uncorrupted signal-to-noise ratio of 45.2 dB.

## Results

The evaluation confirms that DsNA preserves audio signal fidelity, reporting an identical SNR of 45.2 dB for both original and fingerprinted audio samples. Perceptual quality remains virtually unchanged, with the MOS shifting negligibly from 4.51 in the original audio to 4.48 in the fingerprinted version. Transcription metrics show zero degradation, maintaining a Character Error Rate (CER) of 1.8% and a Word Error Rate (WER) of 2.6% across both conditions, proving that the chunking and interleaving process does not compromise linguistic intelligibility.

| Metric | Original Audio | Fingerprinted Audio |
|---|---|---|
| SNR (dB) | 45.2 | 45.2 |
| MOS | 4.51 | 4.48 |
| CER (%) | 1.8 | 1.8 |
| WER (%) | 2.6 | 2.6 |

## Limitations

The current study serves as an initial proof of concept evaluated on a limited scale of 250 samples from a single dataset environment. The framework lacks empirical testing regarding its robustness against real-world audio transformations, such as lossy compression (e.g., MP3/AAC), background noise, channel distortions, and re-recording. Additionally, resilience against malicious adversarial manipulation or intentional signature stripping requires further cryptographic hardening and security analysis.

## Why read this

Researchers and engineers working on audio deepfake detection, provenance, and synthetic speech watermarking should read this to understand a novel waveform-level interleaving approach for self-contained cryptographic signatures. It offers a practical blueprint for embedding identity markers without altering upstream generative model weights or depending on external lookup servers.

## Code

None released (as of this page's `updated` date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

## Applications

Tracing the provenance of synthetic speech in content moderation platforms, preventing deepfake audio dissemination, and embedding copyright verification in enterprise text-to-speech generation tools.

## Institutions / 機構

Oracle

## Related

- [Phoneme-Aware Mamba Watermark: An Active Defense System Against Purified Speech Deepfakes](shao26_interspeech.md) — same problem · relatedness 2.7/3
- [DuraMark: Duration-Embedded Watermarking in LLM-based TTS](mou26_interspeech.md) — same problem · relatedness 2.7/3
- [AudioNoisePrints: Model-free audio watermarking using spatial correlation in flow matching TTS](tse26_interspeech.md) — same problem · relatedness 2.6/3
- [A Training-Free Proactive Defense Against Partial Speech Manipulation via Self-Embedding Steganography](ozer26_interspeech.md) — same problem · relatedness 2.1/3
- [VoxWatermark: A Large-Scale Benchmark for Audio Watermark Detection under Perturbations](sedaghati26_interspeech.md) — same problem · relatedness 2.0/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
