All papers
Paralinguistics & emotionFull-paper digest

SEAM: Shortcut-Aware Real-Time Detection of Scripted vs. Spontaneous Speech for Interview Guardrails

Vsevolod Kovalev, Pranay Manocha

Code & resourcesgithub.com/vsevolod-kovalev/seam

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.6 KB · Ready to paste

Preview copied content

TL;DR — SEAM is a shortcut-aware real-time framework for detecting scripted versus spontaneous speech in job interviews, achieving 0.971 ROC-AUC on external interview data using a compact 23M-parameter DistilHuBERT backbone.

Key contributions

  • Integrates uniform waveform preprocessing, seam-aware sampling, and non-speech noise-bank augmentation to suppress spurious corpus, channel, and boundary shortcuts.
  • Demonstrates that shortcut-prevention components improve external domain transfer despite lowering internal held-out accuracy.
  • Identifies a compact DistilHuBERT configuration (23M parameters, 41.8 MB at INT4) optimized for low-latency real-time deployment.
  • Establishes a rigorous transfer-oriented evaluation protocol using both grouped internal splits and an adversarial external interview benchmark.

Problem

Distinguishing scripted from spontaneous speech is useful for interview guardrails, but prior models often exploit spurious shortcuts like microphone characteristics, room acoustics, codec artifacts, and dataset-specific provenance boundaries rather than true speaking style. This causes strong internal benchmark performance to collapse under slight distribution shifts in real-world deployments. Creating a lightweight, real-time guardrail requires an architecture that is both computationally efficient and explicitly hardened against shortcut learning.

Method

SEAM uses a pretrained DistilHuBERT encoder coupled with a 2-layer MLP classification head (ReLU, dropout p=0.30) trained with BCEWithLogitsLoss using the AdamW optimizer (weight decay 0.01). The pipeline processes audio at mono 16 kHz through a uniform waveform preprocessing chain: DC removal, a 70 Hz high-pass biquad filter, integrated loudness normalization to -23 LUFS, and peak limiting at 0.99 clamped to [-1, 1]. To prevent provenance and channel shortcuts, audio is stored in 10-minute FLAC chunks and sampled using seam-aware constraints (never crossing recording boundaries; padding with non-speech if necessary) combined with a 14-hour non-speech noise bank (room tone, breathing) mixed additively at a sampled SNR over 40-70% of the window.

The training regimen explores shallow adaptation strategies, ultimately selecting 8-second analysis windows and unfreezing only the top transformer layer with constant learning rates of 5e-6 for the encoder and 3e-4 for the head across 3 epochs. For inference, window-level outputs are aggregated via median logit pooling, and post-training quantization is applied to reduce the footprint for live stream deployment.

Experimental setup

The internal dataset comprises 4 English corpora (People's Speech, PodcastFillers, LibriSpeech, Spoken Wikipedia) totaling 240 hours per class, split into grouped 80/10/10 train/eval/test sets across 2,238 total speakers. The external evaluation benchmark consists of 720 proprietary English interview clips balanced across scripted/spontaneous styles and clean/mixed room-microphone conditions. Systems are evaluated using window-level Accuracy and ROC-AUC, implemented with a fixed-budget ablation regime (batch size 256, 4 shards/class, 120 steps) and a full-training regime on NVIDIA A100 and L4 hardware.

Results

Under the full-training regime, the final SEAM model achieves an internal test accuracy of 0.9623 (±0.0068) and ROC-AUC of 0.9766 (±0.0045), alongside an external interview-domain accuracy of 0.9517 (±0.0077) and ROC-AUC of 0.9713 (±0.0039). Ablation studies prove that removing shortcut defenses (noise bank and seam-aware sampling) artificially boosts internal test ROC-AUC from 0.9287 to 0.9557, but drastically crashes external interview transfer ROC-AUC from 0.8991 down to 0.7324.

Regarding window lengths, 8-second windows outperform 2s, 4s, and 12s variants. For adaptation depth, unfreezing only the top transformer layer (tr1, Test AUC 0.9287) surpasses head-only tuning, deeper unfreezing (tr2), and partial CNN unfreezing (cnn2). Post-training quantization successfully compresses the model to 48.74 MB (INT8) and 41.80 MB (INT4) with negligible degradation in external AUC (0.9700 and 0.9743 respectively).

System / ConditionTest AccTest AUCExt AccExt AUC
Baseline (Full SEAM)0.96230.97660.95170.9713
Seam-aware off0.93280.93280.81790.8674
SpecAugment instead0.94050.94050.78070.8202
Noise bank off0.94910.94910.70890.7518
Noise off + Seam off0.95570.95570.68820.7324

Limitations

The framework remains English-first with only partial zero-shot transfer capabilities across other languages. Scriptedness labels remain inherently entangled with genre and recording practices, preventing the model from capturing a complete, pristine representation of speaking style. The evaluation is restricted to a proprietary interview domain and requires broader testing across diverse accents, recording hardware, and acoustic environments.

Why read this

Speech and ML engineers building real-time audio guardrails should read this paper to learn how to prevent self-supervised models from overfitting to dataset shortcuts and channel artifacts.

Code

Applications

Real-time AI interview guardrails, conversational system integrity checks, and speaking-style corpus curation.

Institutions

Symbal AI, Princeton University

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1480