All papers
Paralinguistics & emotionFull-paper digest

What Makes Synthetic Speech Sound Sarcastic? A Prosody-Controlled Perception Study

Zhu Li, Shekhar Nayak, Matt Coler

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.6 KB · Ready to paste

Preview copied content

TL;DR — The paper investigates the causal roles of individual prosodic dimensions in sarcasm perception using neural text-to-speech synthesis, finding that human listeners rely primarily on loudness while foundational multimodal models rely on speech rate.

Key contributions

  • Constructs a fully crossed, orthogonally manipulated stimulus set using neural TTS to isolate speech rate, pitch variation, and loudness.
  • Compares human listener ratings of sarcasm and naturalness against predictions from a large-scale multimodal foundation model (Qwen3-Omni).
  • Reveals a significant divergence in acoustic cue weighting: humans are driven primarily by loudness, whereas the foundation model is driven primarily by speech rate.
  • Provides a validated methodology for using controllable generative speech models as experimental testbeds in psycholinguistics.

Problem

Prior research on sarcasm perception relies heavily on natural speech where acoustic dimensions like pitch, tempo, and intensity co-vary, making it impossible to isolate their individual causal contributions. Context-minimal paradigms and post-hoc acoustic analyses describe group differences but fail to determine which specific vocal features independently bias perception. Furthermore, it remains unknown whether current foundation models process and weigh prosodic cues for sarcasm in the same way human listeners do, highlighting a gap in behavioral alignment.

Method

The study utilized short, semantically neutral English utterances adapted from Bryant and Fox Tree as text inputs for Qwen3-TTS-12Hz-1.7B-CustomVoice. A single synthetic speaker voice was used to eliminate inter-speaker variability. The authors implemented a fully crossed 2 x 2 x 2 factorial design manipulating pitch variation (dynamic vs. flat), loudness (loud vs. soft), and speech rate (fast vs. slow) via natural-language prompting. Sampling temperature was adjusted during autoregressive decoding to control prosodic variability (e.g., lower temperatures for flat conditions). To ensure orthogonality, candidate pool samples (100 per utterance/condition) were extracted and filtered using effect-size contrasts measured by Cohen's d. Final stimuli achieved large effect sizes in target dimensions (pitch d=1.14, loudness d=0.81, duration d=1.76) while keeping non-target dimensions near zero (|d|<0.25). Voice quality metrics (H1-H2, HNR) confirmed no systematic side effects (p>0.05).

For evaluation, 66 participants rated 192 total stimuli (24 utterances x 8 conditions) on 5-point Likert scales for sarcasm and naturalness. For machine perception, Qwen3-Omni was fed the exact same audio waveforms with a fixed prompt instructing it to evaluate prosodic cues across 6 random seeds to reduce stochasticity. Statistical analyses utilized linear mixed-effects models (lme4 package) with fixed effects for the three prosodic factors and random intercepts for participants and items, followed by Tukey-adjusted pairwise comparisons.

Experimental setup

The experiment used 24 short, semantically neutral English text utterances synthesized into 192 total stimuli across 8 prosodic conditions. Human evaluation engaged 66 native or near-native English speakers. Machine evaluation used Qwen3-Omni across 6 random inference seeds. Metrics included 5-point Likert scales for perceived sarcasm and naturalness, evaluated via intraclass correlation coefficients (ICC) and linear mixed-effects models.

Results

For human naturalness, fast stimuli were perceived as more natural than slow (beta=0.09, p<0.001) and soft stimuli were more natural than loud (beta=0.11, p<0.001). For human sarcasm perception, loudness was the dominant main effect (beta=0.29, p<0.05), where loud stimuli received significantly higher sarcasm ratings than soft ones; speech rate and pitch contour showed no significant main effects. Pairwise comparisons confirmed that flat-pitch, loud conditions (e.g., fast flat loud, slow flat loud) were rated significantly more sarcastic than soft counterparts.

In contrast, the foundation model (Qwen3-Omni) exhibited a completely different weighting scheme: it showed a significant main effect of speech rate on sarcasm (beta=0.31, p<0.01), rating slow stimuli as more sarcastic, with no significant main effect for loudness (beta=0.04, p=0.77). Spearman rank correlation revealed no significant overall alignment between human and model ratings (rho=-0.11, p=0.26).

Cue / ConditionHuman Sarcasm Beta (beta)Machine Sarcasm Beta (beta)
Speech Rate (slow)0.0610.313
Pitch Variation (flat)0.1380.132
Loudness (loud)0.2850.035

Limitations

The study used context-free stimuli without rich discourse pragmatics, measuring acoustic bias rather than full communicative intent, which may explain mid-range rating clusters. Many human participants were non-native English speakers, potentially altering cue weighting. Sarcasm was treated as a single monolithic category, ignoring subtypes of verbal irony. Additionally, synthetic manipulation via text prompts, while statistically orthogonal, may not perfectly replicate natural covariance patterns in human speech.

Why read this

Speech and ML researchers building multimodal audio LLMs should read this paper to understand the severe divergence in prosodic cue weighting between humans and foundation models. It offers a concrete blueprint for using controllable neural TTS to rigorously audit speech-language models for human-like pragmatic perception.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Improving the pragmatic alignment of conversational AI agents, speech-to-speech translation systems, and affective text-to-speech synthesis.

Institutions

University of Groningen

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1487