All papers
Paralinguistics & emotionFull-paper digest

Exploring the Effect of the Visual Channel in Vocal Expression of Affect in an Irish (Gaelic) Synthetic Voice

Anna Maria Giovannini, Zihan Wang, Ailbhe Ní Chasaide, Christer Gobl

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.3 KB · Ready to paste

Preview copied content

TL;DR — This paper investigates how visual facial expressions interact with affect-conditioned synthetic voice source parameters in an Irish (Gaelic) TTS voice. While congruent visual contexts did not significantly enhance affect perception compared to voice-only stimuli, they successfully enabled better differentiation among high- and low-activation affective states.

Key contributions

  • Created an affective audio-visual Irish corpus using Voice Source Generator (VSG) modifications and 3D avatar animations (Genesis 9.0) controlled by Facial Action Coding System (FACS) action units.
  • Conducted a rigorous validation test across 19 visual stimuli with 20 subjects to isolate subtle facial expressions meeting a minimum 70% recognition threshold (except for relaxed and interested which met 45%).
  • Evaluated 17 distinct audio-visual stimulus combinations (voice-only, congruent, incongruent) across 31 listeners using mixed-effects ordinal logistic regression models.
  • Demonstrated that while the vocal channel dominates (congruent visuals do not boost overall intensity ratings), incongruent visual cues significantly attenuate perceived affect and visual context aids high/low activation state differentiation.

Problem

Prior work examining affect in synthetic speech has shown that specific voice qualities often map to multiple affective states, leading to overlaps among high- or low-activation states. While multimodal emotion perception is widely studied, the interplay between voice source parameters and visual facial expressions remains underexplored for lesser-resourced languages like Irish. Specifically, it was unclear whether adding visual context would enhance affect recognition or resolve ambiguities inherent in voice-only synthetic cues.

Method

Affective vocal stimuli were produced by modifying a single neutral carrier sentence ("Geobhaimid bád ar an lá sin") generated by a male Irish TTS voice from the ABAIR platform (Kerry dialect). Voice source modifications were performed using the Voice Source Generator (VSG), implementing the aliasing-free Liljencrants-Fant (LF) glottal flow model. The system controlled three primary source parameters: fundamental frequency (f0f_0), excitation strength (EeEe), and the global waveshape parameter RdRd (correlating with vocal tension), alongside PSOLA tempo shifts (10% duration increase for sad, 10% decrease for happy) and a 4% formant frequency increase for happy. Visual stimuli were designed using a pre-rigged 3D male avatar ('Ty' from Daz 3D Genesis 9.0) with facial movements parameterized via the Facial Action Coding System (FACS) to portray angry, happy, sad, bored, relaxed, and interested affects. These were manually lip-synched and combined with the audio using Adobe Photoshop to handle variable frame rates for tempo-shifted utterances.

The perceptual evaluation utilized a 7-point scale (-3 to +3) across three paired forced-choice tests: Happy vs. Sad, Interested vs. Bored, and Angry vs. Relaxed. Listeners evaluated 17 stimulus conditions per test, encompassing voice-only, congruent, and incongruent audio-visual pairings presented in a randomized sequence after a neutral anchor.

Experimental setup

Evaluated using a perception test with 31 human participants possessing varying levels of proficiency in Irish. The stimuli consisted of 17 conditions across 6 affects (angry, happy, sad, bored, relaxed, neutral, plus interested for visual/incongruent pairings). Statistical analysis relied on mixed-effects ordinal logistic regression models fitted in R, treating test responses as ordinal dependent variables with target affect and stimulus type as independent variables.

Results

Voice-only and congruent stimuli achieved very high target identification rates, with happy attaining particularly strong recognition contrary to prior literature trends where it often proved elusive. Congruent visual contexts failed to yield statistically significant increases in perceived affect intensity scores compared to voice-only stimuli, pointing toward vocal channel dominance and potential saturation effects. However, incongruent stimuli caused marked score reductions (e.g., binary identification for happy dropped from 94% down to 58% when paired with sad visuals, with β=−4.13,p<.001\beta = -4.13, p < .001 in the Happy-Sad test), demonstrating clear visual channel sensitivity without triggering a full polarity switch. Furthermore, congruent visual contexts significantly improved differentiation between high-activation affects (p=.038p = .038 for happy vs. angry) and low-activation affects (p=.006p = .006 in Happy-Sad, p=.023p = .023 in Interested-Bored).

Stimulus ConditionHappy-Sad ID RateAngry-Relaxed ID RateInterested-Bored ID Rate
Voice-Only (Happy)High (~94%)--
Congruent (Happy)High (~94%)--
Incongruent (Happy + Sad)Moderate (58%)--
Voice-Only (Angry)-High-
Congruent (Angry)-High-

Limitations

The study is constrained by using a single male speaker/avatar voice and a single carrier sentence in Irish, which may limit generalizability across different phonetic contexts and speaker genders. The visual expressions were intentionally kept subtle to avoid overpowering the voice, which likely contributed to the absence of a statistically significant main effect for visual enhancement. Furthermore, participant proficiency in Irish varied, and the evaluation was restricted to adult native/proficient speakers of a specific dialect.

Why read this

Speech and TTS engineers building expressive avatars or AAC systems for low-resource languages should read this to understand the complex dominance and interaction dynamics between synthetic voice source manipulations and facial animation channels.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Augmentative and Alternative Communication (AAC) systems for non-verbal users, expressive text-to-speech synthesis, and animated conversational virtual agents.

Institutions

Trinity College Dublin

Funding / 經費: Irish Research Council, Department of Rural and Community Development and the Gaeltacht, National Library

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1363