TL;DR — This paper investigates how visual facial expressions interact with affect-conditioned synthetic voice source parameters in an Irish (Gaelic) TTS voice. While congruent visual contexts did not significantly enhance affect perception compared to voice-only stimuli, they successfully enabled better differentiation among high- and low-activation affective states.
Key contributions
- Created an affective audio-visual Irish corpus using Voice Source Generator (VSG) modifications and 3D avatar animations (Genesis 9.0) controlled by Facial Action Coding System (FACS) action units.
- Conducted a rigorous validation test across 19 visual stimuli with 20 subjects to isolate subtle facial expressions meeting a minimum 70% recognition threshold (except for relaxed and interested which met 45%).
- Evaluated 17 distinct audio-visual stimulus combinations (voice-only, congruent, incongruent) across 31 listeners using mixed-effects ordinal logistic regression models.
- Demonstrated that while the vocal channel dominates (congruent visuals do not boost overall intensity ratings), incongruent visual cues significantly attenuate perceived affect and visual context aids high/low activation state differentiation.
Problem
Prior work examining affect in synthetic speech has shown that specific voice qualities often map to multiple affective states, leading to overlaps among high- or low-activation states. While multimodal emotion perception is widely studied, the interplay between voice source parameters and visual facial expressions remains underexplored for lesser-resourced languages like Irish. Specifically, it was unclear whether adding visual context would enhance affect recognition or resolve ambiguities inherent in voice-only synthetic cues.
Method
Affective vocal stimuli were produced by modifying a single neutral carrier sentence ("Geobhaimid bád ar an lá sin") generated by a male Irish TTS voice from the ABAIR platform (Kerry dialect). Voice source modifications were performed using the Voice Source Generator (VSG), implementing the aliasing-free Liljencrants-Fant (LF) glottal flow model. The system controlled three primary source parameters: fundamental frequency (), excitation strength (), and the global waveshape parameter (correlating with vocal tension), alongside PSOLA tempo shifts (10% duration increase for sad, 10% decrease for happy) and a 4% formant frequency increase for happy. Visual stimuli were designed using a pre-rigged 3D male avatar ('Ty' from Daz 3D Genesis 9.0) with facial movements parameterized via the Facial Action Coding System (FACS) to portray angry, happy, sad, bored, relaxed, and interested affects. These were manually lip-synched and combined with the audio using Adobe Photoshop to handle variable frame rates for tempo-shifted utterances.
The perceptual evaluation utilized a 7-point scale (-3 to +3) across three paired forced-choice tests: Happy vs. Sad, Interested vs. Bored, and Angry vs. Relaxed. Listeners evaluated 17 stimulus conditions per test, encompassing voice-only, congruent, and incongruent audio-visual pairings presented in a randomized sequence after a neutral anchor.
Experimental setup
Evaluated using a perception test with 31 human participants possessing varying levels of proficiency in Irish. The stimuli consisted of 17 conditions across 6 affects (angry, happy, sad, bored, relaxed, neutral, plus interested for visual/incongruent pairings). Statistical analysis relied on mixed-effects ordinal logistic regression models fitted in R, treating test responses as ordinal dependent variables with target affect and stimulus type as independent variables.
Results
Voice-only and congruent stimuli achieved very high target identification rates, with happy attaining particularly strong recognition contrary to prior literature trends where it often proved elusive. Congruent visual contexts failed to yield statistically significant increases in perceived affect intensity scores compared to voice-only stimuli, pointing toward vocal channel dominance and potential saturation effects. However, incongruent stimuli caused marked score reductions (e.g., binary identification for happy dropped from 94% down to 58% when paired with sad visuals, with in the Happy-Sad test), demonstrating clear visual channel sensitivity without triggering a full polarity switch. Furthermore, congruent visual contexts significantly improved differentiation between high-activation affects ( for happy vs. angry) and low-activation affects ( in Happy-Sad, in Interested-Bored).
| Stimulus Condition | Happy-Sad ID Rate | Angry-Relaxed ID Rate | Interested-Bored ID Rate |
|---|---|---|---|
| Voice-Only (Happy) | High (~94%) | - | - |
| Congruent (Happy) | High (~94%) | - | - |
| Incongruent (Happy + Sad) | Moderate (58%) | - | - |
| Voice-Only (Angry) | - | High | - |
| Congruent (Angry) | - | High | - |
Limitations
The study is constrained by using a single male speaker/avatar voice and a single carrier sentence in Irish, which may limit generalizability across different phonetic contexts and speaker genders. The visual expressions were intentionally kept subtle to avoid overpowering the voice, which likely contributed to the absence of a statistically significant main effect for visual enhancement. Furthermore, participant proficiency in Irish varied, and the evaluation was restricted to adult native/proficient speakers of a specific dialect.
Why read this
Speech and TTS engineers building expressive avatars or AAC systems for low-resource languages should read this to understand the complex dominance and interaction dynamics between synthetic voice source manipulations and facial animation channels.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Augmentative and Alternative Communication (AAC) systems for non-verbal users, expressive text-to-speech synthesis, and animated conversational virtual agents.
Institutions
Trinity College Dublin
Funding / 經費: Irish Research Council, Department of Rural and Community Development and the Gaeltacht, National Library
Related
- A barrier or a booster? Familiarity effects on Mandarin emotion prosody recognition using AI-powered voice cloning — same problem · relatedness 1.9/3
- Eye and Mouth Cues in Audiovisual Perception of Mandarin Irony: Evidence from Eye-Tracking — same problem · relatedness 1.8/3
- Pā‑Kakare: The First Emotional Speech Database for Te Reo Māori — same problem · relatedness 1.8/3
- MER-Live: An Interactive Browser Demo of Prosody-Driven Multimodal Emotion Recognition — same problem · relatedness 1.8/3
- A Large-Scale Dataset of Listener Impressions of Emotional TTS — same problem · relatedness 1.7/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-1363