---
id: giovannini26_interspeech
title: Exploring the Effect of the Visual Channel in Vocal Expression of Affect
  in an Irish (Gaelic) Synthetic Voice
authors:
  - Anna Maria Giovannini
  - Zihan Wang
  - Ailbhe Ní Chasaide
  - Christer Gobl
year: 2026
doi: 10.21437/Interspeech.2026-1363
isca_url: https://www.isca-archive.org/interspeech_2026/giovannini26_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/giovannini26_interspeech.pdf
session: Multimodal Emotion Recognition
topics:
  - tts
  - paralinguistics
  - evaluation
category: paralinguistics-emotion
institutions:
  - Trinity College Dublin
funding:
  - Irish Research Council
  - Department of Rural and Community Development and the Gaeltacht
  - National Library
code:
  url: ""
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: giovannini26_interspeech
  category: paralinguistics-emotion
  institutions:
    - Trinity College Dublin
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-1363
  pdf: https://www.isca-archive.org/interspeech_2026/giovannini26_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/giovannini26_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/giovannini26_interspeech/markdown.md
---

# Exploring the Effect of the Visual Channel in Vocal Expression of Affect in an Irish (Gaelic) Synthetic Voice

*Anna Maria Giovannini, Zihan Wang, Ailbhe Ní Chasaide, Christer Gobl*

[PDF](https://www.isca-archive.org/interspeech_2026/giovannini26_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/giovannini26_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-1363)

**Category:** `paralinguistics-emotion`

**TL;DR** — This paper investigates how visual facial expressions interact with affect-conditioned synthetic voice source parameters in an Irish (Gaelic) TTS voice. While congruent visual contexts did not significantly enhance affect perception compared to voice-only stimuli, they successfully enabled better differentiation among high- and low-activation affective states.

## Key contributions

- Created an affective audio-visual Irish corpus using Voice Source Generator (VSG) modifications and 3D avatar animations (Genesis 9.0) controlled by Facial Action Coding System (FACS) action units.
- Conducted a rigorous validation test across 19 visual stimuli with 20 subjects to isolate subtle facial expressions meeting a minimum 70% recognition threshold (except for relaxed and interested which met 45%).
- Evaluated 17 distinct audio-visual stimulus combinations (voice-only, congruent, incongruent) across 31 listeners using mixed-effects ordinal logistic regression models.
- Demonstrated that while the vocal channel dominates (congruent visuals do not boost overall intensity ratings), incongruent visual cues significantly attenuate perceived affect and visual context aids high/low activation state differentiation.

## Problem

Prior work examining affect in synthetic speech has shown that specific voice qualities often map to multiple affective states, leading to overlaps among high- or low-activation states. While multimodal emotion perception is widely studied, the interplay between voice source parameters and visual facial expressions remains underexplored for lesser-resourced languages like Irish. Specifically, it was unclear whether adding visual context would enhance affect recognition or resolve ambiguities inherent in voice-only synthetic cues.

## Method

Affective vocal stimuli were produced by modifying a single neutral carrier sentence ("Geobhaimid bád ar an lá sin") generated by a male Irish TTS voice from the ABAIR platform (Kerry dialect). Voice source modifications were performed using the Voice Source Generator (VSG), implementing the aliasing-free Liljencrants-Fant (LF) glottal flow model. The system controlled three primary source parameters: fundamental frequency ($f_0$), excitation strength ($Ee$), and the global waveshape parameter $Rd$ (correlating with vocal tension), alongside PSOLA tempo shifts (10% duration increase for sad, 10% decrease for happy) and a 4% formant frequency increase for happy. Visual stimuli were designed using a pre-rigged 3D male avatar ('Ty' from Daz 3D Genesis 9.0) with facial movements parameterized via the Facial Action Coding System (FACS) to portray angry, happy, sad, bored, relaxed, and interested affects. These were manually lip-synched and combined with the audio using Adobe Photoshop to handle variable frame rates for tempo-shifted utterances.

The perceptual evaluation utilized a 7-point scale (-3 to +3) across three paired forced-choice tests: Happy vs. Sad, Interested vs. Bored, and Angry vs. Relaxed. Listeners evaluated 17 stimulus conditions per test, encompassing voice-only, congruent, and incongruent audio-visual pairings presented in a randomized sequence after a neutral anchor.

## Experimental setup

Evaluated using a perception test with 31 human participants possessing varying levels of proficiency in Irish. The stimuli consisted of 17 conditions across 6 affects (angry, happy, sad, bored, relaxed, neutral, plus interested for visual/incongruent pairings). Statistical analysis relied on mixed-effects ordinal logistic regression models fitted in R, treating test responses as ordinal dependent variables with target affect and stimulus type as independent variables.

## Results

Voice-only and congruent stimuli achieved very high target identification rates, with happy attaining particularly strong recognition contrary to prior literature trends where it often proved elusive. Congruent visual contexts failed to yield statistically significant increases in perceived affect intensity scores compared to voice-only stimuli, pointing toward vocal channel dominance and potential saturation effects. However, incongruent stimuli caused marked score reductions (e.g., binary identification for happy dropped from 94% down to 58% when paired with sad visuals, with $\beta = -4.13, p < .001$ in the Happy-Sad test), demonstrating clear visual channel sensitivity without triggering a full polarity switch. Furthermore, congruent visual contexts significantly improved differentiation between high-activation affects ($p = .038$ for happy vs. angry) and low-activation affects ($p = .006$ in Happy-Sad, $p = .023$ in Interested-Bored).

| Stimulus Condition | Happy-Sad ID Rate | Angry-Relaxed ID Rate | Interested-Bored ID Rate |
| :--- | :--- | :--- | :--- |
| Voice-Only (Happy) | High (~94%) | - | - |
| Congruent (Happy) | High (~94%) | - | - |
| Incongruent (Happy + Sad) | Moderate (58%) | - | - |
| Voice-Only (Angry) | - | High | - |
| Congruent (Angry) | - | High | - |

## Limitations

The study is constrained by using a single male speaker/avatar voice and a single carrier sentence in Irish, which may limit generalizability across different phonetic contexts and speaker genders. The visual expressions were intentionally kept subtle to avoid overpowering the voice, which likely contributed to the absence of a statistically significant main effect for visual enhancement. Furthermore, participant proficiency in Irish varied, and the evaluation was restricted to adult native/proficient speakers of a specific dialect.

## Why read this

Speech and TTS engineers building expressive avatars or AAC systems for low-resource languages should read this to understand the complex dominance and interaction dynamics between synthetic voice source manipulations and facial animation channels.

## Code

None released (as of this page's `updated` date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

## Applications

Augmentative and Alternative Communication (AAC) systems for non-verbal users, expressive text-to-speech synthesis, and animated conversational virtual agents.

## Institutions / 機構

Trinity College Dublin

**Funding / 經費:** Irish Research Council, Department of Rural and Community Development and the Gaeltacht, National Library

## Related

- [A barrier or a booster? Familiarity effects on Mandarin emotion prosody recognition using AI-powered voice cloning](xu26i_interspeech.md) — same problem · relatedness 1.9/3
- [Eye and Mouth Cues in Audiovisual Perception of Mandarin Irony: Evidence from Eye-Tracking](xia26_interspeech.md) — same problem · relatedness 1.8/3
- [Pā‑Kakare: The First Emotional Speech Database for Te Reo Māori](rathnayake26_interspeech.md) — same problem · relatedness 1.8/3
- [MER-Live: An Interactive Browser Demo of Prosody-Driven Multimodal Emotion Recognition](song26h_interspeech.md) — same problem · relatedness 1.8/3
- [A Large-Scale Dataset of Listener Impressions of Emotional TTS](cooper26_interspeech.md) — same problem · relatedness 1.7/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
