---
id: sepanta26_interspeech
title: "From Game-Based Annotation to Representation Probing: Cross-Validated
  Prosodic Speech and Privacy Implications"
authors:
  - Sia Vosh Sepanta
  - Roberto Zamparelli
  - Alessio Brutti
year: 2026
doi: 10.21437/Interspeech.2026-2459
isca_url: https://www.isca-archive.org/interspeech_2026/sepanta26_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/sepanta26_interspeech.pdf
session: Challenges in Speech Data Collection, Curation, and Annotation
topics:
  - speech-llm
  - emotion-recognition
  - dataset
category: paralinguistics-emotion
labels:
  - dataset-or-benchmark-release
institutions:
  - Fondazione Bruno Kessler
  - University of Trento
funding:
  - European Union
code:
  url: ""
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: sepanta26_interspeech
  category: paralinguistics-emotion
  labels:
    - dataset-or-benchmark-release
  institutions:
    - Fondazione Bruno Kessler
    - University of Trento
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-2459
  pdf: https://www.isca-archive.org/interspeech_2026/sepanta26_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/sepanta26_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/sepanta26_interspeech/markdown.md
---

# From Game-Based Annotation to Representation Probing: Cross-Validated Prosodic Speech and Privacy Implications

*Sia Vosh Sepanta, Roberto Zamparelli, Alessio Brutti*

[PDF](https://www.isca-archive.org/interspeech_2026/sepanta26_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/sepanta26_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-2459)

**Category:** `paralinguistics-emotion` · **Labels:** `dataset-or-benchmark-release`

**TL;DR** — This paper introduces the Actor's Challenge (AC), a game-based crowdsourced prosodic speech corpus with built-in human validation, and uses it to evaluate automatic speech emotion recognition (ASER) and privacy leakage of demographic traits in speech-LLM pipelines. Results show that AC yields robust, contextually realistic emotion recognition models while revealing that age and emotion attributes persist in intermediate representations.

## Key contributions

- Developed a crowdsourced Games-With-A-Purpose (GWAP) framework where players alternate as performers and casting directors to collect and self-validate emotional prosody data.
- Created a multilingual prosodic speech corpus (English, Italian, French, German) featuring 1,018 completed recordings and 1,489 evaluations across 7 emotion classes.
- Evaluated self-supervised representations (emotion2vec) on the AC corpus, demonstrating superior cross-corpus generalization compared to traditional acted corpora like Emozionalmente.
- Probed a speech-LLM architecture (Whisper, MEUSLI projector, EuroLLM-1.7B) to expose privacy leakage, showing strong preservation of emotion and speaker age in intermediate hidden states.

## Problem

Traditional emotional speech datasets face severe limitations, including small pools of professional actors that restrict demographic and acoustic diversity, lack of contextual grounding for elicited phrases, and narrow linguistic coverage. These shortcomings cause models trained on canonical datasets to overfit to clean, prototypical studio conditions and fail in real-world scenarios. Furthermore, as speech-to-text and multimodal speech-LLM architectures become ubiquitous, understanding whether and where sensitive biometric traits (like age and emotion) leak through internal representations remains an open security challenge.

## Method

The data collection platform, Actor's Challenge (AC), uses Stanislavskian acting principles where users read neutral target phrases (e.g., 'It's a cappuccino') embedded within specific communicative contexts to elicit authentic emotional prosody. Players rotate between 'Auditioner' (performing) and 'Casting Director' (evaluating and rating performance on a 1-5 Likert scale) roles, providing continuous quality control. Audio is resampled to 16 kHz for downstream tasks.

For Automatic Speech Emotion Recognition (ASER), the frozen self-supervised model emotion2vec base extracts embeddings, which are aggregated via mean pooling, mean-std concatenation, or attention pooling, and fed into lightweight classifiers (logistic regression or a 2-layer MLP). For privacy probing, a SLAM-based speech-LLM pipeline is evaluated using a frozen Whisper large-v3 encoder (Stage E), a trainable linear MEUSLI projector (Stage Z), and EuroLLM-1.7B-Instruct final hidden states restricted to speech-aligned positions (Stage H). Shallow MLPs probe these frozen stages to measure feature recoverability for emotion and speaker age.

## Experimental setup

Evaluations used the AC corpus (English and Italian subsets), RAVDESS (English, professional acted), and Emozionalmente (Italian). Metrics included Accuracy, Macro-F1, AUC, and balanced accuracy (bACC). Probing experiments utilized speaker-disjoint splits and 7-class configuration for age based on Common Voice metadata, and 4-class/7-class settings for emotion.

## Results

On 4-class ASER, AC achieves 0.73 Accuracy / 0.77 Macro-F1 for English and 0.75 Accuracy / 0.73 Macro-F1 for Italian, outperforming Emozionalmente (0.56 Acc / 0.31 Macro-F1) though trailing RAVDESS (0.92 Acc / 0.92 Macro-F1) due to AC's naturalistic variability. In cross-corpus testing, training on combined AC (EN+IT) yields strong transfer to RAVDESS (0.93 Accuracy). In pipeline probing, 7-class emotion accuracy peaks at Stage Z (Projector, 0.41) and Stage E (Encoder, 0.42), up from 0.21 at Stage H. Speaker age is exceptionally recoverable at the Whisper encoder stage (AUC 0.93), but heavily attenuates by the final LLM hidden states (AUC 0.50).

| System / Condition | Accuracy | Macro-F1 |
|---|---|---|
| AC (English) | 0.73 | 0.77 |
| AC (Italian) | 0.75 | 0.73 |
| AC (EN+IT combined) | 0.60 | 0.61 |
| RAVDESS (English) | 0.92 | 0.92 |
| Emozionalmente (Italian) | 0.56 | 0.31 |

## Limitations

The dataset currently has limited volume (1,018 recordings across 240 users) and an uneven distribution favoring emotional prosody over attitudinal prosody. Engagement tends to drop over time due to the cognitive effort required for complex contextual prompts. Furthermore, privacy probing was constrained to frozen backbones, and demographic evaluation was limited primarily to age and emotion.

## Why read this

Speech and ML researchers building multimodal speech-LLMs or emotion recognition systems should read this to understand how crowdsourced GWAP frameworks can generate ecologically valid speech datasets, and to learn how sensitive speaker attributes leak across speech-LLM pipeline stages.

## Code

None released (as of this page's `updated` date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

## Applications

Automatic speech emotion recognition, trustworthy speech-LLM development, privacy-preserving speech representation learning, and crowdsourced affective dataset generation.

## Institutions / 機構

Fondazione Bruno Kessler, University of Trento

**Funding / 經費:** European Union

## Related

- [Collecting Prosody in the Wild: A Content-Controlled, Privacy-First Smartphone Protocol and Empirical Evaluation](koch26_interspeech.md) — shared data / evaluation · relatedness 2.0/3
- [From Single to Multi-Label SER: Dataset and Mamba-Based Fusion Model](tran26c_interspeech.md) — same problem · relatedness 2.0/3
- [Personal Attribute Leakage in Federated Speech Models](alali26_interspeech.md) — same problem · relatedness 2.0/3
- [Prosody-Aware Speech Representations for Emotion Recognition under Pragmatic Ambiguity](park26l_interspeech.md) — same problem · relatedness 2.0/3
- [ERM-MinMaxGAP: Benchmarking and Mitigating Gender Bias in Multilingual Multimodal Speech-LLM Emotion Recognition](pang26b_interspeech.md) — same problem · relatedness 2.0/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
