All papers
Health & clinical speechFull-paper digest

Revisiting Emotion-Based Triage: Evidence from French Emergency Call Data

Elio Stasica, Clément Joly, Amandine Lecomte, Vincent P. Martin, Romain Serizel, Emmanuel Vincent, Tahar Chouihed

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.7 KB · Ready to paste

Preview copied content

TL;DR — This paper evaluates whether speech emotion recognition (SER) adds predictive value to automatic emergency call triage using 250 real French emergency calls. Results show that categorical and dimensional emotions fail to outperform or reliably augment basic metadata (age, sex, and speaker role) in predicting clinically grounded priority levels.

Key contributions

  • Evaluates SER models on real, naturalistic emergency call data (SAMU54) annotated with official 5-level clinical priority codes (P3 to P0).
  • Compares categorical SER (Lajavaness Wav2Vec classifier) and dimensional SER (SpeechDimEmo arousal/valence) using ordinal logistic regression models selected via AIC.
  • Exposes a counter-intuitive inverse relationship where higher caller arousal associates with lower priority levels, heavily modulated by speaker role.
  • Demonstrates that demographic metadata (age, sex) and speaker context (patient vs. family vs. clinician) vastly outperform emotion features for triage prediction.

Problem

Prior speech emotion recognition (SER) research for emergency triage relies heavily on simulated or acted datasets, assumes emotion directly maps to medical urgency without clinical validation, and treats callers as a homogeneous group. These unvalidated assumptions and corpus mismatches between acted data and naturalistic emergency call center (ECC) recordings limit real-world deployment. This work questions whether speech emotion actually correlates with professional dispatcher triage decisions in real clinical workflows.

Method

The study analyzes a balanced dataset of 250 audio recordings (50 per priority level P0-P3) from the SAMU54 emergency center spanning January to December 2024. Main-speaker audio segments are extracted using DiariZen diarization followed by manual verification in Praat, excluding overlapping speech and dispatcher turns. Two public French SER models are used: Lajavaness (a 5-class Wav2Vec categorical classifier outputting pleased, relaxed, neutral, sad, tense) and SpeechDimEmo (estimating frame-level arousal and valence, aggregated via median per recording).

Statistical analysis employs ordinal logistic regression with proportional odds models (using the MASS package in R). Model 1 tests priority against categorical emotion, speaker role, patient age, and patient sex (AIC = 726.2). Model 2 tests priority against median valence, arousal, speaker role, patient age, patient sex, and interaction terms (arousal × speaker role, arousal × valence; AIC = 718.4). Stratified regressions per speaker role are also conducted to isolate effects.

Experimental setup

Evaluated on a controlled subset of 250 real French emergency calls from SAMU54 (50 calls per priority level P0, P1, P2 SNP, P2 AMU, P3), totaling 250 speakers with annotated age, sex, duration (mean 37.0s), and caller type (patient, family, healthcare professional, other). Compared across categorical vs. dimensional SER feature sets within ordinal logistic regression models. Implemented in R 4.5.1 using the MASS and MuMIn packages.

Results

Neither categorical nor dimensional emotion features provided meaningful predictive power over metadata. In Model 1 (categorical), age (β=0.035,p<0.001\beta = 0.035, p < 0.001) and speaker role (other: β=2.13,p<0.001\beta = 2.13, p < 0.001; family: β=1.25,p<0.001\beta = 1.25, p < 0.001; clinician: β=1.21,p=0.022\beta = 1.21, p = 0.022) strongly predicted priority, whereas none of the emotion categories reached statistical significance relative to neutral. The relaxed emotion was never predicted across the entire dataset, and tense was the most frequent category across all priority levels.

In Model 2 (dimensional), arousal and valence main effects were non-significant (p=0.645p = 0.645 and p=0.243p = 0.243). Unexpectedly, higher arousal interacted negatively with family calls (β=−0.56,p=0.059\beta = -0.56, p = 0.059) and other callers (β=−1.00,p=0.024\beta = -1.00, p = 0.024), meaning high caller arousal correlated with lower assigned priority. Age (β=0.035,p<0.001\beta = 0.035, p < 0.001), male sex (β=0.493,p=0.046\beta = 0.493, p = 0.046), and speaker role remained the dominant predictors.

System / Model ConditionPredictors IncludedAICKey Significant Drivers (p<0.05p < 0.05)
Baseline Metadata OnlyAge, Sex, Speaker Role~725Age, Speaker Role (Family/Other/Clinician)
Model 1: Categorical SEREmotion, Age, Sex, Speaker Role726.2Age, Speaker Role (Emotions non-significant)
Model 2: Dimensional SERValence, Arousal, Interactions, Age, Sex, Speaker Role718.4Age, Sex, Speaker Role, Arousal ×\times Other/Valence

Limitations

The study is limited by a small sample size of 250 calls from a single French emergency center (SAMU54), restricting generalizability. P0-level calls lack direct patient speech because patients are unconscious or in cardiac arrest, biasing speaker availability. Audio recordings contain sensitive clinical data and cannot be made publicly available, limiting external replication.

Why read this

Speech and ML researchers building affective computing systems for high-stakes healthcare should read this paper to understand the limitations of treating speech emotion as a proxy for clinical urgency. It offers a cautionary empirical baseline showing that basic metadata heavily outperforms complex emotion features in real ECC environments.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Emergency call center analytics, clinical triage decision-support systems, and robust speech affective computing evaluation.

Institutions

University of Lorraine, CNRS, Inria, CHRU-Nancy, INSERM

Funding / 經費: Grand Est ENACT AI Cluster

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1265