TL;DR — An LLM-based pipeline converts non-verbal acoustic cues from mental health hotline recordings into explicit textual markers (paralinguistic injection) and uses reasoning-enhanced auxiliary supervision, achieving a macro F1 of 0.802 on three-way crisis-level classification.
Key contributions
- Proposed a paralinguistic injection pipeline that translates non-verbal vocal cues (sobbing, trembling voice, low volume) into explicit textual annotations appended to ASR transcripts, bridging the modality gap for text LLMs.
- Introduced reasoning-enhanced training using LLM-generated TAF (Triage Assessment Form) clinical rationales as a multi-task auxiliary loss alongside the classification objective.
- Applied a chunk-based data augmentation strategy (splitting calls into non-overlapping 5-minute segments) to overcome authentic clinical data scarcity, expanding 154 raw calls into over 900 chunks.
- Demonstrated substantial performance gains over direct SpeechLLM fine-tuning, zero-shot LLM prompting, and traditional OpenSMILE acoustic baselines.
Problem
Automated psychological crisis assessment from real-world support hotline audio is challenging due to extreme data scarcity, privacy sensitivities, and the fact that text-only ASR transcripts strip away crucial paralinguistic and affective cues needed for clinical triage. Prior approaches either rely on handcrafted acoustic features (e.g., OpenSMILE) that fail to capture semantic context, zero-shot LLMs that struggle without domain adaptation, or end-to-end SpeechLLMs that underutilize prosody and vocal affect in data-scarce regimes.
Method
The framework operates in a two-step pipeline. First, raw audio is transcribed using Paraformer-zh, and non-verbal paralinguistic cues are extracted using Step-Audio-R1 guided by Triage Assessment Form (TAF) affective domain criteria. These are merged into the ASR output as paralinguistically enriched transcripts (e.g., adding tags like '[Trembling voice, obvious sobbing...; Emotion: Sadness > Anxiety]').
Second, the enriched transcript is fed into a fine-tuned Qwen2.5-7B-Instruct model using LoRA (rank r = 8, alpha = 64 on attention projections) via a multi-task objective. During training, the model undergoes two parallel forward passes: a classification pass predicting the 3-way crisis level (0, 1, or 2) via cross-entropy loss (L_cls), and a generation pass using teacher forcing to output step-by-step TAF-aligned clinical rationales (Affective, Behavioral, Cognitive domain scores) generated by gpt-oss-120b (L_gen). The final objective is L = L_cls + L_gen.
For inference, calls are segmented into contiguous 5-minute chunks to preserve temporal flow, processed by the model, and aggregated using majority voting at the call level. Training uses AdamW with a learning rate of 3e-5, cosine annealing, 10% warmup, effective batch size 16 (batch size 1 with gradient accumulation over 16 steps), and a single NVIDIA A800 GPU.
Experimental setup
Evaluated on a Chinese psychological support hotline dataset comprising 154 authentic calls (~100 hours total, 39.5 minutes per sample on average) annotated into three balanced classes by two expert chief directors (55 no crisis, 42 low crisis, 57 medium-to-high crisis). Evaluated using call-level 5-fold cross-validation with Accuracy and Macro F1-score as primary metrics, comparing against OpenSMILE (eGeMAPSv02 + SVM), zero-shot LLM (gpt-oss-120b with enriched transcript and TAF prompt), and SpeechLLM (Qwen2.5-Omni-7B).
Results
The proposed system achieves a headline Macro F1-score of 0.802 and Accuracy of 0.805, substantially outperforming the zero-shot LLM baseline (F1: 0.371), the OpenSMILE acoustic baseline (F1: 0.471), and the direct SpeechLLM fine-tuning baseline (F1: 0.551). Ablation studies confirm that removing data augmentation causes the steepest drop (Macro F1 decreases by 10.0% to 0.702), omitting paralinguistic injection drops F1 by 4.1% (to 0.761, matching an alternative emotion2vec adapter baseline at 0.760), and dropping the auxiliary reasoning loss results in a 1.7% decrease (to 0.785).
| System / Condition | Accuracy | Macro F1 |
|---|---|---|
| Zero-shot LLM (w/ Enriched Transcript) | 0.455 | 0.371 |
| OpenSMILE Baseline | 0.486 ± 0.053 | 0.471 ± 0.062 |
| SpeechLLM Baseline | 0.564 ± 0.075 | 0.551 ± 0.079 |
| Ours (Full Model) | 0.805 ± 0.061 | 0.802 ± 0.062 |
| w/o Paralinguistic Injection | 0.764 ± 0.089 | 0.761 ± 0.088 |
| w/o Auxiliary Loss | 0.790 ± 0.075 | 0.785 ± 0.075 |
Limitations
The dataset is restricted in scale (154 calls, ~100 hours from a single Chinese hotline center) and language coverage (Mandarin Chinese only), which may limit cross-cultural and cross-dialect generalizability. The auxiliary reasoning targets are synthetically generated by an LLM rather than verified by clinical experts for intermediate TAF domain scores. Evaluation is retrospective and conducted under offline cross-validation rather than prospective deployment settings.
Why read this
Speech and ML researchers tackling low-resource clinical audio tasks will learn how explicit textual injection of paralinguistic features outperforms end-to-end SpeechLLM fine-tuning. It provides a blueprint for combining multi-task LLM reasoning chains with chunk-based data augmentation in safety-critical domains.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Decision support systems for psychological support hotline operators, automated triage triage-assistance tools, and real-time clinical monitoring of mental health distress.
Institutions
Tsinghua University, Peking University Huilongguan Clinical Medical School, WHO Collaborating Centre for Research and Training in Suicide Prevention
Related
- Towards Paradigm-General Suicide Risk Detection via Speech LLM — same problem · relatedness 2.2/3
- Moot-Court: Training-Free Dialectical Reasoning for Depression Detection — same problem · relatedness 2.1/3
- Revisiting Emotion-Based Triage: Evidence from French Emergency Call Data — same problem · relatedness 2.0/3
- Prosody-Aware Speech Representations for Emotion Recognition under Pragmatic Ambiguity — shared technique · relatedness 2.0/3
- Investigating LLMs Behavior in Depression Severity Prediction — same problem · relatedness 1.9/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-997