All papers
Health & clinical speechFull-paper digest

Speech-based Psychological Crisis Assessment using LLMs

Terumi Chiba, Yang Luo, Ziyun Cui, Yongsheng Tong, Chao Zhang

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.7 KB · Ready to paste

Preview copied content

TL;DR — An LLM-based pipeline converts non-verbal acoustic cues from mental health hotline recordings into explicit textual markers (paralinguistic injection) and uses reasoning-enhanced auxiliary supervision, achieving a macro F1 of 0.802 on three-way crisis-level classification.

Key contributions

  • Proposed a paralinguistic injection pipeline that translates non-verbal vocal cues (sobbing, trembling voice, low volume) into explicit textual annotations appended to ASR transcripts, bridging the modality gap for text LLMs.
  • Introduced reasoning-enhanced training using LLM-generated TAF (Triage Assessment Form) clinical rationales as a multi-task auxiliary loss alongside the classification objective.
  • Applied a chunk-based data augmentation strategy (splitting calls into non-overlapping 5-minute segments) to overcome authentic clinical data scarcity, expanding 154 raw calls into over 900 chunks.
  • Demonstrated substantial performance gains over direct SpeechLLM fine-tuning, zero-shot LLM prompting, and traditional OpenSMILE acoustic baselines.

Problem

Automated psychological crisis assessment from real-world support hotline audio is challenging due to extreme data scarcity, privacy sensitivities, and the fact that text-only ASR transcripts strip away crucial paralinguistic and affective cues needed for clinical triage. Prior approaches either rely on handcrafted acoustic features (e.g., OpenSMILE) that fail to capture semantic context, zero-shot LLMs that struggle without domain adaptation, or end-to-end SpeechLLMs that underutilize prosody and vocal affect in data-scarce regimes.

Method

The framework operates in a two-step pipeline. First, raw audio is transcribed using Paraformer-zh, and non-verbal paralinguistic cues are extracted using Step-Audio-R1 guided by Triage Assessment Form (TAF) affective domain criteria. These are merged into the ASR output as paralinguistically enriched transcripts (e.g., adding tags like '[Trembling voice, obvious sobbing...; Emotion: Sadness > Anxiety]').

Second, the enriched transcript is fed into a fine-tuned Qwen2.5-7B-Instruct model using LoRA (rank r = 8, alpha = 64 on attention projections) via a multi-task objective. During training, the model undergoes two parallel forward passes: a classification pass predicting the 3-way crisis level (0, 1, or 2) via cross-entropy loss (L_cls), and a generation pass using teacher forcing to output step-by-step TAF-aligned clinical rationales (Affective, Behavioral, Cognitive domain scores) generated by gpt-oss-120b (L_gen). The final objective is L = L_cls + L_gen.

For inference, calls are segmented into contiguous 5-minute chunks to preserve temporal flow, processed by the model, and aggregated using majority voting at the call level. Training uses AdamW with a learning rate of 3e-5, cosine annealing, 10% warmup, effective batch size 16 (batch size 1 with gradient accumulation over 16 steps), and a single NVIDIA A800 GPU.

Experimental setup

Evaluated on a Chinese psychological support hotline dataset comprising 154 authentic calls (~100 hours total, 39.5 minutes per sample on average) annotated into three balanced classes by two expert chief directors (55 no crisis, 42 low crisis, 57 medium-to-high crisis). Evaluated using call-level 5-fold cross-validation with Accuracy and Macro F1-score as primary metrics, comparing against OpenSMILE (eGeMAPSv02 + SVM), zero-shot LLM (gpt-oss-120b with enriched transcript and TAF prompt), and SpeechLLM (Qwen2.5-Omni-7B).

Results

The proposed system achieves a headline Macro F1-score of 0.802 and Accuracy of 0.805, substantially outperforming the zero-shot LLM baseline (F1: 0.371), the OpenSMILE acoustic baseline (F1: 0.471), and the direct SpeechLLM fine-tuning baseline (F1: 0.551). Ablation studies confirm that removing data augmentation causes the steepest drop (Macro F1 decreases by 10.0% to 0.702), omitting paralinguistic injection drops F1 by 4.1% (to 0.761, matching an alternative emotion2vec adapter baseline at 0.760), and dropping the auxiliary reasoning loss results in a 1.7% decrease (to 0.785).

System / ConditionAccuracyMacro F1
Zero-shot LLM (w/ Enriched Transcript)0.4550.371
OpenSMILE Baseline0.486 ± 0.0530.471 ± 0.062
SpeechLLM Baseline0.564 ± 0.0750.551 ± 0.079
Ours (Full Model)0.805 ± 0.0610.802 ± 0.062
w/o Paralinguistic Injection0.764 ± 0.0890.761 ± 0.088
w/o Auxiliary Loss0.790 ± 0.0750.785 ± 0.075

Limitations

The dataset is restricted in scale (154 calls, ~100 hours from a single Chinese hotline center) and language coverage (Mandarin Chinese only), which may limit cross-cultural and cross-dialect generalizability. The auxiliary reasoning targets are synthetically generated by an LLM rather than verified by clinical experts for intermediate TAF domain scores. Evaluation is retrospective and conducted under offline cross-validation rather than prospective deployment settings.

Why read this

Speech and ML researchers tackling low-resource clinical audio tasks will learn how explicit textual injection of paralinguistic features outperforms end-to-end SpeechLLM fine-tuning. It provides a blueprint for combining multi-task LLM reasoning chains with chunk-based data augmentation in safety-critical domains.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Decision support systems for psychological support hotline operators, automated triage triage-assistance tools, and real-time clinical monitoring of mental health distress.

Institutions

Tsinghua University, Peking University Huilongguan Clinical Medical School, WHO Collaborating Centre for Research and Training in Suicide Prevention

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-997