All papers
Health & clinical speechFull-paper digest

Analyzing Longitudinal Vocal Changes During Cognitive Behavioral Therapy for Hikikomori Patients

Samara S. Leal, Stavros Ntalampiras, Antonio Trabacca, Marcella Bellani, Roberto Sassi

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.4 KB · Ready to paste

Preview copied content

TL;DR — This paper investigates longitudinal vocal changes in hikikomori patients undergoing internet-based cognitive behavioral therapy, demonstrating that session-by-session speech trajectories predict treatment outcomes better than static pre-post comparisons. A feature fusion model combining Wav2vec 2.0 embeddings with traditional Mel-frequency cepstral coefficients and fundamental frequency achieves the highest F1-scores across age and gender groups.

Key contributions

  • Identifies stable longitudinal Mel-frequency cepstral coefficients (MFCC1-4, MFCC6) characterized by consistent temporal trends and low early inter-subject variability.
  • Formulates a longitudinal trajectory modeling pipeline using session-level audio summaries from 35 hikikomori patients across 8 CBT sessions (276.7 total speech hours).
  • Evaluates multiple acoustic representation strategies (MFCCs, Wav2vec 2.0, and their feature/decision fusion) under a Leave-One-Patient-Out (LOPO) cross-validation protocol.
  • Demonstrates that fusing self-supervised deep representations with hand-crafted spectral and prosodic descriptors consistently outperforms single-modality baselines.

Problem

Monitoring treatment response during cognitive behavioral therapy for individuals with prolonged social isolation like hikikomori is difficult because prior methods rely on static classification or cross-sectional snapshots like DAIC-WoZ and MODMA. Furthermore, existing longitudinal studies focus predominantly on pre-post differences or symptom scales alone rather than tracking continuous within-individual temporal speech dynamics. These static formulations miss critical intra-individual fluctuations and fail to capture early signs of treatment response or clinical deterioration.

Method

The study analyzes a dataset of 35 patients comprising 276.7 hours of speech (approx. 59.3 minutes per session) captured across 8 CBT sessions via the AWS Chime SDK (48 kHz, mono). Audio was segmented into 5-second windows, preprocessed using RMS normalization and pre-emphasis filtering, and speaker-diarized via clinician-annotated reference segments and k-means clustering (k=2) to yield a Diarization Error Rate (DER) of 0.1698. Segment-level acoustic features were aggregated into session summaries (S1-S8) and baseline-corrected relative to the first session to isolate within-subject temporal variation.

For RQ1, temporal scores combining longitudinal slope, variability, late-session change, and early inter-subject variance identified the top-5 stable MFCCs. For RQ3, the framework evaluates four acoustic representations (AR): AR1 (all MFCCs), AR2 (selected MFCCs), AR3 (Wav2vec 2.0 only via facebook/wav2vec2-base-960h processed with a 1-layer GRU with 128 hidden units), and AR4 (fusion of Wav2vec 2.0 GRU outputs with session-level MFCC+F0 descriptors processed via an MLP). Models were trained for 10 epochs per fold using BCEWithLogitsLoss with per-fold class weights, optimized using GridSearchCV for MLPs and an F2-oriented threshold calibration to prioritize sensitivity to clinical deterioration.

Experimental setup

The dataset contains 35 patients (24 young adults, 11 adolescents; 57.14% female, 42.86% male) who completed 8 CBT sessions, identified by a Hikikomori Questionnaire score above 42. Models were evaluated using a Leave-One-Patient-Out (LOPO) cross-validation protocol across age and gender strata, compared against various feature configurations (AR1 to AR4), and measured using macro recall mean/std, positive-class recall, specificity, precision, and positive-class F1-score.

Results

Session-level acoustic trajectories successfully revealed distinct speech evolution patterns between patients showing clinical improvement or worsening, outperforming simple pre-post difference correlations. For the young adult cohort, the fusion model (AR4) achieved a positive-class F1-score of 0.37 with a positive recall of 0.70, compared to AR3 (Wav2vec 2.0 only) which achieved an F1 of 0.00 due to zero positive recall. For adolescents, AR4 reached an F1-score of 0.62 and positive recall of 0.71. In gender stratifications, AR4 achieved F1-scores of 0.47 for females and 0.40 for males, consistently outperforming unirepresentational baselines (AR1–AR3) which suffered from poor sensitivity or high false negatives.

CohortRepresentationModelMacro Recall (Mean)SpecificityPositive RecallPositive F1
YAAR1 (All MFCCs)MLP0.250.390.110.15
YAAR3 (Wav2vec2)GRU0.581.000.000.00
YAAR4 (Fusion)GRU+MLP0.520.360.700.37
ADAR1 (All MFCCs)MLP0.320.140.510.55
ADAR3 (Wav2vec2)GRU0.340.250.430.30
ADAR4 (Fusion)GRU+MLP0.720.750.710.62

Limitations

Clinical severity scores were available only at pre- and post-treatment intervals rather than continuously session-by-session, constraining direct supervision of intermediate trajectory changes. The sample size is relatively small (35 patients), limiting broader model generalization and statistical robustness across age and gender subgroups.

Why read this

Researchers and engineers building speech-based digital health tools will learn how to formulate longitudinal, multi-session speech tracking pipelines instead of relying on static pre-post classification. It provides empirical proof that fusing self-supervised representations with handcrafted spectral features bridges the gap between global structure and clinical sensitivity.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Automated telepsychology monitoring tools, remote cognitive behavioral therapy support systems, and objective vocal biomarker tracking for depressive disorders and social withdrawal.

Institutions

University of Milan, Scientific Institute IRCCS E. Medea, University of Verona, Azienda Ospedaliera Universitaria Integrata

Funding / 經費: European Union

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-742