All papers
Paralinguistics & emotionFull-paper digest

Stress Detection Across Daily Activities: A Context-Aware Multimodal Framework with Trajectory and Ambient Speech

Wei-Heng Huang, Woan-Shiuan Chien, Chi-Chun Lee

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.4 KB · Ready to paste

Preview copied content

TL;DR — This paper proposes a context-aware multimodal framework (Trajectory Speech Embedding, TSE) that integrates daily acoustic features and indoor mobility trajectories to detect occupational stress in hospital workers. It improves F1 score by 9.32% and MCC by 0.038 compared to audio-only baselines.

Key contributions

  • Proposes a multimodal stress detection framework (TSE) combining daily audio representations and session-based indoor mobility trajectories.
  • Utilizes a self-supervised next-location prediction pre-training objective on auxiliary trajectory-only data to learn intrinsic spatial relationships across hospital areas.
  • Formulates daily acoustic features via openSMILE LLDs aggregated through 15 statistical functionals to form 360-dimensional daily vectors.
  • Demonstrates through mobility-stratified analysis how spatial trajectory data and acoustic signals symbiotically recover unimodal classification failures in high-dynamic and low-mobility regimes.

Problem

Occupational stress detection in healthcare is predominantly reliant on intrusive physiological wearables tracking heart rate variability (HRV), which limits real-world scaling and continuous monitoring. While speech offers a non-intrusive alternative, acoustic models lack environmental and situational context essential for interpreting stress in high-pressure clinical settings. Prior approaches fail to capture spatial dynamics, such as how stress levels vary across hospital rooms, ward types, and activity routines.

Method

The framework processes daily audio by extracting 24 low-level descriptors (LLDs: 8 prosodic, 16 spectral) using openSMILE, computing 15 statistical functionals per descriptor to yield a 360-dimensional vector. This sequence is passed through a 2-layer Transformer encoder with GELU activations prepended with a learnable [REP] token to generate a daily acoustic embedding. For mobility, indoor Wi-Fi/Bluetooth RSSI signals from 243 devices across 21 nursing units yield minute-level trajectories (1440 mins/day), merged into sessions of location tokens and durations, combined with sinusoidal periodic temporal embeddings.

The trajectory encoder is a 2-layer Transformer pre-trained via self-supervised next-location prediction on 2,426 extra trajectory-only days (achieving 38.85% top-1 and 60.40% top-3 accuracy). The spatial and acoustic daily embeddings are concatenated via a late-fusion strategy into a joint representation vector passed to an MLP classifier. The model is trained end-to-end using a joint objective combining fusion cross-entropy loss and auxiliary supervised classification losses for each modality branch weighted by lambda = 0.5.

Experimental setup

Evaluated on the TILES-2018 dataset containing real-world data from hospital workers over 10 weeks (166 participants with audio, 2,878 days with audio/trajectory, plus 2,426 trajectory-only days). Binarized global mean self-reported stress labels (31.9% positive). Compared against Bi-LSTM, multi-similarity loss (MSLoss), unimodal acoustic Transformer, and unimodal trajectory Transformer baselines. Metrics include Accuracy, Balanced Accuracy (BACC), F1-score, and Matthews Correlation Coefficient (MCC). Implemented with subject-independent 80/10/10 splits, batch size 32, up to 500 epochs with early stopping, AdamW optimizer (trajectory encoder learning rate 1e-5, others 1e-4).

Results

Coordinate TSE achieves the best performance with an MCC of 0.147, BACC of 58.91%, and F1 of 38.81%, compared to the Bi-LSTM baseline MCC of 0.109, audio-only Transformer MCC of 0.074, and MSLoss MCC of 0.051. Stratified mobility analysis reveals peak fusion performance in high-mobility conditions (F1-score 51.5%), where frequent contextual transitions complement acoustics, and in low-mobility conditions where structured location patterns offer robust behavioral cues.

Systems / ConditionsAccF1BACCMCC
MSLoss [21]58.7028.9753.180.051
Bi-LSTM [22]70.1129.4955.860.109
Transformer (Audio)54.6434.2054.450.074
Transformer (Trajectory)55.3635.2355.490.091
TSE54.6438.6558.590.142
Coordinate TSE56.0738.8158.910.147

Limitations

The evaluation is restricted to a single hospital dataset (TILES-2018) limiting generalized demographic and environmental scope. Indoor location tracking relies heavily on hospital-specific infrastructure (RSSI from Wi-Fi/Bluetooth beacons), which may not transfer smoothly to settings lacking dense positioning setups. Day-level aggregation and binary label thresholds discard fine-grained short-term emotional fluctuations.

Why read this

Researchers building ambient intelligence or occupational health monitoring systems will find a clear blueprint for fusing spatial trajectories with speech representations. It provides practical insights into how multi-modal architectures recover from unimodal failures under varying behavioral dynamics.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Automated mental health monitoring, continuous occupational stress tracking for healthcare personnel, and context-aware workplace wellness systems.

Institutions

National Tsing Hua University, National Yang Ming Chiao Tung University

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1262