TL;DR — Moot-Court is a training-free reasoning engine for automated depression detection that uses adversarial multi-agent debate and a dynamic experience library to guide frozen LLMs, achieving an F1-score of 0.8856 on DAIC-WOZ.
Key contributions
- Introduces Moot-Court, a training-free multimodal depression detection framework that outperforms fully fine-tuned and instruction-tuned models.
- Proposes a dialectical reasoning mechanism integrating psychiatric symptom networks into LLM diagnostics via a Prosecutor and Defense Attorney adversarial structure.
- Develops a dynamic, self-evolving experience library (Codex) governed by a reward-based feedback algorithm to constrain reasoning trajectories and suppress hallucinations without weight updates.
- Combines utterance-level symbolic acoustic markers, session-level statistical trajectories, and subject-level descriptive portraits into a unified multimodal blueprint.
Problem
Current automated depression detection methods using LLMs either rely on costly parameter adaptation (like instruction tuning or LoRA) that risks severe overfitting in low-resource clinical settings, or discard non-verbal acoustic biomarkers by operating in text-only mode. Furthermore, black-box neural networks and standard LLMs lack clinical interpretability and suffer from shortcut bias, failing to meet medical requirements for trustworthy explanations. This gap necessitates an objective, interpretable, and training-efficient framework that successfully integrates speech and text.
Method
The framework begins with multimodal blueprinting, translating raw audio and transcripts into three hierarchical tiers: utterance-level acoustic markers (M) combined with text, session-level statistical trajectories (linear regression slopes, distributional extrema, and interruption rates via F_slope, F_dist, F_inter), and subject-level descriptive portraits generated by segmenting audio into 3-minute blocks processed by Qwen-Audio and Qwen-max. These discrete fact nodes V are categorized into External Fields (V_ext), Internal States (V_int), and Evidence Markers (V_evid).
Next, adversarial agents perform dialectical reasoning. The Prosecutor builds a Pathogenic Graph (G_P) using edges that map biomarkers to symptoms (e_ind), trace feedback loops (e_refin), identify external triggers (e_gen, e_trig), and detect clinical paradoxes like smiling depression via contradiction edges (e_cong). Concurrently, the Defense Attorney constructs a Contextual Graph (G_D) using edges for reactive sadness (e_ctx), biological alibis (e_exp), anxiety comorbidity (e_diff), resilience markers (e_inhib), and manic exclusion criteria (e_excl).
In the decision phase, k=5 independent Judges with temperatures ranging from 0.5 to 0.8 evaluate the graphs using retrieved semantic priors from the dynamic Codex (governed by authority alpha, frequency f, utility u, and a newbie bonus delta_new for f < 5). Outcomes trigger retrospective reflection across three game-theoretic scenarios (consensus, system failure, or divergent judgments), updating the Codex via MODIFY, MERGE, and ADD operations driven by a statistical reward function S.
Experimental setup
Evaluated on the DAIC-WOZ dataset (107 training, 35 validation, 47 test splits), which contains semi-clinical interviews with a virtual agent assessing psychological distress. Audio is denoised using ZipEnhancer. Acoustic features are extracted using Librosa (latency, WPS, F0, RMS energy) and Parselmouth (jitter, shimmer). The system uses frozen LLM backbones including DeepSeek V3.2 and Qwen3 series (30B, 80B). Performance is measured using Precision, Recall, and F1-score.
Results
Moot-Court using DeepSeek V3.2 achieves an F1-score of 0.8856, a Recall of 0.9394, and a Precision of 0.8378, outperforming fully supervised convolutional and recurrent architectures, fully fine-tuned SSL-SDD (F1 0.8290), LoRA-tuned GPT-2 (F1 0.8260), and instruction-tuned DepressInstruct (F1 0.8235). Cross-backbone evaluations show that Qwen3-30b achieves an F1 of 0.8234 and Recall of 0.9545 entirely training-free.
Ablation studies reveal that direct vanilla multimodal inference yields an F1 of 0.6296 with a low recall of 0.5152. Adding linear reasoning raises F1 to 0.7143. Introducing dialectical reasoning alone drives precision to 0.8333 but drops recall to 0.6061 due to over-conservatism; the dynamic Codex successfully resolves this bottleneck by boosting recall to 0.9394. Furthermore, removing audio features and relying solely on text drops the F1-score from 0.8856 down to 0.7890, confirming the critical complementarity of acoustic biomarkers.
| Method | Strategy | F1 | Rec. | Pre. |
|---|---|---|---|---|
| SSL-SDD [28] | Full Fine-tuning | .8290 | – | – |
| GPT-2 [29] | LoRA | .8260 | .8590 | .7950 |
| LLM-Landmark [30] | LoRA/P-tuning | .8330 | – | – |
| DepressInstruct [32] | Instruction Tuning | .8235 | .8400 | .8077 |
| (Ours) DeepSeek V3.2 [21] | Training Free | .8856 | .9394 | .8378 |
Limitations
The framework's performance is inherently bounded by the capabilities and token context limits of the underlying foundation models (e.g., DeepSeek and Qwen series). Evaluation is restricted to the DAIC-WOZ dataset, meaning generalizability across diverse languages, cultural contexts, and clinical interview protocols remains unverified. Additionally, reliance on multiple LLM reasoning passes and judge ensembles introduces higher inference latency compared to single-pass feedforward models.
Why read this
Speech and ML researchers building clinical decision-support systems should read this paper to see how structured judicial reasoning and experience libraries can replace costly, overfitting-prone parameter fine-tuning for multimodal tasks.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Automated mental health screening, objective psychiatric risk assessment tools, and interpretable clinical decision-support systems.
Institutions
University of Chinese Academy of Sciences
Funding / 經費: National Natural Science Foundation of China, Key Scientific Research Program of Hangzhou, Natural Science Foundation of Hangzhou, Scientific Research Starting Foundation of Hangzhou Institute for Advanced Study, Zhejiang Provincial Natural Science Foundation of China, Key R&D Program of Zhejiang
Related
- Learning to Attend to Depression-Related Patterns: An Adaptive Cross-Modal Gating Network for Depression Detection — same problem · relatedness 2.9/3
- Label Correction Enhanced Dual-Stream Multiple Instance Learning for Weakly-Supervised Depression Detection in Speech — same problem · relatedness 2.8/3
- Layer-wise Multi-factor Adaptive Disentanglement for Cross-corpus Speech Depression Detection — same problem · relatedness 2.6/3
- Investigating LLMs Behavior in Depression Severity Prediction — same problem · relatedness 2.5/3
- Uncovering Latent Depression Severity for Binary Depression Detection via Advantage-weighting Ranking — same problem · relatedness 2.3/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-3099