All papers
Health & clinical speechFull-paper digest

Moot-Court: Training-Free Dialectical Reasoning for Depression Detection

Yuqing Sun, Jian Zhao, Haoxun Li, Leyuan Qu, Taihao Li

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.5 KB · Ready to paste

Preview copied content

TL;DR — Moot-Court is a training-free reasoning engine for automated depression detection that uses adversarial multi-agent debate and a dynamic experience library to guide frozen LLMs, achieving an F1-score of 0.8856 on DAIC-WOZ.

Key contributions

  • Introduces Moot-Court, a training-free multimodal depression detection framework that outperforms fully fine-tuned and instruction-tuned models.
  • Proposes a dialectical reasoning mechanism integrating psychiatric symptom networks into LLM diagnostics via a Prosecutor and Defense Attorney adversarial structure.
  • Develops a dynamic, self-evolving experience library (Codex) governed by a reward-based feedback algorithm to constrain reasoning trajectories and suppress hallucinations without weight updates.
  • Combines utterance-level symbolic acoustic markers, session-level statistical trajectories, and subject-level descriptive portraits into a unified multimodal blueprint.

Problem

Current automated depression detection methods using LLMs either rely on costly parameter adaptation (like instruction tuning or LoRA) that risks severe overfitting in low-resource clinical settings, or discard non-verbal acoustic biomarkers by operating in text-only mode. Furthermore, black-box neural networks and standard LLMs lack clinical interpretability and suffer from shortcut bias, failing to meet medical requirements for trustworthy explanations. This gap necessitates an objective, interpretable, and training-efficient framework that successfully integrates speech and text.

Method

The framework begins with multimodal blueprinting, translating raw audio and transcripts into three hierarchical tiers: utterance-level acoustic markers (M) combined with text, session-level statistical trajectories (linear regression slopes, distributional extrema, and interruption rates via F_slope, F_dist, F_inter), and subject-level descriptive portraits generated by segmenting audio into 3-minute blocks processed by Qwen-Audio and Qwen-max. These discrete fact nodes V are categorized into External Fields (V_ext), Internal States (V_int), and Evidence Markers (V_evid).

Next, adversarial agents perform dialectical reasoning. The Prosecutor builds a Pathogenic Graph (G_P) using edges that map biomarkers to symptoms (e_ind), trace feedback loops (e_refin), identify external triggers (e_gen, e_trig), and detect clinical paradoxes like smiling depression via contradiction edges (e_cong). Concurrently, the Defense Attorney constructs a Contextual Graph (G_D) using edges for reactive sadness (e_ctx), biological alibis (e_exp), anxiety comorbidity (e_diff), resilience markers (e_inhib), and manic exclusion criteria (e_excl).

In the decision phase, k=5 independent Judges with temperatures ranging from 0.5 to 0.8 evaluate the graphs using retrieved semantic priors from the dynamic Codex (governed by authority alpha, frequency f, utility u, and a newbie bonus delta_new for f < 5). Outcomes trigger retrospective reflection across three game-theoretic scenarios (consensus, system failure, or divergent judgments), updating the Codex via MODIFY, MERGE, and ADD operations driven by a statistical reward function S.

Experimental setup

Evaluated on the DAIC-WOZ dataset (107 training, 35 validation, 47 test splits), which contains semi-clinical interviews with a virtual agent assessing psychological distress. Audio is denoised using ZipEnhancer. Acoustic features are extracted using Librosa (latency, WPS, F0, RMS energy) and Parselmouth (jitter, shimmer). The system uses frozen LLM backbones including DeepSeek V3.2 and Qwen3 series (30B, 80B). Performance is measured using Precision, Recall, and F1-score.

Results

Moot-Court using DeepSeek V3.2 achieves an F1-score of 0.8856, a Recall of 0.9394, and a Precision of 0.8378, outperforming fully supervised convolutional and recurrent architectures, fully fine-tuned SSL-SDD (F1 0.8290), LoRA-tuned GPT-2 (F1 0.8260), and instruction-tuned DepressInstruct (F1 0.8235). Cross-backbone evaluations show that Qwen3-30b achieves an F1 of 0.8234 and Recall of 0.9545 entirely training-free.

Ablation studies reveal that direct vanilla multimodal inference yields an F1 of 0.6296 with a low recall of 0.5152. Adding linear reasoning raises F1 to 0.7143. Introducing dialectical reasoning alone drives precision to 0.8333 but drops recall to 0.6061 due to over-conservatism; the dynamic Codex successfully resolves this bottleneck by boosting recall to 0.9394. Furthermore, removing audio features and relying solely on text drops the F1-score from 0.8856 down to 0.7890, confirming the critical complementarity of acoustic biomarkers.

MethodStrategyF1Rec.Pre.
SSL-SDD [28]Full Fine-tuning.8290––
GPT-2 [29]LoRA.8260.8590.7950
LLM-Landmark [30]LoRA/P-tuning.8330––
DepressInstruct [32]Instruction Tuning.8235.8400.8077
(Ours) DeepSeek V3.2 [21]Training Free.8856.9394.8378

Limitations

The framework's performance is inherently bounded by the capabilities and token context limits of the underlying foundation models (e.g., DeepSeek and Qwen series). Evaluation is restricted to the DAIC-WOZ dataset, meaning generalizability across diverse languages, cultural contexts, and clinical interview protocols remains unverified. Additionally, reliance on multiple LLM reasoning passes and judge ensembles introduces higher inference latency compared to single-pass feedforward models.

Why read this

Speech and ML researchers building clinical decision-support systems should read this paper to see how structured judicial reasoning and experience libraries can replace costly, overfitting-prone parameter fine-tuning for multimodal tasks.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Automated mental health screening, objective psychiatric risk assessment tools, and interpretable clinical decision-support systems.

Institutions

University of Chinese Academy of Sciences

Funding / 經費: National Natural Science Foundation of China, Key Scientific Research Program of Hangzhou, Natural Science Foundation of Hangzhou, Scientific Research Starting Foundation of Hangzhou Institute for Advanced Study, Zhejiang Provincial Natural Science Foundation of China, Key R&D Program of Zhejiang

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-3099