All papers
Speech LLMs & dialogueFull-paper digest

Integrating Facial Generation into Full-Duplex Spoken Dialogue Systems

Jingjing Jiang, Atsumoto Ohashi, Ryuichiro Higashinaka

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.4 KB · Ready to paste

Preview copied content

TL;DR — Moshi-Face extends the Moshi full-duplex spoken dialogue model to jointly process and generate 3D facial expressions and speech simultaneously. Trained on 180 hours of dialogue data, it achieves low-latency audiovisual alignment without degrading the core dialogue quality of the audio-only base model.

Key contributions

  • Constructed a face VQ-VAE codec that maps 3D FLAME facial meshes into 8 discrete face tokens per frame at 12.5 Hz.
  • Integrated a non-causal Face Transformer into Moshi's 7B RQ-Transformer architecture for parallel, non-autoregressive face token generation.
  • Curated a 180-hour 3D audiovisual dialogue dataset (3,400 dialogues) derived from Meta's Seamless Interaction dataset using the VHAP 3D extraction pipeline.
  • Demonstrated successful real-time full-duplex conversational interaction with synchronized lip movements and head motion.

Problem

Prior multimodal conversational systems are strictly constrained to turn-based architectures where only one party speaks at a time, preventing natural phenomena like simultaneous speech, interruptions, and backchanneling. Conversely, modern full-duplex models like Moshi support bidirectional simultaneous voice interaction but remain entirely audio-only, lacking crucial nonverbal cues such as facial expressions, lip movements, and head motion. Developing a model that combines the real-time dynamics of full-duplex audio with synchronized facial generation is essential for natural human-computer interaction, but requires handling complex token synchronization and high-dimensional 3D facial representations.

Method

Moshi-Face builds on top of Moshi's architecture by appending 8 face token streams (in addition to 1 text stream and 2M=16 audio streams for system/user inputs) processed at a 12.5 Hz frame rate. The 3D facial codec uses a VQ-VAE with an encoder that takes FLAME vertex displacements (5,143 vertices, 3D coordinates), downsamples them temporally by factor r=2, and quantizes latents via a codebook of size K=256 and embedding dimension C=128. The VQ-VAE is trained using an L1 reconstruction loss, a quantization loss, and a velocity loss.

The core model retains Moshi's 7B-parameter pretrained RQ-Transformer (which generates hidden states, text, and audio tokens autoregressively) and attaches a non-causal Face Transformer module. At each timestep, the conditioning vector aggregates the hidden state, text embedding, and audio token embeddings, which are then combined with learnable positional embeddings. The Face Transformer applies non-causal self-attention across the N=8 face token positions within the frame, using 8 parallel linear heads to predict face tokens non-autoregressively. The overall objective combines standard Moshi text and audio cross-entropy losses with a face token cross-entropy loss weighted by lambda.

Training proceeds in two stages: first, freezing the RQ-Transformer and training only the Face Transformer for 500 steps (batch size 32, lr 5e-4); second, jointly fine-tuning all components for 1,200 steps (batch size 16) with specific layer learning rates (Temporal 2e-6, Depth 4e-6, Face 1e-5). Teacher forcing is used during training, feeding ground-truth face tokens from the previous timestep into the current queries.

Experimental setup

Experiments use a 180-hour subset of Meta's Seamless Interaction dataset (3,400 dialogues total), split into 70 hours for face codec training and the remaining for dialogue training with 100 unseen dialogues reserved for testing. The face codec is evaluated using Mean Vertex Error (MVE), Lip Vertex Error (LVE), and Perplexity; Moshi-Face is evaluated using LSE-D and LSE-C via SyncNet, UTMOS for speech naturalness, and an LLM-as-a-Judge protocol (GPT-5-mini) across 1-5 scale metrics (Coherence, Naturalness, Relevance, Overall). Baselines include pretrained Moshi, Moshi fine-tuned on the dataset without faces (Moshi-ft), Reconstructed face (upper bound), and Random face (lower bound).

Results

Under teacher-forced evaluation, Moshi-Face achieves a Lip Sync Error Distance (LSE-D) of 8.76 and LSE-C of 0.14, approaching the Reconstructed face upper bound (8.53 / 0.12) and outperforming the Random face lower bound (11.7 / 0.13). In free-run full-duplex self-dialogue generation, Moshi-Face maintains strong synchronization with an LSE-D of 11.0 and LSE-C of 0.16. While speech naturalness (UTMOS) drops slightly from 3.08 (base Moshi) to 1.75 due to domain fine-tuning on the smaller dataset, Moshi-Face achieves the highest semantic coherence (3.79) and strong naturalness (4.52) in LLM-as-a-Judge evaluations, performing comparably to base Moshi overall (3.76 vs 3.85). Ablations confirm that pre-training the Face Transformer and performing joint fine-tuning are both essential for optimal synchronization and semantic performance.

System / ConditionLSE-D (↓)LSE-C (↑)UTMOS (↑)LLMAJ Overall (1-5)
Moshi (Base)––3.083.85
Moshi-ft (Audio-only)––1.693.55
Reconstructed Face (Upper Bound)8.530.12––
Random Face (Lower Bound)11.700.13––
Moshi-Face (Ours, Teacher-forced)8.760.141.753.76
Moshi-Face (Ours, Free-run)11.000.161.753.76

Limitations

The current face codec is non-causal, requiring a small temporal lookahead that prevents truly real-time streaming visual input/output (though the language model backbone is streaming). The evaluation relies on indirect 2D video rendering via SyncNet pipelines for LSE metrics rather than direct human perceptual studies. Furthermore, the system was trained on a restricted 180-hour subset, leading to a minor drop in general acoustic quality (UTMOS) compared to the original massive pre-trained Moshi checkpoints.

Why read this

Researchers and engineers building multimodal conversational agents will find this paper a clear blueprint for scaling full-duplex audio-only language models into the visual domain via discrete token codebooks and parallel transformer heads. It provides concrete design choices for handling multi-stream token synchronization and non-autoregressive face token generation.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Real-time interactive avatars, full-duplex virtual assistants, and embodied conversational AI agents with realistic facial expressions and lip synchronization.

Institutions

Nagoya University

Funding / 經費: JST Moonshot R&D

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-3114