All papers
Paralinguistics & emotionFull-paper digest

MultiEmoVec: Learning Generalised Multimodal Emotion Representation by Momentum Contrast and Multi-task Reconstruction

Junchen Liu, Jesin James, Karan Nathwani, Michael Witbrock

Code & resourcesgithub.com/MaoriEnglish-Codeswitch/MultiEmoVec

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.2 KB · Ready to paste

Preview copied content

TL;DR — MultiEmoVec is an unsupervised pre-training framework for multimodal emotion recognition that combines momentum contrastive learning with dual-task reconstruction and adaptive loss scaling, achieving a 55.31% 7-class accuracy on CMU-MOSEI—a 3% absolute improvement over prior state-of-the-art baselines.

Key contributions

  • Proposes a multi-task unsupervised pre-training framework for multimodal emotion recognition that eliminates the need for expensive emotional annotation during pre-training.
  • Introduces a dual-reconstruction mechanism combining masked feature completion and denoising (via mixup, masking, and Gaussian noise) to handle incomplete and corrupted input modalities.
  • Applies an adaptive loss scaling (ALS) mechanism utilizing exponential moving averages to dynamically balance contrastive learning and multi-task reconstruction losses.
  • Demonstrates strong cross-database transferability, matching fully supervised performance within 3% on IEMOCAP using only frozen representations from CMU-MOSEI pre-training.

Problem

Supervised multimodal emotion recognition (MER) models heavily rely on large-scale annotated datasets that are expensive and difficult to collect. Prior unsupervised methods either rely on shallow clustering approaches that fail to capture subtle semantic ambiguities of emotions or use graph neural networks that incur high computational costs and complex structure construction. Furthermore, emotional databases frequently suffer from missing modalities, noisy signals, and inconsistent recording conditions, which break naive contrastive learning approaches.

Method

MultiEmoVec extracts features using pre-trained backbones: Wav2Vec (base, 512-dim) for speech, MANet (1024-dim) for video, and DeBERTa (large, 1024-dim) for text. These unimodal representations are projected into a shared 512-dimensional space via two-layer MLPs with ReLU activations and layer normalization. To model inter-modal and intra-modal interactions, a 9-embedding design is constructed using 3 unimodal and 6 cross-modal attention embeddings, which are processed by a single-layer transformer encoder.

Unsupervised pre-training is driven by two parallel pathways: masked reconstruction (randomly masking 20% of feature dimensions in a single modality to force cross-modal completion from other modalities) and denoising reconstruction (restoring clean features from inputs corrupted by mixup with parameter 0.5, feature masking, and Gaussian noise with variance 0.05). Concurrently, momentum contrast (MoCo) learns a discriminative embedding space where queries and positive keys are generated by query and momentum-updated key encoders (m=0.999m=0.999), backed by a dynamic queue of 4096 negative samples and optimized via InfoNCE loss with temperature 0.06.

To prevent objective starvation across multi-task training, an adaptive loss scaling mechanism uses the masked completion loss as a reference scale, dynamically normalizing contrastive and denoising losses using an exponential moving average (β=0.9\beta=0.9, scale exponent τ=0.4\tau=0.4, 10-epoch warm-up, and clipping bounds [0.5,2.0][0.5, 2.0]). The resulting frozen pre-trained encoder outputs 1024-dimensional fused representations for downstream lightweight classification.

Experimental setup

Pre-training was conducted on CMU-MOSEI (22,856 video utterances, split 70/10/20 for train/val/test). Downstream fine-tuning and evaluation were performed on CMU-MOSI and IEMOCAP (4-class and 6-class speaker-independent settings). The primary baseline is MGAFR, alongside SURE, CPSPAN, and DCP. Evaluation metrics include 7-class accuracy (Acc-7), binary accuracy (Acc-2), and binary F1 score (BF1). The framework was trained for 80 epochs using the AdamW optimizer (learning rate 1e−41e-4, weight decay 1e−51e-5) on a single NVIDIA RTX 4090 GPU, while the downstream lightweight classifier was trained for 10 epochs using Adam (learning rate 1e−31e-3).

Results

MultiEmoVec achieves an Acc-7 of 55.31%, Acc-2 of 84.10%, and BF1 of 89.08% on CMU-MOSEI, outperforming the state-of-the-art MGAFR baseline (52.24% Acc-7) by 3.07% while reducing trainable parameters by 6.28M down to 16.28M. Ablation studies reveal that combining MoCo with both masked and denoising reconstruction yields the highest performance, whereas removing pre-training entirely drops Acc-7 down to 42.91%. In cross-database evaluation on IEMOCAP under a 6-class setting, MultiEmoVec achieves 55.21% weighted F1, coming within 3% of a fully supervised baseline (58.64%) despite utilizing zero emotion labels during pre-training.

System / ConditionAcc-7Acc-2BF1
SURE51.77%83.75%83.97%
CPSPAN50.68%83.92%84.18%
DCP50.93%83.86%84.08%
MGAFR (Baseline)52.24%83.97%84.34%
MultiEmoVec (Ours)55.31%84.10%89.08%

Limitations

The framework relies heavily on textual semantics for anchor alignment, causing performance to drop significantly when text is absent or noisy (as shown in bi-modal ablation where video+speech drops Acc-7 to 42.86%). The evaluation is restricted to English-language databases (CMU-MOSEI, CMU-MOSI, IEMOCAP), leaving cross-lingual and low-resource generalizability unverified. Furthermore, the reliance on fixed frozen upstream extractors (Wav2Vec, MANet, DeBERTa) bounds the representation ceiling by the quality of these pre-extracted features.

Why read this

Speech and ML researchers working on self-supervised multimodal representation learning or emotion recognition should read this paper to see how momentum contrast can be effectively combined with dual-task reconstruction and adaptive loss scaling for robust feature transfer.

Code

Applications

Unsupervised multimodal representation learning for customer service analytics, mental health monitoring, and interactive conversational agents.

Institutions

University of Auckland, Indian Institute of Technology Jammu

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1563