---
id: liu26m_interspeech
title: "MultiEmoVec: Learning Generalised Multimodal Emotion Representation by
  Momentum Contrast and Multi-task Reconstruction"
authors:
  - Junchen Liu
  - Jesin James
  - Karan Nathwani
  - Michael Witbrock
year: 2026
doi: 10.21437/Interspeech.2026-1563
isca_url: https://www.isca-archive.org/interspeech_2026/liu26m_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/liu26m_interspeech.pdf
session: Multimodal Speech Processing and Speech LLM Systems
topics:
  - paralinguistics
  - emotion-recognition
  - self-supervised
category: paralinguistics-emotion
labels:
  - self-supervised
institutions:
  - University of Auckland
  - Indian Institute of Technology Jammu
code:
  url: https://github.com/MaoriEnglish-Codeswitch/MultiEmoVec
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: liu26m_interspeech
  category: paralinguistics-emotion
  labels:
    - self-supervised
  institutions:
    - University of Auckland
    - Indian Institute of Technology Jammu
  code: https://github.com/MaoriEnglish-Codeswitch/MultiEmoVec
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-1563
  pdf: https://www.isca-archive.org/interspeech_2026/liu26m_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/liu26m_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/liu26m_interspeech/markdown.md
---

# MultiEmoVec: Learning Generalised Multimodal Emotion Representation by Momentum Contrast and Multi-task Reconstruction

*Junchen Liu, Jesin James, Karan Nathwani, Michael Witbrock*

[PDF](https://www.isca-archive.org/interspeech_2026/liu26m_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/liu26m_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-1563)

**Category:** `paralinguistics-emotion` · **Labels:** `self-supervised`

**TL;DR** — MultiEmoVec is an unsupervised pre-training framework for multimodal emotion recognition that combines momentum contrastive learning with dual-task reconstruction and adaptive loss scaling, achieving a 55.31% 7-class accuracy on CMU-MOSEI—a 3% absolute improvement over prior state-of-the-art baselines.

## Key contributions

- Proposes a multi-task unsupervised pre-training framework for multimodal emotion recognition that eliminates the need for expensive emotional annotation during pre-training.
- Introduces a dual-reconstruction mechanism combining masked feature completion and denoising (via mixup, masking, and Gaussian noise) to handle incomplete and corrupted input modalities.
- Applies an adaptive loss scaling (ALS) mechanism utilizing exponential moving averages to dynamically balance contrastive learning and multi-task reconstruction losses.
- Demonstrates strong cross-database transferability, matching fully supervised performance within 3% on IEMOCAP using only frozen representations from CMU-MOSEI pre-training.

## Problem

Supervised multimodal emotion recognition (MER) models heavily rely on large-scale annotated datasets that are expensive and difficult to collect. Prior unsupervised methods either rely on shallow clustering approaches that fail to capture subtle semantic ambiguities of emotions or use graph neural networks that incur high computational costs and complex structure construction. Furthermore, emotional databases frequently suffer from missing modalities, noisy signals, and inconsistent recording conditions, which break naive contrastive learning approaches.

## Method

MultiEmoVec extracts features using pre-trained backbones: Wav2Vec (base, 512-dim) for speech, MANet (1024-dim) for video, and DeBERTa (large, 1024-dim) for text. These unimodal representations are projected into a shared 512-dimensional space via two-layer MLPs with ReLU activations and layer normalization. To model inter-modal and intra-modal interactions, a 9-embedding design is constructed using 3 unimodal and 6 cross-modal attention embeddings, which are processed by a single-layer transformer encoder.

Unsupervised pre-training is driven by two parallel pathways: masked reconstruction (randomly masking 20% of feature dimensions in a single modality to force cross-modal completion from other modalities) and denoising reconstruction (restoring clean features from inputs corrupted by mixup with parameter 0.5, feature masking, and Gaussian noise with variance 0.05). Concurrently, momentum contrast (MoCo) learns a discriminative embedding space where queries and positive keys are generated by query and momentum-updated key encoders ($m=0.999$), backed by a dynamic queue of 4096 negative samples and optimized via InfoNCE loss with temperature 0.06.

To prevent objective starvation across multi-task training, an adaptive loss scaling mechanism uses the masked completion loss as a reference scale, dynamically normalizing contrastive and denoising losses using an exponential moving average ($eta=0.9$, scale exponent $	au=0.4$, 10-epoch warm-up, and clipping bounds $[0.5, 2.0]$). The resulting frozen pre-trained encoder outputs 1024-dimensional fused representations for downstream lightweight classification.

## Experimental setup

Pre-training was conducted on CMU-MOSEI (22,856 video utterances, split 70/10/20 for train/val/test). Downstream fine-tuning and evaluation were performed on CMU-MOSI and IEMOCAP (4-class and 6-class speaker-independent settings). The primary baseline is MGAFR, alongside SURE, CPSPAN, and DCP. Evaluation metrics include 7-class accuracy (Acc-7), binary accuracy (Acc-2), and binary F1 score (BF1). The framework was trained for 80 epochs using the AdamW optimizer (learning rate $1e-4$, weight decay $1e-5$) on a single NVIDIA RTX 4090 GPU, while the downstream lightweight classifier was trained for 10 epochs using Adam (learning rate $1e-3$).

## Results

MultiEmoVec achieves an Acc-7 of 55.31%, Acc-2 of 84.10%, and BF1 of 89.08% on CMU-MOSEI, outperforming the state-of-the-art MGAFR baseline (52.24% Acc-7) by 3.07% while reducing trainable parameters by 6.28M down to 16.28M. Ablation studies reveal that combining MoCo with both masked and denoising reconstruction yields the highest performance, whereas removing pre-training entirely drops Acc-7 down to 42.91%. In cross-database evaluation on IEMOCAP under a 6-class setting, MultiEmoVec achieves 55.21% weighted F1, coming within 3% of a fully supervised baseline (58.64%) despite utilizing zero emotion labels during pre-training.

| System / Condition | Acc-7 | Acc-2 | BF1 |
|---|---|---|---|
| SURE | 51.77% | 83.75% | 83.97% |
| CPSPAN | 50.68% | 83.92% | 84.18% |
| DCP | 50.93% | 83.86% | 84.08% |
| MGAFR (Baseline) | 52.24% | 83.97% | 84.34% |
| MultiEmoVec (Ours) | 55.31% | 84.10% | 89.08% |

## Limitations

The framework relies heavily on textual semantics for anchor alignment, causing performance to drop significantly when text is absent or noisy (as shown in bi-modal ablation where video+speech drops Acc-7 to 42.86%). The evaluation is restricted to English-language databases (CMU-MOSEI, CMU-MOSI, IEMOCAP), leaving cross-lingual and low-resource generalizability unverified. Furthermore, the reliance on fixed frozen upstream extractors (Wav2Vec, MANet, DeBERTa) bounds the representation ceiling by the quality of these pre-extracted features.

## Why read this

Speech and ML researchers working on self-supervised multimodal representation learning or emotion recognition should read this paper to see how momentum contrast can be effectively combined with dual-task reconstruction and adaptive loss scaling for robust feature transfer.

## Code

- https://github.com/MaoriEnglish-Codeswitch/MultiEmoVec

## Applications

Unsupervised multimodal representation learning for customer service analytics, mental health monitoring, and interactive conversational agents.

## Institutions / 機構

University of Auckland, Indian Institute of Technology Jammu

## Related

- [Audio-Visual Feature Reconstruction Pretraining for Noise-Robust Emotion Recognition](parmonangan26_interspeech.md) — same problem · relatedness 2.4/3
- [The Illusion of Balanced Multimodal Sentiment Analysis: Beyond the Limits of Optimization-Based Methods](kaffeza26_interspeech.md) — same problem · relatedness 2.4/3
- [EmoEUS: Uncertainty Supervision for Multimodal Emotion Recognition in Conversation](huang26n_interspeech.md) — same problem · relatedness 2.3/3
- [Segment-wise Embedding based Graph Attention Network for Effective Speech Emotion Recognition](song26c_interspeech.md) — same problem · relatedness 2.3/3
- [Segregate, Refine, Integrate: Decomposing Multimodal Fusion for Sentiment Analysis](filippakopoulos26_interspeech.md) — shared data / evaluation · relatedness 2.3/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
