All papers
Resources & evaluationFull-paper digest

AnimeScore: A Preference-Based Dataset and Framework for Evaluating Anime-Like Speech Style

Joonyong Park, Jerry Li

Code & resourcesgithub.com/sizigi/animescore

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.1 KB · Ready to paste

Preview copied content

TL;DR — AnimeScore is a preference-based dataset and evaluation framework for automated assessment of anime-like speech styles, achieving up to 90.8% AUC using SSL-based ranking models compared to a 69.3% ceiling for handcrafted acoustic features.

Key contributions

  • Constructed a curated dataset of 3,000 Japanese utterances (2,500 train, 500 test) paired with 15,000 pairwise human preference judgments from 187 evaluators.
  • Established a multi-stage filtering pipeline utilizing Qwen3-30B-Instruct, Sidon speech enhancement, Whisper-large-v3, and ECAPA-TDNN clustering to control linguistic, acoustic, and speaker biases.
  • Revealed through acoustic analysis that anime-likeness is driven by controlled resonance shaping (lower formants), prosodic continuity, and dense syllable rates rather than simply high pitch.
  • Demonstrated that SSL backbones (specifically HuBERT) combined with BiLSTM and MLP layers trained on pairwise logistic loss vastly outperform handcrafted acoustic feature baselines.

Problem

Evaluating anime-like speech style currently relies on costly subjective listening tests and lacks a standardized, reproducible automatic evaluation metric. Unlike naturalness or intelligibility, anime-likeness lacks a shared absolute perceptual scale and is multidimensional, rendering traditional Mean Opinion Score (MOS) protocols inconsistent. Developing speech generation systems in this domain is severely bottlenecked by this absence of scalable scoring protocols and automated objective reward signals.

Method

The speech set combines samples from Anim-400k, ReazonSpeech, and Coco-Nut. The data curation pipeline removes linguistic bias by filtering out texts with high anime-subtitle probability using Qwen3-30B-Instruct (keeping scores <= 2), enhances audio using Sidon, applies UTMOS > 3 filtering, and ensures speaker diversity via ECAPA-TDNN embeddings and t-SNE clustering. Sparse comparison pairs (12,500 train, 2,500 test) are built using text and speaker similarity to emphasize cross-corpus contrasts.

The score prediction framework passes input audio through a frozen SSL encoder (wav2vec 2.0, WavLM, HuBERT, or data2vec) to extract frame-level features H. These features are fed into a bidirectional LSTM (BiLSTM), mean-pooled into a fixed-length representation, and mapped to a scalar score s(x) via a multi-layer perceptron (MLP). During training, the network predicts scores for pairs (a, b) and is optimized using the pairwise logistic loss -log sigma(sa - sb) against ground-truth A/B preference outcomes, enabling drop-in model screening or reward shaping for generative models.

Experimental setup

Evaluations used a held-out test set of 2,500 A/B pairs derived from 500 utterances. Evaluators (n=187, predominantly male aged 30s-50s) provided 15,000 total comparative judgments. Models were compared against a multivariate logistic regression baseline using handcrafted acoustic features evaluated via 5-fold cross-validation. Metrics reported include pairwise accuracy, negative log-likelihood (NLL), and ROC-AUC.

Results

Handcrafted acoustic features combined via logistic regression achieved a 69.3% AUC (63.4% accuracy) with dominant predictors including pause ratio and syllable rate. In contrast, frozen SSL backbones trained on the pairwise loss substantially outperformed this baseline. HuBERT achieved the highest performance with 0.3852 NLL, 82.43% accuracy, and 90.82% AUC, followed closely by WavLM at 89.44% AUC, wav2vec 2.0 at 82.47% AUC, and data2vec at 85.80% AUC.

Masked-prediction models (HuBERT, WavLM) consistently surpassed the contrastive model (wav2vec 2.0), capturing paralinguistic and prosodic properties necessary for stylistic evaluation. The models maintained robust performance on out-of-distribution within-corpus test pairs, though the study notes limitations in demographic balance and moderate data scale.

System / BackboneNLLAccuracy (%)AUC (%)
Handcrafted Features (LR)-63.4 ± 1.269.3 ± 1.5
wav2vec 2.00.513974.3082.47
data2vec0.468677.0985.80
WavLM0.428481.0589.44
HuBERT0.385282.4390.82

Limitations

The work is constrained by a moderate data scale (15,000 pairs across 3,000 utterances) and a demographic imbalance among human evaluators (76% male, heavily skewed toward individuals in their 30s to 50s). The scope is strictly limited to Japanese speech styles and does not explore model architecture ablations beyond changing the frozen SSL backbone.

Why read this

Speech and ML researchers building generative anime or character-styled speech models should read this to understand how to replace costly subjective listening tests with an automated, SSL-backed reward signal.

Code

Applications

Automated quality screening for stylized speech generation systems and reinforcement learning reward signals for aligning text-to-speech models toward target voice aesthetics.

Institutions

Spellbrush

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-3025