All papers
Speech synthesisFull-paper digest

Towards Digital Preservation of Efik: TTS for a Low-Resource African Language

Offiong Bassey Edet, Emmanuel Oyo-Ita, Archibong Okon Archibong, David Effanga Bassey, Mbuotidem Sunday Awak

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.2 KB · Ready to paste

Preview copied content

TL;DR — This paper presents the first end-to-end text-to-speech (TTS) study for Efik, a low-resource tonal African language, by introducing a curated 3-hour single-speaker corpus and benchmarking four neural architectures. MMS-TTS emerges as the top-performing model, achieving a mean opinion score (MOS) of 3.80 ± 0.63.

Key contributions

  • Introduces the first documented single-speaker Efik TTS corpus consisting of 2,632 validated utterances totaling approximately 3.08 hours.
  • Provides a comparative evaluation of four distinct neural TTS architectures (VITS, MMS-TTS, SpeechT5, and Orpheus-TTS) under extreme low-resource conditions.
  • Establishes a foundational evaluation benchmark using native speaker evaluations across MOS, Nat-MOS, and A-MOS metrics.
  • Identifies transfer learning via multilingual pretraining (e.g., initializing MMS-TTS with a Yoruba checkpoint) as a key enabler for intelligible low-resource tonal synthesis.

Problem

Efik is a Lower Cross tonal language spoken by 1.5 million native speakers and 3 million second-language speakers in Southeastern Nigeria, yet it remains absent from modern speech technology pipelines. Developing TTS for Efik is hindered by severe data scarcity and the need to accurately model lexical pitch variations, as inadequate tone realization destroys intelligibility. Prior automatic forced alignment using Whisper and XLS-R failed completely due to a lack of pretraining exposure to Lower Cross languages, necessitating manual annotation to construct a reliable training set.

Method

The authors fine-tuned four models: VITS (conditional VAE with normalizing flows), MMS-TTS (multilingual framework leveraging cross-lingual transfer), SpeechT5 (transformer sequence-to-sequence), and Orpheus-TTS (adversarial waveform realism). Because MMS-TTS lacked an Efik checkpoint, it was initialized using a Yoruba checkpoint with vocabulary and embedding extensions for Efik-specific characters like o. and ˜n. VITS and SpeechT5 similarly required updated embeddings, while Orpheus-TTS worked without modification.

Training was conducted on a single NVIDIA A100 GPU using mixed precision and early stopping based on validation loss. Hyperparameters included: VITS trained for 50 epochs (lr=2e-4, batch size 4, Adam); MMS-TTS trained for 50 epochs (lr=2e-5, batch size 16, AdamW); SpeechT5 trained for up to 2,500 epochs (lr=1e-5, batch size 4, 0.1 dropout); and Orpheus-TTS trained for 50 epochs (lr=2e-5, batch size 8). The audio dataset was preprocessed into uncompressed 16 kHz mono WAV files with trailing silences of 60-100 ms preserved to protect sentence-final tonal cues.

Experimental setup

The dataset contains 2,632 utterances (1,975 train, 264 validation, 393 test) summing to 3.08 hours from a single native speaker, drawn from novels, folktales, and educational texts. Evaluation was performed by 5 native Efik speakers rating short clips on a 1-5 scale across MOS (overall naturalness), Nat-MOS (native naturalness), and A-MOS (accent/phonetic preservation). Models were compared against each other as baselines under identical low-resource constraints.

Results

MMS-TTS achieved the headline-leading MOS of 3.80 ± 0.63, Nat-MOS of 3.60 ± 0.56, and A-MOS of 3.04 ± 0.52, demonstrating superior stability in generating continuous speech up to 3 minutes without hallucination. Orpheus-TTS ranked second with an MOS of 3.08 ± 0.48 and Nat-MOS of 2.32 ± 0.46, though it exhibited a foreign European male accent and struggled with tonal nuances. SpeechT5 scored an MOS of 2.48 ± 0.49, maintaining intelligibility only for short sequences under 20-30 seconds before hallucinating. VITS performed the worst with an MOS of 1.08 ± 0.27, completely failing to capture tonal variations or produce intelligible long-form audio due to its heavy reliance on large-scale datasets.

ModelMOSNat-MOSA-MOS
VITS1.08 ± 0.271.04 ± 0.19-
SpeechT52.48 ± 0.491.88 ± 0.511.64 ± 0.48
Orpheus-TTS3.08 ± 0.482.32 ± 0.462.21 ± 0.43
MMS-TTS3.80 ± 0.633.60 ± 0.563.04 ± 0.52

Limitations

The study is restricted to a single-speaker dataset of roughly 3 hours, limiting prosodic variation and robust long-sequence modeling. Rare phonemes like ˜n caused persistent pronunciation failures across all evaluated models, and non-MMS models suffered from foreign accent drift and poor preservation of cultural tonal contours.

Why read this

Speech researchers and engineers working on extremely low-resource, tonal, or underrepresented African languages should read this paper to understand how cross-lingual transfer (such as initializing with Yoruba checkpoints) bridges data gaps where mainstream models like VITS fail.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Digital language preservation, educational tools, and text-to-speech accessibility applications for the Efik-speaking community.

Institutions

University of Cross River State, University of Calabar, ML Collective

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1868