All papers
Speech recognitionFull-paper digest

Beyond Standard Greek: Adapting Whisper for Greek Dialects through Curriculum Multitask Learning

Antigoni Klimi, Dimitrios Damianos, Stavros Bompolas, Vivian Stamou, Stella Markantonatou, Vassilis Katsouros, Georgios Paraskevopoulos

Code & resourcesgithub.com/athena-ilsp/Dialect-Adaptation

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.3 KB · Ready to paste

Preview copied content

TL;DR — This paper proposes a staged multitask curriculum framework for adapting Whisper to low-resource Greek dialects, progressively shifting training from standard-language speech translation to dialect-specific ASR and achieving significant error reductions over standard fine-tuning.

Key contributions

  • A unified four-stage curriculum framework combining domain adaptation (Standard Greek to dialect) and task adaptation (speech translation to ASR).
  • Empirical demonstration of consistent WER and CER improvements across three distinct Greek dialects (Cypriot, Cretan, Messenian) and three Whisper model scales (small, medium, large-v3).
  • Comprehensive ablation studies isolating the critical role of joint translation-ASR multitasking and donor-language pre-alignment data.
  • Release of open-source fine-tuned models and datasets for Greek dialectal speech processing.

Problem

Adapting Automatic Speech Recognition (ASR) to regional dialects results in extremely high zero-shot Word Error Rates (>100% for highly divergent varieties) and a persistent "dialectal tax" where error rates double compared to Standard Modern Greek (SMG) due to orthographic variation and language-contact effects. Prior approaches like naive supervised fine-tuning, pseudo-labeling, or standard multi-task learning fail to control the adaptation trajectory under extreme data scarcity, causing optimization instability. This matters because robust speech technology must support non-standard linguistic varieties rather than exclusively high-resource standard languages.

Method

The proposed approach structures adaptation across task space (translation to ASR) and domain space (Standard Greek to dialect) using a sequential four-stage curriculum. Stage 0 (Foundation Warm-up) trains exclusively on Standard Greek-to-English speech translation (GST) using the Greek Podcast Corpus (GPC-5h subset) to establish stable multilingual acoustic-semantic representations. Stage 1 (Mixed Domain) introduces dialectal speech through a mixed-domain sampling strategy where data are drawn with probability alpha from Standard Greek ASR and 1-alpha from Dialect Speech Translation (DST). Stage 2 (Adaptation Specialization) shifts supervision entirely within the dialect domain, sampling Dialect ASR with probability beta and Dialect ST with probability 1-beta to provide regularizing signals. Stage 3 (Target Domain Convergence) specializes exclusively on Dialect ASR data.

Experiments evaluate Whisper-small, Whisper-medium, and Whisper-large-v3 models. The encoder is frozen during fine-tuning while keeping the convolutional feature extractor trainable. Stage-specific learning rates are decayed progressively across stages S0-S3 (e.g., from 5e-5 down to 5e-6 for Whisper-small). Training uses fixed budgets of 8,000 samples per stage for Cypriot and Cretan, and 2,000 samples for the lower-resource Messenian setting. All transcripts are paired with English translations generated by Llama-Krikri-8B-Instruct and manually reviewed to support the speech-to-text translation auxiliary objectives.

Experimental setup

Evaluated on three Greek dialects: Cypriot Greek (12.9 hours across 1,284 clips from Mozilla Common Voice v23.0), Cretan Greek (1h 21m across 2,589 utterances from archival radio broadcasts), and Messenian Greek (37m across 590 utterances). Baselines include zero-shot Whisper and conventional fine-tuning (FT) with a frozen encoder for a maximum of 2,048 update steps. Metrics reported are Word Error Rate (WER) and Character Error Rate (CER).

Results

On Cypriot Greek with the Whisper-small model, the proposed curriculum reduces WER from 81.27% (zero-shot) and 52.38% (fine-tuned) down to 35.61%. For Whisper-large-v3 on Cypriot Greek, WER drops from 58.33% zero-shot and 33.93% fine-tuned to 24.89%. In the extremely low-resource Messenian setting, Whisper-small improves from 67.49% zero-shot to 30.25% with the curriculum.

Ablations on task supervision demonstrate that removing the joint translation task increases WER significantly (e.g., rising from 35.61% to 47.84% for Whisper-small on Cypriot Greek). Ablations on donor language data show that omitting Standard Greek pre-alignment data (GPC) causes performance degradation primarily in lower-resource settings like Messenian.

ModelSettingCypriot WERCypriot CERCretan WERCretan CERMessenian WERMessenian CER
Whisper-smallZero-Shot81.2749.2193.8865.6967.4946.69
Whisper-smallFine-Tuned52.3832.3351.3718.7541.9719.74
Whisper-smallCurriculum35.6121.0929.518.5330.2512.94
Whisper-large-v3Zero-Shot58.3331.4260.4637.0822.3115.42
Whisper-large-v3Fine-Tuned33.9321.1533.9312.1415.316.87
Whisper-large-v3Curriculum24.8915.3917.695.0414.565.66

Limitations

The evaluation is restricted to three southern Greek dialects, leaving highly divergent contact-influenced varieties (such as Griko or Cappadocian) and Northern Greek varieties unexplored. The approach requires parallel English translations for the curriculum's speech-translation stages, which may increase annotation overhead for new low-resource languages lacking machine translation or LLM support.

Why read this

Speech researchers and engineers tackling low-resource dialect adaptation should read this to see how combining task-level translation supervision with domain-progressive curricula stabilizes large encoder-decoder models like Whisper.

Code

Applications

Building robust regional speech recognition systems for low-resource or dialectal languages where standard ASR models suffer from severe domain mismatch.

Institutions

Athena Research Center

Funding / 經費: European High-Performance Computing Joint Undertaking, Greek Ministry of Digital Governance and Artificial Intelligence

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-2567