All papers
Speech recognitionFull-paper digest

SᴜTRA: Structurally-Unified Tokenization with Root Awareness

Vaibhav Rathore, Siddhant Gole, Dadhichi Telwadkar, Rooshil Bhatia, Maulik Ruparel, Siddharth Surekha, Neha Bhargava

Code & resourcesmo-vaibhavr-43300.github.io/SuTRA

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.7 KB · Ready to paste

Preview copied content

TL;DR — SuTRA is a morphology-aware subword tokenization framework for morphologically rich Indic languages that prevents root-affix fragmentation through script-aware grouping and boundary-penalized merging. It achieves an average +8.08 chrF2 improvement in machine translation and a +34% gain in semantic recoverability for Hindi over standard BPE.

Key contributions

  • Introduces SuTRA, a two-phase tokenization framework combining akshara-level pre-tokenization with boundary-penalized BPE merging to eliminate Morphological Shattering.
  • Releases an LLM-verified gold-standard morphological segmentation dataset spanning ~560,000 unique words across Hindi, Marathi, and Gujarati.
  • Designs a dynamic rigidity constraint with an annealed schedule to prioritize core lexical roots early in training before attaching functional affixes.
  • Demonstrates consistent improvements across morphological alignment, semantic recoverability, orthographic robustness, and downstream machine translation.

Problem

Standard statistical tokenizers like BPE, WordPiece, and Unigram treat text purely as data compression streams, causing severe Morphological Shattering and Semantic Blindness in morphologically rich languages. For Indic scripts (abugidas), off-the-shelf tokenizers frequently split dependent vowels (matras) from base consonants and violate morpheme boundaries created by phonetic fusions like Sandhi. This results in high token fertility ('Indic Tax'), poorly anchored vector representations, and an inability of downstream models to easily recover whole-word semantics.

Method

SuTRA operates in two distinct phases: pre-tokenization and morphology-aware merging. In Phase 1, orthographic rules (Phi) map each word to a sequence of akshara-like units to preserve script atomicity, while a gold lexicon and a fine-tuned sequence-to-sequence model identify forbidden morpheme boundaries for out-of-vocabulary terms. In Phase 2, a BPE-style merging process scores candidate pairs via S(a, b) = f(a, b) * Psi(a, b)^(gamma_t), where f is corpus frequency, Psi in [0, 1] penalizes merges crossing forbidden boundaries, and gamma_t is a dynamic rigidity constraint.

The rigidity constraint gamma_t is annealed exponentially from gamma_start to gamma_end over training. This curriculum forces the tokenizer to prioritize merging continuous lexical roots early in training while gradually relaxing the penalty to handle functional affixes later. Surface-form integrity is prioritized over canonical purity to allow exact, zero-overhead detokenization via simple string concatenation.

Experimental setup

Experiments are conducted on three Indic languages (Hindi, Marathi, and Gujarati) using the IndicCorp corpus and a ~560k word gold-standard morphological lexicon. Baselines include statistical tokenizers (BPE, WordPiece, SentencePiece, Unigram) and morphological tokenizers (SuperBPE, MorphTok). Evaluation metrics include Boundary F1, Fertility Ratio, Word2Vec Linear and MLP R^2 for semantic recoverability, Jaccard Overlap and Root-Affected Distance for robustness, and chrF2/COMET for machine translation using a 3-layer Transformer (L=3, H=4, d_ff=400) trained for 100k updates with a shared 32k vocabulary.

Results

SuTRA achieves top Boundary F1 scores of 0.586 for Hindi and 0.617 for Marathi while maintaining a controlled fertility ratio (1.412 to 1.755). In semantic recoverability probes, SuTRA yields a +34% relative gain in Linear R^2 over BPE for Hindi (0.4464 vs 0.3329) and exceeds 0.50 R^2 with MLP probes for Marathi and Gujarati. In machine translation, SuTRA attains 38.84 chrF2 and 0.6554 COMET on Marathi-to-Hindi, outperforming BPE (36.55 chrF2) and WordPiece (27.18 chrF2). Furthermore, SuTRA drastically reduces root-affected distance under orthographic noise down to 0.038-0.042 compared to 0.228-0.324 for standard BPE.

TokenizerHi Boundary F1Hi FertilityMr Boundary F1Mr FertilityGu Boundary F1Gu Fertility
BPE (ACL'16)0.4821.2850.4701.2250.5911.126
WordPiece0.4111.2140.5271.3000.5961.173
SentencePiece0.4381.3150.0831.1830.5911.156
Unigram0.4391.3100.5071.2560.6691.137
SuperBPE0.0892.5170.0843.0840.0963.113
SuTRA (Ours)0.5861.4120.6171.7550.5841.454

Limitations

Evaluated exclusively on three Indo-Aryan Indic languages (Hindi, Marathi, Gujarati) using a curated gold morphological dataset, meaning its generalizability to non-abugida or non-Indic morphologically rich languages remains unverified. The framework relies on an auxiliary sequence-to-sequence model and LLM-verified lexicons to identify out-of-vocabulary boundaries, introducing pipeline complexity and resource dependencies during vocabulary construction.

Why read this

Researchers and engineers building large language models for morphologically rich or low-resource abugida scripts should read this to understand how integrating lightweight linguistic priors into tokenization can eliminate morphological shattering, improve semantic grounding, and boost translation performance without increasing vocabulary size.

Code

Applications

Large language model pre-training, machine translation, and speech-to-text systems handling morphologically complex Indic languages.

Institutions

Motilal Oswal Financial Services, Indian Institute of Technology Bombay

Funding / 經費: Motilal Oswal Financial Services

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-291