TL;DR — This study investigates Kaxi, a Chinese folk art where a reed serves as the sound source and hand movements act as an external resonator, demonstrating through kinematic, acoustic, and perceptual analyses that hands can effectively articulate intelligible vowels with an average listener identification accuracy of 83.7%.
Key contributions
- Captured 3D kinematic hand landmarks (21 joints, 63 features via MediaPipe) during Kaxi vowel production, revealing a spatial posture map analogous to tongue height and advancement.
- Identified a staircase-like formant tuning strategy where formants fluctuate in discrete zigzag patterns across harmonic tracks (300-700 Hz) to maintain vowel distinctiveness at high fundamental frequencies.
- Evaluated perceptual intelligibility across 24 native listeners using isolated vowel tests and reaction time metrics, showing robust recognition (83.7% overall accuracy) alongside systematic error patterns.
- Provided empirical insights into speech plasticity and the source-filter model, highlighting potential non-invasive applications for clinical speech rehabilitation following glossectomy.
Problem
The classical source-filter model posits that speech is generated by a laryngeal source and shaped by an internal vocal tract filter (tongue, lips). While physical impairments like laryngectomy or glossectomy disrupt this pathway, individuals rely on vocal tract plasticity or electrolaryngeal devices. However, whether an external tool independent of the vocal tract can dynamically shape both pitch and vowel quality remains underexplored, particularly given the acoustic challenges of high fundamental frequencies where widely spaced harmonics hinder formant estimation and reduce intelligibility.
Method
The study analyzes a skilled male Kaxi performer uttering 197 Mandarin monosyllabic words containing five vowels (a, o, e, i, u) read in both normal speech and Kaxi via a reed and hand gestures. Video recordings were processed using the MediaPipe package in Python to track 21 right-hand anatomical landmarks across 3D coordinates, followed by principal component analysis (PCA) with wrist-relative normalization to extract spatial deformation patterns. Acoustic features (F0, F1, F2, F3) were extracted using VoiceSauce and Praat with manually optimized wideband spectrogram settings, and outliers were filtered via Mahalanobis distance with a Chi-square threshold of p = 0.70.
For perceptual evaluation, 24 native Mandarin speakers participated in a vowel identification task using 200 ms vowel segments lengthened to 500 ms via PSOLA, presented across two randomized blocks of 10 repetitions (100 total trials per listener). Mixed-effects logistic regression and linear mixed-effects models (LMM) were fitted with block and target vowel as fixed effects and participant random intercepts to analyze accuracy and log-transformed reaction times (N = 2,195). Random Forest classifiers with 500 trees were trained on acoustic formants using a 20% test split to evaluate acoustic separability between normal speech and Kaxi conditions.
Experimental setup
Data collection involved 197 Mandarin monosyllabic (C)V words (53 a, 11 o, 16 e, 48 i, 69 u) recorded in a sound-attenuated room at 44,100 Hz / 16-bit using a Maono PD200X microphone, Behringer XENYX 302 USB mixer, and XONAR U5 sound card. Perceptual tests involved 24 native listeners (balanced gender, no speech disorders). Baselines compare Kaxi acoustic distributions and listener recognition against normal speech conditions. Notable implementations include Python (3.12), MediaPipe (0.10.21), R (4.5.1), and randomForest package (500 trees).
Results
Random Forest classification test accuracy dropped from 96.4% in normal speech to 84.1% in Kaxi, driven by acoustic compression and high harmonic spacing (f0 of 300-700 Hz). In perceptual identification tasks, overall Kaxi accuracy reached 83.7% (compared to near-perfect normal speech), with vowel 'a' achieving 99.6% accuracy while 'o' dropped to 68.3% and showed high confusion with 'e' (15%) and 'u' (8.3%). Cognitive processing effort, measured via reaction times, increased significantly in Kaxi, with the vowel 'o' requiring 1,522 ms (a 40% slowdown compared to normal speech) and response times for 'i' increasing by 29% to 1,219 ms.
| System / Condition | Overall Accuracy (%) | F1-F2 RF Test Accuracy (%) | Mean Reaction Time (ms) |
|---|---|---|---|
| Normal Speech | ~100.0% | 96.4% | 940 - 1,110 |
| Kaxi External Resonator | 83.7% | 84.1% | 1,219 - 1,522 |
Limitations
The study investigates a single expert Kaxi performer, limiting speaker-dependent generalizability across different hand sizes and anatomical variations. Evaluation was restricted to isolated monosyllabic vowels rather than continuous speech or diphthongs/consonants. The dataset is constrained to five Mandarin vowels, leaving tonal and multi-syllabic contexts unexplored.
Why read this
Speech scientists and rehabilitation engineers should read this to understand non-vocal-tract acoustic filtering and speech plasticity, providing a foundational reference for hand-articulated speech rehabilitation devices.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Non-invasive clinical speech rehabilitation tools for patients with severe vocal tract impairments such as glossectomy or cleft palate.
Institutions
Peking University
Funding / 經費: National Social Science Foundation of China
Related
- Articulatory Dynamics using Physical Vocal-tract Models — same problem · relatedness 1.6/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-1807