All papers
Phonetics & linguisticsFull-paper digest

Can Hands Articulate? Kinematic, Acoustic, and Perceptual Analyses of Vowel Production via External Resonators in Kaxi

Jinyang Ren, Bingliang Zhao, Jiangping Kong, Xiyu Wu

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

11.8 KB · Ready to paste

Preview copied content

TL;DR — This study investigates Kaxi, a Chinese folk art where a reed serves as the sound source and hand movements act as an external resonator, demonstrating through kinematic, acoustic, and perceptual analyses that hands can effectively articulate intelligible vowels with an average listener identification accuracy of 83.7%.

Key contributions

  • Captured 3D kinematic hand landmarks (21 joints, 63 features via MediaPipe) during Kaxi vowel production, revealing a spatial posture map analogous to tongue height and advancement.
  • Identified a staircase-like formant tuning strategy where formants fluctuate in discrete zigzag patterns across harmonic tracks (300-700 Hz) to maintain vowel distinctiveness at high fundamental frequencies.
  • Evaluated perceptual intelligibility across 24 native listeners using isolated vowel tests and reaction time metrics, showing robust recognition (83.7% overall accuracy) alongside systematic error patterns.
  • Provided empirical insights into speech plasticity and the source-filter model, highlighting potential non-invasive applications for clinical speech rehabilitation following glossectomy.

Problem

The classical source-filter model posits that speech is generated by a laryngeal source and shaped by an internal vocal tract filter (tongue, lips). While physical impairments like laryngectomy or glossectomy disrupt this pathway, individuals rely on vocal tract plasticity or electrolaryngeal devices. However, whether an external tool independent of the vocal tract can dynamically shape both pitch and vowel quality remains underexplored, particularly given the acoustic challenges of high fundamental frequencies where widely spaced harmonics hinder formant estimation and reduce intelligibility.

Method

The study analyzes a skilled male Kaxi performer uttering 197 Mandarin monosyllabic words containing five vowels (a, o, e, i, u) read in both normal speech and Kaxi via a reed and hand gestures. Video recordings were processed using the MediaPipe package in Python to track 21 right-hand anatomical landmarks across 3D coordinates, followed by principal component analysis (PCA) with wrist-relative normalization to extract spatial deformation patterns. Acoustic features (F0, F1, F2, F3) were extracted using VoiceSauce and Praat with manually optimized wideband spectrogram settings, and outliers were filtered via Mahalanobis distance with a Chi-square threshold of p = 0.70.

For perceptual evaluation, 24 native Mandarin speakers participated in a vowel identification task using 200 ms vowel segments lengthened to 500 ms via PSOLA, presented across two randomized blocks of 10 repetitions (100 total trials per listener). Mixed-effects logistic regression and linear mixed-effects models (LMM) were fitted with block and target vowel as fixed effects and participant random intercepts to analyze accuracy and log-transformed reaction times (N = 2,195). Random Forest classifiers with 500 trees were trained on acoustic formants using a 20% test split to evaluate acoustic separability between normal speech and Kaxi conditions.

Experimental setup

Data collection involved 197 Mandarin monosyllabic (C)V words (53 a, 11 o, 16 e, 48 i, 69 u) recorded in a sound-attenuated room at 44,100 Hz / 16-bit using a Maono PD200X microphone, Behringer XENYX 302 USB mixer, and XONAR U5 sound card. Perceptual tests involved 24 native listeners (balanced gender, no speech disorders). Baselines compare Kaxi acoustic distributions and listener recognition against normal speech conditions. Notable implementations include Python (3.12), MediaPipe (0.10.21), R (4.5.1), and randomForest package (500 trees).

Results

Random Forest classification test accuracy dropped from 96.4% in normal speech to 84.1% in Kaxi, driven by acoustic compression and high harmonic spacing (f0 of 300-700 Hz). In perceptual identification tasks, overall Kaxi accuracy reached 83.7% (compared to near-perfect normal speech), with vowel 'a' achieving 99.6% accuracy while 'o' dropped to 68.3% and showed high confusion with 'e' (15%) and 'u' (8.3%). Cognitive processing effort, measured via reaction times, increased significantly in Kaxi, with the vowel 'o' requiring 1,522 ms (a 40% slowdown compared to normal speech) and response times for 'i' increasing by 29% to 1,219 ms.

System / ConditionOverall Accuracy (%)F1-F2 RF Test Accuracy (%)Mean Reaction Time (ms)
Normal Speech~100.0%96.4%940 - 1,110
Kaxi External Resonator83.7%84.1%1,219 - 1,522

Limitations

The study investigates a single expert Kaxi performer, limiting speaker-dependent generalizability across different hand sizes and anatomical variations. Evaluation was restricted to isolated monosyllabic vowels rather than continuous speech or diphthongs/consonants. The dataset is constrained to five Mandarin vowels, leaving tonal and multi-syllabic contexts unexplored.

Why read this

Speech scientists and rehabilitation engineers should read this to understand non-vocal-tract acoustic filtering and speech plasticity, providing a foundational reference for hand-articulated speech rehabilitation devices.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Non-invasive clinical speech rehabilitation tools for patients with severe vocal tract impairments such as glossectomy or cleft palate.

Institutions

Peking University

Funding / 經費: National Social Science Foundation of China

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1807