TL;DR — This study investigates how Taiwan Mandarin speakers perceive Korean unreleased stop codas (/p, t, k/), revealing an accuracy hierarchy of /p/ > /t/ > /k/ that is heavily modulated by preceding vowel context but unaffected by intermediate course-based proficiency.
Key contributions
- Identifies a clear perceptual accuracy hierarchy (/p/ > /t/ > /k/) for Korean unreleased stop codas among Taiwan Mandarin-speaking learners.
- Demonstrates that preceding vowel contexts (/i, u, a/) systematically alter both identification accuracy and the direction of place-of-articulation confusions.
- Shows that course-based proficiency (beginner vs. intermediate) yields no significant difference in stop-coda identification accuracy, pointing toward a potential perceptual plateau.
- Provides a direct cross-modal comparison mapping these perceptual confusion matrices against known production error profiles from prior literature.
Problem
Taiwan Mandarin-speaking learners of Korean struggle persistently with perceiving and producing Korean stop codas (/p, t, k/), which are obligatorily unreleased and contrast only in place of articulation. Because Taiwan Mandarin permits only nasal codas (/n, ŋ/) and lacks oral stop codas entirely, learners face a cross-linguistic structural mismatch. This gap causes lexical confusion and forces learners to rely heavily on subphonemic, vowel-conditioned coarticulatory cues rather than established L1 phonological categories.
Method
The experiment utilized a four-alternative forced-choice coda identification task implemented in OpenSesame (v3.3.9). The stimulus set consisted of monosyllabic Korean items crossing three preceding vowels (/i, u, a/) with three stop codas (/p, t, k/) and a filler liquid coda (/l/), yielding 12 unique syllables spoken by one male and one female native Seoul Korean speaker (selected via dictation check from eight candidates). Thirty participants (17 beginners, 13 intermediate learners) listened to tokens over headphones and selected the perceived coda consonant from four on-screen options ('ㄹ' /l/, 'ㅂ' /p/, 'ㄷ' /t/, 'ㄱ' /k/).
Statistical analyses evaluated participant-level accuracy using Friedman tests followed by paired Wilcoxon signed-rank tests with Holm correction, while between-group proficiency differences were tested via Wilcoxon rank-sum tests. To analyze response distributions, vowel-by-response chi-square tests of independence with Holm-adjusted p-values were conducted within each stop-coda target, excluding sparse /l/ filler responses.
Experimental setup
The study evaluated 30 Taiwan Mandarin speakers (17 beginners in TOPIK Level 1-aligned courses, 13 intermediate learners in TOPIK Level 2-4-aligned courses) minoring in Korean at National Chengchi University. Stimuli comprised 24 tokens total (12 syllable types multiplied by 2 native Seoul Korean talkers). Metrics included percentage identification accuracy and response distribution percentages across target-coda by vowel-context cells, analyzed via non-parametric inferential statistics (Friedman, Wilcoxon, and chi-square tests).
Results
Collapsed across vowel contexts, learners demonstrated a significant overall accuracy hierarchy of /p/ > /t/ > /k/ (Friedman χ²(2) = 36.07, p < .001; all pairwise comparisons p ≤ .006). Preceding vowel context significantly modulated response distributions for all targets: /p/ was identified most accurately in /i/ and /u/, but in the /a/ context was frequently misperceived as /t/ (26.7% of responses); /t/ was most accurately identified in /a/ (65.0%) and shifted from being misperceived as /p/ after /i/ to being misperceived as /k/ after /a/ and /u/; and /k/ showed the lowest accuracy overall (ranging from 20.0% to 26.7%), being misidentified primarily as /t/ after /i/ and as /p/ after /a/ and /u/ (chi-square Holm-adjusted p < .05 for all targets). Proficiency group showed no statistically significant effect on accuracy for /p/ (p = .5891), /t/ (p = .2591), or /k/ (p = .5923), indicating no reliable performance boost between beginner and intermediate course levels.
| System / Condition | /i/ Context Accuracy | /a/ Context Accuracy | /u/ Context Accuracy |
|---|---|---|---|
| Taiwan Learners (/p/) | 88.3% | 61.7% | 76.7% |
| Taiwan Learners (/t/) | 41.7% | 65.0% | 61.7% |
| Taiwan Learners (/k/) | 20.0% | 26.7% | 20.0% |
| Seoul Natives ([7], /p/) | 89.6% | 97.9% | 84.7% |
| Seoul Natives ([7], /t/) | 84.7% | 93.1% | 95.8% |
| Seoul Natives ([7], /k/) | 85.4% | 93.8% | 70.8% |
Limitations
The study relies on a cross-sectional design with course-based instructional groupings (beginner vs. intermediate) rather than standardized proficiency exams like TOPIK. The participant sample is restricted to 30 Taiwan Mandarin speakers at a single Taiwanese university, and the stimulus set is limited to two talkers and a subset of vowel contexts (/i, u, a/), limiting broader acoustic generalization.
Why read this
Speech researchers and L2 acquisition engineers should read this paper to understand how cross-linguistic phonological inventory gaps and vowel-to-closure coarticulatory transitions drive asymmetric perceptual confusions in non-native stop codas. It provides concrete evidence of a perceptual plateau in intermediate language learners, underscoring why general exposure fails to resolve subphonemic acoustic ambiguities without targeted perceptual training.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Design of computer-assisted pronunciation training (CAPT) software and targeted perceptual training curricula for L2 Korean learners focusing on vowel-conditioned stop-coda minimal pairs.
Institutions
Seoul National University, National Chengchi University
Related
- Perception of English /iː/–/ɪ/ by Japanese Listeners under Silent-Centre and Devoiced Vowel Conditions — same problem · relatedness 1.8/3
- Lexical stress-conditioned spatiotemporal gestural coordination in L2 English — relatedness 1.7/3
- Vowel Allophony Improves Maximum-Likelihood Classification of Warlpiri Consonants — relatedness 1.7/3
- Categorical Perception of Mandarin Tones in Jingpo Native Speakers — relatedness 1.5/3
- Achieving voicelessness in coda stop contexts: Insights from combined electroglottography and laryngoscopy — relatedness 1.5/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-3292