All papers
Phonetics & linguisticsFull-paper digest

Morphoacoustic Modeling of a Dynamic 3D Vocal Tract Using MRI-Constrained Deformations and FEM Acoustics

Tharinda Piyadasa, Michael Proctor, Tünde Szalay, Joan Glaunès, Amelia Gully, Kirrie Ballard, Emily Kiff, Naeim Sanaei, Sheryl Foster, David Waddington, Craig Jin

Code & resourcesgithub.com/TharindaDilshan/vocal-tract-acoustics-comsol-fem

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

14.0 KB · Ready to paste

Preview copied content

TL;DR — This paper presents a proof-of-concept pipeline combining rtMRI-constrained Large Deformation Diffeomorphic Metric Mapping (LDDMM) with 3D finite element method (FEM) acoustic simulations to model dynamic vocal tract geometries and their time-varying formants during an Australian English VCV sequence. The offset-aligned FEM simulations successfully track acoustic formant trajectories and capture key phonetic features like rhotic F3 lowering, achieving residual dispersions of 78 Hz for F1 and 74 Hz for F2 across landmarks.

Key contributions

  • Integrates volumetric structural MRI of sustained postures with midsagittal real-time MRI (rtMRI) contours via a constrained LDDMM framework to generate smooth 3D intermediate vocal tract geometries across a continuous [5-ô-5] speech transition.
  • Establishes a 3D finite-element acoustic simulation pipeline in COMSOL Multiphysics using head-enclosed watertight geometries, perfectly matched boundary conditions, and a spherical external air domain to evaluate time-varying formants.
  • Performs an empirical acoustic validation comparing 3D FEM-derived resonances against out-of-scanner upright speech measurements and in-scanner audio across five discrete articulatory landmarks.
  • Releases an open-source repository containing the COMSOL acoustic models, minimal example geometries, and documentation to facilitate reproducibility.

Problem

Understanding speech production requires linking dynamic 3D vocal tract geometry to time-varying acoustic responses, yet combining high-resolution volumetric MRI (limited to sustained postures) with real-time 2D midsagittal MRI (limited to continuous speech) remains an open challenge. Simplified 1D acoustic tube models are computationally efficient but inherently neglect critical 3D spatial properties, complex branching, and anterior cavity shaping that dictate acoustic behavior—particularly during complex articulations like rhotics where small geometric variations heavily impact F3 lowering. Meanwhile, fully 3D numerical acoustic simulations lack rigorous validation frameworks to prove that intermediate morphed shapes generated between static targets remain acoustically and anatomically plausible.

Method

The pipeline models speech transitions using two consecutive LDDMM deformations ([5] to [ô] and [ô] to [5]), where a source surface mesh is deformed into a target mesh via a smooth diffeomorphic flow generated by a time-varying velocity field. The deformation is optimized in an RKHS space with a Gaussian kernel scale and varifold-based surface representation, uniquely constrained frame-by-frame by all available rtMRI midsagittal airway contours (42 frames for [5] to [ô] and 23 frames for [ô] to [5]).

For acoustic simulation, ITK-SNAP extracted 3D tissue boundaries from volumetric structural MRI (1.6 x 1.6 x 2.0 mm voxels), which were repaired in Geomagic to be watertight and embedded into an enclosed head and a 0.4 m radius spherical external air domain with a perfectly matched layer boundary condition to simulate free-field radiation. Linear, lossless pressure acoustics frequency-domain simulations were executed in COMSOL Multiphysics using the Helmholtz equation, with vocal tract surfaces treated as rigid, sound-hard walls (thermoviscous boundary layer impedance tests yielded negligible formant shifts < 10 Hz). A uniform normal particle velocity of vn = 1 m/s was applied at the glottal termination, and frequency sweeps from 50 to 3000 Hz with 10 Hz resolution were evaluated at a receiver point 15 cm anterior to the lips to extract F1-F3 spectral peaks.

Experimental setup

Data was collected from an adult female native speaker of Australian English, utilizing 8 mm midsagittal slice rtMRI (FLASH sequence, 72 fps, 0.97 mm^2 in-plane resolution, 16 kHz audio) alongside volumetric structural MRI for 3D target postures. FEM evaluation was performed on five discrete landmark states ([51], [5ô], [ô], [ôa], [52]) sampled from the deformation trajectory and compared against out-of-scanner upright acoustic recordings and in-scanner supine audio, with formants extracted in Praat using the Burg method (30 ms window, 5 formants, 5500 Hz ceiling, 50 Hz pre-emphasis).

Results

Endpoint acoustic checks showed that 3D FEM simulations closely matched sustained out-of-scanner formants for F2 and F3 (e.g., sustained [5:] FEM F2=1400 Hz, F3=2600 Hz vs. upright F2=1358 Hz, F3=2671 Hz), while exhibiting a consistent +110 Hz F1 offset attributed to supine versus upright posture and imaging constraints. After applying a per-formant additive alignment to correct for systematic resonance-formant bias, the FEM points tracked reference landmarks with a residual dispersion (audio minus aligned FEM) of 78 Hz for F1, 74 Hz for F2, and 163 Hz for F3. The higher F3 dispersion was driven by the post-rhotic landmark where the model remained higher than measured values by ~257 Hz, pointing to local geometric differences or missing anterior details like dentition during release.

System / ConditionF1 Residual Dispersion (Hz)F2 Residual Dispersion (Hz)F3 Residual Dispersion (Hz)
Offset-Aligned 3D FEM vs. Out-of-Scanner [5-ô-5]7874163

Limitations

The study is presented as a single-speaker proof-of-concept demonstration, limiting generalizability across speakers, phonetic contexts, and dialect variations. The data pipeline lacks multi-subject evaluations and does not fully model complex thermoviscous boundary layer losses or detailed micro-anatomy such as dentition, which significantly impacts high-frequency transfer functions and F3 behavior. Furthermore, direct temporal synchronization between in-scanner imaging and out-of-scanner acoustic references was not performed, relying instead on landmark-based state alignment.

Why read this

Speech production researchers and computational phoneticians will find this paper a clear technical blueprint for bridging static volumetric MRI targets with dynamic 2D rtMRI and 3D finite-element acoustics. It offers actionable insights into handling systematic MRI-to-audio frequency biases and validating intermediate articulatory morphing trajectories.

Code

Applications

Foundational modeling for physiological speech synthesis, computer animation of talking heads, and clinical assessment of speech motor disorders.

Institutions

University of Sydney, Macquarie University, Universite Paris Cite, University of York, Westmead Hospital

Funding / 經費: Australian Research Council, National Health and Medical Research Council

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1890