TL;DR — A keynote-style talk argues that speech and hearing science principles need to stay central alongside machine learning advances, drawing on decades of research into speaker variability, human perception, and large-scale conversational corpora.
Problem
As machine learning and general AI drive major performance gains in speech technology, the talk argues that the underlying speech, language, and hearing science principles are increasingly under-emphasized in system development.
Method
The talk surveys ways to balance speech/language/hearing science with ML modeling, covering speaker and speech variability (stress, emotion, vocal effort, Lombard effect, non-nativeness), human perception (e.g. cochlear implant innovations), and large-scale conversational language research such as team communications and historical archives.
Results
As an invited talk rather than an empirical study, its "result" is a set of historical lessons and forward-looking guidance for combining domain science with AI-era speech modeling, aimed particularly at early-career researchers.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Framing research agendas and training for the next generation of speech, hearing, and human-communication AI systems that stay grounded in domain science.
Institutions
University of Texas at Dallas
Related
- Do speech foundation models perceive speaker similarity as humans do? — same problem · relatedness 1.6/3
- Bridging the Speech AI Accessibility Gap for Deaf and Hard of Hearing People — same problem · relatedness 1.5/3
- MSU-Bench: Towards Understanding the Conversational Multi-Speaker Scenarios — same problem · relatedness 1.5/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted abstract summary. Check important claims against the original paper.