---
id: hansen26_interspeech
title: "Balancing Speech, Language and Hearing Science with Machine Learning
  Modeling in the Age of AI: “Know your Problem, Data, and Solution”"
authors:
  - John Hansen
year: 2026
isca_url: https://www.isca-archive.org/interspeech_2026/hansen26_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/hansen26_interspeech.pdf
session: "Keynote1 - John Hansen: Balancing Speech, Language and Hearing Science
  with Machine Learning Modeling in the Age of AI: “Know your Problem, Data, and
  Solution”"
topics:
  - evaluation
category: resources-evaluation
institutions:
  - University of Texas at Dallas
code:
  url: ""
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: hansen26_interspeech
  category: resources-evaluation
  institutions:
    - University of Texas at Dallas
  updated: 2026-09-28
  confidence: abstract-only
  source: https://www.isca-archive.org/interspeech_2026/hansen26_interspeech.html
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/hansen26_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/hansen26_interspeech/markdown.md
---

# Balancing Speech, Language and Hearing Science with Machine Learning Modeling in the Age of AI: "Know your Problem, Data, and Solution"

**Category:** `resources-evaluation`

**TL;DR** — A keynote-style talk argues that speech and hearing science principles need to stay central alongside machine learning advances, drawing on decades of research into speaker variability, human perception, and large-scale conversational corpora.

## Problem

As machine learning and general AI drive major performance gains in speech technology, the talk argues that the underlying speech, language, and hearing science principles are increasingly under-emphasized in system development.

## Method

The talk surveys ways to balance speech/language/hearing science with ML modeling, covering speaker and speech variability (stress, emotion, vocal effort, Lombard effect, non-nativeness), human perception (e.g. cochlear implant innovations), and large-scale conversational language research such as team communications and historical archives.

## Results

As an invited talk rather than an empirical study, its "result" is a set of historical lessons and forward-looking guidance for combining domain science with AI-era speech modeling, aimed particularly at early-career researchers.

## Code

None released (as of this page's `updated` date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

## Applications

Framing research agendas and training for the next generation of speech, hearing, and human-communication AI systems that stay grounded in domain science.

## Institutions / 機構

University of Texas at Dallas

## Related

- [Do speech foundation models perceive speaker similarity as humans do?](kishi26_interspeech.md) — same problem · relatedness 1.6/3
- [Bridging the Speech AI Accessibility Gap for Deaf and Hard of Hearing People](glasser26_interspeech.md) — same problem · relatedness 1.5/3
- [MSU-Bench: Towards Understanding the Conversational Multi-Speaker Scenarios](sun26j_interspeech.md) — same problem · relatedness 1.5/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
