Build a spoken interface
Explore how systems recognize speech, generate a voice, hold a conversation, and translate between languages.
Explore 501 papersTHE RESEARCH LANDSCAPE
See how the collection is distributed across 14 primary categories. Follow a field you know, or find a new place to start.
Looking for a starting point? Explore reading routes01 / THE COLLECTION AT A GLANCE
Ordered by paper count.
Each paper appears in one primary category.
| Research category | PapersShare of 1,379 | Distribution, in papers | Resource coverageLinked papers / category total |
|---|---|---|---|
| Speech recognition From speech to text, across languages and listening conditions. | 23016.7% of all papers | 84 / 23036.5%With resources | |
| Speech synthesis Natural, expressive, and controllable speech generation. | 15211.0% of all papers | 99 / 15265.1%With resources | |
| Resources & evaluation Datasets, benchmarks, and better ways to measure progress. | 14410.4% of all papers | 83 / 14457.6%With resources | |
| Enhancement & separation Making voices clearer and separating the sounds around us. | 1329.6% of all papers | 63 / 13247.7%With resources | |
| Phonetics & linguistics How speech is produced, perceived, and structured. | 1269.1% of all papers | 21 / 12616.7%With resources | |
| Speech LLMs & dialogue Spoken interaction, audio language models, and conversational agents. | 1037.5% of all papers | 46 / 10344.7%With resources | |
| Health & clinical speech Speech as a window into health, accessibility, and care. | 967.0% of all papers | 34 / 9635.4%With resources | |
| Deepfakes & security Detecting synthetic speech and building trustworthy voice systems. | 946.8% of all papers | 47 / 9450.0%With resources | |
| Paralinguistics & emotion The emotion, intent, and human signals beyond words. | 876.3% of all papers | 34 / 8739.1%With resources | |
| Speaker recognition Who is speaking: verification, identification, and diarization. | 735.3% of all papers | 24 / 7332.9%With resources | |
| Audio understanding Understanding acoustic scenes, events, music, and multimodal signals. | 533.8% of all papers | 21 / 5339.6%With resources | |
| Speech coding Audio codecs, compression, and efficient speech representations. | 372.7% of all papers | 21 / 3756.8%With resources | |
| Applications & other Speech technology in practice and emerging research directions. | 362.6% of all papers | 9 / 3625.0%With resources | |
| Speech translation Bridging languages through speech-to-text and speech-to-speech translation. | 161.2% of all papers | 10 / 1662.5%With resources | |
| Entire collection | 1,379100% of papers | Browse all papers | 596 / 1,37943.2%With resources |
A resource link can lead to code, a dataset, a demo, or a project page. Coverage describes links recorded in this wiki, not code availability or research quality. Percentages are rounded.
02 / FOLLOW A QUESTION
Each route opens papers from any of its listed categories.
Choose a direction, then refine your search.
Explore how systems recognize speech, generate a voice, hold a conversation, and translate between languages.
Explore 501 papersFollow work on clearer audio, speaker identity, synthetic-speech detection, and efficient representations.
Explore 336 papersConnect the structure of speech with emotion, human communication, and clinical applications.
Explore 309 papers