TL;DR — This study presents the first empirical investigation into the experiences of Australian Aboriginal English speakers using voice technologies, revealing that systemic ASR failures force users to modify their speech to sound non-Indigenous.
Key contributions
- Conducted the first empirical study exploring how speakers of an indigenised local variety of a global language interact with voice technologies.
- Provided qualitative evidence that ASR-powered systems frequently fail for Australian Aboriginal English speakers, who associate performance with race.
- Documented how virtual assistants extend public pressures of linguistic assimilation into private spaces, requiring users to adopt a 'white voice' to gain accurate transcriptions.
- Employed an Indigenous-led research methodology involving First Nations Research Assistants from the Nyungar Nation to ensure cultural safety and collaborative data interpretation.
Problem
Automated speech recognition technologies promise universal inclusion, yet commercial systems consistently perform poorly for minoritised language varieties like African American English or Native American speech. In Australia, colonisation has left Australian Aboriginal English as the primary communicative mode for most Indigenous people, yet this vibrant contact variety remains unsupported by mainstream language technologies. When ASR systems fail to recognise Aboriginal English, users experience frustration, anger, and self-blame regarding their intelligence, which reinforces historical colonial hierarchies and psychological harm.
Method
The study adopted a qualitative research design built on semi-structured, in-person interviews and written responses administered by Indigenous Research Assistants. A total of 39 participants aged 16 to 92 (median age 33, predominantly from the Nyungar Nation alongside over 25 other First Nations groups across Western Australia and the Northern Territory) were interviewed in private homes and relaxed environments. Interviews relied on culturally appropriate conversational styles, such as 'yarning', to accommodate pragmatic norms like avoiding direct questions and mitigating gratuitous concurrence.
Recorded audio data was transcribed, filtering out non-technology topics to protect participant privacy, and evaluated via thematic analysis following inductive and deductive coding procedures. The analytical process involved recurring collaborative discussions between the academic investigators and the Indigenous Research Assistants to review, refine, and contextualize emerging themes within broader frameworks of linguistic minoritisation and decolonial technology development.
Experimental setup
The dataset comprised audio-recorded interviews from 39 Indigenous participants representing over 25 First Nations across Western Australia and the Northern Territory, with a median interview duration of 38 minutes. Qualitative thematic analysis was performed on the resulting transcripts following institutional review board approval and regular consultations with an Indigenous Advisory Committee. The study compared user perceptions across various commercial voice-enabled devices, virtual assistants, in-car navigation systems, and automated telephone lines like Centrelink.
Results
Participants reported widespread frustration and system abandonment due to ASR failing to handle local vocabulary (such as 'dardy') and phonological differences inherent to Australian Aboriginal English. Users frequently blamed themselves for the technology's shortcomings, questioning their intelligence or education level, and reported adopting linguistic code-switching strategies—such as putting on a 'proper white voice' or speaking 'white way'—to achieve accurate performance. The study did not evaluate baseline quantitative word error rates across automated systems, focusing instead on capturing the lived emotional, psychological, and social impacts of misrecognition.
Limitations
The research is geographically bound to Australia and reflects the specific historical legacy of settler colonialism and language loss in that region. Distrust of academic institutions within Indigenous communities may have self-selected participants who felt comfortable engaging with the research team. Because Aboriginal English lacks a standardized orthography, reliance on qualitative thematic analysis of transcribed oral interactions introduces interpretive challenges inherent to re-contextualizing spoken contact varieties.
Why read this
Speech and ML engineers building spoken language technologies and voice assistants should read this paper to understand the real-world psychological and social harms inflicted by algorithmic bias on minoritised language communities. It challenges technical assumptions about universal coverage and highlights the necessity of community-governed dataset curation.
Code
None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.
Applications
Development of inclusive, community-governed speech recognition systems and culturally safe voice technologies for minoritised indigenous language varieties.
Institutions
University of Western Australia, Google
Related
- Speech Technology and Linguistic Diversity — same problem · relatedness 2.0/3
- Decolonizing Linguistic Policies in Automatic Speech Recognition: A Framework for Cross-Culturally Competent Speech AI — same problem · relatedness 2.0/3
- From Academic Tool to Community Infrastructure: A Call for Indigenous Partnership in Speech Data Governance — same problem · relatedness 1.9/3
- Hybrid Continual Learning for Low-Resource Australian Aboriginal Language Identification — same problem · relatedness 1.9/3
- Dialect Bias in Speech Recognition Across 10 Spanish and French Varieties — same problem · relatedness 1.9/3
All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.
AI-assisted full-paper digest. Check important claims against the original paper.
DOI: 10.21437/Interspeech.2026-1595