---
id: dixit26_interspeech
title: "AURA Score: A Metric for Holistic Audio Question Answering Evaluation"
authors:
  - Satvik Dixit
  - Soham Deshmukh
  - Bhiksha Raj
year: 2026
doi: 10.21437/Interspeech.2026-3185
isca_url: https://www.isca-archive.org/interspeech_2026/dixit26_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/dixit26_interspeech.pdf
session: Evaluation of Speech and Audio Analysis
topics:
  - evaluation
  - speech-llm
  - self-supervised
category: resources-evaluation
labels:
  - dataset-or-benchmark-release
institutions:
  - Carnegie Mellon University
funding:
  - National Science Foundation
code:
  url: ""
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: dixit26_interspeech
  category: resources-evaluation
  labels:
    - dataset-or-benchmark-release
  institutions:
    - Carnegie Mellon University
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-3185
  pdf: https://www.isca-archive.org/interspeech_2026/dixit26_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/dixit26_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/dixit26_interspeech/markdown.md
---

# AURA Score: A Metric for Holistic Audio Question Answering Evaluation

*Satvik Dixit, Soham Deshmukh, Bhiksha Raj*

[PDF](https://www.isca-archive.org/interspeech_2026/dixit26_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/dixit26_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-3185)

**Category:** `resources-evaluation` · **Labels:** `dataset-or-benchmark-release`

**TL;DR** — The paper introduces AQEval, a 10k-sample human-annotated benchmark for Audio Question Answering (AQA) evaluation, and proposes AURA, a metric combining LLM reasoning with CLAP audio entailment that outperforms traditional n-gram and captioning metrics.

## Key contributions

- Introduces AQEval, the first human-annotated benchmark specifically for evaluating AQA metrics, containing roughly 10k model responses annotated across 5 human raters for absolute correctness and partial correctness.
- Performs a comprehensive evaluation of legacy NLP and audio captioning metrics (BLEU, METEOR, ROUGE-L, CIDER, SPICE, SPIDER, MACE, FENSE) on AQEval, proving they fail on longer, more complex free-form answers.
- Proposes AURA (Audio Response Assessment), a metric that combines few-shot LLM reasoning with chain-of-thought and an audio entailment module.
- Demonstrates state-of-the-art correlation with human ratings, exceeding the best baseline metric by a factor of 2.2 and outperforming a plain LLM-as-judge baseline by 9.1% overall.

## Problem

As Audio-Language Models (ALMs) shift from closed-form classification to open-ended Audio Question Answering (AQA), researchers have relied on text-based NLG metrics (BLEU, METEOR, ROUGE-L, BERTScore) and audio captioning metrics (FENSE, MACE). These prior approaches fail because they measure mere lexical overlap or surface embedding similarity, remaining entirely question-agnostic. They cannot judge whether a response contextually answers a specific question or aligns with actual audio content, particularly for nuanced, partially correct, or long-form generations.

## Method

AURA computes a holistic score by combining an LLM-based contextual correctness evaluation with an audio grounding module. For the textual component, an LLM (such as Llama 3.1-8B, Gemini 2.5 Pro, Claude Sonnet 3.5, or GPT-4o) is prompted with the question, reference answer, and candidate response, instructed to first output a natural language rationale (Chain-of-Thought) and then rate the answer on a 3-point scale (1 = incorrect, 2 = ambiguous/partially correct, 3 = correct). This category score S_LLM is mapped to 0, 0.5, and 1.

Simultaneously, for audio grounding, the question and response are rewritten into a declarative hypothesis text (h) using an LLM prompt. The hypothesis is embedded via the CLAP text encoder (Et), while the source audio (a) is embedded via the CLAP audio encoder (Ea). The cosine similarity between these embeddings produces an audio entailment score S_AE, thresholded at 0.35.

The final AURA score is calculated as a weighted sum of the normalized LLM score and the audio entailment score: S_AURA = Normalised(S_LLM + w * S_AE), where the entailment weight w is set to 0.1 based on validation ablations. Best configuration uses 3-shot in-context learning with rationalization.

## Experimental setup

Evaluations are conducted on the newly proposed AQEval benchmark, comprising 9,974 entries (8k test, 2k validation) synthesized from ClothoAQA and OpenAQA, utilizing audio clips from Clotho and AudioCaps. Candidate responses are generated by four distinct ALMs: Qwen AudioChat, Audio Flamingo, GAMA, and Qwen2 Audio. Alignment with human judgment is measured via Pearson's rank correlation coefficient (rho). Baselines include BLEU, ROUGE-L, METEOR, CIDER, SPICE, SPIDER, MACE, FENSE, and a zero-shot LLM-as-judge without demonstrations or CoT.

## Results

On aggregate ClothoAQA, AURA achieves a correlation of 72.62 compared to 62.59 for the plain LLM baseline and 31.00 for BLEU. On aggregate OpenAQA, AURA scores 45.44 vs 43.56 for the plain LLM and 17.05 for BLEU. In question-type breakdowns, traditional metrics plummet on medium and long responses (e.g., BLEU drops from 36.92 on words to 17.02 on long responses), whereas AURA maintains high correlation across all lengths, peaking at 61.80 overall. Ablations show that utilizing advanced frontier models like GPT-4o as the core judge pushes overall correlation up to 65.88.

| System / Metric | ClothoAQA Correlation | OpenAQA Correlation | Overall Correlation |
|---|---|---|---|
| BLEU | 31.00 | 17.05 | 23.91 |
| METEOR | 31.65 | 22.64 | 27.86 |
| ROUGE-L | 33.74 | 19.06 | 27.34 |
| FENSE | 23.73 | 21.76 | 17.52 |
| LLM Baseline | 62.59 | 43.56 | 56.64 |
| AURA (Proposed) | 72.62 | 45.44 | 61.80 |

## Limitations

The current audio entailment component relies on zero-shot CLAP models that achieve only around 50% accuracy on standard audio entailment tasks, limiting the marginal gain of the grounding term (w = 0.1). The evaluation is bounded by English-centric datasets (Clotho and AudioCaps) and does not explore multilingual AQA robustness. Furthermore, utilizing frontier LLMs like GPT-4o or Claude Sonnet as the backbone introduces significant computational inference overhead compared to traditional n-gram metrics.

## Why read this

Researchers and engineers building Audio-Language Models or evaluating open-ended audio tasks should read this to adopt a rigorous, human-aligned evaluation metric that goes beyond lexical overlap. It exposes the severe failure modes of traditional NLP metrics on complex audio answers and provides a reproducible blueprint for combining LLM reasoning with audio grounding.

## Code

None released (as of this page's `updated` date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

## Applications

Automated benchmarking of audio-language models, continuous evaluation pipelines for conversational speech assistants, and data quality filtering for speech-text dataset curation.

## Institutions / 機構

Carnegie Mellon University

**Funding / 經費:** National Science Foundation

## Related

- [UG-Bench: A Comprehensive Benchmark for Evaluating Large Audio-Language Models](zhou26c_interspeech.md) — shared data / evaluation · relatedness 2.4/3
- [MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models](yang26c_interspeech.md) — shared data / evaluation · relatedness 2.1/3
- [EChO-Agent: Evidence Chain Orchestration Agent for Audio Reasoning](zhang26t_interspeech.md) — same problem · relatedness 2.1/3
- [Multi-Source Evidence Fusion for Audio Question Answering](olev26_interspeech.md) — same problem · relatedness 2.1/3
- [All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation](foo26_interspeech.md) — shared data / evaluation · relatedness 2.1/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
