All papers
Resources & evaluationFull-paper digest

AURA Score: A Metric for Holistic Audio Question Answering Evaluation

Satvik Dixit, Soham Deshmukh, Bhiksha Raj

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.5 KB · Ready to paste

Preview copied content

TL;DR — The paper introduces AQEval, a 10k-sample human-annotated benchmark for Audio Question Answering (AQA) evaluation, and proposes AURA, a metric combining LLM reasoning with CLAP audio entailment that outperforms traditional n-gram and captioning metrics.

Key contributions

  • Introduces AQEval, the first human-annotated benchmark specifically for evaluating AQA metrics, containing roughly 10k model responses annotated across 5 human raters for absolute correctness and partial correctness.
  • Performs a comprehensive evaluation of legacy NLP and audio captioning metrics (BLEU, METEOR, ROUGE-L, CIDER, SPICE, SPIDER, MACE, FENSE) on AQEval, proving they fail on longer, more complex free-form answers.
  • Proposes AURA (Audio Response Assessment), a metric that combines few-shot LLM reasoning with chain-of-thought and an audio entailment module.
  • Demonstrates state-of-the-art correlation with human ratings, exceeding the best baseline metric by a factor of 2.2 and outperforming a plain LLM-as-judge baseline by 9.1% overall.

Problem

As Audio-Language Models (ALMs) shift from closed-form classification to open-ended Audio Question Answering (AQA), researchers have relied on text-based NLG metrics (BLEU, METEOR, ROUGE-L, BERTScore) and audio captioning metrics (FENSE, MACE). These prior approaches fail because they measure mere lexical overlap or surface embedding similarity, remaining entirely question-agnostic. They cannot judge whether a response contextually answers a specific question or aligns with actual audio content, particularly for nuanced, partially correct, or long-form generations.

Method

AURA computes a holistic score by combining an LLM-based contextual correctness evaluation with an audio grounding module. For the textual component, an LLM (such as Llama 3.1-8B, Gemini 2.5 Pro, Claude Sonnet 3.5, or GPT-4o) is prompted with the question, reference answer, and candidate response, instructed to first output a natural language rationale (Chain-of-Thought) and then rate the answer on a 3-point scale (1 = incorrect, 2 = ambiguous/partially correct, 3 = correct). This category score S_LLM is mapped to 0, 0.5, and 1.

Simultaneously, for audio grounding, the question and response are rewritten into a declarative hypothesis text (h) using an LLM prompt. The hypothesis is embedded via the CLAP text encoder (Et), while the source audio (a) is embedded via the CLAP audio encoder (Ea). The cosine similarity between these embeddings produces an audio entailment score S_AE, thresholded at 0.35.

The final AURA score is calculated as a weighted sum of the normalized LLM score and the audio entailment score: S_AURA = Normalised(S_LLM + w * S_AE), where the entailment weight w is set to 0.1 based on validation ablations. Best configuration uses 3-shot in-context learning with rationalization.

Experimental setup

Evaluations are conducted on the newly proposed AQEval benchmark, comprising 9,974 entries (8k test, 2k validation) synthesized from ClothoAQA and OpenAQA, utilizing audio clips from Clotho and AudioCaps. Candidate responses are generated by four distinct ALMs: Qwen AudioChat, Audio Flamingo, GAMA, and Qwen2 Audio. Alignment with human judgment is measured via Pearson's rank correlation coefficient (rho). Baselines include BLEU, ROUGE-L, METEOR, CIDER, SPICE, SPIDER, MACE, FENSE, and a zero-shot LLM-as-judge without demonstrations or CoT.

Results

On aggregate ClothoAQA, AURA achieves a correlation of 72.62 compared to 62.59 for the plain LLM baseline and 31.00 for BLEU. On aggregate OpenAQA, AURA scores 45.44 vs 43.56 for the plain LLM and 17.05 for BLEU. In question-type breakdowns, traditional metrics plummet on medium and long responses (e.g., BLEU drops from 36.92 on words to 17.02 on long responses), whereas AURA maintains high correlation across all lengths, peaking at 61.80 overall. Ablations show that utilizing advanced frontier models like GPT-4o as the core judge pushes overall correlation up to 65.88.

System / MetricClothoAQA CorrelationOpenAQA CorrelationOverall Correlation
BLEU31.0017.0523.91
METEOR31.6522.6427.86
ROUGE-L33.7419.0627.34
FENSE23.7321.7617.52
LLM Baseline62.5943.5656.64
AURA (Proposed)72.6245.4461.80

Limitations

The current audio entailment component relies on zero-shot CLAP models that achieve only around 50% accuracy on standard audio entailment tasks, limiting the marginal gain of the grounding term (w = 0.1). The evaluation is bounded by English-centric datasets (Clotho and AudioCaps) and does not explore multilingual AQA robustness. Furthermore, utilizing frontier LLMs like GPT-4o or Claude Sonnet as the backbone introduces significant computational inference overhead compared to traditional n-gram metrics.

Why read this

Researchers and engineers building Audio-Language Models or evaluating open-ended audio tasks should read this to adopt a rigorous, human-aligned evaluation metric that goes beyond lexical overlap. It exposes the severe failure modes of traditional NLP metrics on complex audio answers and provides a reproducible blueprint for combining LLM reasoning with audio grounding.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Automated benchmarking of audio-language models, continuous evaluation pipelines for conversational speech assistants, and data quality filtering for speech-text dataset curation.

Institutions

Carnegie Mellon University

Funding / 經費: National Science Foundation

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-3185