---
id: ly26_interspeech
title: "TinyGiantALM: A Compact Audio-Language Model for Intent-Aware Reasoning
  under Resource Constraints"
authors:
  - Vinh-Thuan Ly
year: 2026
doi: 10.21437/Interspeech.2026-491
isca_url: https://www.isca-archive.org/interspeech_2026/ly26_interspeech.html
pdf_url: https://www.isca-archive.org/interspeech_2026/ly26_interspeech.pdf
session: Challenge - Audio Reasoning Challenge
topics:
  - speech-llm
  - self-supervised
  - low-resource
category: speech-llm-dialogue
labels:
  - efficient-on-device
  - self-supervised
institutions:
  - Zalo AI
  - University of Science, VNU-HCM
  - Vietnam National University
code:
  url: ""
  license: ""
open_to_collaboration: false
wiki_frontmatter:
  id: ly26_interspeech
  category: speech-llm-dialogue
  labels:
    - efficient-on-device
    - self-supervised
  institutions:
    - Zalo AI
    - University of Science, VNU-HCM
    - Vietnam National University
  updated: 2026-09-29
  confidence: full-paper
  digest: v2
  source: https://doi.org/10.21437/Interspeech.2026-491
  pdf: https://www.isca-archive.org/interspeech_2026/ly26_interspeech.pdf
wiki_url: https://interspeech-2026-wiki.vercel.app/papers/ly26_interspeech/
markdown_url: https://interspeech-2026-wiki.vercel.app/papers/ly26_interspeech/markdown.md
---

# TinyGiantALM: A Compact Audio-Language Model for Intent-Aware Reasoning under Resource Constraints

*Vinh-Thuan Ly*

[PDF](https://www.isca-archive.org/interspeech_2026/ly26_interspeech.pdf) · [ISCA page](https://www.isca-archive.org/interspeech_2026/ly26_interspeech.html) · [DOI](https://doi.org/10.21437/Interspeech.2026-491)

**Category:** `speech-llm-dialogue` · **Labels:** `efficient-on-device`, `self-supervised`

**TL;DR** — TinyGiantALM is a compact 1.5B audio-language model that uses an instruction-aware feature refinement framework with a query-guided projector and semantic gating to achieve 46.4% zero-shot accuracy on the MMAR benchmark, outperforming much larger 7B–13B models.

## Key contributions

- Proposes a 1.5B efficiency-oriented audio-language model demonstrating that architectural priors can compensate for reduced scale in audio reasoning.
- Introduces a triple-stream acoustic front-end combining Whisper, HTS-AT, and CLAP encoders to capture fine-grained temporal, event-level, and global semantic representations.
- Develops a Query-guided Triple-stream Projector using E-Branchformer blocks, masked mean-pooling user intent cross-attention, and a CLAP-driven semantic gating mechanism.
- Outperforms traditional 7B-13B base lines (e.g., SALMONN-13B, Qwen2-Audio) on complex mixed-modality tasks by over 36%.

## Problem

Current state-of-the-art audio reasoning models rely on massive parameter scaling exceeding 7B to 30B parameters and expensive reinforcement learning, making them unsuitable for resource-constrained edge devices. Prior architectures use linear projectors or passive processing that fail to filter task-relevant information, causing models like Qwen2-Audio and SALMONN to collapse in complex, overlapping multi-source acoustic scenes. This creates a critical gap in achieving deep audio reasoning and signal disentanglement within an edge-friendly footprint.

## Method

TinyGiantALM adopts a multi-rate resampling triple-stream acoustic front-end consisting of Whisper-Large-v3-turbo (16kHz, 732M total encoder parameters across streams), HTS-AT (48kHz audio for event-level perception), and a CLAP encoder for global semantic priors, all sequence-length fixed to N = 300 tokens via adaptive average pooling. The features are processed by a Query-guided Triple-stream Projector initialized with L = 2 E-Branchformer blocks (combining global MHSA branches and 1D depth-wise convolutional local branches with kernel size 17) to map representations to the d_model = 1024 LLM latent space.

To align acoustic features with the user's instruction, a global user intent vector q_intent is derived via masked mean pooling over instruction tokens M_user, serving as the Query in a Multi-Head Cross-Attention mechanism where encoded audio acts as Key and Value. Next, a soft gate g in (0, 1) is computed via a sigmoid projection of the CLAP global anchor, modulating the features via affine scaling (0.5 * g + 0.5) to inject global context while preserving signal integrity. The final refined embeddings are inserted at the <audio> token position into a Qwen3-0.6B language model backbone.

The model is trained on 558,423 instruction-tuning samples from the CoTA dataset using next-token prediction with assistant response masking, optimizing for a Chain-of-Thought format encapsulating Plan, Audio Analysis, Logic, and Summary tags inside <think> blocks. Training runs for 3 epochs on a single NVIDIA A100 GPU with AdamW optimizer (projector lr 1e-4, LLM lr 5e-5), effective batch size of 32, BFloat16 precision with TF32, and maximum sequence lengths of 300 audio frames and 2048 text tokens.

## Experimental setup

Evaluated on the MMAR benchmark across single and mixed modalities (Sound, Music, Speech, and combinations) containing 16 sub-tasks, compared against baselines including Flamingo-2 (3B), LTU-AS (7B), GAMA (7B), Qwen2-Audio (8.4B), SALMONN (13B), Audio-Reasoner (8.4B), Qwen2.5-Omni (7B), and Gemini 2.0 Flash. Notable implementation details include a single NVIDIA A100 GPU, 3 epochs of fine-tuning, an edge-friendly inference memory footprint of 5GB VRAM for the 1.5B parameter system, and evaluation metrics comprising zero-shot MMAR accuracy (%) and intermediate reasoning Rubrics score.

## Results

TinyGiantALM achieves an overall zero-shot MMAR accuracy of 46.4%, outperforming midscale models like Qwen2-Audio (30.0%) and SALMONN (33.2%), and surpassing Audio-Reasoner (36.8%) by 9.6%. In mixed-modality tasks like Mix-Sound-Music, it attains 45.5% accuracy, avoiding the catastrophic collapse seen in 7B-13B baselines which drop to 9.1%. Ablations confirm that combining the inference query and CLAP gate yields a synergistic +8.40% boost over a vanilla baseline (38.00% to 46.40%), with mixed sound-music accuracy jumping by +36.36%. However, a persistent reasoning gap remains compared to 30B+ foundation models, reflected in a lower Rubrics score (23.77% vs >62.00%) due to limited language modeling capacity causing it to omit granular acoustic triggers.

| System / Condition | Size | MMAR Accuracy (%) |
|---|---|---|
| Qwen2-Audio | 8.4B | 30.0 |
| SALMONN | 13B | 33.2 |
| Audio-Reasoner | 8.4B | 36.8 |
| Qwen2.5-Omni | 7B | 56.7 |
| Gemini 2.0 Flash | - | 65.6 |
| TinyGiantALM (Ours) | 1.5B | 46.4 |

## Limitations

The 1.5B model exhibits a clear reasoning gap, scoring lower in intermediate reasoning Rubrics (23.77%) compared to 30B+ models because it struggles to generate exhaustive, multi-step narrative descriptions. The CLAP-driven semantic gating mechanism can introduce noise in overly dense multi-source scenes (Mix All dropping 4.16% compared to the variant without CLAP) and can strip fine physical cues, leading to performance drops in spatial analysis and correlation tasks.

## Why read this

Speech and ML researchers building efficient, edge-friendly audio-language models should read this to understand how architectural priors like intent-aware cross-attention and semantic gating can match or exceed the perception capabilities of models 5x to 8x larger without brute-force scaling.

## Code

None released (as of this page's `updated` date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

## Applications

On-device interactive virtual assistants, resource-constrained audio surveillance systems, and intent-aware edge audio reasoning hardware.

## Institutions / 機構

Zalo AI, University of Science, VNU-HCM, Vietnam National University

## Related

- [VISA: A Visual Information Strengthened Audio-Reasoning System for the Interspeech 2026 ARC Agent Track](tu26b_interspeech.md) — shared data / evaluation · relatedness 2.9/3
- [Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models](li26o_interspeech.md) — shared data / evaluation · relatedness 2.8/3
- [EChO-Agent: Evidence Chain Orchestration Agent for Audio Reasoning](zhang26t_interspeech.md) — shared data / evaluation · relatedness 2.7/3
- [Structured Prompting vs. Self-Training for Audio Reasoning Under Limited Data and Compute: Lessons from Interspeech Audio Reasoning Challenge 2026](noronha26_interspeech.md) — shared data / evaluation · relatedness 2.7/3
- [ALARM: Audio–Language Alignment for Reasoning Models](grinberg26_interspeech.md) — same problem · relatedness 2.5/3

<sub>All 950k paper pairs scored by TypeSafe Jev (`scripts/related/`); relatedness 0 = unrelated … 3 = directly comparable.</sub>
