Investigating Faithfulness in Large Audio Language Models

84d ago · Global · primary source: export.arxiv.org

A new study questions whether the reasoning steps generated by Large Audio Language Models are truly grounded in the audio they process. Researchers have proposed a systematic framework to evaluate the faithfulness of Chain-of-Thought explanations in these multimodal systems, identifying a potential disconnect between fluent reasoning and actual auditory input. Large Audio Language Models, or LALMs, combine audio encoders with pretrained Large Language Models to handle tasks that require both listening and reasoning [1]. While they can produce Chain-of-Thought explanations — step-by-step rationales for their conclusions — the reliability of these chains has not been systematically verified [2]. A research team led by Cem Subakan has introduced a framework to measure this faithfulness, defining three core criteria: the reasoning must be hallucination-free, holistic, and demonstrate attentive listening [2]. In artificial intelligence, a hallucination is a response that contains false or misleading information presented as fact [3]. Such errors pose significant challenges for deploying large language models in high-stakes fields like medical diagnostics and supply chain logistics [3]. The new benchmark applies this concern to the audio domain by introducing interventions on both the audio input and the reasoning chain itself to test whether the model's explanations hold up under scrutiny [2]. The framework was tested on two models, Audio Flamingo 3 and Qwen2.5-Omni [2]. Results pointed to a multimodal disconnect: the generated reasoning often aligned with the model's final prediction but was not always strongly anchored in the actual audio content [2]. The models proved vulnerable to hallucinations and could be swayed by adversarial perturbations, raising questions about their trustworthiness in real-world applications [2]. The work arrives as the broader AI field grapples with the reliability of large language models, which are statistical systems trained on vast text corpora [11]. Unlike symbolic AI, which generally avoids hallucination, these neural network-based models can embed plausible-sounding falsehoods in their outputs [3]. The study extends this line of inquiry into the audio modality, where a model might, for example, describe a sound that is not present in the recording. The paper was submitted to the arXiv preprint repository on September 26, 2025, and revised through four versions, with the latest posted on June 17, 2026 [1]. arXiv, which began in 1991, now hosts over two million e-prints and receives roughly 24,000 new submissions each month, serving as a primary venue for rapid dissemination in computer science and physics [9]. The researchers have also released a benchmarking interface and evaluation results on a dedicated project page [2].

research-papertool-releasemodel-releaseproduct-launchsafety-researchbenchmark

Background sources we checked (10)
  • arxiv.org ↗ Large Audio Language Models (LALMs) integrate audio encoders with pretrained Large Language Models to perform complex multimodal reasoning tasks. While these models can generate Chain-of-Thought (CoT) explanations, the faithfulness of these reasoning chains remains unclear. In th…
  • en.wikipedia.org ↗ In the field of artificial intelligence (AI), a hallucination or artificial hallucination (also called bullshitting, confabulation, or delusion) is a response generated by AI that contains false or misleading information presented as fact. The term draws a loose analogy with huma…
  • en.wikipedia.org ↗ Machine learning (ML) is a field of study in artificial intelligence concerned with the development and study of statistical algorithms that can learn from data and generalize to unseen data, and thus perform tasks without being explicitly programmed. Advances in the field of de…
  • en.wikipedia.org ↗ Artificial intelligence (AI) is the capability of computational systems to perform tasks typically associated with human intelligence, such as learning, reasoning, problem-solving, perception, and decision-making. It is a field of research in engineering, mathematics and computer…
  • info.arxiv.org ↗ arXiv Labs - arXiv info | arXiv e-print repository Skip to content # arXiv Labs Attention arXiv Users: arXiv Labs is pausing new proposals ## What are arXiv Labs? arXiv Labs are a way for the community to contribute new, useful features to arXiv. These integrations are avail…
  • blog.arxiv.org ↗ arXivLabs: a space for community innovation – arXiv blog arXiv has launched a new, formalized framework enabling innovative collaborations with individuals and organizations. “Members of our community want to contribute tools that enhance the arXiv experience, and we val…
  • info.arxiv.org ↗ arXivLabs: Showcase - arXiv info | arXiv e-print repository ... # arXivLabs: Showcase ... arXiv is surrounded by a community of researchers and developers working at the cutting edge of information science and technology. ... While the arXiv team is focused on our core mission—pr…
  • en.wikipedia.org ↗ arXiv (pronounced as "archive"—the X represents the Greek letter chi ⟨χ⟩) is an open-access repository of electronic preprints and postprints (known as e-prints) approved for posting after moderation, but not peer reviewed. It consists of scientific papers in the fields of mathem…
  • en.wikipedia.org ↗ 14 (fourteen) is the natural number following 13 and preceding 15.…
  • en.wikipedia.org ↗ A large language model (LLM) is a type of machine learning model designed for natural language processing tasks such as language generation. LLMs are language models with many parameters, and are trained with self-supervised learning on a vast amount of text.…

Sources

Spot something wrong? Report an issue