The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs
- lab arXiv
- lab arXivLabs
- location Eastern Europe
- model SpeechLLMs
- person Shree Harsha Bokkahalli Satish
A new study finds that Speech Large Language Models exhibit measurable intersectional bias, delivering less helpful responses to speakers with Eastern European accents, especially when the voice is female-presenting, according to research posted on arXiv [1]. The preprint, authored by Shree Harsha Bokkahalli Satish and colleagues, evaluated three SpeechLLMs using 2,880 controlled interactions across six English accents and two gender presentations [1]. The researchers kept linguistic content constant through voice cloning to isolate the effect of vocal characteristics [1]. SpeechLLMs process spoken input directly, retaining cues such as accent and perceived gender that were previously removed in cascaded pipelines, which introduces speaker-identity-dependent variation in responses [1]. The evaluation employed pointwise LLM-judge ratings, pairwise comparisons, and Best-Worst Scaling with human validation [1]. The results showed recurring directional disparities: Eastern European-accented speech received lower helpfulness scores, and the effect was more pronounced for female-presenting voices [1]. Responses remained polite but differed in helpfulness [1]. While LLM judges captured the directional trend of these biases, human evaluators exhibited significantly higher sensitivity, showing stronger accent-level contrasts [1]. The study was submitted to arXiv on March 15, 2026, as a 70 KB file, with a revised 68 KB version posted on June 18, 2026 [1]. arXiv, an open-access repository of electronic preprints, hosts scientific papers in fields including computer science and electrical engineering and has posted more than two million articles since its founding in 1991 [7]. Large language models, the broader category to which SpeechLLMs belong, are machine learning models with many parameters trained on vast amounts of text for natural language processing tasks such as language generation [9]. The new research adds to a growing body of work examining how such models perform across different demographic groups, as artificial intelligence systems are increasingly deployed in applications throughout industry and academia [3].
controversyresearch-paper
Background sources we checked (8)
- arxiv.org ↗ Speech Large Language Models (SpeechLLMs) process spoken input directly, retaining cues such as accent and perceived gender that were previously removed in cascaded pipelines. This introduces speaker identity dependent variation in responses. We present a large-scale intersection…
- en.wikipedia.org ↗ Artificial intelligence is the capability of computational systems to perform tasks that are typically associated with human intelligence, such as learning, reasoning, problem-solving, perception, and decision-making. Artificial intelligence has been used in applications througho…
- info.arxiv.org ↗ arXiv Labs - arXiv info | arXiv e-print repository Skip to content # arXiv Labs Attention arXiv Users: arXiv Labs is pausing new proposals ## What are arXiv Labs? arXiv Labs are a way for the community to contribute new, useful features to arXiv. These integrations are avail…
- blog.arxiv.org ↗ arXivLabs: a space for community innovation – arXiv blog arXiv has launched a new, formalized framework enabling innovative collaborations with individuals and organizations. “Members of our community want to contribute tools that enhance the arXiv experience, and we val…
- info.arxiv.org ↗ arXivLabs: Showcase - arXiv info | arXiv e-print repository ... # arXivLabs: Showcase ... arXiv is surrounded by a community of researchers and developers working at the cutting edge of information science and technology. ... While the arXiv team is focused on our core mission—pr…
- en.wikipedia.org ↗ arXiv (pronounced as "archive"—the X represents the Greek letter chi ⟨χ⟩) is an open-access repository of electronic preprints and postprints (known as e-prints) approved for posting after moderation, but not peer reviewed. It consists of scientific papers in the fields of mathem…
- en.wikipedia.org ↗ 14 (fourteen) is the natural number following 13 and preceding 15.…
- en.wikipedia.org ↗ A large language model (LLM) is a type of machine learning model designed for natural language processing tasks such as language generation. LLMs are language models with many parameters, and are trained with self-supervised learning on a vast amount of text.…
Sources
- export.arxiv.org — The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs ↗