BenchX: Benchmarking AI Models for Cancer Detection and Localization with Demographic and Protocol Biases
- lab arXiv
- lab arXivLabs
- model CatalyzeX Code Finder for Papers
- model DagsHub
- model Gotit.pub
- model Hugging Face
- model ScienceCast
- model alphaXiv
A new open benchmark called BenchX, built from 85,355 CT scans, reveals that state-of-the-art AI models for tumor detection perform poorly on rare patient subgroups, including young, female African Americans, according to a preprint posted to arXiv [1]. The benchmark systematically evaluates 12 tumor-detection AI models across tumor size, location, patient subgroup, and imaging protocol [1]. The research, submitted to the arXiv preprint repository on June 23, 2026, aims to quantify inconsistencies that arise when patient demographics and imaging protocols vary in real-world clinical settings [1]. arXiv, founded in 1991, is an open-access repository that hosts scientific papers in fields including computer science and quantitative biology and has grown to receive roughly 24,000 submissions per month as of late 2024 [6]. The study leveraged large language models to extract and organize subgroup information from clinical data, a method the authors describe as both scalable and reproducible [1]. Large language models are machine learning systems trained on vast amounts of text for tasks such as language generation [8]. The use of such models allowed the researchers to structure the demographic and protocol data needed for the benchmark without manual annotation [1]. The findings show that models optimized for average accuracy falter when confronted with underrepresented populations [1]. The authors note that collecting sufficient annotated data for these rare cases is often impractical, which makes subgroup-level evaluation critical for building more reliable medical AI [1]. The benchmark’s dataset and code have been made openly available to provide a foundation for future work on robust tumor-detection models [1].
research-paperbenchmarkcommentary
Background sources we checked (7)
- arxiv.org ↗ Artificial intelligence (AI) has achieved remarkable success in medical imaging, but it is widely recognized that these models often perform inconsistently across real-world clinical settings. Such inconsistencies occur when patient demographics and imaging protocols vary, for ex…
- info.arxiv.org ↗ arXiv Labs - arXiv info | arXiv e-print repository Skip to content # arXiv Labs Attention arXiv Users: arXiv Labs is pausing new proposals ## What are arXiv Labs? arXiv Labs are a way for the community to contribute new, useful features to arXiv. These integrations are avail…
- blog.arxiv.org ↗ arXivLabs: a space for community innovation – arXiv blog arXiv has launched a new, formalized framework enabling innovative collaborations with individuals and organizations. “Members of our community want to contribute tools that enhance the arXiv experience, and we val…
- info.arxiv.org ↗ arXivLabs: Showcase - arXiv info | arXiv e-print repository ... # arXivLabs: Showcase ... arXiv is surrounded by a community of researchers and developers working at the cutting edge of information science and technology. ... While the arXiv team is focused on our core mission—pr…
- en.wikipedia.org ↗ arXiv (pronounced as "archive"—the X represents the Greek letter chi ⟨χ⟩) is an open-access repository of electronic preprints and postprints (known as e-prints) approved for posting after moderation, but not peer reviewed. It consists of scientific papers in the fields of mathem…
- en.wikipedia.org ↗ 14 (fourteen) is the natural number following 13 and preceding 15.…
- en.wikipedia.org ↗ A large language model (LLM) is a type of machine learning model designed for natural language processing tasks such as language generation. LLMs are language models with many parameters, and are trained with self-supervised learning on a vast amount of text.…