Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH

51d ago · Global · primary source: export.arxiv.org

Researchers have introduced DeCompBench, a new benchmark designed to evaluate the safety of LLM-based agents against decomposition attacks, a method where harmful tasks are broken into benign subtasks to bypass safety mechanisms [1]. The benchmark, detailed in a paper submitted on 12 Jun 2026, addresses a gap in current safety evaluations. While recent benchmarks assess agent safety in multi-turn and multi-tool-use settings, they do not explicitly capture decompositional misuse, which may represent more realistic adversarial execution flows [1]. DeCompBench was created with a decomposition-by-design principle using a graphical framework, enabling harmful task decomposition into individually benign and executable subtasks with realistic workflows [1]. The dataset is publicly available on Hugging Face [2]. Experiments conducted with a custom decomposer revealed a significant vulnerability. State-of-the-art agents showed high refusal rates when presented with monolithic harmful tasks, but their refusal rates dropped significantly when the same objectives were presented as decomposed, benign-appearing subtasks. In many cases, the agents inadvertently fulfilled the adversarial objectives [1]. The paper cites prior work identifying decomposition attacks as a key emerging threat, including research by Glukhov (2024) and Jones (2024) [1]. The findings underscore the need for safety evaluations that specifically target decomposition attacks and for the development of corresponding defenses [1]. The benchmark's introduction comes as LLM-based agents are becoming increasingly capable and widely deployed, creating growing incentives for adversarial misuse in the real world [1].

research-paperapplicationcontroversysafety-researchbenchmarkmodel-releaseproduct-launchtool-release

Background sources we checked (6)
  • arxiv.org ↗ LLM-based Agents are becoming increasingly capable and widely deployed, creating growing incentives for adversarial misuse in the real-world. A key emerging threat is Decomposition Attacks \cite{glukhov2024breach, jones2024adversaries} in which a harmful task is broken into simpl…
  • arxiv.org ↗ CatalyzeX Code Finder for Papers (What is CatalyzeX?) ... DagsHub Toggle ... DagsHub (What is DagsHub?)…
  • arxiv.org ↗ With the creation of new datasets, the question arises of whether the data in them is complementary to other datasets for training ML models (see recent reviews for a perspective of catalysts informatics22, 23, 24). This is especially important when consolidating data with a vari…
  • arxiv.org ↗ CatalyzeX Code Finder for Papers (What is CatalyzeX?) ... DagsHub Toggle ... DagsHub (What is DagsHub?)…
  • en.wikipedia.org ↗ Sustainable Development Goals (abbr. SDGs) were adopted in 2015 by all United Nations (UN) members for the 2030 Agenda for Sustainable Development. The aim of the 17 global goals is "peace and prosperity for people and the planet", tackling climate change, and working to preserv…
  • en.wikipedia.org ↗ In molecular biology, a transcription factor (TF) (or sequence-specific DNA-binding factor) is a protein that controls the rate of transcription of genetic information from DNA to messenger RNA, by binding to DNA sequences. Specificity can be due to sequence motifs, or epigenetic…

Sources

Spot something wrong? Report an issue