ForensicsTok: Forensics-Guided Tokenized Modeling for Image Tampering Localization
- lab arXiv
- lab arXivLabs
- product DagsHub
- product ForensicsTok
- product Hierarchical Expert Fusion
- product Hugging Face
- product MLLMs
- product Token Splatting Decoder
A new method called ForensicsTok reformulates image tampering localization as an autoregressive sequence generation task, aiming to overcome limitations in existing multi-modal large language model approaches, according to research submitted to arXiv in 2026 [1][2]. The work, posted on the preprint server arXiv, proposes a unified architecture that directly generates spatially grounded token sequences for precise mask prediction without intermediary supervision [2]. The researchers introduce a Token Splatting Decoder (TSD) to map tokens to binary masks via codebook-aware code smoothing, a technique designed to mitigate sharp gradients from deterministic detokenizers [2]. A Hierarchical Expert Fusion (HEF) module is also proposed to inject multi-scale features from a forensic expert model, compensating for the lack of forensic priors in standard MLLMs [2]. Experiments across six benchmarks showed that ForensicsTok substantially improves over existing MLLM-based baselines and slightly improves over strong forensic expert baselines, while exhibiting stronger robustness to perturbations [2]. The paper was submitted to arXiv on June 23, 2026 [1]. arXiv, which began on August 14, 1991, is an open-access repository of electronic preprints that are moderated but not peer-reviewed, and as of November 2024 received about 24,000 submissions per month [6]. The research appears on the site with a suite of community-developed tools under the arXivLabs framework, which allows collaborators to build features such as bibliographic explorers and code finders directly on article pages [4][5].
research-papersafety-research
Background sources we checked (7)
- arxiv.org ↗ Multi-modal Large Language Models (MLLMs) offer powerful reasoning for forensic tasks, yet existing approaches utilizing exogenous segmentation decoders often suffer from suboptimal localization. The reliance on stitched pipelines introduces information bottlenecks during backpro…
- info.arxiv.org ↗ arXiv Labs - arXiv info | arXiv e-print repository Skip to content # arXiv Labs Attention arXiv Users: arXiv Labs is pausing new proposals ## What are arXiv Labs? arXiv Labs are a way for the community to contribute new, useful features to arXiv. These integrations are avail…
- blog.arxiv.org ↗ arXivLabs: a space for community innovation – arXiv blog arXiv has launched a new, formalized framework enabling innovative collaborations with individuals and organizations. “Members of our community want to contribute tools that enhance the arXiv experience, and we val…
- info.arxiv.org ↗ arXivLabs: Showcase - arXiv info | arXiv e-print repository ... # arXivLabs: Showcase ... arXiv is surrounded by a community of researchers and developers working at the cutting edge of information science and technology. ... While the arXiv team is focused on our core mission—pr…
- en.wikipedia.org ↗ arXiv (pronounced as "archive"—the X represents the Greek letter chi ⟨χ⟩) is an open-access repository of electronic preprints and postprints (known as e-prints) approved for posting after moderation, but not peer reviewed. It consists of scientific papers in the fields of mathem…
- en.wikipedia.org ↗ 14 (fourteen) is the natural number following 13 and preceding 15.…
- en.wikipedia.org ↗ A large language model (LLM) is a type of machine learning model designed for natural language processing tasks such as language generation. LLMs are language models with many parameters, and are trained with self-supervised learning on a vast amount of text.…