AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning

78d ago · Global · primary source: export.arxiv.org

A new benchmark called Agora tests whether large language models can reason across sprawling workplace document collections, a task that requires finding sparse evidence and reconciling inconsistent terminology. The strongest model evaluated reached only 59.4% accuracy, according to a paper posted to arXiv on June 23, 2026 [1]. The benchmark pairs 362 questions with eight domain collections containing 9,664 authentic documents and 372 million tokens, a scale that exceeds any model's context window and forces deliberate exploration rather than exhaustive scanning [1][2]. Agora was built using an agentic pipeline that combines cross-document task synthesis, obfuscation to prevent data leakage, and difficulty filtering [1][2]. Researchers evaluated eight models on the benchmark and found performance varied notably across domains, with no model exceeding the 59.4% accuracy mark [1][2]. The paper defines archive-grounded reasoning as the ability to locate sparse evidence across large, messy collections of workplace files while reconciling inconsistent terminology, units, and time conventions [1][2]. The work appears on arXiv, the open-access repository of electronic preprints that has hosted scientific papers since August 1991 and now receives roughly 24,000 submissions per month as of November 2024 [6]. The paper's abstract page features arXivLabs integrations, a framework launched in 2020 that allows community collaborators to develop tools such as the Bibliographic Explorer and CORE Recommender, which appear as tabs on article record pages [4][5]. arXivLabs requires partners to adhere to values of openness, community, excellence, and user data privacy, and collaborators receive only minimal, anonymized user data necessary for their features to function [4]. Large language models, the type of system tested on Agora, are machine learning models with many parameters trained on vast amounts of text for natural language processing tasks such as language generation [8]. The Agora benchmark is designed to stress-test these models in a setting that combines archive-groundedness, agentic exploration, and cross-domain coverage — dimensions the authors say existing benchmarks address only in isolation [1][2].

research-paperbenchmarkapplication

Background sources we checked (7)
  • arxiv.org ↗ Large language models are increasingly deployed as agents that reason over documents rather than answer from parametric knowledge. We study archive-grounded reasoning: locating sparse evidence across a large, messy collection of workplace files, reconciling inconsistent terminolo…
  • info.arxiv.org ↗ arXiv Labs - arXiv info | arXiv e-print repository Skip to content # arXiv Labs Attention arXiv Users: arXiv Labs is pausing new proposals ## What are arXiv Labs? arXiv Labs are a way for the community to contribute new, useful features to arXiv. These integrations are avail…
  • blog.arxiv.org ↗ arXivLabs: a space for community innovation – arXiv blog arXiv has launched a new, formalized framework enabling innovative collaborations with individuals and organizations. “Members of our community want to contribute tools that enhance the arXiv experience, and we val…
  • info.arxiv.org ↗ arXivLabs: Showcase - arXiv info | arXiv e-print repository ... # arXivLabs: Showcase ... arXiv is surrounded by a community of researchers and developers working at the cutting edge of information science and technology. ... While the arXiv team is focused on our core mission—pr…
  • en.wikipedia.org ↗ arXiv (pronounced as "archive"—the X represents the Greek letter chi ⟨χ⟩) is an open-access repository of electronic preprints and postprints (known as e-prints) approved for posting after moderation, but not peer reviewed. It consists of scientific papers in the fields of mathem…
  • en.wikipedia.org ↗ 14 (fourteen) is the natural number following 13 and preceding 15.…
  • en.wikipedia.org ↗ A large language model (LLM) is a type of machine learning model designed for natural language processing tasks such as language generation. LLMs are language models with many parameters, and are trained with self-supervised learning on a vast amount of text.…

Sources

Spot something wrong? Report an issue