JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks
- lab arXivLabs
- location arXiv
- model JADE
- person Lanbo Lin
- product BizBench
- product DagsHub
- product HealthBench
- product Hugging Face
A research team has introduced JADE, a two-layer evaluation framework designed to resolve a persistent tension in assessing AI systems on open-ended professional tasks, according to a paper posted on arXiv [1]. The framework, detailed in a submission last revised on 14 June 2026, aims to balance the rigor of static rubrics with the flexibility of dynamic assessment [1]. The authors, including Lanbo Lin, argue that current methods fall short: fixed rubrics cannot accommodate diverse valid responses, while evaluators that use large language models as judges often exhibit instability and bias [2]. Human experts navigate this by applying domain-grounded principles alongside claim-level analysis, a process JADE seeks to replicate [2]. The framework operates in two layers. The first encodes expert knowledge as a predefined set of evaluation skills to provide stable criteria. The second performs report-specific, claim-level evaluation, incorporating an evidence-dependency gating mechanism that invalidates conclusions built on refuted claims [2]. Experiments on BizBench, a business-task benchmark, showed that JADE improved evaluation stability and surfaced critical agent failure modes that holistic LLM-based evaluators missed [2]. The researchers also reported strong alignment with expert-authored rubrics and demonstrated the framework’s transferability to HealthBench, a benchmark spanning 10 medical and professional domains [2]. The initial preprint was submitted on 6 February 2026 at a size of 4,410 KB, with a revised version posted at 2,210 KB [1]. The paper appears on arXiv, an open-access repository that hosts preprints without peer review and has grown to receive about 24,000 submissions per month as of late 2024 [6]. The field of AI evaluation has drawn broader scrutiny as researchers debate the reliability of machine learning systems. The metascience movement, which studies the practices and incentive structures of research itself, has highlighted widespread methodological flaws and a replication crisis across scientific disciplines [4]. These concerns extend to artificial intelligence, where the challenge of aligning advanced systems with human values has prompted warnings from prominent figures. A 2023 statement signed by hundreds of AI experts and public figures declared that mitigating the risk of extinction from AI should be a global priority alongside pandemics and nuclear war [3]. The JADE framework enters this landscape as a targeted effort to make the evaluation of professional AI agents more reproducible and transparent. The code and data for the project are publicly available on GitHub [2].
tool-releasemodel-releaseresearch-paperproduct-launchcontroversysafety-researchbenchmarkapplication
Background sources we checked (7)
- arxiv.org ↗ Evaluating agentic AI on open-ended professional tasks faces a fundamental dilemma between rigor and flexibility. Static rubrics provide rigorous, reproducible assessment but fail to accommodate diverse valid response strategies, while LLM-as-a-judge approaches adapt to individua…
- en.wikipedia.org ↗ Existential risk from artificial intelligence, or AI x-risk, refers to the idea that substantial progress in artificial general intelligence (AGI) and artificial superintelligence (ASI) could lead to human extinction or an irreversible global catastrophe. One argument for the val…
- en.wikipedia.org ↗ Metascience, also known as science of science, is the systematic study of science itself. It analyzes the practices, structures, and outcomes of scientific research, including research design, peer review, publication bias, replication, and incentive structures, with the aim of i…
- en.wikipedia.org ↗ David is a masterpiece of Italian Renaissance sculpture in marble created from 1501 to 1504 by Michelangelo. With a height of 5.17 metres (17 ft 0 in), the David was not only the first colossal marble statue made in the High Renaissance, but also the first since classical antiqu…
- en.wikipedia.org ↗ arXiv (pronounced as "archive"—the X represents the Greek letter chi ⟨χ⟩) is an open-access repository of electronic preprints and postprints (known as e-prints) approved for posting after moderation, but not peer reviewed. It consists of scientific papers in the fields of mathem…
- en.wikipedia.org ↗ 14 (fourteen) is the natural number following 13 and preceding 15.…
- en.wikipedia.org ↗ "Attention Is All You Need" is a 2017 research paper in machine learning authored by eight scientists and engineers working at Google. The paper introduced a new deep learning architecture known as the transformer, based on the attention mechanism proposed in 2014 by Bahdanau et …
Sources
- export.arxiv.org — JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks ↗