Gaming-Resistant Insurance Contracts for Autonomous AI Agents: Strategy-Proof Toll Mechanism Design

85d ago · Global · primary source: export.arxiv.org

A new paper proposes a strategy-proof toll mechanism designed to prevent autonomous AI agents from gaming insurance contracts, extending prior actuarial runtime work by treating the operator as a strategic actor rather than a passive one [1][2]. The paper, submitted to arXiv on 15 June 2026 under the Computer Science and Game Theory category, builds on an earlier framework referred to as Paper A [1]. That prior work defined a time-consistent actuarial runtime that prices each side-effect-bearing action against a contractually fixed safe default and gates execution against a reserve budget, but it assumed a passive operator [2]. The new research makes the operator strategic and characterizes a five-attack space for autonomous AI-agent insurance contracts [2]. Two of those attack surfaces — post-toll safe-default selection and within-boundary action splitting — are already closed by Paper A's minimal-authority and no-splitting clauses [2]. The remaining three require new contract clauses. First, common-control aggregation prevents cross-boundary re-routing from reducing the toll below the boundary potential applied to total exposure [2]. Second, interface failures such as invalid JSON are treated as contract-relevant events rather than safety wins; treating them as zero-toll safe defaults can reward unreliable models, while escalation fees reverse the incentive [2]. The authors validate this interface-compliance theorem on committed cross-model traces from a companion empirical paper [2]. Third, a model-identity menu with a componentwise-minimum penalty schedule makes truthful reporting of the deployed model weakly dominant [2]. The clauses are then composed with Paper A's runtime guarantees to obtain joint incentive compatibility over the full five-attack space [2]. A two-parameter premium family discharges operator individual rationality and weak budget balance at the truthful equilibrium [2]. The result is described as an incentive-compatibility layer for actuarial control of autonomous-agent side effects [2]. The paper appears on arXiv, an open-access repository of electronic preprints that, as of November 2024, receives about 24,000 submissions per month and has surpassed two million articles [6].

research-paperapplicationregulation

Background sources we checked (7)
  • arxiv.org ↗ Paper A defines a time-consistent actuarial runtime that prices each side-effect-bearing action against a contractually fixed safe default and gates execution against a reserve budget. It treats the operator as passive. This paper makes the operator strategic. We characterise a f…
  • info.arxiv.org ↗ arXiv Labs - arXiv info | arXiv e-print repository Skip to content # arXiv Labs Attention arXiv Users: arXiv Labs is pausing new proposals ## What are arXiv Labs? arXiv Labs are a way for the community to contribute new, useful features to arXiv. These integrations are avail…
  • blog.arxiv.org ↗ arXivLabs: a space for community innovation – arXiv blog arXiv has launched a new, formalized framework enabling innovative collaborations with individuals and organizations. “Members of our community want to contribute tools that enhance the arXiv experience, and we val…
  • info.arxiv.org ↗ arXivLabs: Showcase - arXiv info | arXiv e-print repository ... # arXivLabs: Showcase ... arXiv is surrounded by a community of researchers and developers working at the cutting edge of information science and technology. ... While the arXiv team is focused on our core mission—pr…
  • en.wikipedia.org ↗ arXiv (pronounced as "archive"—the X represents the Greek letter chi ⟨χ⟩) is an open-access repository of electronic preprints and postprints (known as e-prints) approved for posting after moderation, but not peer reviewed. It consists of scientific papers in the fields of mathem…
  • en.wikipedia.org ↗ 14 (fourteen) is the natural number following 13 and preceding 15.…
  • en.wikipedia.org ↗ A large language model (LLM) is a type of machine learning model designed for natural language processing tasks such as language generation. LLMs are language models with many parameters, and are trained with self-supervised learning on a vast amount of text.…

Sources

Spot something wrong? Report an issue