Agent trajectories as programs: fingerprinting and programming coding-agent behavior

86d ago · Global · primary source: export.arxiv.org

Multi-source synthesis by The Embedding Report from 2 sources. Every numeric and quoted claim traces to a cited source body (see methodology).

Researchers have introduced new methods for comparing AI agents and understanding their behavior, enabling the identification of unique 'behavioral habits' and the development of procedural representations for agent problem-solving procedures.

The researchers, as reported in a paper submitted on June 15, 2026[1], compared ten agents and found they could be identified by their behavioral habits with 85.7% accuracy when attributing unseen trajectories to the correct agent[1]. They developed procedural representations for agent problem-solving procedures using an emergent vocabulary induction technique. The framework was applied to the SWE-Bench dataset to study the structural distinctness of agent trajectories. Additionally, the researchers introduced ProcGrep, a library for auditing and evaluating agents. Another study proposed a new approach to understanding AI model behavior by analyzing agent trajectories and identifying the 'intent-execution' gap between model assumptions and harness behavior[2]. The researchers developed a simple and customizable harness called 'Simple Strands Agent' (SSA), which was able to reproduce or improve on the pass@1 performance reported by diverse model-provider families on popular agentic benchmarks. An analysis of 138k trajectories generated by SSA revealed model-level differences in problem-solving behavior[2].

applicationbenchmarkmodel-releaseresearch-papertool-releasecommentary

Background sources we checked (9)
  • arxiv.org ↗ Benchmark scores tell you what an agent got right; they do not tell you how it got there. In this work, we introduce methods for comparing agents procedurally in different contexts, where the model, tasks, and approaches vary. We compare ten agents and find that they are identifi…
  • en.wikipedia.org ↗ Total Information Awareness (TIA) was a mass detection program by the United States Information Awareness Office. It operated under this title from February to May 2003 before being renamed Terrorism Information Awareness. Based on the concept of predictive policing, TIA was mean…
  • en.wikipedia.org ↗ The assassination of John F. Kennedy, the 35th president of the United States, on November 22, 1963, has spawned numerous conspiracy theories. These theories allege the involvement of the Central Intelligence Agency (CIA), the Mafia, Vice President Lyndon B. Johnson, Cuban prime …
  • en.wikipedia.org ↗ Algorithmic bias describes systematic and repeatable harmful tendency in a computerized sociotechnical system to create "unfair" outcomes, such as "privileging" one category over another in ways that may or may not be different from the intended function of the algorithm. Bias ca…
  • en.wikipedia.org ↗ These datasets are used in machine learning (ML) research and have been cited in peer-reviewed academic journals. Datasets are an integral part of the field of machine learning. Major advances in this field can result from advances in learning algorithms (such as deep learning), …
  • en.wikipedia.org ↗ Functional magnetic resonance imaging or functional MRI (fMRI) measures brain activity by detecting changes associated with blood flow. This technique relies on the fact that cerebral blood flow and neuronal activation are coupled: When an area of the brain is in use, blood flow …
  • en.wikipedia.org ↗ Hangzhou DeepSeek Artificial Intelligence Basic Technology Research Co., Ltd., doing business as DeepSeek, is a Chinese artificial intelligence (AI) company that develops large language models (LLMs). Based in Hangzhou, Zhejiang, DeepSeek is owned and funded by High-Flyer, a Chin…
  • en.wikipedia.org ↗ Douwe Kiela is a Dutch-American research scientist and entrepreneur working in the field of artificial intelligence with a focus on machine learning and natural language processing. He is a research scientist director at Google DeepMind. He previously co-founded and served as CEO…
  • en.wikipedia.org ↗ A large language model (LLM) is a type of machine learning model designed for natural language processing tasks such as language generation. LLMs are language models with many parameters, and are trained with self-supervised learning on a vast amount of text.…

Sources cited (2)

  1. arxiv.org ↗ E
  2. arxiv.org ↗ E
Spot something wrong? Report an issue