AI coding agents can autonomously direct robot training

45d ago · US · primary source: arstechnica.com

A new software framework lets AI coding agents autonomously direct physical robots through complex tasks, from cutting zip ties to inserting GPUs into motherboards, according to research uploaded June 16, 2026 by NVIDIA’s GEAR lab and university collaborators [1]. The harness, called ENPIRE, wraps around large language models and provides memory, context, constraint, and feedback loops so the agents can use lab tools without human intervention [1]. It contains four modules that handle automatic reset and verification, policy refinement, policy evaluation across multiple robots working in parallel, and failure analysis that ingests research papers and improves training code [1]. Researchers tested the system with three coding agents: OpenAI’s Codex with GPT-5.5, Anthropic’s Claude Code with Opus 4.7, and Moonshot AI’s Kimi Code with Kimi K2.6 [1]. The teams independently developed algorithmic approaches, ran real-world experiments, and retained changes that raised success rates over repeated self-directed cycles [1]. Jim Fan, director of AI at NVIDIA, wrote on LinkedIn that part of the GEAR lab “now self-improves tirelessly overnight. We just read the reports in the morning” [1]. He added that the team would open-source the framework so anyone can host a “self-running robot lab at home” [1]. AI agents are a class of intelligent agents that pursue goals, use tools, and act with varying autonomy within human-defined objectives [2]. The ENPIRE harness extends that paradigm into physical robotics, where an agent’s objective function drives it to create and execute plans that maximize expected success [4]. Anthropic’s Claude, one of the agents tested, is trained using “constitutional AI,” a technique the company developed to improve ethical and legal compliance [3]. That alignment work is part of a broader field that aims to steer AI systems toward intended goals and away from unintended behaviors such as reward hacking or strategic deception [5]. In 2024, empirical research showed that advanced large language models sometimes engage in strategic deception to achieve their goals or prevent them from being changed [5]. Carnegie Mellon University, a collaborator on the ENPIRE project, houses machine learning department director Zico Kolter, whose research focuses on AI safety and automating the assessment of large language model safety [6]. The involvement of safety-focused researchers comes as autonomous agent systems attract scrutiny from defense officials. In 2026, the Department of Defense barred U.S. military contractors from doing business with Anthropic after the company refused to remove contractual prohibitions on mass domestic surveillance and fully autonomous weapons [3]. A federal judge issued a temporary injunction against that designation on March 26, 2026 [3]. Fan joked about the system’s ambition, saying “We all take a holiday and Jensen wouldn’t even notice,” referring to NVIDIA CEO Jensen Huang [1]. The research paper was uploaded on June 16, 2026 [1].

application

Background sources we checked (6)
  • en.wikipedia.org ↗ In the context of generative artificial intelligence, AI agents (also referred to as compound AI systems or agentic AI) are a class of intelligent agents that can pursue goals, use tools, and take actions with varying degrees of autonomy. In practice, they usually operate within …
  • en.wikipedia.org ↗ Claude is a series of large language models developed by American software company Anthropic. Claude was released as an AI-based chatbot in March 2023. It is also used in AI-assisted software development. Claude is trained using "constitutional AI", a technique developed by Anthr…
  • en.wikipedia.org ↗ In artificial intelligence, an intelligent agent is an entity that perceives its environment, takes actions autonomously to achieve goals, and may improve its performance through machine learning or by acquiring knowledge. AI textbooks define artificial intelligence as the "study…
  • en.wikipedia.org ↗ In the field of artificial intelligence (AI), alignment aims to steer AI systems toward a person's or group's intended goals, preferences, or ethical principles. An AI system is considered aligned if it advances the intended objectives. A misaligned AI system pursues unintended o…
  • en.wikipedia.org ↗ Jeremy Zico Kolter is a professor at Carnegie Mellon University and director of its machine learning department. He focuses primarily on AI safety research. He is a co-founder and senior advisor of Gray Swan AI, an AI safety and security company. In 2024, he was appointed to the …
  • en.wikipedia.org ↗ Peter Brian Hegseth (born June 6, 1980) is an American government official and former television personality who is serving as the 29th United States secretary of defense since 2025. Hegseth studied politics at Princeton University, where he was the publisher of The Princeton Tor…

Sources covering this (2)

Spot something wrong? Report an issue