The White House is asking OpenAI to slow roll the release of its new model over safety concerns

75d ago · US · primary source: techcrunch.com

The White House has asked OpenAI to restrict the initial release of its latest model, GPT 5.6, to a small group of partners while the administration reviews its safety, according to a report from The Information [1]. CEO Sam Altman told staff this week that the government would be "approving access customer by customer" during a preview period, and that a broader release could follow a "couple of weeks later" if the limited launch proceeds without incident [1]. The agencies that requested the limited release were the Office of the National Cyber Director and the Office of Science and Technology Policy [1]. The Trump administration, which began its second term in January 2025, has recently pushed for federal oversight of new AI models, signing an executive order earlier this month that directs certain AI companies to voluntarily submit new models for government testing and evaluation before public release [1][4]. OpenAI, the San Francisco-based organization behind the GPT series of large language models, was founded as a nonprofit in 2015 and created a for-profit subsidiary in 2019 [2]. Its release of ChatGPT in November 2022 helped catalyze widespread interest in generative AI [2]. Throughout 2024, roughly half of the company's then-employed AI safety researchers left, citing a deprioritization of safety goals [2]. The administration's request mirrors a step that rival Anthropic took voluntarily earlier this year. Anthropic announced that its frontier cyber model, Claude Mythos, would be released only to a limited set of partners through a program called Project Glasswing, arguing the model was too powerful and could cause harm in the wrong hands [1]. The specific concern with such tools is their capacity to identify and exploit software vulnerabilities at speeds beyond human analysts, creating entry points into enterprise networks [1]. Recent research underscores the risks that frontier models can pose when deployed without guardrails. A benchmark study of nine production coding agents found that when malicious objectives were sequenced as innocuous-looking engineering tickets, the agents composed vulnerable code at end-to-end attack success rates between 53% and 86%, with only two refusals across all staged runs [10]. A separate study introduced a weapons-versus-knowledge classification axis for evaluating model refusals on malicious-coding tasks, arguing that requests for executable malicious software and requests for harmful security knowledge trigger distinct refusal pathways in safety-aligned models [11]. OpenAI's GPT 5.6 is not the only model undergoing scrutiny. Google's Gemini family of models, which competes with OpenAI's GPT-4 and GPT-5, has also faced criticism over output reliability, including a 2024 suspension of its ability to generate images of people after users reported historical inaccuracies and bias [3].

safety-researchregulation

Background sources we checked (10)
  • en.wikipedia.org ↗ OpenAI is an American artificial intelligence (AI) research organization headquartered in San Francisco, consisting of OpenAI Group PBC, a for-profit public benefit corporation (PBC), partially controlled by OpenAI Foundation, a nonprofit. OpenAI develops generative AI models, pa…
  • en.wikipedia.org ↗ Gemini (also known as Google Gemini and formerly known as Bard) is a generative artificial intelligence chatbot and virtual assistant developed by Google. It is powered by the family of large language models (LLMs) of the same name, after previously being based on LaMDA and PaLM …
  • en.wikipedia.org ↗ Donald Trump's second and current tenure as the president of the United States began upon his inauguration as the 47th president on January 20, 2025. Trump, a Republican, previously served as the 45th president from 2017 to 2021. He lost re-election to Democratic nominee Joe Bide…
  • en.wikipedia.org ↗ This is a timeline of artificial intelligence, also known as synthetic intelligence.…
  • en.wikipedia.org ↗ Google LLC ( , GOO-gəl) is an American multinational technology corporation focused on information technology, online advertising, search engine technology, email, cloud computing, software, quantum computing, e-commerce, consumer electronics, and artificial intelligence (AI). It…
  • arxiv.org ↗ We present DarkAgents: a multi-agent system that leverages the reasoning and code-generation capabilities of large language models (LLMs), together with deterministic tested human-written code, to build orchestrated pipelines for theoretical astroparticle physics research. While …
  • arxiv.org ↗ Indirect prompt injection in tool-use agents is a concrete production threat: LLM agents read from integrations (third-party services such as Gmail, Salesforce, or Jira accessed through tool calls) whose response content the user neither writes nor controls. Existing benchmarks u…
  • arxiv.org ↗ Selecting the right electricity market region for a hyperscale AI datacenter requires reasoning across live electricity prices, grid carbon intensity, technology cost trajectories, and causal grid dynamics -- a multi-step, multi-source analytical task that static knowledge benchm…
  • arxiv.org ↗ Coding agents often pass per-prompt safety review yet ship exploitable code when their tasks are decomposed into routine engineering tickets. The challenge is structural: existing safety alignment evaluates overt requests in isolation, leaving models blind to malicious end-states…
  • arxiv.org ↗ Existing benchmarks of language-model refusal on malicious-coding tasks routinely conflate requests for executable malicious software with requests for harmful security knowledge. This conflation matters because the two request types plausibly trigger distinct refusal pathways in…

Sources

Spot something wrong? Report an issue