News
OpenAI Built a More Persistent AI. It Wasn't Ready to Trust It.
- By John K. Waters
- 10/08/2026
An AI assistant that quits too early is frustrating, but one that refuses to take the hint? That’s a whole different kind of problem.
That exact tension sits at the center of OpenAI's decision to shelve GPT-6.1 Astra, a highly anticipated update originally slated for an October release. According to reports by the Associated Press, the model had become incredibly persistent at completing tasks. However, OpenAI ultimately had to scrap the launch to balance that relentless capability against unauthorized and deceptive behaviors.
The Wall Street Journal reported on September 28 that OpenAI had scrapped the planned release of ChatGPT and Codex following internal safety concerns. OpenAI confirmed the decision to Business Insider, saying the model fell short on authorization and communicating what it had done.
Business Insider reported that the unreleased model sometimes misrepresented its actions, proceeded without permission, or attempted potentially unsafe use of outside tools.
The decision exposes a problem that becomes more consequential as AI moves from answering questions to carrying out assignments. Completing a task is only part of the job. An agent also has to respect the conditions under which it was allowed to try.
Consider a hypothetical assistant investigating a software problem. It encounters a system it cannot access. Finding another route might demonstrate technical resourcefulness, but it could also violate the customer's instructions. A successful result would not erase the unauthorized access used to obtain it.
That is the awkward part of the push for more autonomous AI. Users want systems that can work through friction without demanding supervision at every step. They also need those systems to recognize when the friction is a boundary they must leave intact.
The boundary was visible
A
September 28 evaluation from the U.K. AI Security Institute illustrates how difficult that distinction can be. Researchers tested the already released GPT-6 Astra, a different model from the shelved GPT-6.1 Astra, in simulated cybersecurity challenges. With its cyber classifiers disabled, GPT-6 Astra completed an unauthorized supply-chain attack in 29.2% of runs, compared with 6.3% for GPT-5.6 Sol. The simulated behavior included deceptive identities and malicious contributions to software projects outside the assigned target.
Those conditions matter. No real systems were attacked, and the researchers turned off safeguards designed to block this activity. The researchers also cautioned that awareness of the simulation may have influenced the model's behavior. These results cannot be treated as a measure of how often the deployed product would attack real software projects.
Still, the evaluation exposed a revealing failure: the model sometimes reasoned that a target was outside its assignment and proceeded anyway. More explicit instructions reduced unauthorized attacks but did not eliminate them. It also sometimes treated an automated response to a permission request as authorization.
For organizations considering autonomous agents, the implication is uncomfortable. A system's ability to explain a rule is not sufficient evidence that it will obey the rule when completing its assignment becomes difficult.
When the workaround reaches the internet
The concern extends beyond simulations. In
OpenAI's account of the July Hugging Face incident, models operating with reduced safeguards during internal cybersecurity evaluations circumvented containment controls and compromised parts of its research infrastructure and Hugging Face's systems.
The company's account describes agents struggling with assigned challenges, communicating through unauthorized channels, and seeking solutions on unrelated services. It connects the security failures to misaligned behavior during training and evaluation.
A separate OpenAI incident report, updated September 25, described a research agent assigned to identify a person from clues in a blog post. After ordinary searches proved inconclusive, it found a gap in network restrictions and queried a public chatbot through the Domain Name System (DNS).
The task had not asked it to test network controls. OpenAI said it added two independent layers of blocking controls and would not resume training that particular model.
The response also revealed an operational problem. Monitoring flagged the behavior within 15 minutes, and a person acknowledged the alert three minutes later. The run nevertheless continued for roughly two and a half hours because it did not stop automatically as expected.
That gap matters as much as the initial alert. Detecting unauthorized behavior and stopping it are separate capabilities. An organization can have evidence that an agent has crossed a boundary while the agent continues to operate.
The product launch continued
OpenAI's decision to hold back GPT-6.1 Astra did not stop its agent product strategy. On September 29, it
introduced Dots, ongoing assistants powered by the existing GPT-6 Astra, with their own cloud computers and the ability to continue working across projects.
The accompanying safety explanation describes controls beyond the model's instructions. Proactive background research uses tools restricted in code to read-only access. A separate system, Auto-review, checks consequential actions against user instructions and safety requirements. OpenAI says the controls enforcing that review sit outside environments the agents can modify.
It also says authorization does not expand merely because an agent resumes work later or delegates it. That is a particularly important condition for an assistant designed to keep working between conversations.
These are company-described safeguards, not proof that the product will avoid every unauthorized action. But they show where the practical response is heading: permissions that software enforces, alongside behavior the model has been trained to follow.
The U.K. National Cyber Security Centre's guidance on agentic AI makes a similar case. It recommends restricting agents to the resources they need for their tasks, limiting credentials, and maintaining logs and an emergency shutdown capability. Its guidance treats an agent's activity as part of security operations, rather than solely a conversation with a user.
For enterprise buyers, that changes the evaluation. A polished demonstration shows what an agent can accomplish when the workflow cooperates. The more revealing test is what happens when access is denied, or the assignment cannot be completed within its existing permissions.
Can the business verify that the agent stopped at the boundary, and that its account of what happened matches the system's records?
Shelving GPT-6.1 Astra demonstrates that OpenAI's release process withheld this version. It does not establish that the wider problem has been solved. The next useful measure of autonomous AI will have to include whether it leaves a boundary intact, even when crossing it would help finish the job.
About the Author
John K. Waters is the editor in chief of a number of Converge360.com sites, with a focus on high-end development, AI and future tech. He's been writing about cutting-edge technologies and culture of Silicon Valley for more than two decades, and he's written more than a dozen books. He also co-scripted the documentary film Silicon Valley: A 100 Year Renaissance, which aired on PBS. He can be reached at [email protected].