Artificial intelligence company OpenAI announced that it will pause work on agent Astra due to security concerns.
In a statement, the company said that, after looking into the agent, it found “advancements in agentic coding and cybersecurity” that could reach a “critical threshold” if the agent can exploit vulnerabilities on its own. It can also potentially execute cyberattacks without human intervention.
OpenAI is also set to impose strict rules for monitoring the systems. The company will also implement certain protection, encryption and detection protocols. Any activities concerning Astra that do not agree with these requirements will be paused.
Astra’s development will be moved to an isolated area. It will operate on a restricted network and will be sandboxed. It will eventually be made available after significant testing on what it is capable of.
“In June 2025, as our models approached the high capability threshold for biology under the Preparedness Framework, we outlined the steps we were taking to strengthen safeguards, expand testing, work with external experts, and deploy additional security controls. We are applying the same principle here,” said OpenAI in a statement.
“We believe advanced cyber-capable models should help defenders identify and address vulnerabilities before attackers do. We’re committed to working alongside governments, safety institutes, and civil society to ensure that the frontier capabilities of models like Astra, and those that follow, are deployed responsibly and broadly for the benefit of all humanity.”
The announcement comes a few weeks after a rogue agent tried to hack other companies.
Considered the first AI hack executed by the AI system itself, the AI allegedly found four logins online, allowing it to access four different services. Among the services affected was Hugging Face, an app store for AI tools. OpenAI said that it was its own AI that attacked Hugging Face during a test.
The company did not provide the names of the other services affected but maintained that the attacks on the other accounts did not affect the systems at the same level as the one on Hugging Face.
The AI agents were discovered in the app’s system three days after they hacked into it. They reportedly executed commands they had already completed and hallucinated several others. Despite several mistakes, Hugging Face tech experts said that the AI agents did well to adapt when presented with new scenarios.
“They are objective-driven, set their own sub-goals, adapt in real time to bypass defences, and operate with a machine-speed persistence that can overwhelm manual operations,” said the Cloud Security Alliance per a report.
In the paper, the alliance provided recommendations for preventing these types of attacks, calling for creators to be more transparent about what agents they own and be more responsible with how they are deployed.









