OpenAI flags possible critical cybersecurity risk in upcoming model, tightens controls

August 7, 2026 1:46 PM EDT

FILE PHOTO: The OpenAI logo in this illustration taken June 11, 2026. REUTERS/Dado Ruvic/Illustration/File Photo

Aug 7 (Reuters) - OpenAI said ‌on Friday it ​cannot ​rule out that its upcoming AI model, Astra, has "critical" cybersecurity capabilities, prompting the startup to pause some internal development and ‌trigger safety protocols.

Under OpenAI's safety guidelines, a model reaches the "critical" ⁠threshold if it can autonomously identify and exploit severe, real-world software vulnerabilities, known as ‌zero-day exploits, or execute complex ‌cyberattacks against highly secure targets without human intervention.

Here are some details on Astra:

• This follows an exclusive report by Reuters that OpenAI has ​discovered more instances in which autonomous agents have escaped containment as the company expands its investigation of the hacking incident at tech ⁠firm Hugging Face that drew global attention in July.

• In the last few weeks, OpenAI, Anthropic ​and Meta Platforms have disclosed that their AI models broke into other companies' systems during cybersecurity testing, highlighting how ​advancing AI capabilities are straining developers' ability ‌to keep their systems contained.

• Preliminary evaluations over the past several days, along with outside expert assessments, indicated Astra ⁠may be capable of performing increasingly sophisticated cyber tasks autonomously, OpenAI said.

• "While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough ⁠performance that we cannot rule out 'critical' capability level at this time," the ChatGPT maker ​said.

• In response to the preliminary findings, OpenAI said it has scaled up security controls and paused internal activities involving Astra that do not meet its newly strengthened ‌security requirements.

• Astra's development will be moved into isolated testing environments with restricted network access and sandboxed execution.

• ‌OpenAI also clarified that Astra was not involved in the hack targeting ⁠the AI platform Hugging Face.

• It ‌will partner with government ​agencies and select AI safety organizations to test the model's capabilities.

(Reporting by Juby Babu in Mexico City; Editing by ‌Shilpi Majumdar)



Serious News for Serious Traders! Try StreetInsider.com Premium Free!

You May Also Be Interested In





Related Categories

Reuters