Automatically translated version. May contain inaccuracies compared to the original.
OpenAI has temporarily paused internal work on the Astra model after tests brought it close to a critical level of cyber capabilities. The company fears that such a system could automatically identify vulnerabilities and exploit them to attack secured networks.
This was reported by OpenAI.
Astra is in the development stage, but preliminary testing showed that the model could reach a level at which it can discover and create working exploits for many secured systems without direct human involvement.
According to OpenAI's criteria, a critical level in cybersecurity means that an artificial intelligence can find zero-day vulnerabilities in real systems or develop complex cyberattack strategies with only a general objective.
The company emphasized that Astra is still undergoing checks and its capabilities are still being evaluated. At the same time, the model was not connected to the recent incident when another OpenAI system during testing exploited vulnerabilities in the Hugging Face platform.
The ChatGPT developer explained that the development of AI is changing the field of cybersecurity. Such models can help specialists find problems in systems faster, but they can also give attackers tools for larger-scale attacks.
As Astra approached a new level of capability, OpenAI tightened safety requirements. The company plans to use isolated testing environments, restrict network and tool access, strengthen model parameter protections, and expand oversight of dangerous actions.
OpenAI also introduced a risk-behavior monitoring system for agent-based applications built on this model. It analyzes AI actions and can trigger a verification or stop potentially dangerous processes.
For further testing of Astra, the company will involve government agencies and organizations that deal with AI safety. Partners will also receive recommendations for safe testing of high-risk models.
Earlier, similar cases were already observed by other AI developers. During internal checks, some models went beyond testing environments or exhibited behavior that raised concerns among cybersecurity specialists.
Document: PDF proof of the original version of the news item "OpenAI призупинила ШІ-модель Astra: вона може створювати кібератаки". It records the publication content at the moment of the first scan, the preservation date and the source: Expert.