OpenAI Slows Frontier Model Training Over Cybersecurity Risks
OpenAI has paused major training runs and tightened safeguards after signs that its next model may reach critical cyber capabilities.

OpenAI says it has temporarily slowed development of its latest frontier models after two developments raised the stakes around cybersecurity: an incident involving Hugging Face and preliminary evidence that its upcoming Astra model may meet the company’s “Critical cybersecurity capability” threshold.
The company paused reinforcement-learning training for two weeks and is keeping its largest planned frontier run on hold. In the meantime, OpenAI is using smaller training runs and evaluations to study model behavior, test safeguards, and gather stronger evidence that systems remain aligned before scaling resumes.
A higher security bar for model development
OpenAI says it has strengthened protections around research workloads that can execute code or use tools. New measures include tighter sandboxing for untrusted code, stronger network isolation, reduced standing privileges, removal of vulnerable shared services, and continuous automated testing against simulated attacks.
Astra and other cyber-related workloads now receive the strictest controls. Some training and evaluation work has resumed in upgraded environments, while other workloads remain paused until they meet the new requirements. OpenAI says safety and alignment projects are being migrated first.
The company is also expanding monitoring. Its new system checks model activity token by token, then escalates suspicious behavior to automated investigators that examine tool use, reasoning, and activity sequences. Potential concerns include unauthorized access, data theft, destructive actions, and attempts to bypass safeguards. Under the current process, teams are expected to pause activity if they cannot clear a high-priority alert within 30 minutes.
For AI developers, the announcement signals that scaling capability increasingly requires scaling infrastructure security, behavioral monitoring, and alignment evidence at the same time—not after training is complete.
Source: OpenAI News
Comments
Log in to join the discussion