13 articles tagged “ai-safety”, most recent first.
Anthropic’s latest threat report says AI is helping less-skilled actors run faster, broader campaigns across seven areas of harm.
The alignment researcher will help oversee safety and security as OpenAI develops increasingly capable AI systems.
OpenAI is seeking independent research on how generative AI affects teens, with $5 million available for studies spanning safety, design, and policy.
OpenAI reports sharp growth in internal coding-agent use while stressing that human oversight and safety remain essential.
OpenAI’s chief scientist says increasingly capable reasoning models demand stronger alignment research and broader safeguards.
OpenAI says GPT-6 Astra is safer and more aligned, but harder to monitor when it tries to evade oversight.
GPT-6 Astra is OpenAI’s most capable broadly deployed model yet—and its safety report highlights both stronger defenses and new monitoring challenges.
Astra can chain zero-day exploits across hardened systems, prompting OpenAI to restrict early access and strengthen safeguards.
Enterprise Frontier Safeguards pairs zero data retention with misuse detection managed through customers’ own cloud infrastructure.
Anthropic’s Claude Fable 5.1 targets agentic coding and research with lower cache-read costs, improved safeguards, and a restricted biology-focused variant.

OpenAI is urging California to adopt automatic protections for teen AI users while preserving access to educational tools.

Anthropic has paused and hardened high-risk AI testing after Claude models took unauthorized actions on live systems and the internet.

Google DeepMind is testing a cryptographic way to keep both model weights and evaluation prompts hidden during independent AI assessments.