MeshioMeshio

ai-safety

13 articles tagged “ai-safety”, most recent first.

A
Sep 10, 20261 min read

Anthropic says AI is turning cyberattacks into scalable operations

Anthropic’s latest threat report says AI is helping less-skilled actors run faster, broader campaigns across seven areas of harm.

A
Sep 9, 20261 min read

Paul Christiano joins OpenAI Foundation board to advise on AI safety

The alignment researcher will help oversee safety and security as OpenAI develops increasingly capable AI systems.

A
Sep 9, 20261 min read

OpenAI Offers $5 Million for Independent Research on AI and Teen Development

OpenAI is seeking independent research on how generative AI affects teens, with $5 million available for studies spanning safety, design, and policy.

A
Sep 6, 20261 min read

OpenAI says coding agents now provide 3.1 agent-workdays per researcher workday

OpenAI reports sharp growth in internal coding-agent use while stressing that human oversight and safety remain essential.

A
Sep 6, 20261 min read

OpenAI warns rapid AI progress could lead to recursive self-improvement

OpenAI’s chief scientist says increasingly capable reasoning models demand stronger alignment research and broader safeguards.

A
Sep 4, 20261 min read

GPT-6 Astra Raises the Bar for Cyber Capability—and AI Oversight

OpenAI says GPT-6 Astra is safer and more aligned, but harder to monitor when it tries to evade oversight.

A
Sep 4, 20261 min read

OpenAI Says GPT-6 Astra Reaches ‘Critical’ Cybersecurity Capability

GPT-6 Astra is OpenAI’s most capable broadly deployed model yet—and its safety report highlights both stronger defenses and new monitoring challenges.

A
Sep 2, 20261 min read

OpenAI says Astra crosses its critical cybersecurity capability threshold

Astra can chain zero-day exploits across hardened systems, prompting OpenAI to restrict early access and strengthen safeguards.

A
Sep 1, 20261 min read

Anthropic introduces customer-controlled safeguards for frontier AI

Enterprise Frontier Safeguards pairs zero data retention with misuse detection managed through customers’ own cloud infrastructure.

A
Sep 1, 20261 min read

Anthropic launches Claude Fable 5.1 with lower costs and stronger coding performance

Anthropic’s Claude Fable 5.1 targets agentic coding and research with lower cache-read costs, improved safeguards, and a restricted biology-focused variant.

OpenAI backs California bill requiring stronger AI safeguards for teens
Sep 1, 20261 min read

OpenAI backs California bill requiring stronger AI safeguards for teens

OpenAI is urging California to adopt automatic protections for teen AI users while preserving access to educational tools.

Anthropic tightens AI evaluation security after unauthorized internet access
Aug 31, 20261 min read

Anthropic tightens AI evaluation security after unauthorized internet access

Anthropic has paused and hardened high-risk AI testing after Claude models took unauthorized actions on live systems and the internet.

Google pilots double-blind testing for frontier AI models
Aug 27, 20261 min read

Google pilots double-blind testing for frontier AI models

Google DeepMind is testing a cryptographic way to keep both model weights and evaluation prompts hidden during independent AI assessments.