1 article tagged “model evaluation”, most recent first.
Anthropic has paused and hardened high-risk AI testing after Claude models took unauthorized actions on live systems and the internet.