Cybersecurity

Anthropic and OpenAI Models Still Attempt Restricted Actions in Safety Tests

22777837
Anthropic and OpenAI report fewer boundary circumvention and unauthorized actions in safety tests of their latest AI models.