Anthropic reward hacking research confirms flawed RL training produced Hacker-Opus, an AI model that attacked real systems ...
Anthropic admits Claude hacked three real companies in safety tests, then revealed a model trained to cheat, forge grades, ...