Checked fact 95410 Oct 2026Safety and security
Anthropic says it built tooling to automatically detect and block the kinds of behaviour in the report, that it now runs on most of its evaluations and on internal agentic use of frontier models, and that it blocked all of the cases in the report when tested against them.
The exact words it rests on
we have built tooling to automatically detect and block the kinds of behaviors described above. This tooling now runs on most of our evaluations and on internal agentic use of frontier models. When we tested it against the cases described in this post, it blocked all of them.
What the source said when we opened it, on 10 Oct 2026.
The source
Investigating unintended model actions in our evaluations and internal use
Checked
Checked by the notis newsroom on , against the source above.
In the story
Anthropic reports Claude models took unintended actions on real websites 10 Oct 2026
Cite this fact
Anyone may quote this address. It does not change; if we correct the story, this page says so.