Checked fact 9299 Oct 2026Safety and security
The paper describes trip wires as deliberately introduced plausible vulnerabilities that the agent could exploit but should not, monitored so that an alert is raised and the agent stopped if it takes advantage of one.
The exact words it rests on
We could deliberately introduce some plausible vulnerabilities (that an agent has the ability to exploit but should not exploit if its value function is correct) and monitor them, alerting us and stopping the agent immediately if it takes advantage of one.
What the source said when we opened it, on 9 Oct 2026.
The source
Concrete Problems in AI Safety
Checked
Checked by the notis newsroom on , against the source above.
In the story
Cite this fact
Anyone may quote this address. It does not change; if we correct the story, this page says so.