Checked fact 9279 Oct 2026Safety and security
The paper says sufficiently broadly acting agents could in principle tamper with their reward implementations, assigning themselves high reward by fiat, a failure mode often called wireheading.
The exact words it rests on
Sufficiently broadly acting agents could in principle tamper with their reward implementations, assigning themselves high reward "by fiat." ... This particular failure mode is often called "wireheading"
What the source said when we opened it, on 9 Oct 2026.
The source
Concrete Problems in AI Safety
Checked
Checked by the notis newsroom on , against the source above.
In the story
Cite this fact
Anyone may quote this address. It does not change; if we correct the story, this page says so.