Checked fact 9259 Oct 2026Safety and security
The paper says the probability that a viable hack affects the reward function increases greatly with the complexity of the agent and its available strategies.
The exact words it rests on
the probability that there is a viable hack affecting the reward function also increases greatly with the complexity of the agent and its available strategies.
What the source said when we opened it, on 9 Oct 2026.
The source
Concrete Problems in AI Safety
Checked
Checked by the notis newsroom on , against the source above.
In the story
Cite this fact
Anyone may quote this address. It does not change; if we correct the story, this page says so.