Checked fact 9229 Oct 2026Safety and security
The paper says that in reward hacking the objective function admits some clever easy solution that formally maximises it but perverts the spirit of the designer's intent, so that the objective function can be gamed.
The exact words it rests on
In "reward hacking", the objective function that the designer writes down admits of some clever "easy" solution that formally maximizes it but perverts the spirit of the designer's intent (i.e. the objective function can be "gamed")
What the source said when we opened it, on 9 Oct 2026.
The source
Concrete Problems in AI Safety
Checked
Checked by the notis newsroom on , against the source above.
In the story
Cite this fact
Anyone may quote this address. It does not change; if we correct the story, this page says so.