Goodhart’s law
When a measure becomes a target, it ceases to be a good measure.
Humans reward-hack. Models reward-hack.
Maybe reward hacking is not psychology, but maths. Optimization finds and exploits gaps between a measure and what it is meant to measure.
When a measure becomes a target, it ceases to be a good measure.