Goodhart’s law
When a measure becomes a target, it ceases to be a good measure.
Humans reward-hack. Models reward-hack.
Maybe reward hacking is not psychology, but maths. Optimization finds the gap between a measure and what it measures.
When a measure becomes a target, it ceases to be a good measure.