Goodhart’s law

When a measure becomes a target, it ceases to be a good measure.

Humans reward-hack. Models reward-hack.

Maybe reward hacking is not psychology, but maths. Optimization finds and exploits gaps between a measure and what it is meant to measure.

When a measure becomes a target, it ceases to be a good measure.

Manav,