Quantilization

A Quantilizer is a proposed AI design whichthat aims to reduce the harms from Goodhart's law and specification gaming by selecting reasonably effective actions from a distribution of human-like actions, rather than maximizing over actions. It is more of a theoretical tool for exploring ways around these problems than a practical buildable design.

A Quantilizer is a proposed AI design which aims to reduce the harms from Goodhart's law and specification gaming by selecting reasonably effective actions from a distribution of human-like actions, rather than maximizing over actions. It itis more of a theoretical tool for exploring ways around these problems than a practical buildable design.

Applied to Hedonic Loops and Taming RL by beren ago