it includes peoples' death-to-outgroups volitions unmodified [..] whereas CEV (which came first) doesn't
Is there a pointer available to the evidence that an "extrapolation" process a la CEV actually addresses this problem? (Or, if practical, can it be summarized here?)
I've read some but not all of the CEV literature, and I understand that this process intended to solve this problem, but I haven't been able to grasp from that how we know it actually does.
It seems to depend on the idea that if we had world enough and time, we would outgrow things like "death-to-outgroups," and therefore a sufficiently intelligent seed AI tasked with extrapolating what we would want given world enough and time will naturally come up with a CEV that doesn't include such things... perhaps because such things are necessarily instrumental values rather than reflectively stable terminal values, perhaps for other reasons.
But surely there has to be more to it than that, as the "world enough and time" theory seems itself unjustified.
Is there a pointer available to the evidence that an "extrapolation" process a la CEV actually addresses this problem?
I think there's some uncertainty about that, actually. The extrapolation procedure is never really specified in CEV, and I could imagine some extrapolation procedures which probably do eliminate the death-to-outgroups volition, and some extrapolation procedures which don't. So an actual implementation would have a lot of details to fill in, and there are ways of filling in those details which would be bad (but this is true of e...
Link: adarti.blogspot.com/2011/04/review-of-proposals-toward-safe-ai.html