Suggestions for possible general rule:
A: Simulate an argument between the individual at State 1 and State 7. If the individual at State 1 is ultimately convinced, then State 7 is CEV whatever the real State 1 individual thinks. If the individual at State 1 is ultimately unconvinced, it isn't.
If, say, the individual at State is convinced by State 4's values but not by State 7's values (arbitrary choices), then it is extrapolated CEV up to the point where the individual at State 1 would cease to be convinced by the argument even seeing the logical connections.
B: Simulate an argument between the individual at State 1 and the individual at State 7, under the assumption that the individual at State 1 and the individual at State 7 both perfectly follow their own rules for proper argument incorporating appropriate amount of emotion and rationality (by their subjective standards) and getting rid of what they consider to be undue biases. Same rule for further interpretation.
Simulate an argument between the individual at State 1 and State 7. If the individual at State 1 is ultimately convinced
Does this include "convinced by hypnosis", "convinced by brainwashing", "convinced by a clever manipulation" etc.? How will AI tell the difference?
(Maybe "convincing by hypnosis" is considered a standard and ethical method of communication with lesser beings in the society of Stage 7. If a person A is provably more intelligent and rational than a person B, and a person A acts according to general...
My main objection to Coherent Extrapolated Volition (CEV) is the "Extrapolated" part. I don't see any reason to trust the extrapolated volition of humanity - but this isn't just for self centred reasons. I don't see any reason to trust my own extrapolated volition. I think it's perfectly possible that my extrapolated volition would follow some scenario like this:
There are many other ways this could go, maybe ending up as a negative utilitarian or completely indifferent, but that's enough to give the flavour. You might trust the person you want to be, to do the right things. But you can't trust them to want to be the right person - especially several levels in (compare with the argument in this post, and my very old chaining god idea). I'm not claiming that such a value drift is inevitable, just that it's possible - and so I'd want my initial values to dominate when there is a large conflict.
Nor do I give Armstrong 7's values any credit for having originated from mine. Under torture, I'm pretty sure I could be made to accept any system of values whatsoever; there are other ways that would provably alter my values, so I don't see any reason to privilege Armstrong 7's values in this way.
"But," says the objecting strawman, "this is completely different! Armstrong 7's values are the ones that you would reach by following the path you would want to follow anyway! That's where you would get to, if you started out wanting to be more altruistic, had control over you own motivational structure, and grew and learnt and knew more!"
"Thanks for pointing that out," I respond, "now that I know where that ends up, I must make sure to change the path I would want to follow! I'm not sure whether I shouldn't be more altruistic, or avoid touching my motivational structure, or not want to grow or learn or know more. Those all sound pretty good, but if they end up at Armstrong 7, something's going to have to give."