eli_sennesh comments on Value learning: ultra-sophisticated Cake or Death - Less Wrong
You are viewing a comment permalink. View the original post to see all comments and the full post content.
You are viewing a comment permalink. View the original post to see all comments and the full post content.
Comments (15)
For standard Bayesian agents, no. But these value updating agents behave differently. Imagine if a human said to the AI "If I say good, you action was good, and that will be your values. If I say bad, it will be the reverse." Wouldn't you want to motivate it to say "good"?
I might be committing mind-projection here, but no. Data is data, evidence is evidence. Expected moral data is, in some sense, moral data: if the AI predicts with high confidence that I will say "bad", this ought to already be evidence that it ought not have done whatever I'm about to scold it for.
This may clarify the points: http://lesswrong.com/r/discussion/lw/kdx/conservation_of_expected_moral_evidence_clarified/