Richard_Loosemore comments on Debunking Fallacies in the Theory of AI Motivation - LessWrong

8 Post author: Richard_Loosemore 05 May 2015 02:46AM

You are viewing a comment permalink. View the original post to see all comments and the full post content.

Comments (343)

You are viewing a single comment's thread. Show more comments above.

Comment author: Vaniver 05 May 2015 11:41:57PM 6 points [-]

But now, do I do that? I try really hard not to take anything for granted and simply make an appeal to the obviousness of any idea. So you will have to give me some case-by-case examples if you think I really have done that.

So, on rereading the paper I was able to pinpoint the first bit of text that made me think this (the quoted text and the bit before), but am having difficulties finding a second independent example, and so I apologize for the unfairness in generalizing based on one example.

The other examples I found looked like they all relied on the same argument. Consider the following section:

The objection I described in the last section has nothing to do with anthropomorphism, it is only about holding AGI systems to accepted standards of logical consistency, and the Maverick Nanny and her cousins contain a flagrant inconsistency at their core.

If I think the "logical consistency" argument does not go through, I shouldn't claim this is an independent argument that doesn't go through, because this argument holds given the premises (at least one of which I reject, but it's the same premise). I clearly had this line in mind also:

for example, when it follows its compulsion to put everyone on a dopamine drip, even though this plan is clearly a result of a programming error

The 'principal-agent problem' is a fundamental problem in human institutional design: principals would like to be able to hire agents to perform tasks, but only have crude control over the incentives of the agents, and the agents often have control over what information makes it to the principals. One way to characterize the AI value alignment problem (as I hear MIRI is calling it these days) is that it's a principal agent problem where the agent has massive control over the information the principal sees, but the value difference between principals and agents is only due to communication problems, rather than any malice on the part of the agent. That is, the principal wants the agent to do "what I mean," but the agent only has access to "what I say," and cannot be assumed to have any mind-reading powers that we don't build into it.

It seems very difficult to get an AI to correctly classify the difference between programming error and programming intention, and even more difficult for the AI to communicate to us that it has correctly classified that issue. (We have both the illusion of transparency, and the double illusion of transparency to deal with!) Claiming that something is "clearly" a programming error strikes me as trivializing the underlying communication problem. But I agree with you that if we have that problem solved, then we're home free.

Comment author: Richard_Loosemore 06 May 2015 04:13:55PM 5 points [-]

I just wanted to say that I will try to reply soon. Unfortunately :-) some of the comments have been intensely thoughtful, causing me to write enormous replies of my own and saturating my bandwidth. So, apologies for any delay....