It's not necessarily that the AI would have difficulty understanding what "do what humans mean" means, even before being told to do what humans mean.
It just has no reason to obey "do what humans mean" unless we program it to do what humans mean.
"Do what humans mean" is telling the AI to do something that we can currently only specify vaguely. "Figure out what we intend by "do what humans mean", and then do that" is also vaguely specified. It doesn't solve the problem.
It just has no reason to obey "do what humans mean" unless we program it to do what humans mean.
I'm not disputing that this is also a problem, indeed perhaps a harder problem than figuring out what humans mean. In fact there are many failure modes, I was just wondering why people seem to focus in on specifically the fickle genie failure mode to the exclusion of others.
If it's worth saying, but not worth its own post, then it goes here.
Notes for future OT posters:
1. Please add the 'open_thread' tag.
2. Check if there is an active Open Thread before posting a new one. (Immediately before; refresh the list-of-threads page before posting.)
3. Open Threads should start on Monday, and end on Sunday.
4. Unflag the two options "Notify me of new top level comments on this article" and "