it follows its programming, not its knowledge about what its programming is "meant" to be (unless we've successfully programmed in "do what I mean", which is basically the whole of the challenge).
Not necessarily. The instructions to a fully-reflective AI could be more along the lines of “learn what I mean, then do that” or “do what I asked within the constraints of my own unstated principles.” The AI would have an imperative to build a more accurate internal model of your psychology in order to predict the implicit constraints applied to that request, typically by asking you or other trusted humans questions. If you want to take this to a crazy extreme, it is perhaps more probable that the AI would recognize that military campaigns to acquire ore deposits is both outside its recorded experiences and not directly implied by your request. It would then take the prudent step of constructing a passive brain scanner (perhaps developing molecular nanotechnology first, in order to do so), clandestinely scan you while on vacation, and use that knowledge to refine the utility function into something you would be happy with (i.e. not declaring war on humanity).
“learn what I mean, then do that” or “do what I asked within the constraints of my own unstated principles.”
That's just another way of saying "do what I mean". And it doesn't give us the code to implement that.
"Do what I asked within the constraints of my own unstated principles" is a hugely complicated set of instructions, that only seem simple because it's written in English words.
A stub on a point that's come up recently.
If I owned a paperclip factory, and casually told my foreman to improve efficiency while I'm away, and he planned a takeover of the country, aiming to devote its entire economy to paperclip manufacturing (apart from the armament factories he needed to invade neighbouring countries and steal their iron mines)... then I'd conclude that my foreman was an idiot (or being wilfully idiotic). He obviously had no idea what I meant. And if he misunderstood me so egregiously, he's certainly not a threat: he's unlikely to reason his way out of a paper bag, let alone to any position of power.
If I owned a paperclip factory, and casually programmed my superintelligent AI to improve efficiency while I'm away, and it planned a takeover of the country... then I can't conclude that the AI is an idiot. It is following its programming. Unlike a human that behaved the same way, it probably knows exactly what I meant to program in. It just doesn't care: it follows its programming, not its knowledge about what its programming is "meant" to be (unless we've successfully programmed in "do what I mean", which is basically the whole of the challenge). We can't therefore conclude that it's incompetent, unable to understand human reasoning, or likely to fail.
We can't reason by analogy with humans. When AIs behave like idiot savants with respect to their motivations, we can't deduce that they're idiots.