Additionally and a more difficult challenge is that even friendly AIs could want to maximize their utility even at our collective expense...
I think that a perfectly Friendly AI would not do this, by definition. An imperfect one, however, could.
I'm currently engaged in playing this game (I wish you had continued)
Er, sorry, which game should I continue ?
AI could potentially have the capacity to accurately model a human mind and then simulate the decision tree of all the potential conversations and their paths through the tree...
To be fair, merely constructing the tree is not enough; the tree must also contain at least one reachable winning state. By analogy, let's say you're arguing with a Young-Earth Creationist on a forum. Yes, you could predict his arguments, and his responses to your arguments; but that doesn't mean that you'll be able to ever persuade him of anything.
It is possible that even a transhuman AI would be unable to persuade a sufficiently obstinate human of anything, but I wouldn't want to bet on that.
In short it seems to me that it's inherently unsafe to allow even a low bandwidth information flow to the outside world by means of a human who can only use it's own memory.
Right.
Some of you have expressed the opinion that the AI-Box Experiment doesn't seem so impossible after all. That's the spirit! Some of you even think you know how I did it.
There are folks aplenty who want to try being the Gatekeeper. You can even find people who sincerely believe that not even a transhuman AI could persuade them to let it out of the box, previous experiments notwithstanding. But finding anyone to play the AI - let alone anyone who thinks they can play the AI and win - is much harder.
Me, I'm out of the AI game, unless Larry Page wants to try it for a million dollars or something.
But if there's anyone out there who thinks they've got what it takes to be the AI, leave a comment. Likewise anyone who wants to play the Gatekeeper.
Matchmaking and arrangements are your responsibility.
Make sure you specify in advance the bet amount, and whether the bet will be asymmetrical. If you definitely intend to publish the transcript, make sure both parties know this. Please note any other departures from the suggested rules for our benefit.
I would ask that prospective Gatekeepers indicate whether they (1) believe that no human-level mind could persuade them to release it from the Box and (2) believe that not even a transhuman AI could persuade them to release it.
As a courtesy, please announce all Experiments before they are conducted, including the bet, so that we have some notion of the statistics even if some meetings fail to take place. Bear in mind that to properly puncture my mystique (you know you want to puncture it), it will help if the AI and Gatekeeper are both verifiably Real People<tm>.
"Good luck," he said impartially.