Boxing an AI?

tailcalled

Boxing an AI is the idea that you can avoid the problems where an AI destroys the world by not giving it access to the world. For instance, you might give the AI access to the real world only through a chat terminal with a person, called the gatekeeper. This is should, theoretically prevent the AI from doing destructive stuff.

Eliezer has pointed out a problem with boxing AI: the AI might convince its gatekeeper to let it out. In order to prove this, he escaped from a simulated version of an AI box. Twice. That is somewhat unfortunate, because it means testing AI is a bit trickier.

However, I got an idea: why tell the AI it's in a box? Why not hook it up to a sufficiently advanced game, set up the correct reward channels and see what happens? Once you get the basics working, you can add more instances of the AI and see if they cooperate. This lets us adjust their morality until the AIs act sensibly. Then the AIs can't escape from the box because they don't know it's there.

At first glance, I was also skeptical of tailcalled's idea, but now I find I'm starting to warm up to it. Since you didn't ask for a practical proposal, just a concrete one, I give you this:

Implement an AI in Conway's Game of Life.
Don't interact with it in any way.
Limit the computational power the box has, so that if the AI begins engaging in recursive self-improvement, it'll run more and more slowly from our perspective, so we'll have ample time to shut it off. (Of course, from the AI's perspective, time will run as quickly as it always does, since the whole world will slow down with it.)
(optional) Create multiple human-level intelligences in the world (ignoring ethical constraints here), and see how the AI interacts with them. Run the simulation until you are reasonably certain (for a very stringent definition of "reasonably") from the AI's behavior that it is Friendly.
Profit.

The problem with this is that even if you can determine with certainty that an AI is friendly, there is no certainty that it will stay that way. There could be a series of errors as it goes about daily life, each acting as a mutation, serving to evolve the "Friendly" AI into a less friendly one

2Wes_W11y

Hm. That does sound more workable than I had thought.

0tailcalled11y

I would probably only include it as part of a batch of tests and proofs. It would be pretty foolish to rely on only one method to check if something that will destroy the world if it fails works correctly.

3

Boxing an AI?

3

3

3

Boxing an AI?

3

3