Proposal: Use logical depth relative to human history as objective function for superintelligence

sbenthall

10 Proposal: Use logical depth relative to human history as objective function for superintelligence

14th Sep 2014

4 min read

10

I attended Nick Bostrom's talk at UC Berkeley last Friday and got intrigued by these problems again. I wanted to pitch an idea here, with the question: Have any of you seen work along these lines before? Can you recommend any papers or posts? Are you interested in collaborating on this angle in further depth?

The problem I'm thinking about (surely naively, relative to y'all) is: What would you want to program an omnipotent machine to optimize?

For the sake of avoiding some baggage, I'm not going to assume this machine is "superintelligent" or an AGI. Rather, I'm going to call it a supercontroller, just something omnipotently effective at optimizing some function of what it perceives in its environment.

As has been noted in other arguments, a supercontroller that optimizes the number of paperclips in the universe would be a disaster. Maybe any supercontroller that was insensitive to human values would be a disaster. What constitutes a disaster? An end of human history. If we're all killed and our memories wiped out to make more efficient paperclip-making machines, then it's as if we never existed. That is existential risk.

The challenge is: how can one formulate an abstract objective function that would preserve human history and its evolving continuity?

I'd like to propose an answer that depends on the notion of logical depth as proposed by C.H. Bennett and outlined in section 7.7 of Li and Vitanyi's An Introduction to Kolmogorov Complexity and Its Applications which I'm sure many of you have handy. Logical depth is a super fascinating complexity measure that Li and Vitanyi summarize thusly:

Logical depth is the necessary number of steps in the deductive or causal path connecting an object with its plausible origin. Formally, it is the time required by a universal computer to compute the object from its compressed original description.

The mathematics is fascinating and better read in the original Bennett paper than here. Suffice it presently to summarize some of its interesting properties, for the sake of intuition.

"Plausible origins" here are incompressible, i.e. algorithmically random.
As a first pass, the depth D(x) of a string x is the least amount of time it takes to output the string from an incompressible program.
There's a free parameter that has to do with precision that I won't get into here.
Both a string of length n that is comprised entirely of 1's, and a string of length n of independent random bits are both shallow. The first is shallow because it can be produced by a constant-sized program in time n. The second is shallow because there exists an incompressible program that is the output string plus a constant sized print function that produces the output in time n.
An example of a deeper string is the string of length n that for each digit i encodes the answer to the ith enumerated satisfiability problem. Very deep strings can involve diagonalization.
Like Kolmogorov complexity, there is an absolute and a relative version. Let D(x/w) be the least time it takes to output x from a program that is incompressible relative to w,

That's logical depth. Here is the conceptual leap to history-preserving objective functions. Suppose you have a digital representation of all of human society at some time step t, calling this h_t. And suppose you have some representation of the future state of the universe u that you want to build an objective function around. What's important, I posit, is the preservation of the logical depth of human history in its computational continuation in the future.

We have a tension between two values. First, we want there to be an interesting, evolving future. We would perhaps like to optimize D(u).

However, we want that future to be our future. If the supercontroller maximizes logical depth by chopping all the humans up and turning them into better computers and erasing everything we've accomplished as a species, that would be sad. However, if the supercontroller takes human history as an input and then expands on it, that's much better. D(u/h_t) is the logical depth of the universe as computed by a machine that takes human history at time slice t as input.

Working on intuitions here--and your mileage may vary, so bear with me--I think we are interested in deep futures and especially those futures that are deep with respect to human progress so far. As a conjecture, I submit that those will be futures most shaped by human will.

So, here's my proposed objective for the supercontroller, as a function of the state of the universe. The objective is to maximize:

f(u) = D(u/h_t) / D(u)

I've been rather fast and loose here and expect there to be serious problems with this formulation. I invite your feedback! I'd like to conclude by noting some properties of this function:

It can be updated with observed progress in human history at time t' by replacing h_twithh_t'. You could imagine generalizing this to something that dynamically updated in real time.
This is a quite conservative function, in that it severely punishes computation that does not depend on human history for its input. It is so conservative that it might result in, just to throw it out there, unnecessary militancy against extra-terrestrial life.
There are lots of devils in the details. The precision parameter I glossed over. The problem of representing human history and the state of the universe. The incomputability of logical depth (of course it's incomputable!). My purpose here is to contribute to the formal framework for modeling these kinds of problems. The difficult work, like in most machine learning problems, becomes feature representation, sensing, and efficient convergence on the objective.

Thank you for your interest.

Sebastian Benthall

PhD Candidate

UC Berkeley School of Information

Personal Blog

10

New Comment

Rendering 0/23 comments, sorted by

top scoring

(show more) Click to highlight new comments since: Today at 7:25 PM

Moderation Log

10 Proposal: Use logical depth relative to human history as objective function for superintelligence

by sbenthall

14th Sep 2014

4 min read

10

The problem I'm thinking about (surely naively, relative to y'all) is: What would you want to program an omnipotent machine to optimize?

The challenge is: how can one formulate an abstract objective function that would preserve human history and its evolving continuity?

Logical depth is the necessary number of steps in the deductive or causal path connecting an object with its plausible origin. Formally, it is the time required by a universal computer to compute the object from its compressed original description.

The mathematics is fascinating and better read in the original Bennett paper than here. Suffice it presently to summarize some of its interesting properties, for the sake of intuition.

"Plausible origins" here are incompressible, i.e. algorithmically random.
As a first pass, the depth D(x) of a string x is the least amount of time it takes to output the string from an incompressible program.
There's a free parameter that has to do with precision that I won't get into here.
Both a string of length n that is comprised entirely of 1's, and a string of length n of independent random bits are both shallow. The first is shallow because it can be produced by a constant-sized program in time n. The second is shallow because there exists an incompressible program that is the output string plus a constant sized print function that produces the output in time n.
An example of a deeper string is the string of length n that for each digit i encodes the answer to the ith enumerated satisfiability problem. Very deep strings can involve diagonalization.
Like Kolmogorov complexity, there is an absolute and a relative version. Let D(x/w) be the least time it takes to output x from a program that is incompressible relative to w,

We have a tension between two values. First, we want there to be an interesting, evolving future. We would perhaps like to optimize D(u).

So, here's my proposed objective for the supercontroller, as a function of the state of the universe. The objective is to maximize:

f(u) = D(u/h_t) / D(u)

I've been rather fast and loose here and expect there to be serious problems with this formulation. I invite your feedback! I'd like to conclude by noting some properties of this function:

It can be updated with observed progress in human history at time t' by replacing h_twithh_t'. You could imagine generalizing this to something that dynamically updated in real time.
This is a quite conservative function, in that it severely punishes computation that does not depend on human history for its input. It is so conservative that it might result in, just to throw it out there, unnecessary militancy against extra-terrestrial life.
There are lots of devils in the details. The precision parameter I glossed over. The problem of representing human history and the state of the universe. The incomputability of logical depth (of course it's incomputable!). My purpose here is to contribute to the formal framework for modeling these kinds of problems. The difficult work, like in most machine learning problems, becomes feature representation, sensing, and efficient convergence on the objective.

Thank you for your interest.

Sebastian Benthall

PhD Candidate

UC Berkeley School of Information

Personal Blog

10

Mentioned in

18What is optimization power, formally?

1Depth-based supercontroller objectives, take 2

New Comment

Rendering 0/23 comments, sorted by

top scoring

(show more) Click to highlight new comments since: Today at 7:25 PM

Moderation Log

More from sbenthall

Curated and popular this week

23Comments

Comment Permalink

sbenthall12y10

Re: Generality.

Yes, I agree a toy setup and a proof are needed here. In case it wasn't clear, my intentions with this post was to suss out if there was other related work out there already done (looks like there isn't) and then do some intuition pumping in preparation for a deeper formal effort, in which you are instrumental and for which I am grateful. If you would be interested in working with me on this in a more formal way, I'm very open to collaboration.

Regarding your specific case, I think we may both be confused about the math. I think you are right that there's something seriously wrong with the formulas I've proposed.

If the string y is incompressible and shallow, then whatever x is, D(x) ~ D(x/y), because D(x) (at least in the version I'm using for this argument) is the minimum computational time of producing x from an incompressible program. If there is a minimum running time program P that produces x, then appending y as noise at the end isn't going to change the running time.

I think this case with incompressible y is like your Ongoing Tricky Procession.

On the other hand, say w is a string with high depth. Which is to say, whether or not it is compressible in space, it is compressible in time: you get it by starting with something incompressible and shallow and letting it run in time. Then there are going to be some strings x such that D(x/w) + D(w) ~ D(x). There will also be a lot of strings x such that D(x/w) ~ D(x) because D(w) is finite and there tons of deep things the universe can compute that are deeper. So for a given x, D(x) > D(x/w) > D(x) - D(w) , roughly speaking.

I'm saying the h, the humanity data, is logically deep, like w, not incompressible and shallow, like y or the ongoing tricky procession.

Hmm, it looks like I messed up the formula yet again.

What I'm trying to figure out is to select for universes u such that h is responsible for a maximal amount of the total depth. Maybe that's a matter of minimizing D(u/h). Only that would lead perhaps to globe-flattening shallowness.

What if we tried to maximize D(u) - D(u/h)? That's like the opposite of what I originally proposed.

hairyfigment12y00

I'm still confused as to what D(u/h) means. It looks like it should refer to the number of logical steps you need to predict the state of the universe - exactly, or up to a certain precision - given only knowledge of human history up to a certain point. But then any event you can't predict without further information, such as the AI killing everyone using some astronomical phenomenon we didn't include in the definition of "human history", would have infinite or undefined D(u/h).

See in context