OpenAI was founded as a non-profit with a clear commitment to avoiding AI dangers and a unique structure to prevent investor control or excessive pressure from profit motives. This failed. I'm considering launching a commercial venture (possibly named "Prealign") to solve value generalisation and contribute to AI alignment, while avoiding...
So. I previously demonstrated an anthropic impossibility theorem, showing that in Duplicates Sleeping Beauty, there was no possible probability theory that obeyed both the martingale condition and "simple Bayes" in non-anthropic situations. This post will clarify and simplify the result, replacing the martingale with the law of total probability. Let's...
tldr: 1. SIA breaks in many infinite worlds. 2. However, a general insight of what SIA does is that it refuses to pay the Bayes cost for self-location information. It is instead pre-reimbursed for future Bayes costs. 3. For n uniform agent possibilities, the Bayes cost is 1/n and the...
tl;dr 1. There is an impossibility result in anthropic probability: no reasonable probability theory can stay consistent across a duplication event. 2. In particular, it must violate either Bayesian updates-from priors in non-anthropic situations, or the martingale condition – today’s probabilities are expectations of tomorrow’s probabilities. In this post, I’ll...
tl;dr: in situations without duplicates: 1. SIA and standard Bayesian updating are the same thing. 2. Standard Bayesian updating (and hence SIA) is correct. 3. SIA doesn’t have the counter-intuitive conclusions most people think it does. 4. Many of the counter-intuitive parts of anthropic reasoning come exclusively from the choice...
When we get explicit strong generalisation to work (see the first post on the matter and the second) my dream would be to create pre-aligned generalising AIs. Think about the usual conflict between alignment and capabilities, between doing the right thing and doing the easy thing. The standard narrative puts...
A human superpower hidden from even ourselves I though GPT 3.5 was on the verge of Artificial General Intelligence (AGI). It certainly seemed that way – it could combine and extend ideas in ways that were far beyond narrow rigid computing. Sure, it had some flaws, but with its general...