This post argues that fixed-weight models (at least as we understand them today) will a) always be vulnerable to adversarial examples in their concept-spaces, and b) hence will be misaligned, under sufficient optimisation pressure. Boundaries in concept space To serve any purpose whatsoever, an AI will have to draw boundaries...
In the previous post, I presented my theory of change for why value generalisation is vital for AI alignment. Here I'll add the practical part of the argument: given those facts, why explicitly try to do value generalisation, what are the dangers of the approach, how should it be done,...
This is the first part of a theory of change explaining why I'm targeting value generalisation as the path to AI alignment. It presents the definitions and key claims behind the theory, and shows that most AI alignment failure modes are value generalisation failures. Note that a lot of these...
OpenAI was founded as a non-profit with a clear commitment to avoiding AI dangers and a unique structure to prevent investor control or excessive pressure from profit motives. This failed. I'm considering launching a commercial venture (possibly named "Prealign") to solve value generalisation and contribute to AI alignment, while avoiding...
So. I previously demonstrated an anthropic impossibility theorem, showing that in Duplicates Sleeping Beauty, there was no possible probability theory that obeyed both the martingale condition and "simple Bayes" in non-anthropic situations. This post will clarify and simplify the result, replacing the martingale with the law of total probability. Let's...
tldr: 1. SIA breaks in many infinite worlds. 2. However, a general insight of what SIA does is that it refuses to pay the Bayes cost for self-location information. It is instead pre-reimbursed for future Bayes costs. 3. For n uniform agent possibilities, the Bayes cost is 1/n and the...
tl;dr 1. There is an impossibility result in anthropic probability: no reasonable probability theory can stay consistent across a duplication event. 2. In particular, it must violate either Bayesian updates-from priors in non-anthropic situations, or the martingale condition – today’s probabilities are expectations of tomorrow’s probabilities. In this post, I’ll...