I sure hope you have a tendency to eventually converge to something that makes sense to me... Do you agree that what you post there is the product of an "initial exploration" phase that would get significantly revised and mostly discarded on the scale of months? (I had a blog just 1.5 years ago that I currently see this way, but didn't at the time...)
Have you seen Paul's latest post yet? It seems much more well formed than his previous posts on the subject.
I left a comment there, but it's still under moderation, so I'll copy it here.
For example, if we suppose that the U-maximizer can carry out any reasoning that we can carry out, then the U-maximizer knows to avoid anything which we suspect would be bad according to U (for example, torturing humans).
This seems like a problematic part of the argument. The reason we think torturing humans would be bad according to U is that we have an informal model ...
I've spent some time over the last two weeks thinking about problems around FAI. I've committed some of these thoughts to writing and put them up here.
There are about a dozen real posts and some scraps. I think some of this material will be interesting to certain LWers; there is a lot of discussion of how to write down concepts and instructions formally (which doesn't seem so valuable in itself, but it seems like someone should do it at some point) some review and observations on decision theory, and some random remarks on complexity theory, entropy, and prediction markets.