I haven't read any technical papers on goal-system stability; isn't it the case that real-world attempts at that are going to have at least as much of Problem Two as of Problem One about them? ("Internally" -- in the notion of what counts as self-improvement -- if not "externally" in whatever problem(s) the system is trying to solve.) I haven't thought (or read) enough about this for my opinion to have much weight; I could well be completely wrong about it.
Regardless, you're certainly right that Problem One is going to be important as well as Problem Two, and I should have said something like "AI safety is also an instance of Problem Two".
isn't it the case that real-world attempts at that are going to have at least as much of Problem Two as of Problem One about them? ("Internally" -- in the notion of what counts as self-improvement -- if not "externally" in whatever problem(s) the system is trying to solve.) I haven't thought (or read) enough about this for my opinion to have much weight; I could well be completely wrong about it.
Kind of. We expect intuitively that a reasoning system can reason about its own goals and successor-agents. Problem is, that actually requ...
At some point soon, I'm going to attempt to steelman the position of those who reject the AI risk thesis, to see if it can be made solid. Here, I'm just asking if people can link to the most convincing arguments they've found against AI risk.
EDIT: Thanks for all the contribution! Keep them coming...