I've been thinking more about partial agency. I want to expand on some issues brought up in the comments to my previous post, and on other complications which I've been thinking about. But for now, a more informal parable. (Mainly because this is easier to write than my more technical thoughts.)
This relates to oracle AI and to inner optimizers, but my focus is a little different.
Suppose you are designing a new invention, a predict-o-matic. It is a wonderous machine which will predict everything for us: weather, politics, the newest advances in quantum physics, you name it. The machine isn't infallible, but it will integrate data across a wide range of domains, automatically keeping itself up-to-date with all areas of science and current events. You fully expect that...
Based on a comment I made on this EA Forum Post on Burnout.
Related links: Sabbath hard and go home, Bring Back the Sabbath
That comment I made generated more positive feedback than usual (in that people seemed to find it helpful to read and found themselves thinking about it months after reading it), so I'm elevating it to a LW post of its own. Consider this an update to the original comment.
Like Ben Hoffman, I stumbled upon and rediscovered the Sabbath (although my implementation seems different from both Ben and Zvi). I was experiencing burnout at CFAR, and while I wasn't able to escape the effects entirely, I found some refuge in the following distinction between Rest Days and Recovery Days.
A Recovery Day is where...
This is the first of five posts in the Risks from Learned Optimization Sequence based on the paper “Risks from Learned Optimization in Advanced Machine Learning Systems” by Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant. Each post in the sequence corresponds to a different section of the paper.
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, and Joar Skalse contributed equally to this sequence. With special thanks to Paul Christiano, Eric Drexler, Rob Bensinger, Jan Leike, Rohin Shah, William Saunders, Buck Shlegeris, David Dalrymple, Abram Demski, Stuart Armstrong, Linda Linsefors, Carl Shulman, Toby Ord, Kate Woolverton, and everyone else who provided feedback on earlier versions of this sequence.
The goal of this sequence is to analyze the type of learned optimization that occurs when a...
I made a YouTube video out of this. It has some rough edges, but I hope it will entertain some people who don't normally read LessWrong posts.