Good post, but I am curious about your thinking for :
> when you expect that you’re going to run into the same people over and over again for a long time to come, it’s wise to be consistently trustworthy and helpful even if you could get an advantage in the short run by betraying someone; the long-term costs to your reputation, and the costs of making a habit of doing something that’s usually destructive, vastly outweigh any short-term advantage you could gain this one time.
Isn't this the small world case? Following the pattern, my big world example would be: "when there are so many people around that you rarely meet the same person twice, you should scam people as much as you can". Or do I have a different interpretation of small/big world?
Interesting, your small/big-world point reminds me of the mean-field footnote in OP that I felt confused by:
There are zones where the best move is non-consequentialist, in the sense that you do best by following heuristics (or principles, or virtues) that are good regardless of the exact specifics of the situation you’re in. You can neglect your (immediate) action’s (immediate and specific) effect on the environment.[1] [1]: This intro to mean field theory seems to be pointing to a similar concept of a “foreground” particle and a “background”, where the background is made of the interaction of lots of little “foreground” particles, but each particle has negligible impact on the background; you can just model the particle’s behavior as some function of the (assumed approximately fixed) background. But I don’t really know mean field theory in physics so I can’t be sure the analogy is apt.
Here, assuming that I have no effect on the environment is not quite what mean field theory does (at least not in the places where I encountered it in physics): mean field theory assumes "everyone behaves like me. If I follow a principle, the world does too." From that assumption, a natural next step is Kant's 'Act only according to that maxim whereby you can at the same time will that it should become a universal law.'. This does seem related to "the best move is non-consequentialist".
If I had zero effect on the big world (also in the sense that the big world does not treat me any differently if I misbehave) there would be no force that pushes me to cooperate. But with the mean-field worldview, I will assume that there is an effect even if I am not tracking the causality – whether there are repeated interactions or not.
the "world" i was thinking of was the iterated "game" of interacting with people. This world is "big" in time because you expect it to go on for a long time in future. So the immediate impact of your next move in the game is tiny in comparison to long run effects like your reputation or your habits.
And how you behave determines the size of the world you can play in! Are you part of the vast virtuous mutually normative collective of moral agents that is our society? Or, as often happens with smart self-interested sociopaths, does your game end early in jail or shunned or dead.
when you are working on a hard technical problem that’s a long way from solved, and could have both benefits and harms to humanity if it were solved, you don’t really have to worry about “what if we succeed and it’s bad?” or “what if this information gets into the hands of the wrong people?” You’re a long way from winning, worrying about winning too much is silly, and heuristics like “do a good job on important work” and “share scientific knowledge freely” are more trustworthy than galaxy-brained arguments for why in this particular case the utilitarian calculus works out the other way.
I'm not sure why exactly you believe this, can you say more? (Calling the arguments for backfire "galaxy-brained" kinda unduly stacks the deck, I think.)
Heuristics like "doing a good job on important work benefits the world" and "it's beneficial to share scientific knowledge freely" are time-tested (at least over the past several centuries) and have a lot of empirical evidence in their favor. "The Baconian project has been working" is one of the only sweeping historical/social claims I am highly confident in. And you don't even need that to believe a more common-sense thing like "working to solve problems that people find helpful to have solved is beneficial on net." Like, "productivity is good" is something that could have been intuitive in antiquity or prehistory. There's just tons of individual cases, life experience, and logical extrapolation that all point in the same direction.
"THIS time if I invent/improve this technology it'll be harmful" or "THIS time if I disseminate this discovery it'll be harmful" fly in the face of those heuristics and depend on your specific causal argument being correct & your conceptualizations apt to reality. Often it's a fairly conjunctive argument full of lots of beliefs about what people will do, made without specific knowledge of the kinds of people in question. Doesn't mean it's wrong, but it's risky.
Thanks for clarifying!
My worry is, these heuristics are only "time-tested" in that they've (apparently) had good consequences relative to not-super-large-scale goals, in not-super-unfamiliar contexts. Many people working on the kinds of problems you're pointing at do so for impartial altruistic reasons. So I suspect that that track record is only weak evidence that such heuristics will work relative to their goals, in unfamiliar contexts — like, say, ASI takeoff.
"Sharing scientific knowledge freely" contributed to the industrial revolution, for example, without which factory farming wouldn't have happened, nor the development of technologies that pose x-risks. Granted, the implications of these examples are debatable. But they seem prima facie concerning enough to me that calling big-world heuristics "robust to uncertainty" sounds way too strong (for impartial altruistic goals at least).
depend on your specific causal argument being correct & your conceptualizations apt to reality. Often it's a fairly conjunctive argument full of lots of beliefs about what people will do
I don't think the extrapolation from "good on local scales / familiar contexts" to "good on the cosmic scale / very unfamiliar contexts" escapes really conjunctive causal mechanisms. It just hides them at a higher level of abstraction. When making that extrapolation, we're predicting that the same sorts of mechanisms that led to good consequences on the local scales will be at play on the cosmic scale. So we're baking in a lot of implicit beliefs about those mechanisms.
I hesitate to curate this, but am curating it.
I hesitate because this post raises the question "our usual heuristics on how to be a good/effective person don't hold up during extreme circumstances. wat do?". The post then doesn't answer that question.
I'm scared of many people reading this, thinking "I'm an exceptional person during what is maybe the AI Midgame, maybe approaching The Endgame, and indeed yeah the commonsense morality/heuristics don't make sense during these extreme circumstances." I am currently way more worried abut those people jettisoning their morality without thinking it through, than them failing to handle nuanced Midgame/Endgame situations optimally.
The reason I am curating this post is because I think it does a good job of laying out why you should have a lot of commonsense intuitions on how to coordinate and participate in the world, in a way that lays out a different set of mechanistic gears than I've seen laid out before. This is both conceptually interesting, and maybe motivating to some readers who hadn't realized why this was important.
I think we have at least awhile more than all the Big World heuristics just straightforwardly apply, even if the circumstances are starting to feel extreme. In the end, my guess is when you are sufficiently galaxy brained, it turns out that basically you want to just keep applying these intuitions even in various extreme circumstances. But I don't have a succinct argument for that, and would like to have more clear and robust arguments for that, uh, soon.
I'm scared of many people reading this, thinking "I'm an exceptional person during what is maybe the AI Midgame, maybe approaching The Endgame, and indeed yeah the commonsense morality/heuristics don't make sense during these extreme circumstances." I am currently way more worried abut those people jettisoning their morality without thinking it through, than them failing to handle nuanced Midgame/Endgame situations optimally.
This part of the thought process itself feels not very Big-World-pilled to me? Like, in a big world it's generally a good thing to promote interesting/valuable ideas, even if a few people might misinterpret them and go on to do bad things on that basis, or something?
when you expect that you’re going to run into the same people over and over again for a long time to come, it’s wise to be consistently trustworthy and helpful even if you could get an advantage in the short run by betraying someone; the long-term costs to your reputation, and the costs of making a habit of doing something that’s usually destructive, vastly outweigh any short-term advantage you could gain this one time.
Isn't this one more of a "small-world" effect? In a big world, you're unlikely to run into the same people over and over and one-shot dynamics become more prevalent.
This reminds me of when Ada Palmer asked "What would Machiavelli say if you asked him what would happen if Milan suddenly changed from a monarchal duchy to a republic?":
The poli sci students went first: He’d say that it would be very unstable, because the people don’t have a republican tradition, so lots of ambitious families would be tempted to try to take over, so you’d have to get rid of those ambitious families, like the example Livy gives of killing the sons of Brutus in the Roman Republic, and you would have to work hard to get the people passionately invested in the new republican institutions, or they wouldn’t stand by them when the going gets tough or conquerors threaten. It was a great answer. Then my students replied: He’d say it would all depend on whether Cardinal Ascanio Visconti Sforza is or isn’t in the inner circle of the current pope, how badly the Orsini-Colonna feud is raging, whether politics in Florence is stable enough for the Medici to aid Milan’s defenses, and whether Emperor Maximilian is free to defend Milan or too busy dealing with Vladislaus of Hungary. “And I think I’d have something to say about it!” added my fearsome Caterina Sforza; “And me,” added my ominously smiling King Charles. In fact, my class had given a silent answer before anyone spoke, since the instant they heard the phrase, “if Milan became a republic,” all my students had turned as a body to stare at our King Charles with trepidation, with a couple of glances for our Ascanio Visconti Sforza. It was a completely different answer from the other class’s, but the thing that made the moment magical is that both were right.
By contrast, if you are in the endgame of a game of chess, or in an oligopolistic competition, you absolutely need to think about how other players will respond to your actions. You do need to search through a tree of “if I do this, that will happen” and compute that specifically for the specific game state you’re in, and recompute it every time the game state changes.
Since chess is bounded you really can calculate (and in the endgame you're now able to). With oligopolistic competition, your power has increased and the players that matter to you are fewer, but the moves are not bounded -- if Apple and Dell (or whoever) are in an oligopoly but then Apple creates the iPhone...things can change.
Most of the time, for most people, this is not the case!
I mean it is interesting that we are, all of us, now in this case. And I do notice that it causes heuristic problems for even very thoughtful people.
This overlooks one important distinction: most of these scenarios focus on ignoring indirect effects when the direct effects are small enough not to upset the environment and change other parties' behavior. That makes sense, but the direct effects still need to be positive. Your ideal game of chess might not vary much with your opponent's moves, but it does not involve marching your king forwards as fast as possible. Likewise,
- when you are working on a hard technical problem that’s a long way from solved, and could have both benefits and harms to humanity if it were solved, you don’t really have to worry about “what if we succeed and it’s bad?” or “what if this information gets into the hands of the wrong people?” You’re a long way from winning, worrying about winning too much is silly, and heuristics like “do a good job on important work” and “share scientific knowledge freely” are more trustworthy than galaxy-brained arguments for why in this particular case the utilitarian calculus works out the other way.
if you're doing AI research, it might not be significant how other people respond to your incremental progress, in an environment where many others are making incremental progress, but it is still significant whether making progress is the right thing to do. That is a separate discussion and not one that the pattern outlined here pertains to.
I think the chess example is inaccurate, but it doesn't refute the main message (or how I choose to interpret it at the very least).
It's very wrong to say that in chess, in the opening moves, both players do their own little thing and after some development, only then do the cordialities end.
In the opening phase, both players aim to achieve certain advantages and if you let your opponent get all the advantages they want, you won't have a very pleasant middle game.
Openings have a 'logic'. If white plays 1.e4 then black doesn't want white to get the full center and thus plays a move that prevents that, while also grabbing for center. This is the logic of 1.e5. Now white can't have the whole center and so focuses on developing their pieces. Maybe they could do it in a way that would let them keep the initiative and so plays 2.Nf3, getting the Knight out while attacking the e5 pawn and thus again forcing black to respond etc.
If you don't understand your opening's 'logic' and play against someone who does, it's very easy to get disadvantaged early on, possibly several moves before even noticing you're on the backfoot.
On the other hand, what is true, is the slow, gradual, incremental nature of 'advantage' in chess. Unless both opponents decide to go for a very sharp and complicated game (I'm assuming good players btw), then progress is going to be very slow and players won't really be thinking about checkmate until much much later.
This is the nature of Strategy vs Tactics. Tactics are about identifying a short, exact sequence of moves that leads to an immediate concrete advantage.
Strategy is essentially about developing a situation that makes it easier/likelier for there to be tactics for you to execute.
Consider the following situations:
when you are a small, growing startup in a big market, standard advice is not to worry too much about your competitors or try to do anything adversarial “against” them, but just to focus on growing and providing value to your own customers.
when you are a small trader in a big market, you don’t need to worry about your trades shifting the market price or revealing information to your competitors; in many contexts, your optimal strategy is simply to bid your true price, buying when an asset is cheaper than your “happy price” and selling when it’s more expensive.
when you are in the early stages of a game, often your best strategy is to grow your “resources” (like developing your pieces in chess, trying to control more territory and have more value on the board), following a pattern that’s mostly independent of what the other players are doing and gets you more of something that’s valuable across many possible game states.
when you are a species whose resource needs are much smaller than the carrying capacity of your environment, you are r-selected; your fitness is maximized by just having as many offspring as possible and not “worrying” about running out of resources.
when you expect that you’re going to run into the same people over and over again for a long time to come, it’s wise to be consistently trustworthy and helpful even if you could get an advantage in the short run by betraying someone; the long-term costs to your reputation, and the costs of making a habit of doing something that’s usually destructive, vastly outweigh any short-term advantage you could gain this one time.
when you are working on a hard technical problem that’s a long way from solved, and could have both benefits and harms to humanity if it were solved, you don’t really have to worry about “what if we succeed and it’s bad?” or “what if this information gets into the hands of the wrong people?” You’re a long way from winning, worrying about winning too much is silly, and heuristics like “do a good job on important work” and “share scientific knowledge freely” are more trustworthy than galaxy-brained arguments for why in this particular case the utilitarian calculus works out the other way.
These kinds of intuitions are what I’d call “big-world” heuristics.
What they have in common is that the world is “big” enough relative to you — the market is much bigger than you, the problem is much bigger than your progress on it, the game is very far from being over, the environment is much bigger than your resource consumption, etc — that your best strategy for achieving your goals is not to explicitly calculate how your action will affect the environment (including how your actions will affect other agents’ actions.) There are zones where the best move is non-consequentialist, in the sense that you do best by following heuristics (or principles, or virtues) that are good regardless of the exact specifics of the situation you’re in. You can neglect your (immediate) action’s (immediate and specific) effect on the environment.1
By contrast, if you are in the endgame of a game of chess, or in an oligopolistic competition, you absolutely need to think about how other players will respond to your actions. You do need to search through a tree of “if I do this, that will happen” and compute that specifically for the specific game state you’re in, and recompute it every time the game state changes.
Big-world intuitions feel a bit like “play fair, mind your own business, cultivate your own garden, do your best, and it’ll all more-or-less work out for the best in the long run”. Or like “keep your eyes on your own paper”, don’t try to manipulate people, just follow the same path you would if you were alone; whether you’re Robinson Crusoe alone on an island or one anonymous citizen in a big city, your “job” is pretty much the same, in that you need to work to take care of your own needs.
This came to mind because I noticed that I pretty much always rely on big-world intuitions. I don’t really know what to do about questions like “what if we win too much and it’s bad” or “what if I, personally, have the power to shape society, how would I choose to set it up”. And I don’t really even know where to begin with chains of adversarial strategic thinking like “if I do this, she’ll do that, so I’ll do this…” I’m always thinking of myself as one participant among many, with a small share of power/influence, working on problems hard enough that the only thing I have to worry about is doing the best job I can, and trying to follow universally sound principles/heuristics.
The nice thing about sticking to principles/heuristics is that it’s robust to uncertainty. What if you’ve misread the situation? What if you’re not as big a deal as you think you are? Behave in a way that’s usually for the best, even so.
Except, of course, if you really are super powerful, super close to winning, super close to the “endgame” in some sense, or in a really bizarre situation where the consequences of following generally-good heuristics happen to be very bad. Most of the time, for most people, this is not the case! But it is disquieting to notice that none of my grounding intuitions or usual ways of thinking about what “healthy” or “prudent” or “ethical” look like, are fit for such situations.
This intro to mean field theory seems to be pointing to a similar concept of a “foreground” particle and a “background”, where the background is made of the interaction of lots of little “foreground” particles, but each particle has negligible impact on the background; you can just model the particle’s behavior as some function of the (assumed approximately fixed) background. But I don’t really know mean field theory in physics so I can’t be sure the analogy is apt.