Sam Altman and Dario Amodei are the public faces of AI safety and that's poisoning the well
Normies just smell a rat whenever those two bring up AI safety. I think when the AI safety community gets a word in edge-wise (like you're arguing with your friends or maybe a tweet goes viral), we usually say something like
"it could kill everyone. but China could also kill everyone, so now more than ever we have to supercharge OpenAI and Anthropic."
and we just sound like we're "one of them."
I feel like I should be hearing "shut it down" about as often as I hear "bro...
Less_raichu seems to be saying that those two say that thing and that they're the loudest spokespersons for ai safety.
One metric I've been using for AI impressiveness is how much I could do if I was the only one in the world with them. Like, if you gave me Sonnet 5, and nobody else in the world has LLMs or know they exist, what's that look like? I'd look like an insanely good engineer, right?
What kind of engineer knows every coding language on github, runs through features the sprint planning meeting thought would take a few days in half an hour, and can read thousands of lines of code in a minute? We agree in 2016 if you 'saturated' leetcode you were great, right?
So. Ima...
I looked up EpochAI's list of benchmarks. The very poor ECI of Grok 4.5 as opposed to GPT-5.5 seems to be a result of xAI not caring about math. Additionally, Grok 4.6 seems to be closer to Opus 4.8 with respect to ARC-AGI-3, if not outright outperforming Opus, as the actual ARC-AGI-3 score suggests. Does it mean that xAI is less behind than we think? I guess that Zvi will have to write something like "Grok 4.6 is three, not six, mounths behind. Stop xAI to hell!"
P.S. The same issue seems to apply to Meta's Muse Spark, except that it has even less evaluate...
Also, SpaceX has caught up to Meta in credibly having enough compute in 2027-2028 to stay in the game, if either of them can assemble a functional model development team. As Google illustrates, it's not easy to do that (even when you have some of the best people), but as OpenAI and Anthropic illustrate, it's not so difficult that it can't be replicated. Muse Sparks are probably small enough models that their non-frontier performance doesn't count as evidence that they're not well-made (and that a Mythos-sized Muse model won't have Mythos-level capabilities...
i wish i could just exchange money for specific goods and services: specifically, having ubers waiting for me. right now there’s this whole thing where it’s kind of a dick move to leave drivers waiting because it cuts into their profits, so they have some chance of just cancelling the ride, and/or leaving a bad rating which makes it harder for me to get rides consistently in the future. but a big part of the reason why i sometimes want drivers to wait for me is i’m in a hurry and uber wait times have substantial variance, so i want to call early to be safe...
much more expensive and less convenient than uber (usually need to book much further ahead of time); also, overkill for a 30 minute ride because you can usually only book a min of 2 hours or something.
It is brought up quite often, but I have not (yet!) heard a satisfactory answer to the question: "why are so many leading people in advocating AI risk also investing in developing LLMs and AI-dependent ecosystem?"
It seems to me, the position (in simple form) holds:
I have read that writers have a compulsion to write. They can't not write. Writing is like breathing to them.
As I said, for some humans writing gives them a real utility. This is a benefit of human writing so encouraging it has some utility.
How could they not be curious if current AI models can write good fiction? How could they not wonder if these new creatures we've created, who breathe words the way we breathe air, can write excellently?
So your argument is that the utility is that it will provide some satisfaction to the curiosity of Gwern, Alexander Wa...
TL;DR(-ish): I've been thinking about a possible principle/heuristic that I'm provisionally calling "the Aristarchean Principle" (after Aristarchus of Samos). The principle states that the facts of the universe need not "be sensible" or "make sense" from the perspective of what seems "intuitively" "sensible" or "makes sense" to us.
I'm looking for a "nice" formulation of the principle that "nicely" excludes pathologies such as denying the validity of any reasoning.
I'm open to the possibility that, upon further thought, there's no "there" there and there is ...
The universe only needs to resemble our intuition locally, because our intuition comes from the part of the universe we are in. On different scales or in different places, there is no need for it to seem sensible.
Perhaps it is that simple.
I just saw a post about mindreading tech. Has there been/is there going to be any interest among the LW community to consider the likelihood of such technology being within reach or if it's even already here and well-developped through Manhattan style projects?
It's a very murky subject and I'm curious how LWers would approach it.
Coalitional motte-and-bailey has both motte-defenders and bailey-promoters as distinct groups. The motte people can tacitly disagree with the bailey claims, avoid motte-and-bailey dynamics within their castle, regard the bailey people with some distaste, and dislike when some of their own start presenting at the bailey events. At the same time, the bailey people advance their wild claims while gesturing at the motte arguments (which they often don't understand) and the prestige of the authors of the arguments, which directs attention and public renown to t...
Further to Zvi's post on the podcast ...
https://www.lesswrong.com/posts/BZW8CeAHHJ52EvwYt/on-dwarkesh-patel-s-podcast-with-ryan-greenblatt
Here is a transcript summary for reading in < 10 minutes ...
Recursive Self-Improvement & Timelines
Dwarkesh Patel: Today I’m chatting with Ryan Greenblatt, Chief Scientist at Redwood Research. Let's talk about recursive self-improvement: the idea that once we build human-level intelligences, they quickly slingshot toward superintelligences more competent than top experts across every field. Historically, I’ve been ...
This is a follow-up question to my post The Open Problems of the AI Alignment Field and their Cruxes.
Why do you think there has been little (visible?) research on CEV since it was posed?
Because we at least need to point the AI at something.
There are really two questions:
The field overview looked much more at verifying that a system actually does T than on these questions.
Pointing it at the CEV conditioned on mankind having the ability to do so is close to a governance problem
I think I mostly agree now that what remains of CEV today is mostly a governance problem. CEV looks only like a technical problem (and thus as a separa...
I found this list of links interesting and valuable,
Original Research on Less Wrong by lukeprog, 30th Oct 2012 https://www.lesswrong.com/posts/jTkmEGWM4dJAfE62W/original-research-on-less-wrong
It's a large list of original-ish posts concerning philosophy, decision theory, AI architectures, mathematical logic, ethics. Most of the posts are old enough to be forgotten.
Are there newer such lists, newer than 14 years old?
https://www.astralcodexten.com/p/your-book-review-the-escape-artist
Reading about the unwillingness of Jews and Allies to believe that the holocaust was real, while it was happening, and even in some cases while being herded into the camps, was interesting and depressing.
(The review mentions that too.)
Reading what mathematicians are going through from the sidelines, mathematicians are now (or will soon be) going through what Ted Chiang called the Evolution of Human Science in his 2000 short story, where normal "human science" is increasingly superseded by "metascience" -- the constant effort of interpreting the scientific results of metahumans/LLMs.
Perhaps this will hit all of us soon.
In case your August wasn't cyberpunk enough ...
‘Cyber privateers’: Trump issues order allowing US companies to hack overseas groups under certain conditions
https://www.cnn.com/2026/08/13/politics/cyber-privateers-trump-order-overseas-groups-hacking
EXPANDING CAPABILITIES TO COMBAT TRANSNATIONAL CYBER-ENABLED CRIME
My short story, You're Absolutely Right (SPOILERS in block) tried to imagine a future where a frontier lab had some real (and significant!) safeguards beyond what current frontier labs had
but it still didn't matter anyway.
For various reasons, I think it's likely that methods that involve continual learning (i.e. modifying the weights during deployment) will come online soon. Here are some implications for safety:
My impression from what discussions can be found in the literature is that the idea isn't ready, so full weight updating being secretly ready requires a greater conspiracy than RLVR did (where many plausible paths were visible before DeepSeek R1 demonstrated that GRPO with chains of thought is sufficient). Experiments with mostly frozen weights sometimes get something useful and not too broken, and true recurrence isn't too dissimilar from the practical standpoint from shallow recurrent states of SSMs and such in hybrid attention architectures, so it's pla...
Chemists shrink gallium nitride, the material behind LED lighting, into nanocrystals
Molten-salt method from UChicago and Argonne could unlock durable materials for printed electronics, flexible devices
I needed a more detailed sketch of algorithmic condensation so I could build on it, so here's what I've got. I've done a proof for two concrete cases of the objectivity theorem, which I find much easier to understand than keeping track of the indices in the fully general proof.
A latent string model is a set of observations
I didn't mean to imply that it was in the paper as I have it, that's why I said based on instead of from, sorry if that was confusing. It's a presentational thing and it's fine to just use
I don't see how there's an implicit optimality assumption unless you're trying to read quantities at every step in the proof as a program length. I agree that there's something aesthetically worse about a proof that routes through something less nicely interpretable.
Effectively multi-tasking with AIs seems to require a state of mind that is inimical to deep work / thinking. When I'm switching between AI panes I'm constantly distracted, disoriented, frazzled. I get tired out having to pick up and put down large amounts of context each time, etc. etc. I've observed that doing large volumes of work with AI coding agents mostly destroys my ability to do deep thinking for the rest of the day.
I don't know what the right way is to resolve this. What's worked reasonably well for me so far is to do most of my deep thinking in...
For me the solution is to split the work into chunks that take e.g. 3h-12h and the model can do fully autonomously. So you can just start it and not look at what it's doing. Ideally, fire several in parallel and check on them only e.g. 3 times a day. I think the "deep work mode" is destroyed only when all the time there might be something requiring your attention, and it's on you to make sure there's nothing : )
So e.g. there is a giant difference between working with an AI that needs your attention 2 times per hour for 30 seconds and one that never needs your attention at all.
GLM 5.3 is scary. It appears to have relatively strong cyber capabilities: not Mythos level, but it seems like it could cause significant harm.