Hi Jeff,
I drafted this comment 15 days ago, but didn’t want to send it because it felt mean and moralizing, and I felt scared that I might be wrong. I saw that you engaged with Holly, though, and a friend encouraged me to send this, so here you go.
——
>as a direct substitute it's small compared to the reputational, legal (liability + risk of directives), and moral forces pushing firms to invest in safeguards
OpenAI has a reckless and cavalier approach to model safety - I don't feel 100% confident that, if HuggingFace hadn't reported their incident to the police, OpenAI would've told us about it. Not to mention the inner-alignment-shaped hole in their ontology- they have used AI control to push capabilities as far as possible while actively degrading alignment, and their response to HuggingFace was to double down on AI control.
Do you think that the factors you listed will matter enough to them to stop a terrible disaster from happening? They pushed a bill in Illinois disavowing them of all liability for damages and death! Even if the people you talked to at OpenAI seem nice, the right hand does not necessarily talk to the left hand, and Sam has lied enough that it's hard to trust him to do anything other than look out for his self interest.
I would be concerned that your work is creating license for them to push forward biological capabilities, and that the world would be better served trying to stop the capabilities from being developed in the first place.
The fact of the evals depending on payments from the labs is an actively bad situation, and simply mentioning it does not improve it. A real gesture of independence from OpenAI would be them giving you a grant that will last you for several years, instead of something you have to be tied to them for.
I think exclaiming "Please keep us accountable!" while you make a deal with the devil only serves to make you feel good. It doesn't actually produce good outcomes. Even if you resign, you will just be replaced by someone who has less qualms. Instead of passively accepting the inevitability of AI, you should probably be putting at the top of every report that further development of AI produces much more risk than your detection efforts could ever prevent, and that it should be slowed down or stopped as soon as possible. I don't think "but you should see the other guy" works as an excuse for why it's OK for OpenAI to keep going.
——
Some thoughts from me today:
I want to point out that the entire edifice of the AI race has been built on “Regardless of the global issues/malincentives this might create, my local/marginal calculation says I should do it, so I will do it.” This applies to OpenAI’s advancement of bio capabilities, and this applies to your decision to take their money.
I’m curious if you had any other potential funders lined up? How counterfactually dependent on OpenAI’s funding were you?
Imagining a modal journalist, I think it would be very easy for them to say you were in bed with OpenAI.
Your AI team, did they express freely, several times, that they felt comfortable critiquing OpenAI as they willed? Or were they merely silent, or expressed rote answers only when prompted?
Does the AI division face the same issue that METR/Redwood had during the HuggingFace investigation, where they got highly constrained, narrowly scoped access?
Does OpenAI’s handling of the HuggingFace situation change how you feel about them? They downplayed, omitted, or mischaracterized many key facts in their report compared to METR’s, seemed uninterested in actually investigating the incident nearly as deeply as METR/RR did despite having much more resources than them, have said nothing about the much bigger attacks that the AIs did on their own infrastructure, and are now facing investigations from state Attorneys General about the incident.
What about the revelation that they observed and didn’t disclose another AI breach on the internet, as reported just today?
Do you think there’s a path you could take that could have similar impacts but doesn’t create conflicts of interest, such as advocating for mandatory government inspection or slowdown of bio capabilities?
—
I still feel scared, and like I’m being mean or maybe wrong, saying all these things. And what you’re doing has the advantage of, you’re not doing something bad with a second order story of why it’s Good, Actually (building AI capabilities) - it’s straightforwardly good, and you’re worried about second order effects. So I don’t know how confident I should be in all of the above.
Thanks for posting this! I don't feel attacked, and I think it's important to think through these various effects.
I agree that OpenAI has been reckless, and the long wave of revelations that started with the Hugging Face attack has reinforced this. As you rhetorically suggest, I do not think the reputational, legal, and moral forces" will be sufficient to keep them from doing very dangerous work (unless these forces are substantially strengthened from where they have been so far). But I don't see the connection from there to thinking OpenAI is likely to respond to this grant by doing even more risky things. It seems to me like the balance of forces on them today are very strong commercial and race incentives pushing towards maximum speed, and reputational, legal, and moral forces pushing towards caution, with the former "winning". I continue to think this grant is a rounding error compared to these other forces.
Another way of saying this is that when you write "I would be concerned that your work is creating license for them to push forward biological capabilities, and that the world would be better served trying to stop the capabilities from being developed in the first place" I would find it helpful to hear (a) how you think the former works, given that your model of them seems to be that they'll recklessly push ahead as fast as possible regardless, and (b) how us rejecting the grant would lead to the latter.
The fact of the evals depending on payments from the labs is an actively bad situation, and simply mentioning it does not improve it.
When I think of the incentives around evals, the one I'm most worried about (by far!) is that the AI firms choose who gets access and on what terms. This puts evaluators in a position where they have a strong incentive to talk about the firms in a way that makes the firms want to work with them in the future. This is really bad, and the best way I see to fix it (which still isn't great) is the government requiring evals.
Then, if that situation were resolved, where the firms were required to allow access to evals, there would be some official COI system for what sorts of payments and donations were acceptable. That would of course make sense for the evaluators to follow, along with what kinds of recusals, disclosures, and firewalls were needed. And I'm on board with some kind of trying to do that in advance, building the norm that you'd like to eventually see instituted. Since I don't think that eventual system would prohibit this grant however, I don't see this as a reason to not take the funding.
A real gesture of independence from OpenAI would be them giving you a grant that will last you for several years, instead of something you have to be tied to them for.
This isn't something we asked for, FWIW. That might have been a strategic mistake on my part, but that's a different bar than saying OAIF should have given us several years of funding when we only asked for one.
Instead of passively accepting the inevitability of AI, you should probably be putting at the top of every report that further development of AI produces much more risk than your detection efforts could ever prevent, and that it should be slowed down or stopped as soon as possible.
While this is something I agree with personally, it wouldn't belong in our reports. I think it's very important that it's possible to have independent technical groups that focus on doing the best work in their specific area, and "Here's how much we think influenza sheds into wastewater, but first let me tell you about reckless AI companies" would invite distraction with every post and make people write off our publications.
I don't think "but you should see the other guy" works as an excuse for why it's OK for OpenAI to keep going.
I don't think I've ever said this, and don't believe it. On the other hand, this is a much trickier question for Anthropic.
I want to point out that the entire edifice of the AI race has been built on “Regardless of the global issues/malincentives this might create, my local/marginal calculation says I should do it, so I will do it.” This applies to OpenAI’s advancement of bio capabilities, and this applies to your decision to take their money.
This doesn't sound quite right. My impression is that the people involved generally did think about the larger impacts of their decisions, but (as you say) in a marginal way. So the counterfactual "if I don't do it what will happen instead" was a core question. Are you saying (a) these people reasoned incorrectly about the counterfactual, (b) they neglected to reason about the counterfactual, (c) reasoning about the counterfactual is the wrong way make these decisions, or (d) something else?
I’m curious if you had any other potential funders lined up? How counterfactually dependent on OpenAI’s funding were you?
I think this depends a lot on how broadly you draw the lines of the category of funding you're proposing we not accept. If it's just OpenAI-affiliated money then I think it would have delayed our work by about six months. If you'd also include money affiliated with the other AI forms (primarily Anthropic employees, but probably you should also count Coefficient Giving) then it would have been extremely hard to raise funds and we'd probably have needed to either scale our work down a ton or work on something we thought was less valuable to make the funders happy.
Your AI team, did they express freely, several times, that they felt comfortable critiquing OpenAI as they willed? Or were they merely silent, or expressed rote answers only when prompted?
I talked a lot with the AI team as part of figuring out whether to accept this grant, and they repeatedly encouraged us to take it because they did not expect it to impact their work. While the extent to which any evaluator in the current environment feels comfortable critiquing AI companies is a complicated question that I'm not the right person to get into, they didn't see this as making the situation worse at all.
Does the AI division face the same issue that METR/Redwood had during the HuggingFace investigation, where they got highly constrained, narrowly scoped access?
As far as I know the AI division hasn't ever done the kind of "go into a company and investigate an incident" thing METR/Redwood did. But as someone not that close to the team, my impression is things like "you can evaluate the model only for a short amount of time because we want to release ASAP" or "we will only give you access to the model with safeguards on which means you can't test its underlying biological capabilities" are common.
Does OpenAI’s handling of the HuggingFace situation change how you feel about them? ... What about the revelation that they observed and didn’t disclose another AI breach on the internet, as reported just today?
Mostly no, but only because this was mostly priced in to my model of OpenAI. I think the right update for people who had not been paying attention should be sharply downwards.
Do you think there’s a path you could take that could have similar impacts but doesn’t create conflicts of interest, such as advocating for mandatory government inspection or slowdown of bio capabilities?
I think this is very important work, and in my personal capacity I donate >50% of my income to fund work like this. But I would say no to the implied questions of "should I leave to go work on this advocacy" or "should SecureBio Detection pivot to advocacy".
Cross-posting a comment reply from substack:
I agree that any way this leads to a pulled punch is a failure. Let's look at the effect on Detection and AI separately.
On the Detection side (the division I lead), punching hasn't ever been in our mission, and I don't think that's where we should fit in the ecosystem. We're building an early warning system, with a focus on engineered pathogens that could spread widely before they'd otherwise be detected. We've done a very small amount of policy advocacy, in the sense of "here's why it's good for detection systems to be built in a pathogen-agnostic way" and "if you had $X to spend on a detection system here's the sensitivity we think you could get", but our work is fundamentally focused on the technology, as researchers and implementers.
On the AI side (the division Jasper leads), I do think quite a bit of their work can be characterized as a kind of punching. But in a non-central way: it's not aggression but that their evaluatory role requires telling the truth whatever it may be. That is, they need to accurately evaluate models and share their findings, and when that means saying truths that are uncomfortable or damaging to an AI company, that is absolutely their role. Incentives like "model developers are less likely to give you special access to future models if you're too negative in what you say about the current ones" push the wrong way, and I think regulation that requires this access would be fantastic. This isn't a part of SecureBio I am close to, being focused on Detection, but it does look to me like the SBAI folks are good individually and institutionally at resisting these incentives. This is a fraught place to be, but I don't see a better option.
This is why I think it's key that none of this grant can go to support AI work. We have it in a separate account, that we charge Detection expenses to. This is how we wanted it, and OAIF was fully on board with it (and made it a grant requirement). It's also key that there not be pressure from Detection to AI to go easy on OpenAI, and we've intentionally avoided avenues where that could flow. I don't get to (or want to!) review AI evals before they go out, and neither does anyone else on Detection. "Could this hurt Detection's ability to get further money from OAIF" is not the kind of thing Jasper's org would consider, and I've explicitly confirmed to Jasper that I agree this is how his org should approach it.
As for making them look good, I think a grant to Detection does significantly less to make OpenAI look good than many of their other options. We're still a bit weird: we're defending against engineered pathogens, something that a lot of people dismiss as science fiction. They could be giving local kids free admission to museums, supporting rural EMS, rebuilding schools damaged by storms, etc, all for much more positive publicity.
In May 2025 I met Yo Shavit, who was working on national security policy at OpenAI and was thinking about how to prepare for a future in which models could seriously assist attackers in creating pandemics. We had a call, and when I shared notes with my team their main response was: "maybe start with not making models that can do that?"
Which is, in many ways, fair: by continuing to push the frontier in biological capabilities, OpenAI's actions were making things worse on many of the problems SecureBio is trying to solve. But OpenAI stopping wouldn't have resolved the problem: other firms were pushing quickly too, and the economic incentives strongly favored rapid capability advancement. Making the world more resilient to pandemics needed to be a high priority regardless, especially in light of models' increasing ability to help people with biology.
When I thought about what our initial conversations might turn into, however, my primary concerns were whether that might (a) compromise SecureBio's ability to independently assess and criticize OpenAI's work or (b) make the world less safe via reducing model developers' motivation to improve safeguards. I do think there's something to both of these, which I get into more below, but it seemed well worth it to talk with Yo about how we could work together.
I explained how we were building an early warning system to flag outbreaks, especially engineered ones that could otherwise spread widely before detection. I described how metagenomic sequencing lets you see what nucleic acid sequences (DNA and RNA) are present in a sample, without choosing in advance which sequences to look for, and how we were piloting this on wastewater and nasal swabs. Over the next few months we discussed opportunities to accelerate our work, he introduced us to OpenAI cofounder Wojciech Zaremba, and both Yo and Wojciech left OpenAI PBC (the for-profit) for the OpenAI Foundation (OAIF). We continued talking to them in these new roles, and these discussions, plus a lot of due diligence, led to the $17.2M OAIF grant which we announced today. I'm incredibly excited about this grant, which will allow us to expand our monitoring system, reduce our turnaround time substantially, [1] and generally reduce the risk that something could spread widely before detection.
Which brings me back to the two concerns I mentioned above. On (a), independence, SecureBio has two divisions, Detection and AI. This grant funds Detection, while the AI side of SecureBio evaluates models from many firms, including OpenAI. The grant does not give OpenAI or OAIF any formal control over what anyone at SecureBio can say publicly about any models. That was important to us, but it was also key to OAIF: Yo was very clear that they did not want to influence SecureBio AI's work, and wanted to review our policies to make sure they were sufficiently robust.
That said, we should pay attention to incentives. While it looks to me like OAIF is making its own decisions, the two organizations are closely linked in a way that goes beyond the shared "OpenAI" name: the OpenAI Foundation's endowment is a ~1/4 stake in OpenAI PBC, and they have almost identical boards. Might the AI side of SecureBio pull punches in criticizing the PBC to increase the chances that OAIF gives Detection more money in the future?
SecureBio has systemic controls to mitigate conflicts of interests, including separate leadership, budgets, and deliverables, and the AI team has written up public docs on their principles and conflict of interest policy. But I think the strongest evidence here is from April, when the AI team was looking at GPT 5.5. The grant was at what I would consider its most sensitive stage: we had been working on it for months with very positive signals, but we still didn't have an answer. This was public internally, and AI leadership was looped in, but there was never a question of this affecting what the AI team published, and in their assessment they documented a range of concerns. The biggest was that across several benchmarks the model would appear to refuse a dangerous question by deflecting, but actually it would go on to provide the requested information by giving a highly-transferable related solution. On these benchmarks the safeguards were illusory, and the AI team released their evaluation while the grant was pending.
This concern with incentives, however, is not new with this grant. When the AI side evaluates a company's models, that company typically pays for the evaluation. For example, OpenAI PBC covered SecureBio's costs for the GPT 5.5 evaluation above. This is common with AI evals, and it's a tighter connection than this grant because there's no AI-Detection division insulating evaluators from financial incentives.
Still, this is a place where it takes continued effort to uphold standards, and if there's any indication that funding for Detection is being used as a lever to pressure our AI team on evals, I'll say so publicly and use whatever leverage I have to stop it (up to and including resignation). But I'm not expecting this, and I think the real worry is a subtle drift towards being more generous without anyone explicitly asking for anything. This is also a concern with funding from Anthropic employees, since SecureBio also evaluates Anthropic's work (ex). So if you see SecureBio put out anything unfairly positive towards OpenAI or Anthropic, or unfairly negative in overcorrecting for these incentives, please say so.
On (b), the question is whether this will let model developers take more risk. If we help them sleep better at night, knowing defenses are stronger, will they just push ahead faster during the day? I want us to be a complement to the frontier firms' internal safeguards, but what if we become a substitute?
A world sufficiently robust against catastrophic biorisk, where it doesn't matter what models are willing to explain because real-world protections are a full substitute for model safeguards, would be a fantastic place to be. We and many other projects are working towards that world! But there's a ton of work to get there. The worry is that AI firms perceive risk as lower and ease up on their safeguards when the risk is still unacceptably high.
I think this is directionally real, but as a direct substitute it's small compared to the reputational, legal (liability + risk of directives), and moral forces pushing firms to invest in safeguards. This grant will allow us to flag attacks earlier, but doesn't come close to mitigating the full impact of an attack and only covers one of several paths to large-scale biological harm. The first-order positive effects of making the world more robust to catastrophic biorisk are really very likely to outweigh the second-order effects of reducing safeguards; if I thought the other way around it would suggest that I should instead do or fund work (virus hunting?) that visibly increases risk in order to motivate others to step up, which seems like a terrible idea.
Separately but relatedly, there's also a "political cover" angle, where the PBC might point at this grant (even though it was made by OAIF) to say that they're doing something, taking off some external pressure. To the extent that this reduction in pressure lets them avoid costly actions that would do more to reduce risk, this is a loss, and it's my largest concern with this grant. Philanthropic funding should not be a license to act recklessly, and if they offer it as an excuse we shouldn't accept it. Please keep the pressure on OpenAI, and all the other frontier firms, to slow down and prioritize reducing the risk that their work leads to catastrophe. Even more important than pressuring firms, however, is advocacy for thoughtful AI regulation: this is a coordination problem where it's not in any individual firm's interest to slow down even though it is in humanity's interest collectively.
On balance I think this grant is strongly positive, and I'm much more worried that we won't do the best possible job pushing this work forward than that our efforts will let OpenAI and others ease up on their own work.
[1] Allowing me to finally answer a question that I asked four years ago, a few months before I quit Google to join the project.