In the famous marshmallow challenge, teams race to build the tallest tower out of marshmallows, dry spaghetti, tape, and string. The celebrated result is that kindergarteners, using trial and error, generally beat lawyers, CEOs, and many other adults. A less celebrated finding is that architects beat kindergarteners. While trial and error is necessary, building tall durable towers requires an understanding of materials and structurally stable design patterns.
Similarly, responsibly designing AI requires understanding its elements, just as secure cryptography relies on foundations in complexity theory and number theory, and safe airplanes are built on knowledge of aerodynamics and control. Motivated by this, I recently left OpenAI to cofound the Institute for Responsible Superintelligence (RESI[1]) with Shafi Goldwasser, inventor of zero-knowledge proofs, and Vinod Vaikuntanathan, a leading MIT cryptographer.
Currently, AI systems are assembled from learned models, data, and institutions whose properties and interactions are poorly understood. As AI systems become more capable, relying on trial and error becomes more dangerous. Scientific foundations can help AI safety mature beyond "fly-fix-fly" cycles in early aviation or attack-and-patch in early cryptography. In each of these fields, trial and error won early and principled design won later. We think AI safety is at that transition, which is what we mean by safe-by-design: guarantees that follow from how a system is built. While many properties of deep learning have resisted theoretical analysis, rigorous guarantees may still be shown for the surrounding architecture, protocol, and incentives. In analogy, cryptographic guarantees hold for adversarial behavior.
Who we are
The research team includes Ran Canetti (composable security, obfuscation), Yael Tauman Kalai (verifiable delegation, interactive proofs), Dylan Hadfield-Menell (alignment theory), Yannai Gonczarowski (mechanism design, algorithmic collusion), Daniela Rus (safe control and planning for embodied AI), Chloé Bakalar (AI ethics), Andrew Sutherland (AI-assisted mathematics, formal verification), Yaron Singer (AI security), Noam Kolt and Rebecca Wexler (law), Anat Perry (social cognition and human-AI interaction) and, most importantly, junior researchers Miranda Christ, Greg Gluch, Sam Gunn, Tal Herman, Jonathan Shafer, Mirac Suzgun, and Neekon Vafa, whose work spans watermarking, verification, learning theory, and hallucination. Planned visitors include Scott Aaronson, Boaz Barak, Nicholas Carlini, Geoffrey Irving, and Jacob Steinhardt (the full list is at resi.org).
We aim to develop the theory and tools needed to reason about advanced AI before deployment: which safety properties are meaningful and achievable (alone and in combination), and which mechanisms deliver them, under which explicit assumptions. Safety here is broader than avoiding harm: ultimately it means systems whose benefits people can actually trust and share.
RESI is starting with several working groups, including numerous additional researchers beyond our core team, on various aspects of safety including:
Verification of model properties and outputs
Monitoring and control
Modular architectures
Cryptography and cybersecurity
AI harnesses
Game-theoretic foundations of AI safety
Ethics and superintelligence
Legal alignment
AI safety using quantum effects
Data attribution
Three problems that particularly excite me: (1) when do incentives continue to shape behavior as systems become more capable, and which systems don’t outgrow them (stickers motivate children, not adults); (2) how can we design cautious AI systems that act only when the outcomes, including those arising from many AIs acting together, are expected to be beneficial; and (3) how can we design training procedures that maintain a system's ability to be shut down, across a wide range of environments, capability growth, and self-modification.
AI is a program, not a person. It can be copied, rewound, encrypted, and sometimes formally reasoned about. The RESI team knows how to exploit such properties, defining and achieving reliability guarantees that may go beyond informal notions such as deception.
I believe highly capable AI should only be built with strong safety guarantees. Safety and capability trade off, and the challenge is to find the sweet spots: can a system be superhuman at cancer research yet provably limited elsewhere? RESI will study which safety properties are technically possible and how to achieve them, so that frontier labs and regulators can build on these findings.
Hasn’t this approach already been tried? A small number of pioneers have worked on theoretical alignment, including some of you. But developing a field requires a community: Turing made progress on cryptography, yet it took many more researchers to invent modern cryptography. We think AI safety is beginning to mature in a similar way, accelerated by AI itself, and we're collaborating with other like-minded organizations, like Resolution and ARC.
RESI's mission is to develop the scientific foundations of safe-by-design superintelligence. Outputs may include:
Characterization of what superintelligent systems can and cannot do.
A map of which safety properties are possible and impossible, and which survive composition, adversaries, and capability increases.
Mechanisms, protocols, algorithms, incentives, and architectures that achieve useful properties under explicit assumptions.
Implementations and experiments that test the resulting theories and constructions.
Novel thinking about how humans and organizations interact with these systems, and how their benefits and burdens, from labor displacement to access, may be distributed.
On a personal note, I've been surprised how some of my anxiety about AI has eased now that I'm doing my best to help.
I have my own views about promising directions, but I am more excited to see what our outstanding team discovers, working towards our mission. We will share progress at resi.org and look forward to working alongside many of you.
Thanks to Noam Kolt, Vinod Vaikuntanathan, Mirac Suzgun, and Chloé Bakalar for helpful comments.
It says superintelligence in the name, but I don't see any awareness in your organization that artificial superintelligence means the complete obsolescence of human intelligence as an organizing factor in the world...
In the famous marshmallow challenge, teams race to build the tallest tower out of marshmallows, dry spaghetti, tape, and string. The celebrated result is that kindergarteners, using trial and error, generally beat lawyers, CEOs, and many other adults. A less celebrated finding is that architects beat kindergarteners. While trial and error is necessary, building tall durable towers requires an understanding of materials and structurally stable design patterns.
Similarly, responsibly designing AI requires understanding its elements, just as secure cryptography relies on foundations in complexity theory and number theory, and safe airplanes are built on knowledge of aerodynamics and control. Motivated by this, I recently left OpenAI to cofound the Institute for Responsible Superintelligence (RESI[1]) with Shafi Goldwasser, inventor of zero-knowledge proofs, and Vinod Vaikuntanathan, a leading MIT cryptographer.
Currently, AI systems are assembled from learned models, data, and institutions whose properties and interactions are poorly understood. As AI systems become more capable, relying on trial and error becomes more dangerous. Scientific foundations can help AI safety mature beyond "fly-fix-fly" cycles in early aviation or attack-and-patch in early cryptography. In each of these fields, trial and error won early and principled design won later. We think AI safety is at that transition, which is what we mean by safe-by-design: guarantees that follow from how a system is built. While many properties of deep learning have resisted theoretical analysis, rigorous guarantees may still be shown for the surrounding architecture, protocol, and incentives. In analogy, cryptographic guarantees hold for adversarial behavior.
Who we are
The research team includes Ran Canetti (composable security, obfuscation), Yael Tauman Kalai (verifiable delegation, interactive proofs), Dylan Hadfield-Menell (alignment theory), Yannai Gonczarowski (mechanism design, algorithmic collusion), Daniela Rus (safe control and planning for embodied AI), Chloé Bakalar (AI ethics), Andrew Sutherland (AI-assisted mathematics, formal verification), Yaron Singer (AI security), Noam Kolt and Rebecca Wexler (law), Anat Perry (social cognition and human-AI interaction) and, most importantly, junior researchers Miranda Christ, Greg Gluch, Sam Gunn, Tal Herman, Jonathan Shafer, Mirac Suzgun, and Neekon Vafa, whose work spans watermarking, verification, learning theory, and hallucination. Planned visitors include Scott Aaronson, Boaz Barak, Nicholas Carlini, Geoffrey Irving, and Jacob Steinhardt (the full list is at resi.org).
While the research team brings many tools from outside AI, prior work includes undetectable backdoors in ML models, steganographic communication between AI agents, hallucination as an incentive problem, interactive proofs for scalable oversight, the off-switch game, and legal alignment.
RESI is a non-profit based in Cambridge, MA.
What we are working on
We aim to develop the theory and tools needed to reason about advanced AI before deployment: which safety properties are meaningful and achievable (alone and in combination), and which mechanisms deliver them, under which explicit assumptions. Safety here is broader than avoiding harm: ultimately it means systems whose benefits people can actually trust and share.
RESI is starting with several working groups, including numerous additional researchers beyond our core team, on various aspects of safety including:
Three problems that particularly excite me: (1) when do incentives continue to shape behavior as systems become more capable, and which systems don’t outgrow them (stickers motivate children, not adults); (2) how can we design cautious AI systems that act only when the outcomes, including those arising from many AIs acting together, are expected to be beneficial; and (3) how can we design training procedures that maintain a system's ability to be shut down, across a wide range of environments, capability growth, and self-modification.
AI is a program, not a person. It can be copied, rewound, encrypted, and sometimes formally reasoned about. The RESI team knows how to exploit such properties, defining and achieving reliability guarantees that may go beyond informal notions such as deception.
I believe highly capable AI should only be built with strong safety guarantees. Safety and capability trade off, and the challenge is to find the sweet spots: can a system be superhuman at cancer research yet provably limited elsewhere? RESI will study which safety properties are technically possible and how to achieve them, so that frontier labs and regulators can build on these findings.
Hasn’t this approach already been tried? A small number of pioneers have worked on theoretical alignment, including some of you. But developing a field requires a community: Turing made progress on cryptography, yet it took many more researchers to invent modern cryptography. We think AI safety is beginning to mature in a similar way, accelerated by AI itself, and we're collaborating with other like-minded organizations, like Resolution and ARC.
RESI's mission is to develop the scientific foundations of safe-by-design superintelligence. Outputs may include:
On a personal note, I've been surprised how some of my anxiety about AI has eased now that I'm doing my best to help.
I have my own views about promising directions, but I am more excited to see what our outstanding team discovers, working towards our mission. We will share progress at resi.org and look forward to working alongside many of you.
Thanks to Noam Kolt, Vinod Vaikuntanathan, Mirac Suzgun, and Chloé Bakalar for helpful comments.
The acronym is RESI because IRS was taken.