x

LESSWRONG
LW

Anca Dragan — LessWrong

Anca Dragan

Anca Dragan

Message

213

2y

Anca Dragan

AGI Safety and Alignment at Google DeepMind: A Summary of Recent Work

We wanted to share a recap of our recent outputs with the AF community. Below, we fill in some details about what we have been working on, what motivated us to do it, and how we thought about its importance. We hope that this will help people build off things...

222Aug 20, 2024

Message

213 karma

Member for 2 years

AGI Safety and Alignment at Google DeepMind: A Summary of Recent Work

Rohin Shah, Seb Farquhar, Anca Dragan+ 0 more

Rohin Shah, Seb Farquhar, Anca Dragan

1y

We wanted to share a recap of our recent outputs with the AF community. Below, we fill in some details about what we have been working on, what motivated us to do it, and how we thought about its importance. We hope that this will help people build off things we have done and see how their work fits with ours.

Who are we?

We’re the main team at Google DeepMind working on technical approaches to existential risk from AI systems. Since our last post, we’ve evolved into the AGI Safety & Alignment team, which we think of as AGI Alignment (with subteams like mechanistic interpretability, scalable oversight, etc.), and Frontier Safety (working on the Frontier... (read 2668 more words →)

33

222