Josh Engels

How transparent is DiffusionGemma (and why it matters)

Authors: Joshua Engels*, Callum McDougall*, Bilal Chughtai*, Janos Kramar, Senthoran Rajamanoharan, Cindy Wu, Arthur Conmy, Asic Q Chen, Jean Tarbouriech, Min Ma, Brendan O'Donoghue+, João Gabriel Lopes de Oliveira+, Rohin Shah+, Neel Nanda+ *Primary Contributor +Advising Paper here: https://arxiv.org/abs/2606.20560 Overview In a recent collaboration between the GDM interpretability team and...

Jun 2041

Why Do Naive SFT Filters For Safety Properties Fail?

This is the fourth in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The third post can be found here. Since SFT is the cause for many safety relevant properties, a natural strategy is to filter out rollouts from...

Jun 1450

SFT Drives Gemini’s Safety Properties

This is the third in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The second post can be found here. In this short post, we describe a surprising finding: most safety relevant properties in Gemini seem to be caused...

Jun 1369

Building and evaluating model diffing agents

by bilalchughtai, Josh Engels, and Neel Nanda

This is the second in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The first post can be found here. TL;DR * It is possible to build extremely simple agents that reliably find interesting behavioural differences between distinct models....

Jun 1261

[paper] Training on Documents About Monitoring Leads to CoT Obfuscation

by Reilly Haskins, bilalchughtai, and Josh Engels

Authors: Reilly Haskins*, Bilal Chughtai**, Joshua Engels** * primary contributor ** advice and mentorship This is the updated version of our earlier preliminary results post, covering the final results from our paper. The paper extends our preliminary work to eight models, a harder agentic task, CoT controllability analysis, and RL...

May 2731

Test your best methods on our hard CoT interp tasks

by daria, Riya Tyagi, Josh Engels, and Neel Nanda

Authors: Daria Ivanova, Riya Tyagi, Josh Engels, Neel Nanda Daria and Riya are co-first authors. This work was done during Neel Nanda’s MATS 9.0. Claude helped write code and suggest edits for this post. Most of our tasks fall in 3 categories: predicting future actions, detecting the effect of an...

Mar 2658

Training on Documents About Monitoring Leads To CoT Obfuscation

by Reilly Haskins, bilalchughtai, and Josh Engels

Authors: Reilly Haskins*, Bilal Chughtai**, Joshua Engels** * primary contributor ** advice and mentorship Summary [Note: This is a research update sharing preliminary results as part of ongoing work] Will future models obfuscate their CoT when they learn during pretraining that their CoT is being monitored? We investigate this question...

Mar 1865

Josh Engels

Josh Engels

A Pragmatic Vision for Interpretability

Negative Results on Group SAEs

Steering RL Training: Benchmarking Interventions Against Reward Hacking

SFT Drives Gemini’s Safety Properties

Josh Engels

A Pragmatic Vision for Interpretability

Negative Results on Group SAEs

Steering RL Training: Benchmarking Interventions Against Reward Hacking

SFT Drives Gemini’s Safety Properties

How transparent is DiffusionGemma (and why it matters)

Why Do Naive SFT Filters For Safety Properties Fail?

SFT Drives Gemini’s Safety Properties

Building and evaluating model diffing agents

[paper] Training on Documents About Monitoring Leads to CoT Obfuscation

Test your best methods on our hard CoT interp tasks

Training on Documents About Monitoring Leads To CoT Obfuscation