x

Eric Breck

Message

52

1y

Finding Features Causally Upstream of Refusal

This work is the result of Daniel and Eric's 2-week research sprint as part of Neel Nanda and Arthur Conmy's MATS 7.0 training phase. Andy was the TA during the research sprint. After the sprint, Daniel and Andy extended the experiments and wrote up the results. A notebook that contains...

Jan 14, 202555

Eric Breck

Subscribe

Message

52

1y

Eric Breck

Finding Features Causally Upstream of Refusal

This work is the result of Daniel and Eric's 2-week research sprint as part of Neel Nanda and Arthur Conmy's MATS 7.0 training phase. Andy was the TA during the research sprint. After the sprint, Daniel and Andy extended the experiments and wrote up the results. A notebook that contains...

Jan 14, 202555

Finding Features Causally Upstream of Refusal

Daniel Lee

Daniel Lee, Eric Breck, Andy Arditi

1y

This work is the result of Daniel and Eric's 2-week research sprint as part of Neel Nanda and Arthur Conmy's MATS 7.0 training phase. Andy was the TA during the research sprint. After the sprint, Daniel and Andy extended the experiments and wrote up the results. A notebook that contains all the analyses is available here.

Summary

Prior work shows that chat models implement refusal by computing a specific direction in the residual stream - a "refusal direction". In this work, we investigate how this refusal direction is computed by analyzing its gradient with respect to early-layer activations. This simple approach discovers interpretable features that are both causally upstream of refusal and contextually relevant... (read 3594 more words →)

6

55

LESSWRONG
LW

LESSWRONG
LW

Eric Breck

Eric Breck

Eric Breck

Finding Features Causally Upstream of Refusal

Eric Breck

Eric Breck

Eric Breck

Finding Features Causally Upstream of Refusal

Summary