x
prefilling an emergently misaligned model with reasoning traces that produced misaligned answers increases misalignment rates by ~8% but reasoning traces that produce misaligned answers aren't detectable through text monitoring — LessWrong