What is the right way to do a ban on neuralese and/or training against CoTs?
1. Right now, the single most important load-bearing argument in alignment Risk Reports is that models still lack the covert capabilities necessary to do sophisticated sabotage attempts in such a way that they wouldn't get caught by auditing and/or monitoring. That argument currently depends heavily on CoT monitorability: current models still struggle to do very sophisticated sabotage attempts without discussing the details in their CoT, both due to the sabotage tasks being difficult enough that it's hard for them to do it without TTC, as well as that models are still just not that good at controlling the contents of their CoT.
2. In my view, that is the main reason to care about neuralese and training against CoTs: I think the extent to which those things are bad pretty much entirely flows through the extent to which they make it easier for models doing sophisticated sabotage attempts to hide them from auditing and monitoring that is looking at their CoTs. Furthermore, the arguments in favor of those things mostly say that they won't degrade our auditing and monitoring: maybe process-based oversight (which could include training against the CoT) will make models more monitorable, because they'll be more incentivized to lay out all their cognition in their CoTs; and maybe neuralese won't be a problem for auditing because we'll have good interpretability techniques that let us audit the neuralese—but if those things are true, we should be able to demonstrate that they are true by reference to auditing and monitoring evals.
3. Thus, here is a proposal that I think both proponents and skeptics of CoT monitorability should be able to agree on:
1. Labs should agree to only[1] internally deploy models that, even with good elicitation[2], still fail to complete sophisticated sabotage tasks[3] without being caught by a monitor, where the monitor is allowed to use whatever interpretability t