Rauno Arike

Angles of attack for continual learning safety

This is the fourth post in the sequence Implications of Continual Learning for LLM Agents. Summary Continual learning is a capability that largely doesn’t exist yet in LLMs. We first want to acknowledge that this may make it difficult to identify tractable angles of attack for making CL safer: it...

Jun 1647

How might continual learning affect safety and alignment?

This is the third post in our sequence Implications of Continual Learning for LLM Agents. Summary We argue that continual learning (CL) has two major potential safety implications: it may enable changes to LLM goals and values after deployment, and it eliminates the last-mover advantage held by current safety interventions....

Jun 1359

What's Continual Learning, and Why Might We Expect To See It In Advanced LLM Agents?

by RohanS, Rauno Arike, Owen Terry, Achu Menon, Zhijing Jin, Francis Rhys Ward, and Seth Herd

This is the second post in the sequence Implications of Continual Learning for LLM Agents. Summary We say that an agent is a continual learner if it undergoes persistent updates during deployment. That’s more-or-less a binary criterion, but there are several other components to being good at continual learning that...

Jun 1228

Implications of Continual Learning for LLM Agents: Introduction

by RohanS, Rauno Arike, Owen Terry, Achu Menon, Zhijing Jin, Francis Rhys Ward, and Seth Herd

Many people think that continual learning (CL) is a key missing capability of LLM systems, and we think its development could have huge implications for the capabilities and safety of AI agents. Despite this, several important questions about CL remain underexplored: * What counts as continual learning? Through what pathways...

Jun 1248

Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models

by Anders Cairns Woodruff, Francis Rhys Ward, Dewi Gould, Rauno Arike, Jason R Brown, Jo Jiao, wlanderson, ariana_azarbal, harrymayne, Patrick Leask, Twm Stone, Josh Hills, Ida Caspary, and Shubhorup Biswas

(see full author list at the end) About a year ago, METR showed that the length of tasks frontier models can reliably complete doubles every few months. A related safety-relevant question is this: what length of tasks can models complete without any chain of thought (CoT)? We investigate in our...

Jun 10249

A List of Research Directions in Character Training

Thanks to Rohan Subramani, Ariana Azarbal, and Shubhorup Biswas for proposing some of the ideas and helping develop them during a sprint. Thanks to Rohan and Ariana for comments on a draft. Thanks also to Kei Nishimura-Gasparian, Paul Colognese, and Francis Rhys Ward for conversations that inspired some of the...

Mar 1947

[Paper] How does information access affect LLM monitors' ability to detect sabotage?

TL;DR We evaluate LLM monitors in three AI control environments: SHADE-Arena, MLE-Sabotage, and BigCodeBench-Sabotage. We find that monitors with access to less information often outperform monitors with access to the full sequence of reasoning blocks and tool calls, a phenomenon we call the less-is-more effect for automated monitors. We follow...

Feb 1126

Rauno Arike

Rauno Arike

Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models

Hidden Reasoning in LLMs: A Taxonomy

How might continual learning affect safety and alignment?

13 Arguments About a Transition to Neuralese AIs

Rauno Arike

Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models

Hidden Reasoning in LLMs: A Taxonomy

How might continual learning affect safety and alignment?

13 Arguments About a Transition to Neuralese AIs

Angles of attack for continual learning safety

How might continual learning affect safety and alignment?

What's Continual Learning, and Why Might We Expect To See It In Advanced LLM Agents?

Implications of Continual Learning for LLM Agents: Introduction

Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models

A List of Research Directions in Character Training

[Paper] How does information access affect LLM monitors' ability to detect sabotage?