Research

Alignment Research

The Arcadia Alignment team is a group of researchers based in London. We are currently working on understanding how training can shape model behaviour, extending debate to fuzzy alignment tasks, and scaling alignment research with automation.


We work on conceptual and empirical alignment science: what does it mean for AI models to be aligned and can we produce techniques for ensuring this?

BG Image

Research Topics

Model Motivations

Two models can behave identically while pursuing different goals. We study the motivations a model develops, and how training shapes them.

Model Motivations

Two models can behave identically while pursuing different goals. We study the motivations a model develops, and how training shapes them.

Model Motivations

Two models can behave identically while pursuing different goals. We study the motivations a model develops, and how training shapes them.

Scalable Oversight

We are studying debate protocols and how to apply them towards aligning models we might not be able to control.

Scalable Oversight

We are studying debate protocols and how to apply them towards aligning models we might not be able to control.

Scalable Oversight

We are studying debate protocols and how to apply them towards aligning models we might not be able to control.

Automated Alignment

We are building tools for automating our alignment research and understanding how the models can help us help them.

Automated Alignment

We are building tools for automating our alignment research and understanding how the models can help us help them.

Automated Alignment

We are building tools for automating our alignment research and understanding how the models can help us help them.

Our research

Stress-Testing Alignment Midtraining

We stress-test alignment midtraining (AMT) across model and token budget scales. Our results suggest that midtraining cannot tackle the hard problems of AI alignment—namely distributional shift and reward underspecification in the presence of imperfect data.

Stress-Testing Alignment Midtraining

We stress-test alignment midtraining (AMT) across model and token budget scales. Our results suggest that midtraining cannot tackle the hard problems of AI alignment—namely distributional shift and reward underspecification in the presence of imperfect data.

Stress-Testing Alignment Midtraining

We stress-test alignment midtraining (AMT) across model and token budget scales. Our results suggest that midtraining cannot tackle the hard problems of AI alignment—namely distributional shift and reward underspecification in the presence of imperfect data.

scroll for more

scroll for more

scroll for more

Mailing list

Stay updated on our work