Careers

Researcher — Alignment Team

Rolling applications

Arcadia Impact’s alignment team is an AI safety research group in London. We have a generous compute allowance and aim to produce large-scale, open investigations of frontier alignment techniques. Since starting four months ago, we’ve realised that we have the capacity to expand the scope of our work. We are therefore excited about hard-working, conscientious, and agentic people joining our team.

This role is for joining our team directly as a researcher; it is intended for people already doing full-time work in AI safety research.

If you are currently finishing a safety fellowship (MATS, LASR, Astra, Anthropic fellows, Pivotal, etc.) it might make more sense to apply for our fellowship position. If you have questions about which role is the right fit for you, please reach out to andrew@arcadiaimpact.org.

We will be assessing applications on a rolling basis.

Current research topics

Although the labs have a wealth of knowledge regarding prosaic alignment techniques, the third-party ecosystem only sees glimpses of how these work [1, 2, 3, 4, 5]. Our researchers are trying to distill knowledge around empirical alignment and verify when these approaches work and fail. Our work is currently centered on three research threads:

  1. Model motivations. Labs shape their models' motivational structure through methods like constitutional training, deliberative alignment, and perhaps other techniques we don’t know about. We think it would be good to reproduce these methods at scale and test the assumptions around them. We put out some preliminary work on model organisms, and have recently published a larger project stress-testing midtraining.

  2. Empirics of scalable oversight. We’re trying to understand how to effectively oversee models on domains which rely on intuition and heuristics. To this end, we’ve put out a blog post on debate in verifiable settings and have recently published another blog post where we transfer these learnings to fuzzy tasks. After all, these are the settings where it is most important to have effective scalable oversight for AGI/ASI-level models.

  3. Understanding automated alignment and RSI dynamics. Models will increasingly do the alignment research which gets fed back into successors. On top of this, they will generate and filter the data for themselves to train on. This likely makes misalignment easier to introduce and harder to notice. We try to understand the consequences and failure modes of automated alignment research by automating much of our own work and analyzing the conditions which lead these runs to go off the rails (see our blogpost on this topic).

How we work

Below, is a set of non-exhaustive principles which are representative of our research approach:

  • It is important to iterate fast and stay at the frontier; but it is also important to not move so fast that one sacrifices rigor. We think this is a difficult balance to strike and are constantly re-assessing our posture here.

  • One of our goals is to understand the consequences of widespread automation of alignment research. It seems likely that, with each passing month, the modal AI safety researcher will spend more of their time interpreting work that the models have done for them. Within this spirit, we lean into automation ourselves so that our threat modeling is grounded in realistic examples.

  • Now more than ever, research needs to serve a clear purpose and come with well-thought-out routes to impact. It seems that the work being excellent is a necessary but insufficient condition for the work to be valuable.

  • Part of the job is running experiments at the scale where the results are relevant for frontier models. In practice, this means our projects aim to be applicable to large scales fairly early in the research process: generating synthetic corpora of hundreds of millions of tokens, training and evaluating 100B+ parameter models, and running multi-agent automated research jobs for days.

Who we are looking for

In general, we are looking to hire people who can be responsible stewards of good ideas. Below is a list of qualities we think represent this ideal:

  • Constantly trying to understand what’s going on with frontier AI developments

  • Having principled theories about where misalignment comes from and what it means to avoid it

  • Being thoughtful about the consequences of your research

  • Aware of the urgency of the current moment

  • Conscientious and easy to work with

In terms of formal qualifications, we are open to a wide variety of backgrounds. Our primary goal is to hire researchers who can own an impactful AI alignment agenda. The primary requirements are:

  • Being able to iterate fast and effectively deploy many agents under you;

  • Attention to detail and rigor when performing experiments;

  • The ability to effectively network and position your work within the current needs of the AI safety community;

A PhD is common on our team but not required.

The role

You’d join as a technical researcher. At the start, you'll work closely with a few members of the team on projects that are already underway. Over the course of your first 2-3 months, the plan is for you to transition towards owning your own workstream and research agenda.

Logistics and process

  • Location: In-person at LISA, London. We are in what was previously the Bluedot room (the nicest room in LISA 😎)

  • We will sponsor visas if you’re moving to the UK for this role.

  • Compensation: Full-time researcher salaries start at £120K. We also include benefits such as help moving to London, a yearly £5K productivity bonus (coaching, conferences, etc.) and standard UK benefits.

  • Process: Application → work test (~4 hours; we’ll pay £200 for the work + cover up to £200 for your compute) → interview(s) → 1-week paid work trial → full-time offer.

    If you have any questions, please reach out to moritz@arcadiaimpact.org

The team

We are an equal opportunity employer and are eager to consider candidates from a variety of backgrounds.