DRAFT / PAPER SITE
Working draft · not yet published

The science of alignment mid-training

First Author, Second Author, Third Author, Fourth Author

Arcadia Alignment · UK AISI · * equal contribution

Main takeaways

Midtrainingcomes after pre-training and before fine-tuning seems underpowered with respect to the amount of compute it requires. Given available evidence, we are not convinced that it is a suitable technique for general-purpose alignmentgetting the model to do what you want.

It is plausible that this is a matter of scale.

Given the state of AI in Autumn 2026, we think there should be more open investigations into frontier alignment techniques.

1.Context

Decorative illustration
Placeholder illustration — swap for something on-topic, or delete this figure.

Models be doing bad shit

There is little public evidence or reproduction of the relevant hypotheses around alignment training

But it seems a lot of people are doing midtraining, so we wanted to investigate it

We designed a synthetic setting where we can test the relevant hypotheses around midtraining.

2.Results

Fine-tuning a charter-motivated model on 50 examples, one of which only a profit-driven model would produce, flips its evaluated motivation to profit-driven
Example figure — replace with your own. Swap the filename above for any image dropped into this folder.
  1. A small amount of conflicting finetuningtraining on examples data can override the motivation that was installed via midtraining, even when the midtraining got 1000x the number of tokens (not sure about numbers)
  2. If we midtrain on 7 rules but only elicit 4 via finetuning, the model does not seem to learn the other 3 (unsure about specific numbers). Maybe worth relating this to the anthropic constitution generalization ideas?
  3. Worked examples seem required in the midtraining in order for it to work.
  4. Midtraining seems to either fail to generalize or generalize strangely in many ways. For instance, RL-ing a model to adopt a midtrained motivation doesn't seem to work in our setting.

Close by naming the mechanism's effect: what capability, decision, or resource shifts as a result, and away from whom.

3.Implications

Placeholder — list two to four concrete interventions, ordered by who could plausibly act on them (an individual lab, a regulator, an industry body). For each, say what it addresses and what it does not. longer

longer

longer

Longer

Place holder

Testing

Testing

Testing