Model Raising

Models today are trained first and aligned later, often resulting in shallow alignment. Model Raising instead proposes to shape the assistant and its values from the start of pretraining, alongside its knowledge and capabilities.

Paper
Synthetic Persona Pretraining: Alignment from Token Zero
Installing the desired assistant persona during pretraining, not after it.
Paper
Tracing Persona Vectors Through LLM Pretraining
Persona vectors form within the first 0.22% of pretraining and survive alignment.
Position Paper
From Model Training to Model Raising
Why alignment must be woven into training from the start, not fine-tuned in afterwards.
Position Paper
The AI Alignment Paradox
If done without care, the better we align AI models, the easier we may make it to realign them.
Paper
Tandem Training for Language Models
Training powerful models so they remain understandable to smaller models or humans.
Editorial
A Narrowing Window to Understand AI
The time to build AI that we can understand and align is now.