Episode
"Alignment Midtraining Cracks Under Pressure" by J Bostock, sidbaines, Daniel Tan, draganover, ma-rmartinez
- Published
- Sep 23, 2026
- Duration seconds
- 874
- Processing state
not_requested
Actions
POST https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/alignment-midtraining-cracks-under-pressure-by-j-bostock-sidbaines-daniel-tan-draganover-ma-rmartinez/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/lesswrong-curated-popular-5643401/alignment-midtraining-cracks-under-pressure-by-j-bostock-sidbaines-daniel-tan-draganover-ma-rmartinez.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
TL;DR We stress-test alignment midtraining (AMT) across model and token budget scales. Our results suggest that midtraining cannot tackle the hard problems of AI alignment—namely distributional shift and reward underspecification in the presence of imperfect data. For instance, we test whether midtrained motivations are robust to finetuning which elicits competing motivations. In our setting, 190 million tokens of midtrained motivations are overpowered by a relatively tiny amount (~50 th...