Episode

"Alignment Midtraining Cracks Under Pressure" by J Bostock, sidbaines, Daniel Tan, draganover, ma-rmartinez

Podcast
LessWrong (Curated & Popular)
Published
Sep 23, 2026
Duration seconds
874
Processing state
not_requested
Canonical source
https://www.buzzsprout.com/2037297/episodes/19854741-alignment-midtraining-cracks-under-pressure-by-j-bostock-sidbaines-daniel-tan-draganover-ma-rmartinez.mp3
Audio
https://www.buzzsprout.com/2037297/episodes/19854741-alignment-midtraining-cracks-under-pressure-by-j-bostock-sidbaines-daniel-tan-draganover-ma-rmartinez.mp3
JSON
/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/alignment-midtraining-cracks-under-pressure-by-j-bostock-sidbaines-daniel-tan-draganover-ma-rmartinez
Markdown
/podcast/lesswrong-curated-popular-5643401/alignment-midtraining-cracks-under-pressure-by-j-bostock-sidbaines-daniel-tan-draganover-ma-rmartinez.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/lesswrong-curated-popular-5643401/episodes/alignment-midtraining-cracks-under-pressure-by-j-bostock-sidbaines-daniel-tan-draganover-ma-rmartinez/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/lesswrong-curated-popular-5643401/alignment-midtraining-cracks-under-pressure-by-j-bostock-sidbaines-daniel-tan-draganover-ma-rmartinez.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

TL;DR We stress-test alignment midtraining (AMT) across model and token budget scales. Our results suggest that midtraining cannot tackle the hard problems of AI alignment—namely distributional shift and reward underspecification in the presence of imperfect data. For instance, we test whether midtrained motivations are robust to finetuning which elicits competing motivations. In our setting, 190 million tokens of midtrained motivations are overpowered by a relatively tiny amount (~50 th...