Episode
🔬Scaling Past Informal AI - Carina Hong, Axiom Math
- Published
- Jun 3, 2026
- Duration seconds
- 5584
- Processing state
not_requested- Canonical source
- https://www.latent.space/p/axiom
Actions
POST https://stenobird.com/v1/public/podcasts/latent-space-ai-engineer/episodes/scaling-past-informal-ai-carina-hong-axiom-math/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/latent-space-ai-engineer/scaling-past-informal-ai-carina-hong-axiom-math.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
In 2025, seven-month-old startup Axiom solved all 12 of the problems Putnam exam (scoring 8/12 in the time limit) a prestigious undergraduate math exam. The 12/12 score is better than the top undergraduates (110/120) and the closest AI system that reported a result (DeepSeek 103/120), although it is unclear what the people and other systems would have scored with more time. Nonetheless, the Putnam exam is legendary for its difficulty, with the median score typically being 0 or 1 points. Taken by itself, this seems like a minor feather in the cap of AI; one of a long series of accomplishments by AI systems in elite competitions with humans, starting with Deep Blue beating Kasparov. Fast forward to mid-2026, and Claude Code is eating the world. In 2024 Anthropic’s bet on code and enterprise looked like a more pragmatic niche play vs. OpenAI’s better models and massive consume scale. Today, Amodei’s all in bet on acceleration via code (images and video be damned) seems prescient. Despite Anthropic’s growing momentum, however, Axiom CEO Carina Hong sees coding ability as a necessary but not sufficient milestone on the path to AGI. Code arguably pushes the jagged frontier to the point of super intelligence in some domains outside of coding , but there are surprising gaps (link) that Carina believes will bottleneck AI progress. (Stats on math benchmarks). The informal bottleneck “Verified AI” sounds like eating broccoli (footnote: I actually love broccoli, but then again, I also believe strongly in Test Driven Development, so ¯\ (ツ) /¯ ) and paying taxes, but to Axiom it means something very different. “Verification to me is about scaling brilliance, compounding brilliance,” Carina told us. It actually took a while for me to understand what she means by this. It sounded like…