Episode

AI Is Learning at the Wrong Level of Abstraction — Matthieu Wyart

Podcast
Machine Learning Street Talk (MLST)
Published
Aug 10, 2026
Duration seconds
4736
Processing state
not_requested
Canonical source
https://podcasters.spotify.com/pod/show/machinelearningstreettalk/episodes/AI-Is-Learning-at-the-Wrong-Level-of-Abstraction--Matthieu-Wyart-e3n81u2
Audio
https://anchor.fm/s/1e4a0eac/podcast/play/124044674/https%3A%2F%2Fd3ctxlq1ktw2nl.cloudfront.net%2Fstaging%2F2026-7-10%2F429598212-44100-2-df596d8a748df.mp3
JSON
/v1/public/podcasts/machine-learning-street-talk/episodes/ai-is-learning-at-the-wrong-level-of-abstraction-matthieu-wyart
Markdown
/podcast/machine-learning-street-talk/ai-is-learning-at-the-wrong-level-of-abstraction-matthieu-wyart.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/machine-learning-street-talk/episodes/ai-is-learning-at-the-wrong-level-of-abstraction-matthieu-wyart/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/machine-learning-street-talk/ai-is-learning-at-the-wrong-level-of-abstraction-matthieu-wyart.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

This episode is sponsored by Notion. Learn more about Notion's Developer Platform today at https://notion.com/mlstWhy can deep networks discover abstractions that shallow models miss? Statistical physicist Matthieu Wyart joins Tim Scarfe to argue that the answer lies in the hidden hierarchy of data. Language and images are built from parts within parts; depth lets a network recover those coarse-grained variables and escape the curse of dimensionality.The conversation moves from jamming transitions and rough loss surfaces to Chomsky, context-free grammars and machine creativity. Wyart explains why next-token prediction can still recover compositional structure, where current systems fall short of genuine scientific invention, and why predicting latent representations rather than raw tokens could make learning far more sample-efficient.They also examine diffusion models, neural scaling laws and the limits of physics-inspired theory. The final question is on a personal note: if mistakes are the price of leaving the beaten path, how much scientific risk is worth taking?---TIMESTAMPS:00:00:00 Can machines learn abstractions from data?00:02:00 Notion agentic workspace00:02:49 From statistical physics to machine learning00:06:40 What physics can explain about learning00:16:37 From Carnot to Chomsky bulldozer00:21:21 How deep networks recover hidden hierarchies00:32:43 Where machine creativity still falls short00:40:48 How deep nets escape the curse of dimensionality00:52:19 Why predict latents instead of tokens01:02:49 The sample-efficiency case for latent prediction01:08:31 Diffusion, scaling laws and text entropy01:16:40 The scientists we learn from and the mistakes we make---REFERENCES:person:[00:00:43] Noam Chomskyhttps://linguistics.mit.edu/user/chomsky/tool:[00:02:08] N…