Episode

AI, Reasoning or Rambling?

Podcast
muckrAIkers
Published
Jul 14, 2025
Duration seconds
4268
Processing state
not_requested
Canonical source
https://kairos.fm/muckraikers/e015
Audio
https://op3.dev/e/media.transistor.fm/3ac97807/0cedd748.mp3
JSON
/v1/public/podcasts/muckraikers-7026051/episodes/ai-reasoning-or-rambling
Markdown
/podcast/muckraikers-7026051/ai-reasoning-or-rambling.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/muckraikers-7026051/episodes/ai-reasoning-or-rambling/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/muckraikers-7026051/ai-reasoning-or-rambling.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

In this episode, we redefine AI's "reasoning" as mere rambling , exposing the "illusion of thinking" and "Potemkin understanding" in current models. We contrast the classical definition of reasoning (requiring logic and consistency) with Big Tech's new version, which is a generic statement about information processing. We explain how Large Rambling Models generate extensive, often irrelevant, rambling traces that appear to improve benchmarks, largely due to best-of-N sampling and benchmark gaming. Words and definitions actually matter! Carelessness leads to misplaced investments and an overestimation of systems that are currently just surprisingly useful autocorrects. (00:00) - Intro (00:40) - OBB update and Meta's talent acquisition (03:09) - What are rambling models? (04:25) - Definitions and polarization (09:50) - Logic and consistency (17:00) - Why does this matter? (21:40) - More likely explanations (35:05) - The "illusion of thinking" and task complexity (39:07) - "Potemkin understanding" and surface-level recall (50:00) - Benchmark gaming and best-of-n sampling (55:40) - Costs and limitations (58:24) - Claude's anecdote and the Vending Bench (01:03:05) - Definitional switch and implications (01:10:18) - Outro Links Apple paper - The Illusion of Thinking ICML 2025 paper - Potemkin Understanding in Large Language Models Preprint - Large Language Monkeys: Scaling Inference Compute with Repeated Sampling Theoretical understanding Max M. Schlereth Manuscript - The limits of AGI part II Preprint - (How) Do Reasoning Models Reason? Preprint - A Little Depth Goes a Long Way: The Expressive Power of Log-Depth Transformers NeurIPS 2024 paper - How Far Can Transformers Reason? The Globality Barrier and Inductive Scratchpad Empirical explanations Preprint - How Do Large Lan…