# Radically Better Reasoning: Elicit's Andreas Stuhlmüller & Jungwon Byun on World Models for Research Page: https://stenobird.com/podcast/the-cognitive-revolution/radically-better-reasoning-elicit-s-andreas-stuhlm-ller-jungwon-byun-on-world-models-for-research Text version: https://stenobird.com/podcast/the-cognitive-revolution/radically-better-reasoning-elicit-s-andreas-stuhlm-ller-jungwon-byun-on-world-models-for-research.md Podcast: ["The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis](https://stenobird.com/podcast/the-cognitive-revolution) Published: 2026-06-17T20:11:22+00:00 Episode link: https://www.cognitiverevolution.ai/radically-better-reasoning-elicit-s-andreas-stuhlmuller-jungwon-byun-on-world-models-for-research/ Audio file: https://pdst.fm/e/mgln.ai/e/1113/pscrb.fm/rss/p/traffic.megaphone.fm/RINTP9702631647.mp3 Processing state: processed JSON: https://stenobird.com/v1/public/podcasts/the-cognitive-revolution/episodes/radically-better-reasoning-elicit-s-andreas-stuhlm-ller-jungwon-byun-on-world-models-for-research Duration seconds: 6370 ## Resource Elicit founders discuss moving beyond simple LLM outputs toward structured 'world models' that enable verifiable scientific reasoning. They explore how domain-specific reasoning primitives and process supervision can prevent the opacity of modern frontier models. ## Highlights - Main idea: Using a Domain Specific Language (DSL) to define reasoning primitives allows frontier models to execute structured, guaranteed workflows - Practical takeaway: Implementing process supervision—rewarding the quality of steps rather than just the final answer—is essential for high-stakes decision support - Failure mode: Relying on 'neuralese' or uninterpretable chain-of-thought can lead to unverifiable claims in critical fields like toxicology or drug discovery - Technical approach: Developing 'world models' as heterogeneous, evolving representations of knowledge to enable causal and counterfactual analysis - Operational insight: Automating software engineering via systems like 'The Line' can enable high-velocity development, delivering dozens of code changes weekly ## Topics World Models, Process Supervision, Scientific Machine Learning, Reasoning Primitives, Life Sciences AI, Domain Specific Languages, Verifiable AI, Causal Inference ## Chapters - 1:00 — Reasoning Primitives and Microservices: How Elicit uses a DSL to create structured, verifiable reasoning workflows using discrete microservices. - 9:00 — The Importance of Process Supervision: Moving from simple output evaluation to inspecting the entire reasoning path to ensure reliability. - 17:00 — AI in Life Sciences: Applying evidence-based reasoning to drug target ranking and clinical research workflows. - 25:00 — Handling Conflicting Evidence: Strategies for evaluating claims when scientific literature presents contradictory results. - 33:00 — The Need for Intermediate Reasoning Layers: Discussing why complex tasks require specialized layers of reasoning beyond standard LLM prompting. - 41:00 — Verifiable Conclusions in Biology: The challenge of providing proofs and traceable evidence for low-level biological insights. - 49:00 — Building Explicit World Models: Moving away from massive context windows toward structured, interpretable representations of knowledge. ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/the-cognitive-revolution/episodes/radically-better-reasoning-elicit-s-andreas-stuhlm-ller-jungwon-byun-on-world-models-for-research/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/the-cognitive-revolution/radically-better-reasoning-elicit-s-andreas-stuhlm-ller-jungwon-byun-on-world-models-for-research.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.