# Mastering Reasoning LLMs: Decoding AI's Complex Problem-Solving Strategies Page: https://stenobird.com/podcast/smart-enterprises-ai-frontiers-7077810/mastering-reasoning-llms-decoding-ai-s-complex-problem-solving-strategies Text version: https://stenobird.com/podcast/smart-enterprises-ai-frontiers-7077810/mastering-reasoning-llms-decoding-ai-s-complex-problem-solving-strategies.md Podcast: [Smart Enterprises: AI Frontiers](https://stenobird.com/podcast/smart-enterprises-ai-frontiers-7077810) Published: 2025-07-29T23:16:52+00:00 Episode link: https://podcasters.spotify.com/pod/show/ali-mehedi0/episodes/Mastering-Reasoning-LLMs-Decoding-AIs-Complex-Problem-Solving-Strategies-e3672qo Audio file: https://anchor.fm/s/fcb30f54/podcast/play/106187032/https%3A%2F%2Fd3ctxlq1ktw2nl.cloudfront.net%2Fstaging%2F2025-6-29%2F404790272-44100-2-848f624d0493d.m4a Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/smart-enterprises-ai-frontiers-7077810/episodes/mastering-reasoning-llms-decoding-ai-s-complex-problem-solving-strategies Duration seconds: 2023 ## Resource Join us for an insightful exploration into the world of Reasoning LLMs , drawing on the expertise of Sebastian Raschka, PhD. This episode demystifies how Large Language Models (LLMs) are being refined to excel at complex tasks that require intermediate steps , such as solving puzzles, advanced mathematics, and challenging coding problems, moving beyond simple factual question-answering. We'll uncover the four main approaches currently used to build and improve these specialised reasoning capabilities : Inference-time scaling : Discover how techniques like Chain-of-Thought (CoT) prompting encourage LLMs to generate intermediate reasoning steps, mimicking a 'thought process' and often leading to more accurate results on more complex problems. This approach increases computational resources during inference, making it more expensive. Pure Reinforcement Learning (RL) : Learn about the surprising emergence of reasoning behaviour from pure reinforcement learning, as demonstrated by DeepSeek-R1-Zero. This model was trained exclusively with RL, without an initial supervised fine-tuning (SFT) stage, using accuracy and format rewards to develop basic reasoning skills. Supervised Fine-tuning (SFT) + Reinforcement Learning (RL) : Understand this key approach for building high-performance reasoning models , exemplified by DeepSeek's flagship R1 model. This method refines models with additional SFT stages and further RL training, building upon "cold-started" pure RL models. Pure SFT and Distillation : Explore how smaller, more efficient reasoning models can be created by instruction fine-tuning them on high-quality SFT data generated by larger, stronger LLMs. This approach is particularly attractive for creating models that are cheaper to run and can operate on lower-end h… ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/smart-enterprises-ai-frontiers-7077810/episodes/mastering-reasoning-llms-decoding-ai-s-complex-problem-solving-strategies/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/smart-enterprises-ai-frontiers-7077810/mastering-reasoning-llms-decoding-ai-s-complex-problem-solving-strategies.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.