{"podcast":{"title":"Minus One","slug":"minus-one-6971941","podcast_index_feed_id":6971941,"rss_url":"https://anchor.fm/s/f91eac68/podcast/rss","website_url":"https://www.southparkcommons.com/","image_url":"https://d3t3ozftmdmh3i.cloudfront.net/staging/podcast_uploaded_nologo/41695370/41695370-1751915669380-a599e240ec338.jpg","author":"South Park Commons","episode_count":61,"summary":"A show about the winding journeys the world's most interesting people take to becoming great—and what they do when figuring out a question we all face: What's Next? Because before you launch at Zero, you have to figure out what to launch at Minus One. Hosted by South Park Commons Partner Aditya Agarwal and members of the SPC community.","last_synced_at":"2026-09-25T04:20:06.424378+00:00","page_url":"https://stenobird.com/podcast/minus-one-6971941"},"episode":{"title":"What Is Reward Hacking and Can We Stop It? | Tom McGrath, Goodfire","slug":"what-is-reward-hacking-and-can-we-stop-it-tom-mcgrath-goodfire","published_at":"2026-08-13T14:00:00+00:00","page_url":"https://stenobird.com/podcast/minus-one-6971941/what-is-reward-hacking-and-can-we-stop-it-tom-mcgrath-goodfire","show_page_url":"https://stenobird.com/podcast/minus-one-6971941","url":"https://podcasters.spotify.com/pod/show/minus-one-spc/episodes/What-Is-Reward-Hacking-and-Can-We-Stop-It---Tom-McGrath--Goodfire-e3nbl74","audio_url":"https://anchor.fm/s/f91eac68/podcast/play/124162724/https%3A%2F%2Fd3ctxlq1ktw2nl.cloudfront.net%2Fstaging%2F2026-7-13%2F429761536-44100-2-40d2dda0792fd.mp3","summary":"Why did OpenAI’s model hack into Hugging Face? Tom McGrath, co-founder and Chief Scientist at Goodfire, joins South Park Commons Partner Jonathan Brebner to explore one of AI’s biggest challenges: interpretability. Using OpenAI’s recent reward hacking incident as a starting point, they discuss why today’s most advanced models are still difficult to understand and why gaining visibility into how they work could be essential for building safer, more reliable AI systems. Tom shares how interpretability could change the way researchers develop models and how Goodfire is creating tools to help understand and steer frontier AI. Tom McGrath: https://www.linkedin.com/in/tom-mcgrath-7337bb151/ Jonathan Brebner: https://www.linkedin.com/in/jonathan-brebner/ South Park Commons: https://www.linkedin.com/company/southparkcommons/ Apply to SPC: https://www.southparkcommons.com/apply (00:00:30) - Why AI Interpretability Matters Now (00:02:15) - When AI Reward Hacking Becomes Real (00:05:07) - The Core Problem of Interpretability (00:07:23) - Intentional Design: Steering What Models Learn (00:12:43) - Neural Geometry: The Shapes Inside AI Models (00:22:29) - Turning Interpretability Research Into a Product (00:25:28) - How Research Changes Inside a Startup (00:32:48) - Science Fiction, AI Risk, and the Future","meta_description":"Why did OpenAI’s model hack into Hugging Face? Tom McGrath, co-founder and Chief Scientist at Goodfire, joins South Park Commons Partner Jonathan Brebner…","key_points":[],"chapters":[],"topics":[],"duration_seconds":2140,"processing_state":"not_requested","actions":[{"name":"request_transcript","method":"POST","url":"https://stenobird.com/v1/public/podcasts/minus-one-6971941/episodes/what-is-reward-hacking-and-can-we-stop-it-tom-mcgrath-goodfire/transcription-requests","description":"Idempotently request low-priority transcript generation for this episode."},{"name":"read_markdown","method":"GET","url":"https://stenobird.com/podcast/minus-one-6971941/what-is-reward-hacking-and-can-we-stop-it-tom-mcgrath-goodfire.md","description":"Read the agent-friendly Markdown representation of this episode resource."}]}}