Episode

#335 Sriram Raghavan: Why IBM Is Betting Everything on Small AI Models

Podcast
Eye On A.I.
Published
Apr 19, 2026
Duration seconds
3621
Processing state
not_requested
Canonical source
https://aneyeonai.libsyn.com/335-sriram-raghavan-why-ibm-is-betting-everything-on-small-ai-models
Audio
https://dts.podtrac.com/redirect.mp3/pscrb.fm/rss/p/pdst.fm/e/arttrk.com/p/VI4CS/prfx.byspotify.com/e/mgln.ai/e/p58841/pscrb.fm/rss/p/traffic.libsyn.com/secure/aneyeonai/Eye_on_AI_340_-_IBM_-_Sri_Raghavan_audio_.mp3?dest-id=727317
JSON
/v1/public/podcasts/eye-on-ai/episodes/335-sriram-raghavan-why-ibm-is-betting-everything-on-small-ai-models
Markdown
/podcast/eye-on-ai/335-sriram-raghavan-why-ibm-is-betting-everything-on-small-ai-models.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/eye-on-ai/episodes/335-sriram-raghavan-why-ibm-is-betting-everything-on-small-ai-models/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/eye-on-ai/335-sriram-raghavan-why-ibm-is-betting-everything-on-small-ai-models.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

Why IBM Is Betting Everything on Small AI Models In this episode of Eye on AI, Craig Smith sits down with Sriram Raghavan, Vice President of AI at IBM Research, to explore one of the most important debates in enterprise AI right now. Do you actually need a massive model to get world class results? IBM's answer is no, and Sriram breaks down exactly why. Sriram explains why IBM chose to train its Granite models directly using reinforcement learning rather than distilling from larger models like most of the industry. The reason goes beyond performance. It comes down to data lineage, safety alignment, and a belief that small, efficient models are the only sustainable path for enterprises running AI across hybrid cloud environments. We get into the full technical stack behind that bet. How data quality has replaced model size as the real competitive advantage. Why parameter count is becoming the wrong metric entirely. How IBM's inference time scaling techniques allow an 8 billion parameter model to match the performance of GPT-4o and Claude 3.5 on code and math benchmarks. And why IBM is pioneering a new concept called Generative Computing, which treats AI models not as prompt receivers but as programmable computing elements with runtimes, modular LoRA adapters, and proper programming abstractions. Sriram also shares where IBM Research is headed next, including breakthroughs in continuous learning, agent orchestration, and making unstructured enterprise data actually usable at scale. Subscribe for more conversations with the people building the future of AI and emerging technology. Stay Updated: Craig Smith on X: https://x.com/craigss Eye on A.I. on X: https://x.com/EyeOn_AI (00:00) Why IBM Skips Distillation and Trains Small Models Directly (04:50) Did We Even Need Giant A…