Episode

The Next Frontier of AI Video Is Control

Podcast
The a16z Show
Published
Sep 17, 2026
Duration seconds
2372
Processing state
not_requested
Canonical source
https://a16z.simplecast.com/episodes/the-next-frontier-of-ai-video-is-control-RmfmPBwf
Audio
https://mgln.ai/e/1344/afp-848985-injected.calisto.simplecastaudio.com/3f86df7b-51c6-4101-88a2-550dba782de8/episodes/3cafbcc2-e9d6-40ae-adcb-adee4b50432e/audio/128/default.mp3?aid=rss_feed&awCollectionId=3f86df7b-51c6-4101-88a2-550dba782de8&awEpisodeId=3cafbcc2-e9d6-40ae-adcb-adee4b50432e&feed=JGE3yC0V
JSON
/v1/public/podcasts/the-a16z-show-436525/episodes/the-next-frontier-of-ai-video-is-control
Markdown
/podcast/the-a16z-show-436525/the-next-frontier-of-ai-video-is-control.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/the-a16z-show-436525/episodes/the-next-frontier-of-ai-video-is-control/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/the-a16z-show-436525/the-next-frontier-of-ai-video-is-control.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

a16z General Partner Jennifer Li sits down with fal co-founder Gorkem Yurtseven and Head of Engineering Batuhan Taskaya to discuss what changes when generative video becomes fast enough to run in real time. They unpack the technical work behind H3 Max, fal’s post-trained version of MiniMax’s open-weight video model, and how combining model post-training with systems and hardware optimization significantly reduced generation time while maintaining quality. That speed has enabled experiments with continuous video, including streams that can remember previous scenes and respond to new directions while they’re running. They also discuss why the next challenge may be less about speed and more about control, from camera movement and lighting to characters, motion, and lip sync. And they explore what those capabilities could mean for professional creative workflows, where artists and studios need predictable tools rather than simply generating a video from a prompt.