# LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing Page: https://stenobird.com/podcast/daily-paper-cast-7079649/liveedit-towards-real-time-diffusion-based-streaming-video-editing Text version: https://stenobird.com/podcast/daily-paper-cast-7079649/liveedit-towards-real-time-diffusion-based-streaming-video-editing.md Podcast: [Daily Paper Cast](https://stenobird.com/podcast/daily-paper-cast-7079649) Published: 2026-07-01T04:26:00+00:00 Episode link: https://share.transistor.fm/s/f9ed41cb Audio file: https://media.transistor.fm/f9ed41cb/70405b84.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/liveedit-towards-real-time-diffusion-based-streaming-video-editing Duration seconds: 1398 ## Resource 🤗 Upvotes: 72 | cs.CV Authors: Xinyu Wang, Chongbo Zhao, Fangneng Zhan, Yue Ma Title: LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing Arxiv: http://arxiv.org/abs/2606.26740v1 Abstract: Streaming video editing has made rapid progress, yet practical deployment is still limited by two core issues: maintaining stable backgrounds and non-edited regions over time, and achieving the low latency required for real-time interactive scenarios. Meanwhile, recent streaming video generation methods are mostly developed for synthesis and cannot be directly applied to editing due to the strict preservation requirement and region-specific control. In this work, we present a novel streaming video editing framework that performs causal, frame-by-frame editing with strong content preservation and real-time responsiveness. Our key design is a three-stage distillation pipeline that progressively transfers editing capability from a powerful bidirectional foundation model to an efficient unidirectional streaming editor, enabling stable long-horizon edits without sacrificing visual fidelity. To further support real-time deployment, we introduce an AR-oriented mask cache that reuses region-related computation across frames, substantially reducing redundant processing and accelerating inference. Finally, we establish a dedicated benchmark for streaming video editing. Extensive evaluations demonstrate that our method achieves state-of-the-art visual quality among streaming baselines while drastically boosting inference speed to 12.66 FPS, making it suitable for interactive and augmented reality applications. ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/liveedit-towards-real-time-diffusion-based-streaming-video-editing/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/daily-paper-cast-7079649/liveedit-towards-real-time-diffusion-based-streaming-video-editing.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.