# RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation Page: https://stenobird.com/podcast/daily-paper-cast-7079649/rynnworld-4d-4d-embodied-world-models-for-robotic-manipulation Text version: https://stenobird.com/podcast/daily-paper-cast-7079649/rynnworld-4d-4d-embodied-world-models-for-robotic-manipulation.md Podcast: [Daily Paper Cast](https://stenobird.com/podcast/daily-paper-cast-7079649) Published: 2026-07-09T04:18:50+00:00 Episode link: https://share.transistor.fm/s/74cad849 Audio file: https://media.transistor.fm/74cad849/a75883ba.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/rynnworld-4d-4d-embodied-world-models-for-robotic-manipulation Duration seconds: 1516 ## Resource 🤗 Upvotes: 74 | cs.RO Authors: Haoyu Zhao, Xingyue Zhao, Siteng Huang, Xin Li, Deli Zhao, Zhongyu Li Title: RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation Arxiv: http://arxiv.org/abs/2607.06559v1 Abstract: Robotic manipulation in the open world requires not only recognizing what a scene looks like, but also anticipating how its 3D structure moves under interaction. We argue that synchronized RGB, depth, and optical flow, namely RGB-DF, provide a physically grounded representation that captures the underlying 4D dynamics of a scene. Compared to 2D pixel videos, this multi-modal synergy aligns visual appearance with geometric structure and temporal motion, creating a representation space significantly closer to the low-level end-effector actions demanded by robotic systems, thereby narrowing the gap between world prediction and policy learning. Building on this insight, we introduce RynnWorld-4D, a generative model that co-produces future RGB frames, depth maps, and optical flow from a single RGB-D image and a language instruction within one unified diffusion process. This 4D world model features a tri-branch architecture that integrates cross-modal attention with frame-wise 3D RoPE, ensuring that appearance, geometry, and motion evolve consistently. To supply training data at scale, we curate Rynn4DDataset 1.0, a massive dataset of over 254.4 million frames across egocentric human and robotic manipulation videos with high-quality pseudo-labels for depth and optical flow. We further propose RynnWorld-4D-Policy, an inverse dynamics head that consumes the internal 4D representations of RynnWorld-4D in a single forward pass, bypassing expensive multi-step denoising, to output robot actions in a closed-loop manner. Experiments show that RynnWorld-4D produces… ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/rynnworld-4d-4d-embodied-world-models-for-robotic-manipulation/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/daily-paper-cast-7079649/rynnworld-4d-4d-embodied-world-models-for-robotic-manipulation.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.