# PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation Page: https://stenobird.com/podcast/daily-paper-cast-7079649/physisforcing-physics-reinforced-world-simulator-for-robotic-manipulation Text version: https://stenobird.com/podcast/daily-paper-cast-7079649/physisforcing-physics-reinforced-world-simulator-for-robotic-manipulation.md Podcast: [Daily Paper Cast](https://stenobird.com/podcast/daily-paper-cast-7079649) Published: 2026-06-30T03:45:16+00:00 Episode link: https://share.transistor.fm/s/8100e4d4 Audio file: https://media.transistor.fm/8100e4d4/7a1c51c7.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/physisforcing-physics-reinforced-world-simulator-for-robotic-manipulation Duration seconds: 1278 ## Resource 🤗 Upvotes: 42 | cs.CV, cs.AI, cs.RO Authors: Peiwen Zhang, Yufan Deng, Shangkun Sun, Juncheng Ma, Duomin Wang, Jonas Du, Zilin Pan, Ye Huang, Hao Liang, Songyan Huang, Ruihua Zhang, Enze Xie, Ming-Yu Liu, Daquan Zhou Title: PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation Arxiv: http://arxiv.org/abs/2606.28128v1 Abstract: Video generation models have emerged as a promising paradigm for embodied world simulation. However, both general-domain video generators and robot-specific data fine-tuned models can still produce physically implausible manipulations, including discontinuous motion trajectories and inconsistent robot-object interactions, which limits their reliability as world simulators. Through extensive experiments, we find that such physical instability mainly arises from two factors: deformation of moving objects and implausible spatio-temporal correlations among interacting entities, particularly during contact. Building on this observation, we propose PhysisForcing, a scalable training framework that strengthens physical consistency by focusing supervision on physics-informative regions through joint optimization of pixel-level and semantic-level features. The framework consists of a pixel-level trajectory alignment loss, which supervises DiT features using reference point trajectories, and a semantic-level relational alignment loss, which aligns DiT features with inter-region relations extracted from a frozen video understanding encoder. Extensive experiments on R-Bench, PAI-Bench, and EZS-Bench show that PhysisForcing consistently improves embodied video generation over strong baselines, improving the Wan2.2-I2V-A14B and Cosmos3-Nano base models on R-Bench by 22.3\% and 9.2\% (7.1\% and 3.7\% over vanilla finetuning), with the Cosmo… ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/physisforcing-physics-reinforced-world-simulator-for-robotic-manipulation/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/daily-paper-cast-7079649/physisforcing-physics-reinforced-world-simulator-for-robotic-manipulation.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.