Episode
4D Human-Scene Reconstruction from Low-Overlap Captures
- Podcast
- Daily Paper Cast
- Published
- Jul 15, 2026
- Duration seconds
- 1196
- Processing state
not_requested- Canonical source
- https://share.transistor.fm/s/c3142205
Actions
POST https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/4d-human-scene-reconstruction-from-low-overlap-captures/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/daily-paper-cast-7079649/4d-human-scene-reconstruction-from-low-overlap-captures.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
🤗 Upvotes: 41 | cs.CV Authors: Minhyuk Hwang, Sangmin Kim, Seunguk Do, Daneul Kim, Jaesik Park Title: 4D Human-Scene Reconstruction from Low-Overlap Captures Arxiv: http://arxiv.org/abs/2607.09125v1 Abstract: Existing volumetric capture of dynamic human performance achieves high fidelity with dense camera arrays. However, in real-world scenarios, only a handful of low-overlap cameras are available, which degrades the output quality and leaves large areas unobserved. Recent 4D reconstruction methods have focused on low-overlap settings, yet they still produce noticeable artifacts in under-observed regions. Video diffusion models have emerged as another option, but they show geometrically inconsistent results for humans. To address these limitations, we propose StudioRecon, a pipeline that reconstructs 4D human scenes from sparse, low-overlap cameras by decoupling background and humans. We densify background supervision by synthesizing hundreds of camera-controlled novel views with a video diffusion model. We also robustly initialize deformable Gaussian humans with cross-view identity association and triangulated multi-view keypoint fitting. Finally, our recursive enhancement module with motion-adaptive consistency injection harmonizes the composed output, thereby further avoiding remaining artifacts. We achieve state-of-the-art novel view synthesis across four real-world datasets and demonstrate applications such as novel trajectory rendering and human replacement.