Episode
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
- Podcast
- Daily Paper Cast
- Published
- Jul 30, 2026
- Duration seconds
- 1157
- Processing state
not_requested- Canonical source
- https://share.transistor.fm/s/77f7d058
Actions
POST https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/hifi-umi-learning-deployable-manipulation-policies-from-high-fidelity-umi-data-alone/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/daily-paper-cast-7079649/hifi-umi-learning-deployable-manipulation-policies-from-high-fidelity-umi-data-alone.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
🤗 Upvotes: 137 | cs.RO, cs.CV, cs.LG Authors: Simple AI, :, Yuteng Wei, Jinming Ma, Jiawei Wang, Weitao Zhou, Yushen Zuo, Ke Rui, Minglei Li, Jinhao Zhang, Zhikang Pan, Xiang Wang, Haoran Jia, Huan Du, Zicheng Zeng, Jun Ma, Guiyu Qin, Di Zhang, Xiaofei Li Title: HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone Arxiv: http://arxiv.org/abs/2607.25895v1 Abstract: Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capture scales readily, and current practice uses the resulting data mainly for pre-training, adding a small real-robot "anchor" at post-training. We ask whether raising the fidelity of robot-free UMI data, rather than shrinking the real-robot fraction, can remove that anchor. We present HiFi-UMI, a portable UMI data-production system co-designed for trajectory accuracy, inter-gripper relative pose, synchronization, and field of view: head-mounted offline stereo-inertial SLAM, native rather than reconstructed relative pose, a shared microsecond GPIO trigger, and two wide-angle cameras per hand covering ~200 degrees. It reaches 3 mm workspace-local end-effector accuracy without external tracking infrastructure. Using this corpus, we demonstrate zero-robot post-training: a policy post-trained solely on HiFi-UMI demonstrations deploys directly on a real robot and matches in-domain teleoperation across three backbones spanning the vision-language-action and world-action-model families, with success-rate differences of -2.5, +3.1, and -0.6 percentage points on StarVLA-QwenPI, OpenPI-pi_0.5, and LingBot-VA; the strongest policy reaches 85% on a precision insertion task, even though the tele…