# HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone Page: https://stenobird.com/podcast/daily-paper-cast-7079649/hifi-umi-learning-deployable-manipulation-policies-from-high-fidelity-umi-data-alone Text version: https://stenobird.com/podcast/daily-paper-cast-7079649/hifi-umi-learning-deployable-manipulation-policies-from-high-fidelity-umi-data-alone.md Podcast: [Daily Paper Cast](https://stenobird.com/podcast/daily-paper-cast-7079649) Published: 2026-07-30T04:03:40+00:00 Episode link: https://share.transistor.fm/s/77f7d058 Audio file: https://media.transistor.fm/77f7d058/a8b29d89.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/hifi-umi-learning-deployable-manipulation-policies-from-high-fidelity-umi-data-alone Duration seconds: 1157 ## Resource 🤗 Upvotes: 137 | cs.RO, cs.CV, cs.LG Authors: Simple AI, :, Yuteng Wei, Jinming Ma, Jiawei Wang, Weitao Zhou, Yushen Zuo, Ke Rui, Minglei Li, Jinhao Zhang, Zhikang Pan, Xiang Wang, Haoran Jia, Huan Du, Zicheng Zeng, Jun Ma, Guiyu Qin, Di Zhang, Xiaofei Li Title: HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone Arxiv: http://arxiv.org/abs/2607.25895v1 Abstract: Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capture scales readily, and current practice uses the resulting data mainly for pre-training, adding a small real-robot "anchor" at post-training. We ask whether raising the fidelity of robot-free UMI data, rather than shrinking the real-robot fraction, can remove that anchor. We present HiFi-UMI, a portable UMI data-production system co-designed for trajectory accuracy, inter-gripper relative pose, synchronization, and field of view: head-mounted offline stereo-inertial SLAM, native rather than reconstructed relative pose, a shared microsecond GPIO trigger, and two wide-angle cameras per hand covering ~200 degrees. It reaches 3 mm workspace-local end-effector accuracy without external tracking infrastructure. Using this corpus, we demonstrate zero-robot post-training: a policy post-trained solely on HiFi-UMI demonstrations deploys directly on a real robot and matches in-domain teleoperation across three backbones spanning the vision-language-action and world-action-model families, with success-rate differences of -2.5, +3.1, and -0.6 percentage points on StarVLA-QwenPI, OpenPI-pi_0.5, and LingBot-VA; the strongest policy reaches 85% on a precision insertion task, even though the tele… ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/hifi-umi-learning-deployable-manipulation-policies-from-high-fidelity-umi-data-alone/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/daily-paper-cast-7079649/hifi-umi-learning-deployable-manipulation-policies-from-high-fidelity-umi-data-alone.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.