Episode

Data Pyramid for Embodied Manipulation

Podcast
Daily Paper Cast
Published
Jul 29, 2026
Duration seconds
1391
Processing state
not_requested
Canonical source
https://share.transistor.fm/s/c02a0306
Audio
https://media.transistor.fm/c02a0306/e2b2523b.mp3
JSON
/v1/public/podcasts/daily-paper-cast-7079649/episodes/data-pyramid-for-embodied-manipulation
Markdown
/podcast/daily-paper-cast-7079649/data-pyramid-for-embodied-manipulation.md

Actions

  • POST https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/data-pyramid-for-embodied-manipulation/transcription-requests
    Idempotently request low-priority transcript generation for this episode.
  • GET https://stenobird.com/podcast/daily-paper-cast-7079649/data-pyramid-for-embodied-manipulation.md
    Read the agent-friendly Markdown representation of this episode resource.

Summary

🤗 Upvotes: 32 | cs.RO, cs.CV Authors: Yifan Ye, Yankai Fu, Yaoxu Lv, Bohan Hou, Jun Cen, Lingdong Kong, Duo Zheng, Tianxing Chen, Jiaming Liu, Ziang Cao, Yunfan Lou, Wei Chow, Xian Sun, Yingshuo Wang, Kuangzhi Ge, Xiaowei Chi, Xidong Zhang, Zhibo Pang, Yiwu Zhong, Sirui Han, Zhihe Lu, Weihao Yuan, Qifeng Chen, Michael Yu Wang, Yao Mu, Ziwei Liu, Jianfei Yang, Ping Luo, Shanghang Zhang Title: Data Pyramid for Embodied Manipulation Arxiv: http://arxiv.org/abs/2607.24744v1 Abstract: Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data ecosystem as a "pyramid" spanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data. We organize the pyramid around the tension between scalability and robot alignment, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelity. We then analyze recent embodied foundation models through the lens of their data recipes, examining how different sources are selected, aligned, and mixed during pretraining. For embodied brain models, vision-language-action models, and world-action models alike, we relate data composition to capabilities in perception, reasoning, planning, action generation, and world prediction. We close by discussing six open challenges: building large-scale tactile datasets, collecting failure and recovery data, developing scalable data-collection pipelines, aligning actions across embodiments, leveraging egocentr…