# Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On Page: https://stenobird.com/podcast/daily-paper-cast-7079649/oxygen-tryon-fashion-native-foundation-model-for-any-item-virtual-try-on Text version: https://stenobird.com/podcast/daily-paper-cast-7079649/oxygen-tryon-fashion-native-foundation-model-for-any-item-virtual-try-on.md Podcast: [Daily Paper Cast](https://stenobird.com/podcast/daily-paper-cast-7079649) Published: 2026-07-29T03:20:23+00:00 Episode link: https://share.transistor.fm/s/c4cd3577 Audio file: https://media.transistor.fm/c4cd3577/974c3359.mp3 Processing state: not_requested JSON: https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/oxygen-tryon-fashion-native-foundation-model-for-any-item-virtual-try-on Duration seconds: 1397 ## Resource 🤗 Upvotes: 22 | cs.CV Authors: Yong Liu, Xiaolong Fu, Zihang Xu, Wen Xue, Xueheng Li, Lin Song, Yuan Zhang, Chuyang Zhao, Haoyang Huang, Nan Duan, Yipeng Sun, Yan Li, Simiu Gu Title: Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On Arxiv: http://arxiv.org/abs/2607.21694v1 Abstract: We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-purpose image editor, Oxygen-TryOn is fashion-native, built for try-on through a dedicated data engine and try-on-specific training. Given one or more reference items (clean product shots or in-the-wild worn-on photos) and a single target subject image, it synthesizes a photorealistic image of the subject wearing the items across virtually any fashion category. Prior systems handle a single garment category in a studio setting, and recent multi-reference methods remain garment-centric; in contrast, Oxygen-TryOn supports diverse items and scenarios, including full- and half-body views, a variable number of references, and free multi-item composition, while faithfully preserving both subject identity and item appearance. Instead of mask-based inpainting, we reformulate try-on as a multi-reference, understanding-driven generation task. We build a data engine that collects, manufactures, annotates, and filters high-quality try-on data at scale, and design a three-stage recipe of continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL). The RL stage uses a hybrid reward combining an in-house try-on reward model with a proprietary, rubric-guided general-purpose model, jointly supervising fine-grained consistency and instruction-level quality. It also follows general editing instructions (e.g., pose changes) in the same pass.… ## Actions - request_transcript: `POST https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/oxygen-tryon-fashion-native-foundation-model-for-any-item-virtual-try-on/transcription-requests` — Idempotently request low-priority transcript generation for this episode. - read_markdown: `GET https://stenobird.com/podcast/daily-paper-cast-7079649/oxygen-tryon-fashion-native-foundation-model-for-any-item-virtual-try-on.md` — Read the agent-friendly Markdown representation of this episode resource. A page view does not enqueue transcription. Agents should invoke `request_transcript` explicitly when they need this episode processed. ## Transcript Full transcripts are not published on public pages unless there is a clear rights basis.