Garments, outerwear, accessories, shoes, bags, and more; from a single item to free multi-item outfits, with natural layering and occlusion.
Oxygen AIGC Group & Joy Future Academy, JD
Fashion-Native FoundationModel for Any-Item Virtual Try-On
A unified model for photorealistic, multi-reference virtual try-on across any fashion items.
01 / Abstract
We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-purpose image editor through prompting, which tends to hallucinate garment details, drift on subject identity, and break down on fine-grained texture, Oxygen-TryOn is fashion-native: it is purpose-built for try-on, powered by a dedicated data engine and try-on-specific training. Given one or more reference items, provided either as clean product shots or as in-the-wild photos of someone already wearing them, together with a single target subject image, the model synthesizes a photorealistic image of that subject wearing the referenced items, spanning virtually any fashion category: clothing, outerwear, accessories, footwear, bags, and beyond. Most prior systems handle only a single garment category in a controlled studio setting, and even recent multi-reference methods remain garment-centric. In contrast, Oxygen-TryOn supports diverse items and real-world scenarios, including full-body and half-body views, a variable number of references, and free multi-item composition, while faithfully preserving both the subject’s identity and each reference item’s fine-grained appearance, from textures and prints to logos and structural silhouettes. Instead of casting try-on as mask-based inpainting, we reformulate it as a multi-reference, understanding-driven generation task: the model reasons about what each item is and how it should be worn, handling occlusion, deformation, and layering rather than merely filling a predefined region, which in turn yields stronger generalization to unseen items and compositions. To unlock this capability, we build a dedicated data engine that collects, manufactures, annotates, and filters high-quality try-on data at scale, and we design a three-stage training recipe comprising continued pre-training (CPT), large-scale supervised fine-tuning (SFT), and reinforcement learning (RL). The RL stage is driven by a hybrid reward mechanism that integrates an in-house try-on reward model with a proprietary, rubric-guided general-purpose model, jointly supervising fine-grained consistency and instruction-level quality. As a by-product, the same model still follows general editing instructions (e.g., pose changes) within a single pass, with no task switching. Across public benchmarks and our in-house Oxygen-TryOn Bench, Oxygen-TryOn achieves state-of-the-art consistency and realism on single-item try-on and leads on multi-item try-on, matching or surpassing both leading proprietary systems (e.g., Nano Banana Pro, GPT-Image-2, Seedream5 Lite) and the strongest open-source models (e.g., FLUX.2). To the best of our knowledge, it is the first system to deliver any-item, multi-reference try-on at this level of fidelity. We further detail the recipe behind the model, namely the data engine and the CPT–SFT–RL pipeline, to make our design choices transparent and reusable.
Clean product shots or in-the-wild worn-on photos; full- or half-body subjects with a variable number of references.
Both the subject’s identity and each referenced item’s appearance, preserved intact.
General instruction-based edits, such as pose changes, within the same generation pass, with no second model or pass.
Stylized 3D avatars, illustrated characters, statues, or posters — dressed while respecting the original style and geometry.
Best-in-class single-item consistency and realism, surpassing Nano Banana Pro, GPT-Image-2, Seedream5 Lite, and FLUX.2.
02 / Results














































