Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories
- URL: http://arxiv.org/abs/2603.14153v1
- Date: Sat, 14 Mar 2026 23:30:32 GMT
- Title: Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories
- Abstract summary: Garments2Look is the first large-scale multimodal dataset for outfit-level VTON.<n>It comprises 80K many-garments-to-one-look pairs across 40 major categories and 300+ fine-grained subcategories.<n>To balance authenticity and diversity, we propose a synthesis pipeline.
- Score: 27.58214524973654
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Virtual try-on (VTON) has advanced single-garment visualization, yet real-world fashion centers on full outfits with multiple garments, accessories, fine-grained categories, layering, and diverse styling, remaining beyond current VTON systems. Existing datasets are category-limited and lack outfit diversity. We introduce Garments2Look, the first large-scale multimodal dataset for outfit-level VTON, comprising 80K many-garments-to-one-look pairs across 40 major categories and 300+ fine-grained subcategories. Each pair includes an outfit with 3-12 reference garment images (Average 4.48), a model image wearing the outfit, and detailed item and try-on textual annotations. To balance authenticity and diversity, we propose a synthesis pipeline. It involves heuristically constructing outfit lists before generating try-on results, with the entire process subjected to strict automated filtering and human validation to ensure data quality. To probe task difficulty, we adapt SOTA VTON methods and general-purpose image editing models to establish baselines. Results show current methods struggle to try on complete outfits seamlessly and to infer correct layering and styling, leading to misalignment and artifacts.
Related papers
- Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On [50.93194436640413]
Oxygen-TryOn is a unified foundation model for any-item virtual try-on.<n>It is built for try-on through a dedicated data engine and try-on-specific training.
arXiv Detail & Related papers (2026-07-23T17:45:56Z) - FashionStylist: An Expert Knowledge-enhanced Multimodal Dataset for Fashion Understanding [14.029623403884179]
We introduce FashionStylist, an expert-annotated benchmark for holistic and expert-level fashion understanding.<n>FashionStylist provides professionally grounded annotations at both the item and outfit levels.<n>It supports three representative tasks: outfit-to-item grounding, outfit completion, and outfit evaluation.
arXiv Detail & Related papers (2026-04-10T12:03:55Z) - Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off [91.61470961531889]
We introduce the Dress Editing dataset (Dress-ED), the first large-scale benchmark that unifies VTON, VTOFF, and text-guided garment editing.<n>Dress-ED comprises over 146k verified quadruplets spanning three garment categories and seven edit types.
arXiv Detail & Related papers (2026-03-23T22:12:40Z) - MV-Fashion: Towards Enabling Virtual Try-On and Size Estimation with Multi-View Paired Data [33.49074848598509]
MV-Fashion is a large-scale, multi-view video dataset engineered for domain-specific fashion analysis.<n>It features 3,273 sequences from 80 diverse subjects wearing 3-10 outfits each.<n>A core contribution is a rich data representation that includes pixel-level semantic annotations.
arXiv Detail & Related papers (2026-03-09T09:28:15Z) - Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals [76.96387718150542]
We present Text-Enhanced MUlti-category Virtual Try-Off (TEMU-VTOFF)<n>Our architecture is designed to receive garment information from multiple modalities like images, text, and masks to work in a multi-category setting.<n> Experiments on VITON-HD and Dress Code datasets show that TEMU-VTOFF sets a new state-of-the-art on the VTOFF task.
arXiv Detail & Related papers (2025-05-27T11:47:51Z) - COutfitGAN: Learning to Synthesize Compatible Outfits Supervised by Silhouette Masks and Fashion Styles [23.301719420997927]
We propose the new task of generating complementary and compatible fashion items based on an arbitrary number of given fashion items.<n>In particular, given some fashion items that can make up an outfit, the aim of this paper is to synthesize photo-realistic images of other, complementary, fashion items that are compatible with the given ones.<n>To achieve this, we propose an outfit generation framework, referred to as COutfitGAN, which includes a pyramid style extractor, an outfit generator, a UNet-based real/fake discriminator, and a collocation discriminator.
arXiv Detail & Related papers (2025-02-12T03:32:28Z) - Learning to Synthesize Compatible Fashion Items Using Semantic Alignment and Collocation Classification: An Outfit Generation Framework [59.09707044733695]
We propose a novel outfit generation framework, i.e., OutfitGAN, with the aim of synthesizing an entire outfit.<n> OutfitGAN includes a semantic alignment module, which is responsible for characterizing the mapping correspondence between the existing fashion items and the synthesized ones.<n>In order to evaluate the performance of our proposed models, we built a large-scale dataset consisting of 20,000 fashion outfits.
arXiv Detail & Related papers (2025-02-05T12:13:53Z) - IMAGDressing-v1: Customizable Virtual Dressing [58.44155202253754]
IMAGDressing-v1 is a virtual dressing task that generates freely editable human images with fixed garments and optional conditions.
IMAGDressing-v1 incorporates a garment UNet that captures semantic features from CLIP and texture features from VAE.
We present a hybrid attention module, including a frozen self-attention and a trainable cross-attention, to integrate garment features from the garment UNet into a frozen denoising UNet.
arXiv Detail & Related papers (2024-07-17T16:26:30Z) - MV-VTON: Multi-View Virtual Try-On with Diffusion Models [91.71150387151042]
The goal of image-based virtual try-on is to generate an image of the target person naturally wearing the given clothing.<n>Existing methods solely focus on the frontal try-on using the frontal clothing.<n>We introduce Multi-View Virtual Try-ON (MV-VTON), which aims to reconstruct the dressing results from multiple views using the given clothes.
arXiv Detail & Related papers (2024-04-26T12:27:57Z) - VICTOR: Visual Incompatibility Detection with Transformers and
Fashion-specific contrastive pre-training [18.753508811614644]
Visual InCompatibility TransfORmer (VICTOR) is optimized for two tasks: 1) overall compatibility as regression and 2) the detection of mismatching items.
We build upon the Polyvore outfit benchmark to generate partially mismatching outfits, creating a new dataset termed Polyvore-MISFITs.
A series of ablation and comparative analyses show that the proposed architecture can compete and even surpass the current state-of-the-art on Polyvore datasets.
arXiv Detail & Related papers (2022-07-27T11:18:55Z) - Semi-Supervised Visual Representation Learning for Fashion Compatibility [17.893627646979038]
We propose a semi-supervised learning approach to create pseudo-positive and pseudo-negative outfits on the fly during training.
For each labeled outfit in a training batch, we obtain a pseudo-outfit by matching each item in the labeled outfit with unlabeled items.
We conduct extensive experiments on Polyvore, Polyvore-D and our newly created large-scale Fashion Outfits datasets.
arXiv Detail & Related papers (2021-09-16T15:35:38Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.