Guide

IDM-VTON: How to Try It Free, Run It Locally or in ComfyUI, and What the Licence Allows

Checked 6 Oct 2026

IDM-VTON is an ECCV 2024 virtual try-on model by Yisol Choi et al. (KAIST, OMNIOUS.AI). Its official Hugging Face demo is free and running today, but the code and weights are CC BY-NC-SA 4.0, so commercial use is not allowed.

GitHub stars
5,214
HF downloads, all time
1.61M
Official free demo
Running
VRAM, ComfyUI node
≥16 GB

The authors publish no VRAM figure; 16 GB is the requirement stated by the TemryL ComfyUI node. Resolution is 768×1024.

IDM-VTON: How to Try It Free, Run It Locally or in ComfyUI, and What the Licence Allows

IDM-VTON is the open virtual try-on model most people meet first, because its free demo works and its Hugging Face page has passed 1.6 million downloads. This guide covers what it is, what its licence actually allows, and the three ways to run it: in the browser, on your own GPU and in ComfyUI. Every number comes from the GitHub API, the Hugging Face API, arXiv or the project's own README, read on 6 October 2026. Where the authors do not document something, VRAM in particular, we say so.

What is IDM-VTON?

IDM-VTON is a diffusion model for image-based virtual try-on. It takes a photo of a person and a photo of a garment and generates the person wearing the garment, keeping the garment's print, logo and texture. The name comes from the paper, Improving Diffusion Models for Authentic Virtual Try-on in the Wild.

Non-commercial Space running ECCV 2024 Last commit Mar 2025
Paper
arXiv 2403.05139, first posted 8 March 2024, accepted at ECCV 2024
Authors
Yisol Choi, Sangkyung Kwak, Kyungmin Lee, Hyungwon Choi, Jinwoo Shin
Affiliations
KAIST and OMNIOUS.AI (project page)
Architecture
Two UNets built on SDXL inpainting: TryonNet generates the image, GarmentNet encodes the garment's low-level detail, and an IP-Adapter carries its high-level semantics
Base model
stable-diffusion-xl-1.0-inpainting-0.1 + IP-Adapter (SDXL, ViT-H)
Resolution
768×1024 (inference scripts and demo)
Garments
Upper body (VITON-HD checkpoint); upper body, lower body and dresses (DressCode checkpoint)
Training data
VITON-HD and DressCode
Training code
Yes, train_xl.py
Library
diffusers (pinned to 0.25.0), Gradio 4.24.0 for the demo

On 6 October 2026 the GitHub repository had 5,214 stars and 822 forks. Its last commit was on 7 March 2025, so the code is stable rather than actively developed. The main Hugging Face checkpoint had 35,306 downloads in the previous 30 days and 1,610,976 in total, which makes it the most-downloaded open try-on model we track in our comparison of 19 open-source try-on models.

IDM-VTON licence: can you use it commercially?

No. The README says: “The codes and checkpoints in this repository are under the CC BY-NC-SA 4.0 license.” The Hugging Face model and Space carry the same tag. GitHub shows the licence as “Other” only because there is no standard LICENSE file.

LayerLicenceWhat it means for you
CodeCC BY-NC-SA 4.0Use and modify it for non-commercial work, credit the authors, share changes under the same licence
WeightsCC BY-NC-SA 4.0Same terms. No commercial product, no paid service
Base modelSDXL inpainting 0.1 + IP-AdapterPermissive on their own, but they do not lift the NC terms of IDM-VTON's weights
Training dataVITON-HD (CC BY-NC 4.0), DressCode (non-commercial academic)Even a retrained copy would have the same data question
ComfyUI nodeGPL-3.0 (TemryL)Covers the node's code only, not the model

People have asked how to get a commercial licence in the model's Hugging Face discussion since October 2024. When we checked, the thread was still open with no reply from the authors. Hosting it behind an API, as some third-party sites do, does not change the licence. This is a summary of what the licence text says, not legal advice.

How to try IDM-VTON free on Hugging Face

The official Space is huggingface.co/spaces/yisol/IDM-VTON. On 6 October 2026 the Hugging Face API reported it as RUNNING on ZeroGPU (A10G-class hardware), with 2,137 likes. ZeroGPU is free, but each account gets a daily GPU quota, per Hugging Face's ZeroGPU docs:

AccountDaily GPU quotaQueue priority
Not logged in2 minutesLow
Free account5 minutesMedium
PRO40 minutes (extensible)Highest
  1. Upload the person photo in the “Human” panel. It is an image editor, so you can paint a mask with the pen if you need to.
  2. Keep “Use auto-generated mask” ticked for tops. The demo's code builds this mask for the upper body only.
  3. Tick “Use auto-crop & resizing” if your photo is not already a 3:4 portrait. The demo crops to the person, resizes to 768×1024 and scales the result back.
  4. Upload the garment and type a short description, for example “Short Sleeve Round Neck T-shirts”, as the placeholder suggests.
  5. Press Try-on. Under Advanced Settings you can change denoising steps (20–40, default 30) and the seed (default 42).

The output is still under CC BY-NC-SA 4.0 terms. The demo is for evaluating the model, not for producing product images.

How to run IDM-VTON locally

Requirements. The repo's environment.yaml pins Python 3.10, PyTorch 2.0.1 with CUDA 11.8, diffusers 0.25.0, transformers 4.36.2, accelerate 0.25.0, Gradio 4.24.0 and onnxruntime 1.16.2, so in practice you need Linux and an NVIDIA GPU. The authors do not state a VRAM figure; the ComfyUI port asks for at least 16 GB.

1. Clone and create the environment (from the README):

git clone https://github.com/yisol/IDM-VTON.git
cd IDM-VTON
conda env create -f environment.yaml
conda activate idm

2. Download the preprocessing checkpoints from the Space's ckpt folder and place them like this:

ckpt
|-- densepose/model_final_162be9.pkl
|-- humanparsing/parsing_atr.onnx
|-- humanparsing/parsing_lip.onnx
|-- openpose/ckpts/body_pose_model.pth

3. Start the local Gradio demo:

python gradio_demo/app.py

4. Or run batch inference on VITON-HD, the dataset the paper reports on:

accelerate launch inference.py \
    --width 768 --height 1024 --num_inference_steps 30 \
    --output_dir "result" --unpaired \
    --data_dir "DATA_DIR" --seed 42 \
    --test_batch_size 2 --guidance_scale 2.0

For DressCode, use inference_dc.py with --category "upper_body", "lower_body" or "dresses". The dataset folders need image, image-densepose, agnostic-mask and cloth. The authors made their DensePose maps with detectron2 and share pre-computed DensePose images and garment captions for DressCode.

Training. The repo includes it: download the SDXL IP-Adapter (ip-adapter-plus_sdxl_vit-h.bin) and its image encoder from h94/IP-Adapter, then run accelerate launch train_xl.py --gradient_checkpointing --use_8bit_adam --output_dir=result --train_batch_size=6 --data_dir=DATA_DIR, or sh train_xl.sh.

Watch out: the stack is from early 2024. Newer PyTorch, diffusers or CUDA versions are not what the authors tested, so keep the pinned versions in their own conda environment. The VITON-HD and DressCode datasets have their own non-commercial licences and are downloaded separately.

IDM-VTON in ComfyUI

There is no official ComfyUI node. These are the community projects with real adoption, checked through the GitHub API on 6 October 2026:

ProjectWhat it is★Last commitLicence
TemryL/ComfyUI-IDM-VTONThe main ComfyUI node. In ComfyUI Manager; downloads the yisol/IDM-VTON weights; needs ≥16 GB VRAM (README)594Aug 2024GPL-3.0 (node)
MoJIeAIGC/IDMVTON_CNChinese-language one-click install package (not ComfyUI)117Sep 2025None stated
camenduru/IDM-VTON-jupyterJupyter / Colab notebook22Aug 2024None stated

Installing TemryL's node. In ComfyUI Manager, search for ComfyUI-IDM-VTON and check the author is TemryL. Or install it by hand:

cd custom_nodes
git clone https://github.com/TemryL/ComfyUI-IDM-VTON.git
cd ComfyUI-IDM-VTON
python install.py

The example workflow uses ComfyUI Segment Anything for the garment mask and ControlNet Auxiliary Preprocessors for DensePose, so install both. The node has not been updated since August 2024. If a current ComfyUI breaks it, pin an older ComfyUI in a separate install rather than patching your main one.

IDM-VTON quality tips and limits

These follow from how the code works, not from a benchmark we ran.

  • Feed it 3:4. Everything runs at 768×1024. A landscape or square photo gets squashed unless you use auto-crop.
  • Write the garment description. The demo turns it into the prompt “model is wearing …”, so “long sleeve striped cotton shirt” gives the model more to work with than an empty box.
  • Mask by hand for anything but tops. The demo's auto-mask is upper-body only. Lower body and dresses need the DressCode checkpoint locally, or a painted mask in the demo.
  • Try another seed. The demo fixes the seed at 42, so pressing Try-on again with the same inputs gives essentially the same image. Change the seed to get a different attempt.
  • Check the mask first. The demo shows the masked image next to the result. The mask and the DensePose map are both automatic guesses, so when an arm crosses the garment or the clothing is layered, a wrong result usually starts with a wrong mask.
  • One garment per pass. For a full outfit you run it twice: top, then bottom.

IDM-VTON vs CatVTON vs OOTDiffusion vs Leffa

The three open models people most often weigh against it, checked the same day. Full detail on all 19 models is in our open-source virtual try-on comparison.

ModelVenueLicenceFree official SpaceDocumented VRAMGitHub ★ComfyUI
IDM-VTONECCV 2024Non-commercialRunningNot by authors (≥16 GB for the ComfyUI node)5,214Community node, 594★
CatVTONICLR 2025Non-commercialRuntime error~8 GB at 1024×768, bf161,863Official workflow
OOTDiffusionAAAI 2025Non-commercialRunningNot documented6,604Community node, 471★
LeffaCVPR 2025Grey area (MIT, NC data)RunningNot documented1,681Community node, 241★

Swipe the table sideways for VRAM, stars and ComfyUI.

Pick CatVTON if your GPU has 8 GB. Pick OOTDiffusion if you want automatic masks for upper, lower and dress categories. Pick Leffa if you also need pose transfer from the same codebase. None of the four is cleared for commercial use. If you need that, the comparison covers the two options that are: FASHN VTON v1.5 (Apache-2.0) and Qwen-Image-Edit with an Apache-2.0 try-on LoRA.

When a hosted try-on tool is the better fit

IDM-VTON is a good way to see what try-on can do and a solid research baseline. It is the wrong tool for product pages, ads or marketplace listings, because its licence rules out commercial use and the demo's quota rules out volume.

Fashio's virtual try-on is a hosted tool for that job: upload a garment photo, pick a model, and you get on-model images you can use commercially. There is nothing to install and no mask to paint. A try-on costs 4 credits. Plans are Starter at $14 a month for 50 credits, Plus at $29.99 for 160 and Pro at $49.99 for 500, which works out at roughly $0.40 to $1.12 per try-on. Annual plans start with a 3-day free trial, and developers can use the same credits through the Fashio API.

IDM-VTON FAQ

What is IDM-VTON used for?

IDM-VTON is a virtual try-on model: you give it a photo of a person and a photo of a garment, and it generates the person wearing that garment. It was published at ECCV 2024 by Yisol Choi and co-authors from KAIST and OMNIOUS.AI. People use it to test try-on in the official Hugging Face demo, as a research baseline, and inside ComfyUI workflows.

Is IDM-VTON free to use?

It is free of charge, but not free for every purpose. The code and checkpoints are licensed CC BY-NC-SA 4.0, so you can use them at no cost for research and personal, non-commercial projects, with attribution, and anything you derive from them must carry the same licence. Commercial use is not allowed under that licence. The official Hugging Face Space is free to use within Hugging Face's daily ZeroGPU quota.

Can I use IDM-VTON commercially?

Not under its public licence. The README puts both the code and the checkpoints under CC BY-NC-SA 4.0, and the model was trained on VITON-HD and DressCode, which are also licensed for non-commercial research. Users have asked about a commercial licence on the model's Hugging Face discussion board since October 2024, and when we checked on 6 October 2026 there was no answer from the authors there. Running it through a third-party API does not change the licence.

How much VRAM does IDM-VTON need?

The authors do not publish a figure. The most-used ComfyUI node, TemryL's ComfyUI-IDM-VTON, says its implementation needs a GPU with at least 16 GB of VRAM. The official demo loads the models in float16 and runs at 768×1024.

Is the IDM-VTON Hugging Face demo free?

Yes. The official Space, yisol/IDM-VTON, was running on 6 October 2026 on ZeroGPU, Hugging Face's shared GPU pool. Each account gets a daily GPU quota: 2 minutes without logging in, 5 minutes with a free account and 40 minutes with PRO, according to Hugging Face's ZeroGPU documentation. When the quota runs out or the queue is long, the Space shows an error or makes you wait.

Does IDM-VTON work in ComfyUI?

Yes, through community custom nodes. The main one is TemryL/ComfyUI-IDM-VTON (594 GitHub stars, GPL-3.0 node licence, last commit August 2024). It is listed in ComfyUI Manager, downloads the weights from yisol/IDM-VTON, and uses Segment Anything for the mask and ControlNet Auxiliary Preprocessors for DensePose. The weights stay CC BY-NC-SA 4.0 whatever the node's licence says.

Can IDM-VTON try on pants and dresses?

The model can. There is a separate DressCode checkpoint (yisol/IDM-VTON-DC) and the DressCode inference script takes upper_body, lower_body or dresses as a category. The official demo, however, generates its automatic mask for the upper body only. For trousers, skirts or dresses in the demo, untick auto-masking and paint the mask yourself.

Is idmvton.com the official IDM-VTON site?

It is not linked from the official sources. The GitHub README links to the project page idm-vton.github.io, the arXiv paper, the Hugging Face Space and the Hugging Face model, and nothing else. Treat any other site offering IDM-VTON online as a third party, and remember that the non-commercial licence applies to the model whoever hosts it.

How we checked

On 6 October 2026 we read the repository through the GitHub REST API (stars, forks, last push), the model and Space through the Hugging Face API (licence tag, likes, downloads, Space status and hardware), and the paper through the arXiv export API. Requirements and commands are copied from the README and environment.yaml; demo behaviour comes from the Space's app.py. The ZeroGPU quotas are from Hugging Face's documentation. We did not benchmark output quality and we show no images generated by the model. Stars and Space status change, so we re-check monthly.

Share
Keep Reading

Ready to take control of
your fashion content?

Use Fashio AI to put your garments on models, without a photoshoot.

Plans from $14 a month · 4 credits per image
Fashion Model Walking