CatVTON is the open virtual try-on model people reach for when they have a normal graphics card: the authors document about 8 GB of VRAM at 1024×768, and they ship their own ComfyUI workflow. This guide covers what it is, what its licence allows, why the official Hugging Face demo is down, and how to run it on your own GPU or in ComfyUI. Every number comes from the GitHub API, the Hugging Face API, arXiv or the project's own README and code, read on 7 October 2026. Where the authors do not document something, we say so.
- Just want to try itThe official Space is down today. Run it locally, or try IDM-VTON's Space, which is running
- Have an 8 GB NVIDIA GPURun it locally: conda, Python 3.9, PyTorch 2.4.0, one command
- Work in ComfyUIThe official ComfyUI release in the repo, or chflame163's wrapper
- Need images you can sell withNot CatVTON: it is non-commercial. See the commercial options
What is CatVTON?
CatVTON is a diffusion model for image-based virtual try-on. You give it a photo of a person and a photo of a garment, and it generates the person wearing the garment. The name comes from the paper, CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models, and describes the trick: instead of adding a second network to encode the garment, as IDM-VTON and OOTDiffusion do, it concatenates the person and garment images and feeds them through one Stable Diffusion 1.5 inpainting UNet. That keeps the model small and the training cheap.
- Paper
- arXiv 2407.15886, first posted 21 July 2024, v2 February 2025, “Accepted by ICLR 2025”
- Authors
- Zheng Chong, Xiao Dong, Haoxiang Li, Shiyue Zhang, Wenqing Zhang, Xujie Zhang, Hanqing Zhao, Dongmei Jiang, Xiaodan Liang
- Affiliations
- Sun Yat-Sen University, Pixocial Technology, Peng Cheng Laboratory, SIAT (project page)
- Architecture
- A single SD 1.5 inpainting UNet with person and garment concatenated; only the attention layers are trained
- Size
- 899.06M parameters in total, 49.57M trainable
- VRAM
- About 8 GB at 1024×768 in bf16 (README)
- Resolution
- 768×1024 in the demo app; checkpoints trained at 512 (VITON-HD, DressCode) and 1024 (mixed data)
- Garments
- Upper, lower and overall (full-body / dresses)
- Masks
- Automatic, from DensePose and SCHP human parsing, or painted by hand. A separate mask-free checkpoint exists
- Library
- Modified from diffusers; Gradio 4.41.0 for the local app
On 7 October 2026 the GitHub repository had 1,865 stars, 239 forks and 72 open issues. Its last commit was on 16 December 2025, a merged fix to the huggingface_hub requirement, so the project is maintained but no longer developed. Note that the default branch is called edited, not main: the main branch holds an older README. The main Hugging Face checkpoint had 5,948 downloads in the previous 30 days and 61,718 in total. For context, IDM-VTON's has passed 1.6 million; our comparison of 19 open-source try-on models has the full picture.
CatVTON licence: can you use it commercially?
No. The README says: “All the materials, including code, checkpoints, and demo, are made available under the Creative Commons BY-NC-SA 4.0 license.” The repository's LICENSE file is the full text of that licence (GitHub shows it as “Other” because it is not a software licence), and the Hugging Face model and Space carry the same cc-by-nc-sa-4.0 tag. The demo app's own header says: “This demo and our weights are only for Non-commercial Use.”
| Layer | Licence | What it means for you |
|---|---|---|
| Code | CC BY-NC-SA 4.0 | Use and modify it for non-commercial work, credit the authors, share changes under the same licence |
| Weights (zhengchong/CatVTON) | CC BY-NC-SA 4.0 | Same terms. No commercial product, no paid service |
| Mask-free weights | Research-only terms in the model card | Non-commercial scientific research only; no commercial use of the model or its generated content; no sharing the model |
| FLUX LoRA | CC BY-NC-SA 4.0, plus FLUX.1 [dev] Non-Commercial License | Needs FLUX.1-Fill-dev, which is gated and non-commercial itself |
| Base model | SD 1.5 inpainting (CreativeML OpenRAIL-M) | Permissive on its own, but it does not lift the NC terms of CatVTON's weights |
| Training data | VITON-HD (CC BY-NC 4.0), DressCode (non-commercial academic) | Even a retrained copy would have the same data question |
Wrapping it in ComfyUI or putting it behind an API does not change any of this. A node's licence covers the node's code, not the weights it loads. This is a summary of what the licence texts say, not legal advice.
CatVTON on Hugging Face: why the demo is down
The official Space is zhengchong/CatVTON on Hugging Face. On 7 October 2026 the Hugging Face API reported it as RUNTIME_ERROR, the same as the day before. The Space requests ZeroGPU hardware and was last updated on 26 May 2026 (Gradio 5.49.1).
The error log explains it. The app crashes at start-up in FluxTryOnPipeline.from_pretrained, while loading black-forest-labs/FLUX.1-Fill-dev for its CatVTON-FLUX tab: “Cannot access gated repo … Access to model black-forest-labs/FLUX.1-Fill-dev is restricted.” Because the FLUX model fails, the whole app fails, including the Stable Diffusion 1.5 version that has nothing to do with FLUX. Duplicates do not escape it: Ffnjjjch/CatVTON, the copy that ranks first in Google for “catvton hugging face”, also showed RUNTIME_ERROR.
| Where | Status on 7 Oct 2026 | Note |
|---|---|---|
| zhengchong/CatVTON (official) | Runtime error | Fails loading gated FLUX.1-Fill-dev |
| Ffnjjjch/CatVTON (duplicate) | Runtime error | Same code |
| 120.76.142.206:8888 (authors' own server) | No response | The README said in November 2024 it would be taken offline |
| xiaozaa/catvton-flux-try-on (community FLUX version) | Runtime error | Not by the CatVTON authors |
The weights on Hugging Face are fine. zhengchong/CatVTON holds the attention checkpoints in three folders, vitonhd-16k-512, dresscode-16k-512 and mix-48k-1024 (the one the demo app loads), plus DensePose and SCHP for automatic masks and flux-lora for the FLUX version. The local app downloads all of it for you.
How to run CatVTON locally
Requirements. The README asks for a conda environment with Python 3.9.0. The current requirements.txt pins PyTorch 2.4.0, torchvision 0.19.0, accelerate 0.31.0, transformers 4.46.3 and Gradio 4.41.0, installs diffusers from GitHub, and needs peft 0.17 or newer and huggingface_hub 0.34 or newer. The app runs on cuda, so you need an NVIDIA GPU; about 8 GB of VRAM is enough in bf16.
1. Clone and create the environment (from the README):
git clone https://github.com/Zheng-Chong/CatVTON.git
cd CatVTON
conda create -n catvton python==3.9.0
conda activate catvton
pip install -r requirements.txt
2. Start the local Gradio app. Checkpoints download automatically from Hugging Face on the first run:
CUDA_VISIBLE_DEVICES=0 python app.py \
--output_dir="resource/demo/output" \
--mixed_precision="bf16" \
--allow_tf32
The README: “When using bf16 precision, generating results with a resolution of 1024x768 only requires about 8G VRAM.” The app loads the base model from booksforcharlie/stable-diffusion-inpainting, a copy, because a comment in app.py notes that RunwayML deleted the original repository. The repo also contains app_flux.py for the FLUX version and app_p2p.py, which uses the mask-free checkpoint.
3. Or run batch inference on VITON-HD or DressCode, the datasets the paper reports on:
CUDA_VISIBLE_DEVICES=0 python inference.py \
--dataset [dresscode | vitonhd] \
--data_root_path <path> \
--output_dir <path> \
--dataloader_num_workers 8 \
--batch_size 8 \
--seed 555 \
--mixed_precision [no | fp16 | bf16] \
--allow_tf32 \
--repaint \
--eval_pair
For DressCode, the repo includes preprocess_agnostic_mask.py to build the agnostic masks first, and eval.py computes the paper's metrics on your results. Both datasets are downloaded separately and have their own non-commercial licences.
Watch out: clone the default edited branch. Older guides tell you to cd CatVTON-main and follow an INSTALL.md that is not on the default branch. Detectron2 and DensePose are now bundled in the repo, so you do not install them separately.
CatVTON in ComfyUI
CatVTON is one of the few research try-on models with an official ComfyUI integration. The authors released it on 27 July 2024 as a separate download, because the code structure did not fit the main repository.
- Install the requirements for both CatVTON and ComfyUI in the same Python environment.
- Download ComfyUI-CatVTON.zip from the ComfyUI release and unzip it into ComfyUI's
custom_nodesfolder. - Start ComfyUI.
- Drag catvton_workflow.json onto the canvas. On the first run the weights download automatically, which the README says “usually” takes dozens of minutes.
Windows users are pointed to issue #8. The official workflow, like the app, uses DensePose and SCHP to make the mask automatically. Community nodes, checked through the GitHub API on 7 October 2026:
| Project | What it is | ★ | Last commit | Licence |
|---|---|---|---|---|
| chflame163/ComfyUI_CatVTON_Wrapper | The most-starred wrapper. Fixes cropping for odd aspect ratios, adds mask_grow; recommends ≥6 GB VRAM; models from Baidu or Google Drive into ComfyUI/models/CatVTON | 371 | Jan 2025 | None stated |
| pzc163/Comfyui-CatVTON | A modified copy of the official node | 174 | Oct 2024 | None stated |
| lujiazho/ComfyUI-CatvtonFluxWrapper | Wraps the community catvton-flux model; needs a Hugging Face token with FLUX.1-Fill-dev access | 93 | Dec 2024 | None stated |
Installing chflame163's wrapper by hand:
cd ComfyUI/custom_nodes
git clone https://github.com/chflame163/ComfyUI_CatVTON_Wrapper.git
cd ComfyUI_CatVTON_Wrapper
pip install -r requirements.txt
Use the Python that runs your ComfyUI (for the portable build, python_embeded\python.exe -s -m pip install -r requirements.txt). Its node takes the image, a mask, a reference garment image, mask_grow, precision, seed, steps and CFG. If the garment style comes out wrong, the README suggests adjusting mask_grow. None of these nodes has been updated since January 2025, so if a current ComfyUI breaks one, keep an older ComfyUI install for it.
CatVTON-FLUX, the mask-free model and CatV2TON
Searches for “catvton flux” mix up two different projects, so here is what each one is.
| Variant | By | What it is | Licence |
|---|---|---|---|
| CatVTON-FLUX (official) | CatVTON authors, Dec 2024 | A 37.4M LoRA for FLUX.1-Fill-dev, in the flux-lora folder; run with app_flux.py. The README called the app code “not a stable version” | CC BY-NC-SA 4.0 + FLUX.1 [dev] non-commercial |
| catvton-flux (community) | nftblackmagic / xiaozaa | A separate try-on model that combines CatVTON's approach with the FLUX Fill model; 614★, last commit Mar 2025. Its full fine-tune needs 2×H100 80 GB (README) | Code MIT; weights CC BY-NC 2.0 / CC BY-NC-SA 2.0 |
| CatVTON-MaskFree | CatVTON authors, Oct 2024 | Try-on without a mask, built on InstructPix2Pix in app_p2p.py. The card asks for name, affiliation and an institutional email | Research-only terms; no sharing |
| CatV2TON | Zheng Chong et al., Feb 2025 | DiT-based image and video try-on with temporal concatenation (arXiv 2501.11325); 256 and 512 weights; 245★ | CC BY-NC-ND 4.0 |
None of the four is cleared for commercial use. CatV2TON is the strictest: the ND (“no derivatives”) term means you may not distribute modified versions at all.
CatVTON quality tips and limits
These come from the demo app's own hints and code, not from a benchmark we ran.
- Pick the right cloth type. The app offers upper, lower and overall and builds the mask from that choice. A dress with “upper” selected only replaces the top.
- Paint the mask when the auto-mask is wrong. A mask drawn with the brush on the person image takes priority over the automatic one. Set “Show Type” to input & mask & result (the default) to see which area was repainted.
- Turn CFG down if colours look loud. The app notes that CFG is “highly correlated with saturation”. The default is 2.5 on a 0–7.5 scale.
- Add steps for detail. Inference steps run from 10 to 100, default 50. More steps “may enhance details”, at the cost of time.
- Change the seed. The seed is fixed at 42 by default, so the same inputs give the same image. A new seed “may improve pseudo-shadow”, and the app warns that its NSFW safety checker can block normal results, which a different seed usually fixes.
- Feed it 3:4. The person photo is cropped and the garment padded to 768×1024. A landscape photo loses its edges.
- The garment can be a person. The app has “reference person” examples: the condition image can be someone already wearing the garment, not only a flat product shot.
CatVTON vs IDM-VTON vs OOTDiffusion vs Leffa
The open models people most often weigh against it, checked the same day. Full detail on all 19 is in our open-source virtual try-on comparison.
| Model | Venue | Licence | Free official Space | Documented VRAM | GitHub ★ | ComfyUI |
|---|---|---|---|---|---|---|
| CatVTON | ICLR 2025 | Non-commercial | Runtime error | ~8 GB at 1024×768, bf16 | 1,865 | Official workflow |
| IDM-VTON | ECCV 2024 | Non-commercial | Running | Not by authors (≥16 GB for the ComfyUI node) | 5,215 | Community node |
| OOTDiffusion | AAAI 2025 | Non-commercial | Running | Not documented | 6,604 | Community node |
| Leffa | CVPR 2025 | Grey area (MIT, NC data) | Running | Not documented | 1,681 | Community node |
Swipe the table sideways for VRAM, stars and ComfyUI.
Pick CatVTON if your GPU has 8 GB, or if you want a workflow maintained by the authors rather than a community port. Pick IDM-VTON if you want to try one in the browser today; our IDM-VTON guide covers its demo and setup. Pick OOTDiffusion for automatic masks across upper, lower and dress categories, and Leffa if you also need pose transfer. None of the four is cleared for commercial use. If you need that, the comparison covers the two open options that are: FASHN VTON v1.5 (Apache-2.0, also about 8 GB, demo running) and Qwen-Image-Edit with an Apache-2.0 try-on LoRA.
When a hosted try-on tool is the better fit
CatVTON is the easiest research try-on model to run on your own card, and a good way to learn how try-on works. It is the wrong tool for product pages, ads or marketplace listings: its licence rules out commercial use, its official demo is down, and you maintain the GPU, the masks and the Python environment yourself.
Fashio's virtual try-on is a hosted tool for that job: upload a garment photo, pick a model, and you get on-model images you can use commercially. There is nothing to install and no mask to paint. A try-on costs 4 credits. Plans are Starter at $14 a month for 50 credits, Plus at $29.99 for 160 and Pro at $49.99 for 500, which works out at roughly $0.40 to $1.12 per try-on. A new web account starts with no credits; annual plans begin with a 3-day free trial. Developers can use the same credits through the Fashio API. See pricing for the current plans.
CatVTON FAQ
What is CatVTON?
CatVTON is an open virtual try-on diffusion model: give it a photo of a person and a photo of a garment, and it generates the person wearing that garment. It was written by Zheng Chong and co-authors from Sun Yat-Sen University, Pixocial Technology, Peng Cheng Laboratory and SIAT, and accepted at ICLR 2025. Its idea is to concatenate the person and garment images and run them through a single Stable Diffusion 1.5 inpainting UNet, which keeps it small: 899.06M parameters in total, 49.57M of them trainable.
Is CatVTON free for commercial use?
No. The README states that all the materials, including code, checkpoints and demo, are under CC BY-NC-SA 4.0, and the LICENSE file in the repository is the full text of that licence. You can use it free of charge for research and personal, non-commercial work, with credit, and anything you build on it must use the same licence. The mask-free checkpoint is stricter: its model card restricts it to non-commercial scientific research, forbids commercial use of its generated content and forbids sharing the model.
Why is the CatVTON Hugging Face demo not working?
On 7 October 2026 the official Space, zhengchong/CatVTON, reported RUNTIME_ERROR. Its error log shows the app failing at start-up while loading FLUX.1-Fill-dev for the CatVTON-FLUX tab: that base model is gated, and the Space received a 401 error. The Space was in the same state on 6 October. The authors' older self-hosted demo at 120.76.142.206:8888 did not respond when we tried it. To use CatVTON today, run it locally or in ComfyUI.
How much VRAM does CatVTON need?
About 8 GB. The README says that in bf16 precision, generating at 1024×768 only requires about 8 GB of VRAM, and the repository description says under 8 GB. That figure is for the Stable Diffusion 1.5 version. The chflame163 ComfyUI wrapper recommends an NVIDIA GPU with 6 GB or more. The FLUX-based variant needs the much larger FLUX.1-Fill-dev model, and the authors give no VRAM figure for it.
How do I use CatVTON in ComfyUI?
The authors publish an official ComfyUI release on the repository's Releases page, tagged ComfyUI. Download ComfyUI-CatVTON.zip, unzip it into ComfyUI's custom_nodes folder, install CatVTON's requirements, start ComfyUI and drag catvton_workflow.json onto the canvas. The weights download automatically on the first run, which the README says usually takes dozens of minutes. Community alternatives include chflame163's ComfyUI_CatVTON_Wrapper (371 GitHub stars) and pzc163's Comfyui-CatVTON (174).
What is CatVTON-FLUX?
The name covers two different things. The official one is a 37.4M-parameter LoRA for FLUX.1-Fill-dev released by the CatVTON authors in December 2024; its weights sit in the flux-lora folder of the zhengchong/CatVTON Hugging Face repo and app_flux.py runs it. The other is catvton-flux by nftblackmagic, a separate community project inspired by CatVTON, with MIT-licensed code and weights published by xiaozaa under CC BY-NC 2.0 and CC BY-NC-SA 2.0. Both depend on FLUX.1-Fill-dev, which has its own non-commercial licence.
Is there an official CatVTON API?
No. The README links to the GitHub code, the Hugging Face checkpoints, the Space and the project page, and does not offer a hosted API. Any service selling CatVTON through an API is a third party, and the CC BY-NC-SA 4.0 licence still applies to the model whoever hosts it.
What is the CatVTON Wrapper?
ComfyUI_CatVTON_Wrapper is a community ComfyUI node by chflame163 that wraps the official CatVTON code. Its README says it fixes cropping for input images with different proportions and recommends an NVIDIA GPU with 6 GB of VRAM or more. You download the model files from its Baidu or Google Drive links into ComfyUI/models/CatVTON. On 7 October 2026 it had 371 GitHub stars and its last commit was on 1 January 2025. It states no licence of its own and refers you to the original project's licence.
How we checked
On 7 October 2026 we read the repositories through the GitHub REST API (stars, forks, open issues, last commit, releases, default branch), the models and Spaces through the Hugging Face API (licence tags, gating terms, likes, downloads, file tree, Space status and error log), and the paper through the arXiv export API. Commands and requirements are copied from the README and requirements.txt on the default branch; demo behaviour and defaults come from app.py, app_flux.py and app_p2p.py. The related questions come from Google's results for “catvton”, “catvton comfyui” and “catvton hugging face”. We did not benchmark output quality and we show no images generated by the model. Stars and Space status change, so we re-check monthly.






