15 papers
WearWow: Native 2K Multi-Garment Virtual Try-On via Adaptive Token Packing and Preference Alignment
Xujie Zhang, Runyan Du, Song Chang +7
Synthesizing native 2K multi-garment virtual try-on is a formidable frontier in digital fashion, critically bottlenecked by two fundamental limitations: the O(N^2) memory explosion…
StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling
Jianing Peng, Mengyu Wang, Henghui Ding +6
Multi-reference image generation aims to synthesize images by integrating attributes from multiple reference images under textual instructions. As the number of references increase…
Learning Stochastic Bridges for Video Object Removal via Video-to-Video Translation
Zijie Lou, Xiangwei Feng, Jiaxin Wang +7
Existing video object removal methods predominantly rely on diffusion models following a noise-to-data paradigm, where generation starts from uninformative Gaussian noise. This app…
The 1st PortraitCraft Challenge: A CVPR 2026 Workshop Competition on Portrait Composition Understanding and Generation
Zijie Lou, Youyun Tang, Xiaochao Qu +40
This paper presents an overview of the inaugural PortraitCraft Challenge, held as one of the official competitions at CVPR 2026. The challenge focuses on portrait composition under…
FlowSeg: Dynamic Semantic Guidance for LLM-Conditioned Segmentation
Zekang Zhang, Guangyu Gao, Youyun Tang +7
LLM-conditioned segmentation has recently advanced rapidly by coupling large language models with iterative mask generation frameworks. However, we identify a persistent failure mo…
Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Editing via In-Context Learning
Hongxi Li, Tong Wang, Chengjing Wu +6
Scene text editing aims to modify text in a target region of an image while preserving surrounding background style and texture. Existing methods rely solely on image background in…