Publications (5)
OmniDiT: Extending Diffusion Transformer to Omni-VTON Framework
Weixuan Zeng, Pengcheng Wei, Huaiqing Wang +8
Despite the rapid advancement of Virtual Try-On (VTON) and Try-Off (VTOFF) technologies, existing VTON methods face challenges with fine-grained detail preservation, generalization…
DetailAnywhere: Fashion Detail Generation via Cross-Modal Feature Alignment Distillation
Zijun Li, Yimin Zhou, Jia Sun +12
Diffusion-based generative AI has achieved remarkable success in e-commerce applications such as virtual try-on, poster generation, and product background synthesis. However, when…
EVLM: An Efficient Vision-Language Model for Visual Understanding
Kaibing Chen, Dong Shen, Hanwen Zhong +14
In the field of multi-modal language models, the majority of methods are built on an architecture similar to LLaVA. These models use a single-layer ViT feature as a visual prompt,…
CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating
Jiyuan Wang, Huan Ouyang, Jiuzhou Lin +15
In this paper, we propose Concentrate and Concentrate (CaC), a coarse-to-fine anomaly reward model based on Vision-Language Models. During inference, it first conducts a global tem…
Improving Crowded Object Detection via Copy-Paste
Jiangfan Deng, Dewen Fan, Xiaosong Qiu +1
Crowdedness caused by overlapping among similar objects is a ubiquitous challenge in the field of 2D visual object detection. In this paper, we first underline two main effects of…