collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2025

FastFit: Accelerating Multi-Reference Virtual Try-On via Cacheable Diffusion Models

Zheng Chong, Yanwei Lei, Shiyue Zhang +7

Despite its great potential, virtual try-on technology is hindered from real-world application by two major challenges: the inability of current methods to support multi-reference…

cs.CV2025

RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving

Zhijian Huang, Chengjian Feng, Feng Yan +5

Large Multimodal Models (LMMs) have demonstrated exceptional comprehension and interpretation capabilities in Autonomous Driving (AD) by incorporating large language models. Despit…

cs.CV2025

CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models

Zheng Chong, Xiao Dong, Haoxiang Li +6

Virtual try-on methods based on diffusion models achieve realistic effects but often require additional encoding modules, a large number of training parameters, and complex preproc…

cs.CV2025

ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions

Shiyue Zhang, Zheng Chong, Xi Lu +6

Building on the success of diffusion models, significant advancements have been made in multimodal image generation tasks. Among these, human image generation has emerged as a prom…

cs.CV2025

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation

Zheng Chong, Wenqing Zhang, Shiyue Zhang +6

Virtual try-on (VTON) technology has gained attention due to its potential to transform online retail by enabling realistic clothing visualization of images and videos. However, mo…

cs.CV2024

OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion

Hao Wang, Pengzhen Ren, Zequn Jie +8

Open-vocabulary detection is a challenging task due to the requirement of detecting objects based on class names, including those not encountered during training. Existing methods…