collaborators

9 papers

cs.CV2026

ReCap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning

Haonan Jia, Shichao Dong, Zenghui Sun +7

Reinforcement Learning (RL) has demonstrated significant gains in image captioning, yet it is still limited in encouraging Large Vision-Language Models (LVLMs) to explore novel rea…

cs.CV2026

iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance

Jun Zheng, Zhengze Xu, Mengting Chen +6

Video Virtual Try-On (VVT) aims to seamlessly replace a garment on a person in a video with a new one. While existing methods have made significant strides in maintaining temporal…

cs.CV2026

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization

Quanjian Song, Yefeng Shen, Mengting Chen +5

Human-centric video customization, particularly at the garment level, has shown significant commercial value. However, existing approaches cannot support low-latency and interactiv…

cs.CV2026

Improved Baselines with Representation Autoencoders

Jaskirat Singh, Boyang Zheng, Zongze Wu +3

Representation Autoencoders (RAE) replace traditional VAE with pretrained vision encoders. In this paper, we systematically investigate several design choices and find three insigh…

cs.CV2026

Edit-GRPO: A Locality-Preserving Policy Optimization Framework for Image Editing

Shaodong Xu, Zexian Li, Zhendong Wang +5

A fundamental challenge in image editing lies in preserving spatial locality: edits should improve targeted content without inadvertently altering surrounding regions. However, mos…

cs.CV2026

Deep Pre-Alignment for VLMs

Tianyu Yu, Kechen Fang, Zihao Wan +5

Most Vision Language Models (VLMs) directly map outputs from ViT encoders to the LLM via a lightweight projector. While effective, recent analysis suggests this architecture suffer…