9 papers
ReCap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning
Haonan Jia, Shichao Dong, Zenghui Sun +7
Reinforcement Learning (RL) has demonstrated significant gains in image captioning, yet it is still limited in encouraging Large Vision-Language Models (LVLMs) to explore novel rea…
iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance
Jun Zheng, Zhengze Xu, Mengting Chen +6
Video Virtual Try-On (VVT) aims to seamlessly replace a garment on a person in a video with a new one. While existing methods have made significant strides in maintaining temporal…
FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization
Quanjian Song, Yefeng Shen, Mengting Chen +5
Human-centric video customization, particularly at the garment level, has shown significant commercial value. However, existing approaches cannot support low-latency and interactiv…
Improved Baselines with Representation Autoencoders
Jaskirat Singh, Boyang Zheng, Zongze Wu +3
Representation Autoencoders (RAE) replace traditional VAE with pretrained vision encoders. In this paper, we systematically investigate several design choices and find three insigh…
Edit-GRPO: A Locality-Preserving Policy Optimization Framework for Image Editing
Shaodong Xu, Zexian Li, Zhendong Wang +5
A fundamental challenge in image editing lies in preserving spatial locality: edits should improve targeted content without inadvertently altering surrounding regions. However, mos…
Deep Pre-Alignment for VLMs
Tianyu Yu, Kechen Fang, Zihao Wan +5
Most Vision Language Models (VLMs) directly map outputs from ViT encoders to the LLM via a lightweight projector. While effective, recent analysis suggests this architecture suffer…