4 papers
-WM: A Unified Video-Action World Model for Robotic Manipulation
Pengfei Zhou, Shengcong Chen, Di Chen +17
Robotic manipulation requires models that generate executable actions while anticipating and evaluating their future consequences before physical execution. We present -World…
ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning
Yifan Li, Yingda Yin, Lingting Zhu +4
Reasoning-centric video object segmentation is an inherently complex task: the query often refers to dynamics, causality, and temporal interactions, rather than static appearances.…
Repositioning the Subject within Image
Yikai Wang, Chenjie Cao, Ke Fan +4
Current image manipulation primarily centers on static manipulation, such as replacing specific regions within an image or altering its overall style. In this paper, we introduce a…
Unified Lexical Representation for Interpretable Visual-Language Alignment
Yifan Li, Yikai Wang, Yanwei Fu +3
Visual-Language Alignment (VLA) has gained a lot of attention since CLIP's groundbreaking work. Although CLIP performs well, the typical direct latent feature alignment lacks clari…