3 papers
cs.CV2025
Unified Vision-Language-Action Model
Yuqi Wang, Xinghang Li, Wenxuan Wang +5
Vision-language-action models (VLAs) have garnered significant attention for their potential in advancing robotic manipulation. However, previous approaches predominantly rely on t…
cs.CV2025
Towards Unified Referring Expression Segmentation Across Omni-Level Visual Target Granularities
Jing Liu, Wenxuan Wang, Yisi Zhang +5
Referring expression segmentation (RES) aims at segmenting the entities' masks that match the descriptive language expression. While traditional RES methods primarily address objec…
cs.CV2025
Image Difference Grounding with Natural Language
Wenxuan Wang, Zijia Zhao, Yisi Zhang +4
Visual grounding (VG) typically focuses on locating regions of interest within an image using natural language, and most existing VG methods are limited to single-image interpretat…