1 citations · 1 across the 4 of their papers we have counts for
6 papers · 1 filter
Unified Vision-Language-Action Model
Yuqi Wang, Xinghang Li, Wenxuan Wang +5
Vision-language-action models (VLAs) have garnered significant attention for their potential in advancing robotic manipulation. However, previous approaches predominantly rely on t…
Towards Unified Referring Expression Segmentation Across Omni-Level Visual Target Granularities
Jing Liu, Wenxuan Wang, Yisi Zhang +5
Referring expression segmentation (RES) aims at segmenting the entities' masks that match the descriptive language expression. While traditional RES methods primarily address objec…
Image Difference Grounding with Natural Language
Wenxuan Wang, Zijia Zhao, Yisi Zhang +4
Visual grounding (VG) typically focuses on locating regions of interest within an image using natural language, and most existing VG methods are limited to single-image interpretat…
Diffusion Feedback Helps CLIP See Better
Wenxuan Wang, Quan Sun, Fan Zhang +3
Contrastive Language-Image Pre-training (CLIP), which excels at abstracting open-world representations across domains and modalities, has become a foundation for a variety of visio…
Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with Human Intentions
Wenxuan Wang, Yisi Zhang, Xingjian He +4
Visual grounding (VG) aims at locating the foreground entities that match the given natural language expressions. Previous datasets and methods for classic VG task mainly rely on t…
Unveiling Parts Beyond Objects:Towards Finer-Granularity Referring Expression Segmentation
Wenxuan Wang, Tongtian Yue, Yisi Zhang +4
Referring expression segmentation (RES) aims at segmenting the foreground masks of the entities that match the descriptive natural language expression. Previous datasets and method…