5 papers
Unified Vision-Language-Action Model
Yuqi Wang, Xinghang Li, Wenxuan Wang +5
Vision-language-action models (VLAs) have garnered significant attention for their potential in advancing robotic manipulation. However, previous approaches predominantly rely on t…
End-to-End Vision Tokenizer Tuning
Wenxuan Wang, Fan Zhang, Yufeng Cui +5
Existing vision tokenization isolates the optimization of vision tokenizers from downstream training, implicitly assuming the visual tokens can generalize well across various tasks…
Towards Unified Referring Expression Segmentation Across Omni-Level Visual Target Granularities
Jing Liu, Wenxuan Wang, Yisi Zhang +5
Referring expression segmentation (RES) aims at segmenting the entities' masks that match the descriptive language expression. While traditional RES methods primarily address objec…
Image Difference Grounding with Natural Language
Wenxuan Wang, Zijia Zhao, Yisi Zhang +4
Visual grounding (VG) typically focuses on locating regions of interest within an image using natural language, and most existing VG methods are limited to single-image interpretat…
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
Haiwen Diao, Xiaotong Li, Yufeng Cui +6
Existing encoder-free vision-language models (VLMs) are rapidly narrowing the performance gap with their encoder-based counterparts, highlighting the promising potential for unifie…