2 papers
cs.CV2026
Thinker: A vision-language foundation model for embodied intelligence
Baiyu Pan, Daqin Luo, Junpeng Yang +4
When large vision-language models are applied to the field of robotics, they encounter problems that are simple for humans yet error-prone for models. Such issues include confusion…
cs.CV2025
ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model
Kunyang Han, Yibo Hu, Mengxue Qu +3
Advances in CLIP and large multimodal models (LMMs) have enabled open-vocabulary and free-text segmentation, yet existing models still require predefined category prompts, limiting…