3 papers
cs.CV2025
FUSE-RSVLM: Feature Fusion Vision-Language Model for Remote Sensing
Yunkai Dang, Donghao Wang, Jiacheng Yang +7
Large vision-language models (VLMs) exhibit strong performance across various tasks. However, these VLMs encounter significant challenges when applied to the remote sensing domain…
cs.CV2025
HOLODECK 2.0: Vision-Language-Guided 3D World Generation with Editing
Zixuan Bian, Ruohan Ren, Yue Yang +1
3D scene generation plays a crucial role in gaming, artistic creation, virtual reality, and many other domains. However, current 3D scene design still relies heavily on extensive m…
cs.CV2024
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Matt Deitke, Christopher Clark, Sangho Lee +47
Today's most advanced vision-language models (VLMs) remain proprietary. The strongest open-weight models rely heavily on synthetic data from proprietary VLMs to achieve good perfor…