3 papers
cs.CV2026
Improving Large Vision-Language Models' Understanding for Flow Field Data
Xiaomei Zhang, Hanyu Zheng, Xiangyu Zhu +4
Large Vision-Language Models (LVLMs) have shown impressive capabilities across a range of tasks that integrate visual and textual understanding, such as image captioning and visual…
cs.CV2025
Top-Down Guidance for Learning Object-Centric Representations
Junhong Zou, Xiangyu Zhu, Zhaoxiang Zhang +1
Humans' innate ability to decompose scenes into objects allows for efficient understanding, predicting, and planning. In light of this, Object-Centric Learning (OCL) attempts to en…
cs.CV2024
SimCMF: A Simple Cross-modal Fine-tuning Strategy from Vision Foundation Models to Any Imaging Modality
Chenyang Lei, Liyi Chen, Jun Cen +5
Foundation models like ChatGPT and Sora that are trained on a huge scale of data have made a revolutionary social impact. However, it is extremely challenging for sensors in many d…