4 papers
ChemVLR: Prioritizing Reasoning in Perception for Chemical Vision-Language Understanding
Xuanle Zhao, Xinyuan Cai, Xiang Cheng +2
While Vision-Language Models (VLMs) have demonstrated significant potential in chemical visual understanding, current models are predominantly optimized for direct visual question-…
Adaptive Runge-Kutta Dynamics for Spatiotemporal Prediction
Xuanle Zhao, Yue Sun, Ziyi Wang +2
Spatiotemporal prediction is important in solving natural problems and processing video frames, especially in weather forecasting and human action recognition. Recent advances atte…
TinyChemVL: Advancing Chemical Vision-Language Models via Efficient Visual Token Reduction and Complex Reaction Tasks
Xuanle Zhao, Shuxin Zeng, Xinyuan Cai +4
While Vision Language Models (VLMs) have demonstrated remarkable capabilities in general visual understanding, their application in the chemical domain has been limited, with previ…
Revisiting Continual Semantic Segmentation with Pre-trained Vision Models
Duzhen Zhang, Yong Ren, Wei Cong +9
Continual Semantic Segmentation (CSS) seeks to incrementally learn to segment novel classes while preserving knowledge of previously encountered ones. Recent advancements in CSS ha…