activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams

Weihao Bo, Shan Zhang, Yanpeng Sun +7

Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration. For example, OpenAI Prism is a free workspace for scientific wri…

cs.CV2025

Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs

Yanpeng Sun, Shan Zhang, Wei Tang +5

Diagrams represent a form of visual language that encodes abstract concepts and relationships through structured symbols and their spatial arrangements. Unlike natural images, they…

cs.CV2025

Hierarchical Process Reward Models are Symbolic Vision Learners

Shan Zhang, Aotian Chen, Kai Zou +3

Symbolic computer vision represents diagrams through explicit logical rules and structured representations, enabling interpretable understanding in machine vision. This requires fu…

cs.CV2025

Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs

Shan Zhang, Aotian Chen, Yanpeng Sun +6

Current multimodal large language models (MLLMs) often underperform on mathematical problem-solving tasks that require fine-grained visual understanding. The limitation is largely…

cs.CV2024

RS-GPT4V: A Unified Multimodal Instruction-Following Dataset for Remote Sensing Image Understanding

Linrui Xu, Ling Zhao, Wang Guo +5

The remote sensing image intelligence understanding model is undergoing a new profound paradigm shift which has been promoted by multi-modal large language model (MLLM), i.e. from…