7 papers
Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams
Weihao Bo, Shan Zhang, Yanpeng Sun +7
Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration. For example, OpenAI Prism is a free workspace for scientific wri…
Artemis: Structured Visual Reasoning for Perception Policy Learning
Wei Tang, Yanpeng Sun, Shan Zhang +6
Recent reinforcement-learning frameworks for visual perception policy usually incorporate intermediate reasoning chains expressed in natural language. Empirical observations indica…
Agentic Learner with Grow-and-Refine Multimodal Semantic Memory
Weihao Bo, Shan Zhang, Yanpeng Sun +9
MLLMs exhibit strong reasoning on isolated queries, yet they operate de novo -- solving each problem independently and often repeating the same mistakes. Existing memory-augmented…
SSP-SAM: SAM with Semantic-Spatial Prompt for Referring Expression Segmentation
Wei Tang, Xuejing Liu, Yanpeng Sun +1
The Segment Anything Model (SAM) excels at general image segmentation but has limited ability to understand natural language, which restricts its direct application in Referring Ex…
Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
Yanpeng Sun, Shan Zhang, Wei Tang +5
Diagrams represent a form of visual language that encodes abstract concepts and relationships through structured symbols and their spatial arrangements. Unlike natural images, they…
Visual Position Prompt for MLLM based Visual Grounding
Wei Tang, Yanpeng Sun, Qinying Gu +1
Although Multimodal Large Language Models (MLLMs) excel at various image-related tasks, they encounter challenges in precisely aligning coordinates with spatial information within…