activity
20242026
collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2026

Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning

Pengyu Li, Zhitao Gao, Lingling Zhang +4

Unified multimodal models (UMMs) interleave generated ''visual thoughts'' (VTs) with text reasoning to improve spatial tasks. This incurs roughly an order-of-magnitude inference co…

cs.CV2026

ChartAct: A Benchmark for Dynamic Chart Understanding

Muye Huang, Lin Wu, Lingling Zhang +5

Charts are widely used to present complex data for analysis and decision making. Existing chart understanding benchmarks mainly focus on static charts, but real-world charts are of…

cs.CV2026

Neurodynamics-Driven Coupled Neural P Systems for Multi-Focus Image Fusion

Bo Li, Yunkuo Lei, Tingting Bao +5

Multi-focus image fusion (MFIF) is a crucial technique in image processing, with a key challenge being the generation of decision maps with precise boundaries. However, traditional…

cs.CV2026

SketchVL: Policy Optimization via Fine-Grained Credit Assignment for Chart Understanding and More

Muye Huang, Lingling Zhang, Yifei Li +2

Charts are high-density visual carriers of complex data and medium for information extraction and analysis. Due to the need for precise and complex visual reasoning, automated char…

cs.CV2025

CoFFT: Chain of Foresight-Focus Thought for Visual Language Models

Xinyu Zhang, Yuxuan Dong, Lingling Zhang +5

Despite significant advances in Vision Language Models (VLMs), they remain constrained by the complexity and redundancy of visual input. When images contain large amounts of irrele…

cs.CV2025

Diagram-Driven Course Questions Generation

Xinyu Zhang, Lingling Zhang, Yanrui Wu +6

Visual Question Generation (VQG) research focuses predominantly on natural images while neglecting the diagram, which is a critical component in educational materials. To meet the…