collaborators

7 papers

cs.CV2025

See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning

Shuoshuo Zhang, Yizhen Zhang, Jingjing Fu +4

Large vision-language models (VLMs) often benefit from intermediate visual cues, either injected via external tools or generated as latent visual tokens during reasoning, but these…

cs.CV2025

PixelCraft: A Multi-Agent System for High-Fidelity Visual Reasoning on Structured Images

Shuoshuo Zhang, Zijian Li, Yizhen Zhang +6

Structured images (e.g., charts and geometric diagrams) remain challenging for multimodal large language models (MLLMs), as perceptual slips can cascade into erroneous conclusions.…

eess.IV2025

AURAD: Anatomy-Pathology Unified Radiology Synthesis with Progressive Representations

Shuhan Ding, Jingjing Fu, Yu Gu +6

Medical image synthesis has become an essential strategy for augmenting datasets and improving model generalization in data-scarce clinical settings. However, fine-grained and cont…

cs.AI2025

The Illusion of Readiness in Health AI

Yu Gu, Jingjing Fu, Xiaodong Liu +29

Large language models have demonstrated remarkable performance in a wide range of medical benchmarks. Yet underneath the seemingly promising results lie salient growth areas, espec…

cs.IR2025

OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval

Wei Yang, Jingjing Fu, Rui Wang +3

Vision-language retrieval-augmented generation (RAG) has become an effective approach for tackling Knowledge-Based Visual Question Answering (KB-VQA), which requires external knowl…

cs.CV2025

Chain of Functions: A Programmatic Pipeline for Fine-Grained Chart Reasoning Data

Zijian Li, Jingjing Fu, Lei Song +3

Visual reasoning is crucial for multimodal large language models (MLLMs) to address complex chart queries, yet high-quality rationale data remains scarce. Existing methods leverage…