collaborators

7 papers

cs.CV2026

What Makes Synthetic Data Effective in Image Segmentation

Jinjin Zhang, Xiefan Guo, Yizhou Jin +2

Driven by rapid advances in large-scale generative models, synthetic data has emerged as a promising solution for visual understanding. While modern diffusion models achieve remark…

cs.CV2026

CoVFT: Context-aware Visual Fine-tuning for Multimodal Large Language Models

Nan Zhou, Huiqun Wang, Yaoyan Zheng +1

Multimodal large language models (MLLMs) achieve remarkable progress in cross-modal perception and reasoning, yet a fundamental question remains unresolved: should the vision encod…

cs.CV2026

A multimodal vision foundation model for generalizable knee pathology

Kang Yu, Dingyu Wang, Zimu Yuan +8

Musculoskeletal disorders represent a leading cause of global disability, creating an urgent demand for precise interpretation of medical imaging. Current artificial intelligence (…

cs.CV2025

The Illusion of Clinical Reasoning: A Benchmark Reveals the Pervasive Gap in Vision-Language Models for Clinical Competency

Dingyu Wang, Zimu Yuan, Jiajun Liu +5

Background: The rapid integration of foundation models into clinical practice and public health necessitates a rigorous evaluation of their true clinical reasoning capabilities bey…

cs.CV2025

Implicit Modeling for Transferability Estimation of Vision Foundation Models

Yaoyan Zheng, Huiqun Wang, Nan Zhou +1

Transferability estimation identifies the best pre-trained models for downstream tasks without incurring the high computational cost of full fine-tuning. This capability facilitate…

cs.CV2025

ForCenNet: Foreground-Centric Network for Document Image Rectification

Peng Cai, Qiang Li, Kaicheng Yang +6

Document image rectification aims to eliminate geometric deformation in photographed documents to facilitate text recognition. However, existing methods often neglect the significa…