5 papers
A multimodal vision foundation model for generalizable knee pathology
Kang Yu, Dingyu Wang, Zimu Yuan +8
Musculoskeletal disorders represent a leading cause of global disability, creating an urgent demand for precise interpretation of medical imaging. Current artificial intelligence (…
The Illusion of Clinical Reasoning: A Benchmark Reveals the Pervasive Gap in Vision-Language Models for Clinical Competency
Dingyu Wang, Zimu Yuan, Jiajun Liu +5
Background: The rapid integration of foundation models into clinical practice and public health necessitates a rigorous evaluation of their true clinical reasoning capabilities bey…
Implicit Modeling for Transferability Estimation of Vision Foundation Models
Yaoyan Zheng, Huiqun Wang, Nan Zhou +1
Transferability estimation identifies the best pre-trained models for downstream tasks without incurring the high computational cost of full fine-tuning. This capability facilitate…
ForCenNet: Foreground-Centric Network for Document Image Rectification
Peng Cai, Qiang Li, Kaicheng Yang +6
Document image rectification aims to eliminate geometric deformation in photographed documents to facilitate text recognition. However, existing methods often neglect the significa…
RefCut: Interactive Segmentation with Reference Guidance
Zheng Lin, Nan Zhou, Chen-Xi Du +2
Interactive segmentation aims to segment the specified target on the image with positive and negative clicks from users. Interactive ambiguity is a crucial issue in this field, whi…