activity
20242026
collaborators

5 papers

cs.CV2026

CoVFT: Context-aware Visual Fine-tuning for Multimodal Large Language Models

Nan Zhou, Huiqun Wang, Yaoyan Zheng +1

Multimodal large language models (MLLMs) achieve remarkable progress in cross-modal perception and reasoning, yet a fundamental question remains unresolved: should the vision encod…

cs.CV2026

A multimodal vision foundation model for generalizable knee pathology

Kang Yu, Dingyu Wang, Zimu Yuan +8

Musculoskeletal disorders represent a leading cause of global disability, creating an urgent demand for precise interpretation of medical imaging. Current artificial intelligence (…

cs.CV2025

Implicit Modeling for Transferability Estimation of Vision Foundation Models

Yaoyan Zheng, Huiqun Wang, Nan Zhou +1

Transferability estimation identifies the best pre-trained models for downstream tasks without incurring the high computational cost of full fine-tuning. This capability facilitate…

cs.CV2024

Deep Common Feature Mining for Efficient Video Semantic Segmentation

Yaoyan Zheng, Hongyu Yang, Di Huang

Recent advancements in video semantic segmentation have made substantial progress by exploiting temporal correlations. Nevertheless, persistent challenges, including redundant comp…

cs.CV2024

Crowd-SAM: SAM as a Smart Annotator for Object Detection in Crowded Scenes

Zhi Cai, Yingjie Gao, Yaoyan Zheng +2

In computer vision, object detection is an important task that finds its application in many scenarios. However, obtaining extensive labels can be challenging, especially in crowde…