activity
20242026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

The 1st PortraitCraft Challenge: A CVPR 2026 Workshop Competition on Portrait Composition Understanding and Generation

Zijie Lou, Youyun Tang, Xiaochao Qu +40

This paper presents an overview of the inaugural PortraitCraft Challenge, held as one of the official competitions at CVPR 2026. The challenge focuses on portrait composition under…

cs.CV2026

Lost in the Hype: Revealing and Dissecting the Performance Degradation of Medical Multimodal Large Language Models in Image Classification

Xun Zhu, Fanbin Mo, Xi Chen +6

The rise of multimodal large language models (MLLMs) has sparked an unprecedented wave of applications in the field of medical imaging analysis. However, as one of the earliest and…

cs.CV2025

Med-2E3: A 2D-Enhanced 3D Medical Multimodal Large Language Model

Yiming Shi, Xun Zhu, Kaiwen Wang +4

3D medical image analysis is essential for modern healthcare, yet traditional task-specific models are inadequate due to limited generalizability across diverse clinical scenarios.…

cs.CV2025

MedM-VL: What Makes a Good Medical LVLM?

Yiming Shi, Shaoshuai Yang, Xun Zhu +4

Medical image analysis is essential in modern healthcare. Deep learning has redirected research focus toward complex medical multimodal tasks, including report generation and visua…

cs.CV2025

Enhancing Multi-task Learning Capability of Medical Generalist Foundation Model via Image-centric Multi-annotation Data

Xun Zhu, Fanbin Mo, Zheng Zhang +6

The emergence of medical generalist foundation models has revolutionized conventional task-specific model development paradigms, aiming to better handle multiple tasks through join…

cs.CV2024

Uni-Med: A Unified Medical Generalist Foundation Model For Multi-Task Learning Via Connector-MoE

Xun Zhu, Ying Hu, Fanbin Mo +2

Multi-modal large language models (MLLMs) have shown impressive capabilities as a general-purpose interface for various visual and linguistic tasks. However, building a unified MLL…