collaborators

7 papers

cs.AI2026

FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models

Kaili Zheng, Kaiwen Wang, Xun Zhu +3

Fitness Action Quality Assessment (AQA) is important for intelligent sports training, yet the capabilities of Multimodal Large Language Models (MLLMs) in this setting remain undere…

cs.CL2026

ReportQA: QA-Based Radiology Report Evaluation

Yiming Shi, Shaoshuai Yang, Xi Chen +10

Radiology report evaluation is essential for advancing automated report generation. Natural language generation metrics have limited clinical relevance. Clinical efficacy (CE) metr…

cs.CV2026

The 1st PortraitCraft Challenge: A CVPR 2026 Workshop Competition on Portrait Composition Understanding and Generation

Zijie Lou, Youyun Tang, Xiaochao Qu +40

This paper presents an overview of the inaugural PortraitCraft Challenge, held as one of the official competitions at CVPR 2026. The challenge focuses on portrait composition under…

cs.CV2026

InterMesh: Explicit Interaction-Aware End-to-End Multi-Person Human Mesh Recovery

Kaili Zheng, Kaiwen Wang, Xun Zhu +2

Humans constantly interact with their surroundings. Existing end-to-end multi-person human mesh recovery methods, typically based on the DETR framework, capture inter-human relatio…

cs.CV2026

Lost in the Hype: Revealing and Dissecting the Performance Degradation of Medical Multimodal Large Language Models in Image Classification

Xun Zhu, Fanbin Mo, Xi Chen +6

The rise of multimodal large language models (MLLMs) has sparked an unprecedented wave of applications in the field of medical imaging analysis. However, as one of the earliest and…

cs.CV2025

MedM-VL: What Makes a Good Medical LVLM?

Yiming Shi, Shaoshuai Yang, Xun Zhu +4

Medical image analysis is essential in modern healthcare. Deep learning has redirected research focus toward complex medical multimodal tasks, including report generation and visua…