activity
20242026
collaborators

12 papers

cs.CV2026

MotionMERGE: A Multi-granular Framework for Human Motion Editing, Reasoning, Generation, and Explanation

Bizhu Wu, Jinheng Xie, Wenting Chen +5

Recent motion-language models unify tasks like comprehension and generation but operate at a coarse granularity, lacking fine-grained understanding and nuanced control over body pa…

cs.CV2026

Adaptive Texture-aware Masking for Self-Supervised Learning in 3D Dental CBCT Analysis

Xinquan Yang, Jianfeng Ren, Xuguang Li +4

Cone Beam Computed Tomography (CBCT) is pivotal for 3D diagnostic imaging in dentistry. However, the development of robust AI models for volumetric analysis is often constrained by…

cs.CV2026

X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic Diagnosis

Gui Wang, Zehao Zhong, YongSong Zhou +6

Despite significant progress in Multi-modal Large Language Models (MLLMs), their clinical reasoning capacity for multi-modal diagnosis remains largely unexamined. Current benchmark…

cs.CV2026

SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark

Gui Wang, YongSong Zhou, Kaijun Deng +4

Fine-grained spatiotemporal reasoning on surgical videos is critical, yet the capabilities of Multi-modal Large Language Models (MLLMs) in this domain remain largely unexplored. To…

cs.AI2026

DIRCR: Dual-Inference Rule-Contrastive Reasoning for Solving RAVENs

Jiachen Zhang, Chengtai Li, Jianfeng Ren +3

Abstract visual reasoning remains challenging as existing methods often prioritize either global context or local row-wise relations, failing to integrate both, and lack intermedia…

cs.AI2026

Preference-Agile Multi-Objective Optimization for Real-time Vehicle Dispatching

Jiahuan Jin, Wenhao Zhao, Rong Qu +4

Multi-objective optimization (MOO) has been widely studied in literature because of its versatility in human-centered decision making in real-life applications. Recently, demand fo…