1 citations · 1 across the 10 of their papers we have counts for
12 papers
XPos3R: Cross-Modal Transformer for Intraoperative 2D/3D Registration
Shiyan Su, Ruyi Zha, Hongdong Li +2
Intraoperative 2D/3D registration, which aligns live X-ray images with preoperative volumes, is essential for image-guided interventions. Previous regression-based methods suffer f…
Learning Compositional Spatio-Temporal Video Grounding with Synthetic Curriculum
Xingjian Wang, Shijian Wang, Yibo Wang +4
Despite the impressive progress of recent MLLMs on spatio-temporal video grounding (STVG), existing evaluations and training data focus primarily on simple queries. They largely ov…
CALM-AH: An ABAW11-Calibrated Multimodal Ensemble with Reliability-Gated Multi-Expert Consensus for Video-Level Ambivalence and Hesitancy Recognition
Wenzhuo Sun, Mingjian Liang, Richard Attfield +3
Ambivalence and hesitancy (A/H) are subtle behavioural states that may be expressed through language, voice, facial activity, and other non-verbal cues. The ABAW11 A/H Video Recogn…
Beyond Binary Success: A Diagnostic Meta-Evaluation Framework for Fine-Grained Manipulation
He-Yang Xu, Pengyuan Zhang, Zongyuan Ge +5
Fine-grained manipulation marks a regime where global scene context no longer suffices, and success hinges on the tight coupling of local attribute grounding, high-fidelity spatial…
MuSEAgent: A Multimodal Reasoning Agent with Stateful Experiences
Shijian Wang, Jiarui Jin, Runhao Fu +11
Research agents have recently achieved significant progress in information seeking and synthesis across heterogeneous textual and visual sources. In this paper, we introduce MuSEAg…
Project Imaging-X: A Survey of 1000+ Open-Access Medical Imaging Datasets for Foundation Model Development
Zhongying Deng, Cheng Tang, Ziyan Huang +124
Foundation models have demonstrated remarkable success across diverse domains and tasks, primarily due to the thrive of large-scale, diverse, and high-quality datasets. However, in…