collaborators

7 papers

cs.CV2026

Multimodal Causal-Driven Representation Learning for Generalizable Medical Image Segmentation

Xusheng Liang, Lihua Zhou, Nianxin Li +8

Vision-Language Models (VLMs), such as CLIP, have demonstrated remarkable zero-shot capabilities in various computer vision tasks. However, their application to medical imaging rem…

cs.CV2026

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos

Jinlin Wu, Felix Holm, Chuxi Chen +17

While foundation models have advanced surgical video analysis, current approaches rely predominantly on pixel-level reconstruction objectives that waste model capacity on low-level…

cs.CV2025

BFSM: 3D Bidirectional Face-Skull Morphable Model

Zidu Wang, Meng Xu, Miao Xu +5

Building a joint face-skull morphable model holds great potential for applications such as remote diagnostics, surgical planning, medical education, and physically based facial sim…

cs.CV2025

Towards Realistic Hand-Object Interaction with Gravity-Field Based Diffusion Bridge

Miao Xu, Xiangyu Zhu, Xusheng Liang +3

Existing reconstruction or hand-object pose estimation methods are capable of producing coarse interaction states. However, due to the complex and diverse geometry of both human ha…

cs.CV2025

Pose-RFT: Enhancing MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning

Bao Li, Xiaomei Zhang, Miao Xu +3

Generating 3D human poses from multimodal inputs such as images or text requires models to capture both rich spatial and semantic correspondences. While pose-specific multimodal la…

cs.LG2025

Deep Multi-Manifold Transformation Based Multivariate Time Series Fault Detection

Hong Liu, Xiuxiu Qiu, Yiming Shi +3

Unsupervised fault detection in multivariate time series plays a vital role in ensuring the stable operation of complex systems. Traditional methods often assume that normal data f…