collaborators

20 papers

cs.CV2026

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models

Hsiang-Wei Huang, Junbin Lu, Kuang-Ming Chen +3

Vision-Language Models (VLMs) achieve strong performance on spatial question answering benchmarks, yet it remains unclear whether such gains reflect genuine spatial intelligence. W…

cs.CV2026

RAM: Recover Any 3D Human Motion in-the-Wild

Sen Jia, Ning Zhu, Jinqin Zhong +4

RAM incorporates a motion-aware semantic tracker with adaptive Kalman filtering to achieve robust identity association under severe occlusions and dynamic interactions. A memory-au…

cs.CV2026

3DSceneEditor: Controllable 3D Scene Editing with Gaussian Splatting

Ziyang Yan, Yihua Shao, Minwen Liao +7

The creation of 3D scenes has traditionally been both labor-intensive and costly, requiring designers to meticulously configure 3D assets and environments. Recent advancements in g…

cs.CV2026

Detector-in-the-Loop Tracking: Active Memory Rectification for Stable Glottic Opening Localization

Huayu Wang, Bahaa Alattar, Cheng-Yen Yang +5

Temporal stability in glottic opening localization remains challenging due to the complementary weaknesses of single-frame detectors and foundation-model trackers: the former lacks…

cs.CV2025

UniHPR: Unified Human Pose Representation via Singular Value Contrastive Learning

Zhongyu Jiang, Wenhao Chai, Lei Li +3

In recent years, there has been a growing interest in developing effective alignment pipelines to generate unified representations from different modalities for multi-modal fusion…

cs.CL2025

Attention Consistency for LLMs Explanation

Tian Lan, Jinyuan Xu, Xue He +2

Understanding the decision-making processes of large language models (LLMs) is essential for their trustworthy development and deployment. However, current interpretability methods…