2 papers
cs.CV2025
Where, Not What: Compelling Video LLMs to Learn Geometric Causality for 3D-Grounding
Yutong Zhong
Multimodal 3D grounding has garnered considerable interest in Vision-Language Models (VLMs) \cite{yin2025spatial} for advancing spatial reasoning in complex environments. However,…
cs.CV2025
Learning Pyramid-structured Long-range Dependencies for 3D Human Pose Estimation
Mingjie Wei, Xuemei Xie, Yutong Zhong +1
Action coordination in human structure is indispensable for the spatial constraints of 2D joints to recover 3D pose. Usually, action coordination is represented as a long-range dep…