3 papers
cs.CV2026
Vision-language models lag human performance on physical dynamics and intent reasoning
Tianjun Gu, Jingyu Gong, Zhizhong Zhang +4
Spatial intelligence is central to embodied cognition, yet contemporary AI systems still struggle to reason about physical interactions in open-world human environments. Despite st…
cs.CV2025
Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy
Gong Jingyu, Tong Kunkun, Chen Zhuoran +5
Human motion synthesis in 3D scenes relies heavily on scene comprehension, while current methods focus mainly on scene structure but ignore the semantic understanding. In this pape…
cs.CV2024
Fusion-then-Distillation: Toward Cross-modal Positive Distillation for Domain Adaptive 3D Semantic Segmentation
Yao Wu, Mingwei Xing, Yachao Zhang +2
In cross-modal unsupervised domain adaptation, a model trained on source-domain data (e.g., synthetic) is adapted to target-domain data (e.g., real-world) without access to target…