4 papers
Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models
Xinyi Xie, Zican Hu, Zhanyu Liu +7
Vision-Language-Action (VLA) models acquire broad embodied capabilities through large-scale pretraining, yet their generalization remains far more fragile than that of LLMs and VLM…
Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds
Jiaming Bian, Bingliang Li, Yuehao Wu +5
As embodied AI and world models increasingly operate in dynamic 3D environments, visual perception must move beyond passively interpreting given observations toward actively decidi…
Social Structure Matters in 3D Human-Human Interaction Generation
Zhongju Wang, Beier Wang, Yatao Bian +6
Although text-to-motion generation has achieved strong progress in synthesizing realistic single-person motions from language, extending it to text-driven 3D human-human interactio…
HOT: Hierarchical Hourglass Tokenizer for Efficient Video Pose Transformers
Wenhao Li, Mengyuan Liu, Hong Liu +3
Transformers have been successfully applied in the field of video-based 3D human pose estimation. However, the high computational costs of these video pose transformers (VPTs) make…