5 papers
2K Retrofit: Entropy-Guided Efficient Sparse Refinement for High-Resolution 3D Geometry Prediction
Tianbao Zhang, Zhenyu Liang, Zhenbo Song +7
High-resolution geometric prediction is essential for robust perception in autonomous driving, robotics, and AR/MR, but current foundation models are fundamentally limited by their…
VAMPO: Policy Optimization for Improving Visual Dynamics in Video Action Models
Zirui Ge, Pengxiang Ding, Baohua Yin +16
Video action models are an appealing foundation for Vision--Language--Action systems because they can learn visual dynamics from large-scale video data and transfer this knowledge…
HieroAction: Hierarchically Guided VLM for Fine-Grained Action Analysis
Junhao Wu, Xiuer Gu, Zhiying Li +6
Evaluating human actions with clear and detailed feedback is important in areas such as sports, healthcare, and robotics, where decisions rely not only on final outcomes but also o…
Pose-RFT: Enhancing MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning
Bao Li, Xiaomei Zhang, Miao Xu +3
Generating 3D human poses from multimodal inputs such as images or text requires models to capture both rich spatial and semantic correspondences. While pose-specific multimodal la…
Moderating the Generalization of Score-based Generative Model
Wan Jiang, He Wang, Xin Zhang +4
Score-based Generative Models (SGMs) have demonstrated remarkable generalization abilities, e.g. generating unseen, but natural data. However, the greater the generalization power,…