6 papers
ViP-Rig: Visual-Prompted Controllable Rigging
Zihan Qin, Mingze Sun, Yifan Mao +6
Rigging is inherently task-dependent because the same mesh may require different skeletons and deformation behaviors across animation tasks. In practice, artists often inspect an i…
3D Generation for Embodied AI and Robotic Simulation: A Survey
Tianwei Ye, Yifan Mao, Minwen Liao +6
Embodied AI and robotic systems increasingly depend on scalable, diverse, and physically grounded 3D content for simulation-based training and real-world deployment. While 3D gener…
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
Inclusion AI, :, Fudong Wang +12
Recent advancements in Multimodal Large Language Models (MLLMs), particularly through Reinforcement Learning with Verifiable Rewards (RLVR), have significantly enhanced their reaso…
The Fourth Monocular Depth Estimation Challenge
Anton Obukhov, Matteo Poggi, Fabio Tosi +54
This paper presents the results of the fourth edition of the Monocular Depth Estimation Challenge (MDEC), which focuses on zero-shot generalization to the SYNS-Patches benchmark, a…
The RoboDrive Challenge: Drive Anytime Anywhere in Any Condition
Lingdong Kong, Shaoyuan Xie, Hanjiang Hu +88
In the realm of autonomous driving, robust perception under out-of-distribution conditions is paramount for the safe deployment of vehicles. Challenges such as adverse weather, sen…
DINO-SD: Champion Solution for ICRA 2024 RoboDepth Challenge
Yifan Mao, Ming Li, Jian Liu +7
Surround-view depth estimation is a crucial task aims to acquire the depth maps of the surrounding views. It has many applications in real world scenarios such as autonomous drivin…