6 papers
VLGA: Vision-Language-Geometry-Action Models for Autonomous Driving
Jin Yao, Dhruva Dixith Kurra, Tom Lampo +3
Vision-language-action (VLA) models can describe scenes and reason about them in language, yet still struggle to ground their actions in the dense 3D world around them. Existing ap…
Generalizable and Actionable Parts Pose Estimation with Symmetry Annotation-Free Learning Strategy
Wenxiao Chen, Xueyu Yuan, Liu Liu +2
Urgently needed generalizable robot object interaction and manipulation requires high-quality Cross-Category object perception. As a pioneer of this area, Generalizable and Actiona…
MotionAnymesh: Physics-Grounded Articulation for Simulation-Ready Digital Twins
WenBo Xu, Liu Liu, Li Zhang +2
Converting static 3D meshes into interactable articulated assets is crucial for embodied AI and robotic simulation. However, existing zero-shot pipelines struggle with complex asse…
SynHLMA:Synthesizing Hand Language Manipulation for Articulated Object with Discrete Human Object Interaction Representation
Wang zhi, Yuyan Liu, Liu Liu +3
Generating hand grasps with language instructions is a widely studied topic that benefits from embodied AI and VR/AR applications. While transferring into hand articulatied object…
KineDiff3D: Kinematic-Aware Diffusion for Category-Level Articulated Object Shape Reconstruction and Generation
WenBo Xu, Liu Liu, Li Zhang +4
Articulated objects, such as laptops and drawers, exhibit significant challenges for 3D reconstruction and pose estimation due to their multi-part geometries and variable joint con…
Distilling Textual Priors from LLM to Efficient Image Fusion
Ran Zhang, Xuanhua He, Ke Cao +4
Multi-modality image fusion aims to synthesize a single, comprehensive image from multiple source inputs. Traditional approaches, such as CNNs and GANs, offer efficiency but strugg…