6 papers
-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills
Siyao Xiao, Yuhong Zhang, Zhifang Liu +9
Current Vision-Language-Action (VLA) models predominantly rely on end-to-end fine-tuning. While effective, this paradigm compromises the inherent generalization capabilities of Vis…
RoboWheel: A Data Engine from Real-World Human Demonstrations for Cross-Embodiment Robotic Learning
Yuhong Zhang, Zihan Gao, Shengpeng Li +12
We introduce Robowheel, a data engine that converts human hand object interaction (HOI) videos into training-ready supervision for cross morphology robotic learning. From monocular…
Motion2Motion: Cross-topology Motion Transfer with Sparse Correspondence
Ling-Hao Chen, Yuhong Zhang, Zixin Yin +5
This work studies the challenge of transfer animations between characters whose skeletal topologies differ substantially. While many techniques have advanced retargeting techniques…
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
Tianhe Ren, Yihao Chen, Qing Jiang +17
In this paper, we introduce DINO-X, which is a unified object-centric vision model developed by IDEA Research with the best open-world object detection performance to date. DINO-X…
HumanMM: Global Human Motion Recovery from Multi-shot Videos
Yuhong Zhang, Guanlin Wu, Ling-Hao Chen +8
In this paper, we present a novel framework designed to reconstruct long-sequence 3D human motion in the world coordinates from in-the-wild videos with multiple shot transitions. S…
Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset
Yuhong Zhang, Jing Lin, Ailing Zeng +7
In this paper, we introduce Motion-X++, a large-scale multimodal 3D expressive whole-body human motion dataset. Existing motion datasets predominantly capture body-only poses, lack…