5 papers
InstructMove: A Text-Indispensable Benchmark for Instruction-Following Manipulation
Mengao Zhao, Ziang Li, Chaodong Huang +15
Vision-language-action (VLA) models have made general-purpose robot manipulation increasingly plausible by conditioning robot actions on natural-language instructions. A key test o…
HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory
Xiaolin Zhou, Liu Liu, Tingyang Xiao +9
LLM agents follow a practical execution loop in digital environments: they reason over structured states, invoke tools, inspect feedback, and revise actions. Extending this loop to…
ManiSplat: Manipulation Trajectory Synthesis from Monocular Video via Decoupled 3D Gaussian Splatting
Wenhao Hu, Haonan Zhou, Liu Liu +5
Reconstructing dynamic and interactive 3D scenes from real-world observations remains a fundamental challenge in computer vision and robotics. While recent advances in 3D Gaussian…
HoloBrain-0 Technical Report
Xuewu Lin, Tianwei Lin, Yun Du +12
In this work, we introduce HoloBrain-0, a comprehensive Vision-Language-Action (VLA) framework that bridges the gap between foundation model research and reliable real-world robot…
FineGrasp: Towards Robust Grasping for Delicate Objects
Yun Du, Mengao Zhao, Tianwei Lin +3
Recent advancements in robotic grasping have led to its integration as a core module in many manipulation systems. For instance, language-driven semantic segmentation enables the g…