3 papers
cs.AI2026
How Foundational Skills Influence VLM-based Embodied Agents:A Native Perspective
Bo Peng, Pi Bu, Keyu Pan +7
Recent advances in vision-language models (VLMs) have shown promise for human-level embodied intelligence. However, existing benchmarks for VLM-driven embodied agents often rely on…
cs.CV2025
GeoPurify: A Data-Efficient Geometric Distillation Framework for Open-Vocabulary 3D Segmentation
Weijia Dou, Xu Zhang, Yi Bin +5
Recent attempts to transfer features from 2D Vision-Language Models (VLMs) to 3D semantic segmentation expose a persistent trade-off. Directly projecting 2D features into 3D yields…
cs.RO2025
AnchorDP3: 3D Affordance Guided Sparse Diffusion Policy for Robotic Manipulation
Ziyan Zhao, Ke Fan, He-Yang Xu +5
We present AnchorDP3, a diffusion policy framework for dual-arm robotic manipulation that achieves state-of-the-art performance in highly randomized environments. AnchorDP3 integra…