12 papers
Representation Handoffs for OpenArm-Based Laboratory Mobile Manipulation
Yang Shen, Chonghao Cheng, Ziyi Zhao +6
Open-source robotics and foundation models have lowered the barrier to embodied AI, yet language-guided laboratory automation still requires reliable alignment from instructions an…
Anchor3R: Streaming 3D Reconstruction with Transient Anchors for Long-Horizon Visual Mapping
Peilin Tao, Chong Cheng, Yuansen Du +8
Long-horizon online visual mapping is a core capability for robot perception, requiring continuous camera-motion and scene-geometry estimation from visual streams under bounded mem…
HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction
Chong Cheng, Peilin Tao, Nanjie Yao +9
Online 3D reconstruction requires estimating camera pose and scene geometry under strict causal and bounded-memory constraints. Existing methods often suffer from drift, jitter, or…
LongStream: Long-Sequence Streaming Autoregressive Visual Geometry
Chong Cheng, Xianda Chen, Tao Xie +5
Long-sequence streaming 3D reconstruction remains a significant open challenge. Existing autoregressive models often fail when processing long sequences because they anchor poses t…
VGGT4D: Mining Motion Cues in Visual Geometry Transformers for 4D Scene Reconstruction
Yu Hu, Chong Cheng, Sicheng Yu +2
Reconstructing dynamic 4D scenes is challenging, as it requires robust disentanglement of dynamic objects from the static background. While 3D foundation models like VGGT provide a…
Navigation with VLM framework: Towards Going to Any Language
Zecheng Yin, Chonghao Cheng, and Yao Guo +1
Navigating towards fully open language goals and exploring open scenes in an intelligent way have always raised significant challenges. Recently, Vision Language Models (VLMs) have…