From the 1 of 10 linked papers with an AI index.
10 papers
Traj-VLN: Learning Pixel-Space Interaction via Autoregressive Trajectory Generation
Changfei Fu, Guangcheng Chen, Wenjun Xu +3
The paper introduces a method that fine‑tunes vision‑language models to predict sequences of pixel coordinates, enabling an embodied agent to follow natural language navigation ins…
NavIsaacLab: Generating Realistic Crowd via Parallel Robot Learning for Benchmarking Human-aware Navigation
Bingyi Xia, Han Bao, Jingyu Zhu +6
Robot autonomous navigation that accounts for surrounding human activities is crucial for ensuring both safety and natural human-robot interaction in real-world environments shared…
Can Single-View Mesh Reconstruction Generalize to Robot Camera Rotation?
Yu Zhan, Guangcheng Chen, Hanjing Ye +4
Single-view mesh reconstruction predicts object meshes and spatial layouts from a single observation, making it attractive for fast robot spatial reasoning and real-to-sim digital…
UniviewVLA: A Unified Multiview Vision-Language-Action Model with World Modeling
Tao Xu, Runhao Zhang, Zhijian Huang +7
Occluded tasks remain a bottleneck in robot manipulation. Existing solutions either deploy additional physical cameras requiring training-inference camera parity, or rely on explic…
FLM-Occ: Feed-forward Likelihood Maximization for Efficient Indoor Occupancy Prediction
Guangcheng Chen, Lihuang Fang, Huaqi Tao +3
Recent indoor occupancy prediction methods adopt Gaussian primitives as a sparse 3D representation for computational efficiency. However, their training relies on voxel classificat…
Enhancing Glass Surface Reconstruction via Depth Prior for Robot Navigation
Jiamin Zheng, Jingwen Yu, Guangcheng Chen +1
Indoor robot navigation is often compromised by glass surfaces, which severely corrupt depth sensor measurements. While foundation models like Depth Anything 3 provide excellent ge…