7 papers
Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution
Hanyi Zhang, Khang Nguyen, Charith Munasinghe +10
The paper introduces GCA-Bench, a new benchmark for evaluating robotic grasping in complex, multi-step scenarios that require scene-level reasoning and semantic constraints, and as…
RoboDesign1M: A Large-scale Dataset for Robot Design Understanding
Tri Le, Toan Nguyen, Quang Tran +6
The paper presents RoboDesign1M, a million‑sample multimodal dataset of robot designs collected from scientific literature, and shows its usefulness for tasks such as design image…
MAGS-SLAM: Monocular Multi-Agent Gaussian Splatting SLAM for Geometrically and Photometrically Consistent Reconstruction
Zhihao Cao, Qi Shao, Shuhao Zhai +4
Collaborative photorealistic 3D reconstruction from multiple agents enables rapid large-scale scene capture for virtual production and cooperative multi-robot exploration. While re…
AeroScene: Progressive Scene Synthesis for Aerial Robotics
Nghia Vu, Tuong Do, Dzung Tran +6
Generative models have shown substantial impact across multiple domains, their potential for scene synthesis remains underexplored in robotics. This gap is more evident in drone si…
SIGMA: A Physics-Based Benchmark for Gas Chimney Understanding in Seismic Images
Bao Truong, Quang Nguyen, Baoru Huang +6
Seismic images reconstruct subsurface reflectivity from field recordings, guiding exploration and reservoir monitoring. Gas chimneys are vertical anomalies caused by subsurface flu…
Humanoid Agent via Embodied Chain-of-Action Reasoning with Multimodal Foundation Models for Zero-Shot Loco-Manipulation
Congcong Wen, Geeta Chandra Raju Bethala, Yu Hao +8
Humanoid loco-manipulation, which integrates whole-body locomotion with dexterous manipulation, remains a fundamental challenge in robotics. Beyond whole-body coordination and bala…