works on

From the 2 of 8 linked papers with an AI index.

activity
20242026
collaborators

17 papers

cs.RO2026

Gripper-aware Vision Language Action Models

Hanyi Zhang, Zihong Luo, Tianyu Li +16

Vision language action models (VLAs) have advanced general purpose robotic grasping and manipulation by enabling robots to interpret visual observations and natural language instru…

cs.RO2026

Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution

Hanyi Zhang, Khang Nguyen, Charith Munasinghe +10

The paper introduces GCA-Bench, a new benchmark for evaluating robotic grasping in complex, multi-step scenarios that require scene-level reasoning and semantic constraints, and as…

cs.RO2026

RoboDesign1M: A Large-scale Dataset for Robot Design Understanding

Tri Le, Toan Nguyen, Quang Tran +6

The paper presents RoboDesign1M, a million‑sample multimodal dataset of robot designs collected from scientific literature, and shows its usefulness for tasks such as design image…

cs.RO2026

MAGS-SLAM: Monocular Multi-Agent Gaussian Splatting SLAM for Geometrically and Photometrically Consistent Reconstruction

Zhihao Cao, Qi Shao, Shuhao Zhai +4

Collaborative photorealistic 3D reconstruction from multiple agents enables rapid large-scale scene capture for virtual production and cooperative multi-robot exploration. While re…

cs.RO2026

AeroScene: Progressive Scene Synthesis for Aerial Robotics

Nghia Vu, Tuong Do, Dzung Tran +6

Generative models have shown substantial impact across multiple domains, their potential for scene synthesis remains underexplored in robotics. This gap is more evident in drone si…

cs.CV2026

SIGMA: A Physics-Based Benchmark for Gas Chimney Understanding in Seismic Images

Bao Truong, Quang Nguyen, Baoru Huang +6

Seismic images reconstruct subsurface reflectivity from field recordings, guiding exploration and reservoir monitoring. Gas chimneys are vertical anomalies caused by subsurface flu…