From the 2 of 15 linked papers with an AI index.
15 papers
Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution
Hanyi Zhang, Khang Nguyen, Charith Munasinghe +10
The paper introduces GCA-Bench, a new benchmark for evaluating robotic grasping in complex, multi-step scenarios that require scene-level reasoning and semantic constraints, and as…
RoboDesign1M: A Large-scale Dataset for Robot Design Understanding
Tri Le, Toan Nguyen, Quang Tran +6
The paper presents RoboDesign1M, a million‑sample multimodal dataset of robot designs collected from scientific literature, and shows its usefulness for tasks such as design image…
SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models
Jiesong Lian, Zixiang Zhou, Ruizhe Zhong +6
Recent video diffusion models (VDMs) synthesize visually convincing clips, yet still drop entities, mis-bind attributes, and weaken the interactions specified in the prompt. Repres…
CoMo3R-SLAM: Collaborative Monocular Dense SLAM with Learned 3D Reconstruction Priors for Outdoor Multi-Agent Systems
Zhihao Cao, Qi Shao, Shuhao Zhai +4
Collaborative dense SLAM is essential for multi-robot teams to achieve scalable and consistent 3D perception across large-scale outdoor environments. Existing systems typically dep…
MAGS-SLAM: Monocular Multi-Agent Gaussian Splatting SLAM for Geometrically and Photometrically Consistent Reconstruction
Zhihao Cao, Qi Shao, Shuhao Zhai +4
Collaborative photorealistic 3D reconstruction from multiple agents enables rapid large-scale scene capture for virtual production and cooperative multi-robot exploration. While re…
AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers
Nghia Vu, Tuong Do, Khang Nguyen +8
Affordance learning is a complex challenge in many applications, where existing approaches primarily focus on the geometric structures, visual knowledge, and affordance labels of o…