From the 1 of 9 linked papers with an AI index.
9 papers
RoboDesign1M: A Large-scale Dataset for Robot Design Understanding
Tri Le, Toan Nguyen, Quang Tran +6
The paper presents RoboDesign1M, a million‑sample multimodal dataset of robot designs collected from scientific literature, and shows its usefulness for tasks such as design image…
CodeGraphVLP: Code-as-Planner Meets Semantic-Graph State for Non-Markovian Vision-Language-Action Models
Khoa Vo, Sieu Tran, Taisei Hanyu +8
Vision-Language-Action (VLA) models promise generalist robot manipulation, but are typically trained and deployed as short-horizon policies that assume the latest observation is su…
Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
Khoa Vo, Taisei Hanyu, Yuki Ikebe +8
Recent Vision-Language-Action (VLA) models have made impressive progress toward general-purpose robotic manipulation by post-training large Vision-Language Models (VLMs) for action…
Learning Human Motion with Temporally Conditional Mamba
Quang Nguyen, Tri Le, Baoru Huang +4
Learning human motion based on a time-dependent input signal presents a challenging yet impactful task with various applications. The goal of this task is to generate or estimate h…
EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba
Quang Nguyen, Nhat Le, Baoru Huang +6
Estimating human dance motion is a challenging task with various industrial applications. Recently, many efforts have focused on predicting human dance motion using either egocentr…
GraspMAS: Zero-Shot Language-driven Grasp Detection with Multi-Agent System
Quang Nguyen, Tri Le, Huy Nguyen +5
Language-driven grasp detection has the potential to revolutionize human-robot interaction by allowing robots to understand and execute grasping tasks based on natural language com…