From the 1 of 6 linked papers with an AI index.
6 papers
Speech2Grasp: Data-Efficient Transfer of Text-Conditioned Grasp Detection to Speech in Humanoid Robots
Hung Nguyen, Kim Nhat Minh Nguyen, Van Duc Vu +6
The paper presents Speech2Grasp, a method that efficiently adapts a text‑conditioned grasp‑detection model to work directly with spoken commands, improving performance and speed on…
3D Point World Models: Point Completion Enables More Accurate Dynamics Learning
Skand Peri, Hung Nguyen, Chanho Kim +2
Learning predictive models of the world enables robotic control through planning, potentially allowing robots to improvise solutions on new tasks. However, large video-based dynami…
RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis
Minh-Loi Nguyen, Nghiem Tuong Diep, Hung Khang Nguyen +10
Recent advances in robot world models enable synthetic video generation for embodied prediction and planning. However, evaluating these videos is challenging: visually realistic ou…
DiffusionAnything: End-to-End In-context Diffusion Learning for Unified Navigation and Pre-Grasp Motion
Iana Zhura, Yara Mahmoud, Jeffrin Sam +4
Efficiently predicting motion plans directly from vision remains a fundamental challenge in robotics, where planning typically requires explicit goal specification and task-specifi…
SafeHumanoid: VLM-RAG-driven Control of Upper Body Impedance for Humanoid Robot
Yara Mahmoud, Jeffrin Sam, Nguyen Khang +6
Safe and trustworthy Human Robot Interaction (HRI) requires robots not only to complete tasks but also to regulate impedance and speed according to scene context and human proximit…
PhysicalAgent: Towards General Cognitive Robotics with Foundation World Models
Artem Lykov, Jeffrin Sam, Hung Khang Nguyen +6
We introduce PhysicalAgent, an agentic framework for robotic manipulation that integrates iterative reasoning, diffusion-based video generation, and closed-loop execution. Given a…