works on

From the 1 of 11 linked papers with an AI index.

collaborators

11 papers

cs.RO2026

AgenticFocus: Object-Preserving Mixed Reality Synthesis from Human FPV Video for Dexterous Humanoid Learning

Iaroslav Kolomiets, Miguel Altamirano Cabrera, Artem Lykov +6

The paper presents AgenticFocus, a mixed-reality pipeline that turns ordinary first-person human videos into robot-ready demonstrations by reconstructing hidden object geometry, co…

cs.RO2026

Action Agent: Agentic Video Generation Meets Flow-Constrained Diffusion

Jeffrin Sam, Nguyen Khang, Yara Mahmoud +2

We present Action Agent, a two-stage framework that unifies agentic navigation video generation with flow-constrained diffusion control for multi-embodiment robot navigation. In St…

cs.RO2026

GenerativeMPC: VLM-RAG-guided Whole-Body MPC with Virtual Impedance for Bimanual Mobile Manipulation

Marcelino Julio Fernando, Miguel Altamirano Cabrera, Jeffrin Sam +3

Bimanual mobile manipulation requires a seamless integration between high-level semantic reasoning and safe, compliant physical interaction - a challenge that end-to-end models app…

cs.RO2026

DiffusionAnything: End-to-End In-context Diffusion Learning for Unified Navigation and Pre-Grasp Motion

Iana Zhura, Yara Mahmoud, Jeffrin Sam +4

Efficiently predicting motion plans directly from vision remains a fundamental challenge in robotics, where planning typically requires explicit goal specification and task-specifi…

cs.RO2026

HumanoidVLM: Vision-Language-Guided Impedance Control for Contact-Rich Humanoid Manipulation

Yara Mahmoud, Yasheerah Yaqoot, Miguel Altamirano Cabrera +1

Humanoid robots must adapt their contact behavior to diverse objects and tasks, yet most controllers rely on fixed, hand-tuned impedance gains and gripper settings. This paper intr…

cs.HC2026

LLM-Glasses: GenAI-driven Glasses with Haptic Feedback for Navigation of Visually Impaired People

Issatay Tokmurziyev, Miguel Altamirano Cabrera, Muhammad Haris Khan +2

LLM-Glasses is a wearable navigation system which assists visually impaired people by utilizing YOLO-World object detection, GPT-4o-based reasoning, and haptic feedback for real-ti…