From the 1 of 13 linked papers with an AI index.
13 papers
AgenticFocus: Object-Preserving Mixed Reality Synthesis from Human FPV Video for Dexterous Humanoid Learning
Iaroslav Kolomiets, Miguel Altamirano Cabrera, Artem Lykov +6
The paper presents AgenticFocus, a mixed-reality pipeline that turns ordinary first-person human videos into robot-ready demonstrations by reconstructing hidden object geometry, co…
Output-Level Regularization Eliminates the Seed Lottery in Single-GPU VLA Fine-Tuning
Jeffrin Sam, Dzmitry Tsetserukou
Fine-tuning a vision-language-action model (VLA-JEPA) on a single GPU should be simple: load a pretrained checkpoint, run training, deploy. There is a hidden danger. Run the same f…
Action Agent: Agentic Video Generation Meets Flow-Constrained Diffusion
Jeffrin Sam, Nguyen Khang, Yara Mahmoud +2
We present Action Agent, a two-stage framework that unifies agentic navigation video generation with flow-constrained diffusion control for multi-embodiment robot navigation. In St…
GenerativeMPC: VLM-RAG-guided Whole-Body MPC with Virtual Impedance for Bimanual Mobile Manipulation
Marcelino Julio Fernando, Miguel Altamirano Cabrera, Jeffrin Sam +3
Bimanual mobile manipulation requires a seamless integration between high-level semantic reasoning and safe, compliant physical interaction - a challenge that end-to-end models app…
DiffusionAnything: End-to-End In-context Diffusion Learning for Unified Navigation and Pre-Grasp Motion
Iana Zhura, Yara Mahmoud, Jeffrin Sam +4
Efficiently predicting motion plans directly from vision remains a fundamental challenge in robotics, where planning typically requires explicit goal specification and task-specifi…
DreamToNav: Generalizable Navigation for Robots via Generative Video Planning
Valerii Serpiva, Jeffrin Sam, Chidera Simon +4
We present DreamToNav, a novel autonomous robot framework that uses generative video models to enable intuitive, human-in-the-loop control. Instead of relying on rigid waypoint nav…