4 papers
GenerativeMPC: VLM-RAG-guided Whole-Body MPC with Virtual Impedance for Bimanual Mobile Manipulation
Marcelino Julio Fernando, Miguel Altamirano Cabrera, Jeffrin Sam +3
Bimanual mobile manipulation requires a seamless integration between high-level semantic reasoning and safe, compliant physical interaction - a challenge that end-to-end models app…
HapticVLA: Contact-Rich Manipulation via Vision-Language-Action Model without Inference-Time Tactile Sensing
Konstantin Gubernatorov, Mikhail Sannikov, Ilya Mikhalchuk +8
Tactile sensing is a crucial capability for Vision-Language-Action (VLA) architectures, as it enables dexterous and safe manipulation in contact-rich tasks. However, reliance on de…
AnywhereVLA: Language-Conditioned Exploration and Mobile Manipulation
Konstantin Gubernatorov, Artem Voronov, Roman Voronov +4
We address natural language pick-and-place in unseen, unpredictable indoor environments with AnywhereVLA, a modular framework for mobile manipulation. A user text prompt serves as…
VDT-Auto: End-to-end Autonomous Driving with VLM-Guided Diffusion Transformers
Ziang Guo, Konstantin Gubernatorov, Selamawit Asfaw +2
In autonomous driving, dynamic environment and corner cases pose significant challenges to the robustness of ego vehicle's decision-making. To address these challenges, commencing…