9 papers
HapticVLA: Contact-Rich Manipulation via Vision-Language-Action Model without Inference-Time Tactile Sensing
Konstantin Gubernatorov, Mikhail Sannikov, Ilya Mikhalchuk +8
Tactile sensing is a crucial capability for Vision-Language-Action (VLA) architectures, as it enables dexterous and safe manipulation in contact-rich tasks. However, reliance on de…
AnywhereVLA: Language-Conditioned Exploration and Mobile Manipulation
Konstantin Gubernatorov, Artem Voronov, Roman Voronov +4
We address natural language pick-and-place in unseen, unpredictable indoor environments with AnywhereVLA, a modular framework for mobile manipulation. A user text prompt serves as…
SwarmVLM: VLM-Guided Impedance Control for Autonomous Navigation of Heterogeneous Robots in Dynamic Warehousing
Malaika Zafar, Roohan Ahmed Khan, Faryal Batool +5
With the growing demand for efficient logistics, unmanned aerial vehicles (UAVs) are increasingly being paired with automated guided vehicles (AGVs). While UAVs offer the ability t…
METDrive: Multi-modal End-to-end Autonomous Driving with Temporal Guidance
Ziang Guo, Xinhao Lin, Zakhar Yagudin +4
Multi-modal end-to-end autonomous driving has shown promising advancements in recent work. By embedding more modalities into end-to-end networks, the system's understanding of both…
VDT-Auto: End-to-end Autonomous Driving with VLM-Guided Diffusion Transformers
Ziang Guo, Konstantin Gubernatorov, Selamawit Asfaw +2
In autonomous driving, dynamic environment and corner cases pose significant challenges to the robustness of ego vehicle's decision-making. To address these challenges, commencing…
VLM-Auto: VLM-based Autonomous Driving Assistant with Human-like Behavior and Understanding for Complex Road Scenes
Ziang Guo, Zakhar Yagudin, Artem Lykov +2
Recent research on Large Language Models for autonomous driving shows promise in planning and control. However, high computational demands and hallucinations still challenge accura…