3 papers
cs.HC2026
SpaceVLA: Spatially Grounded VLA for Robotic Manipulation with User-Authored Grasp and Place Anchors
Daniia Zinniatullina, Iaroslav Kolomiets, Mikhail Konenkov +2
Vision-language-action (VLA) models follow language commands but often lack explicit spatial intent for manipulation. We present Visual Intent Anchors, an XR pipeline that lets use…
cs.RO2026
AgenticFocus: Object-Preserving Mixed Reality Synthesis from Human FPV Video for Dexterous Humanoid Learning
Iaroslav Kolomiets, Miguel Altamirano Cabrera, Artem Lykov +6
Human egocentric video is a scalable supervision source for humanoid policy learning, but current pipelines struggle with hand-object occlusion, oversimplified motion, or specializ…
cs.RO2026
VersualRL: Closed-Loop Verbal Reinforcement Learning with Visual Execution Feedback for Task-Level Robot Planning
Dmitrii Plotnikov, Iaroslav Kolomiets, Dmitrii Maliukov +9
We introduce VersualRL, a closed-loop framework for task-level robot planning that uses visual execution feedback to iteratively refine executable Behavior Trees through structured…