works on

From the 1 of 24 linked papers with an AI index.

collaborators

24 papers

cs.HC2026

SpaceVLA: Spatially Grounded VLA for Robotic Manipulation with User-Authored Grasp and Place Anchors

Daniia Zinniatullina, Iaroslav Kolomiets, Mikhail Konenkov +2

Vision-language-action (VLA) models follow language commands but often lack explicit spatial intent for manipulation. We present Visual Intent Anchors, an XR pipeline that lets use…

cs.RO2026

GeminiPainter's sequence-formed pipeline comprised of perception, cognition, planning, and action stages

Miguel Altamirano Cabrera, Aleksey Fedoseev, Iana Zhura +1

We present an autonomous robotic portrait-generation system combining real-time face detection, AI-based sketch generation, and robotic drawing. The system captures video frames, e…

cs.RO2026

ORCESTRA: VLM-driven Visual Robot programming in Mixed Reality

Ivan Snegirev, Elizaveta Semenyakina, Mikhail Konenkov +3

ORCESTRA is a mixed-reality system for programming robot digital twins through no-code waypoint teaching and language-guided control. In a passthrough mixed-reality workspace, user…

cs.RO2026

OmniAI: A Surface-Adaptive Aerial Projection Interface for Human--Drone Interaction

Nikita Kuzmin, Yuhua Jin, Georgii Demianchuk +6

Drones in human environments often lack spatially grounded in- terfaces for situated communication. We present OmniAI, an em- bodied aerial agent that supports surface-adaptive int…

cs.RO2026

AgenticFocus: Object-Preserving Mixed Reality Synthesis from Human FPV Video for Dexterous Humanoid Learning

Iaroslav Kolomiets, Miguel Altamirano Cabrera, Artem Lykov +6

The paper presents AgenticFocus, a mixed-reality pipeline that turns ordinary first-person human videos into robot-ready demonstrations by reconstructing hidden object geometry, co…

cs.CV2026

AnythingReality: Robust Online Gaussian Splatting SLAM for Open-Vocabulary VR Scene Exploration

Timofei Kozlov, Dmitrii Maliukov, Andrey Marchenko +2

We present a novel integrated architecture for robust online 3D Gaussian splatting, real-time VR exploration, and speech-driven Vision-Language-Model interaction. Unlike methods as…