activity
20242026
collaborators

8 papers

cs.CV2026

OSMa-Bench++: Toward Open-Ended Benchmarking of Semantic Mapping for Manipulation with Prompt-Generated Synthetic Scenes

Regina Kurkova, Maxim Popov, Sergey Kolyubin

Semantic mapping methods are increasingly used as intermediate scene representations for downstream robotic reasoning and manipulation, yet their evaluation is still largely tied t…

cs.CV2026

R5DGS: Semantic-Aware 4D Gaussian Splatting with Rigid Body Constraints for Efficient Dynamic Scene Reconstruction

Denis Gridusov, Maxim Popov, Sergey Kolyubin

Reconstructing and predicting dynamic 3D scenes from multi-view videos is a foundational task for robotics, AR/VR, and digital twins. Recent physics-informed Gaussian Splatting met…

cs.CV2026

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models

Cuong Huynh, Maxim Popov, Denis Gridusov +1

3D Visual Grounding (3DVG) is an essential capability for embodied AI, requiring agents to localize objects in 3D scenes based on natural language descriptions. Recent zero-shot me…

cs.CV2026

RADIO-ViPE: Online Tightly Coupled Multi-Modal Fusion for Open-Vocabulary Semantic SLAM in Dynamic Environments

Zaid Nasser, Mikhail Iumanov, Tianhao Li +3

We present RADIO-ViPE (Reduce All Domains Into One -- Video Pose Engine), an online semantic SLAM system that enables geometry-aware open-vocabulary grounding, associating arbitrar…

cs.CV2026

OSMa-Bench: Evaluating Open Semantic Mapping Under Varying Lighting Conditions

Maxim Popov, Regina Kurkova, Mikhail Iumanov +2

Open Semantic Mapping (OSM) is a key technology in robotic perception, combining semantic segmentation and SLAM techniques. This paper introduces a dynamically configurable and hig…

cs.CV2025

KM-ViPE: Online Tightly Coupled Vision-Language-Geometry Fusion for Open-Vocabulary Semantic SLAM

Zaid Nasser, Mikhail Iumanov, Tianhao Li +7

We present KM-ViPE (Knowledge Mapping Video Pose Engine), a real-time open-vocabulary SLAM framework for uncalibrated monocular cameras in dynamic environments. Unlike systems requ…