Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
FocusGraph: Graph-Structured Frame Selection for Embodied Long Video Question Answering
Tatiana Zemskova, Solomon Andryushenko, Ilya Obrubov +4
The ability to understand long videos is vital for embodied intelligent agents, because their effectiveness depends on how well they can accumulate, organize, and leverage long-hor…
cs.CV2025
KM-ViPE: Online Tightly Coupled Vision-Language-Geometry Fusion for Open-Vocabulary Semantic SLAM
Zaid Nasser, Mikhail Iumanov, Tianhao Li +7
We present KM-ViPE (Knowledge Mapping Video Pose Engine), a real-time open-vocabulary SLAM framework for uncalibrated monocular cameras in dynamic environments. Unlike systems requ…