4 papers
OVSegDT: Segmenting Transformer for Open-Vocabulary Object Goal Navigation
Tatiana Zemskova, Aleksei Staroverov, Dmitry Yudin +1
Open-vocabulary Object Goal Navigation requires an embodied agent to reach objects described by free-form language, including categories never seen during training. Existing end-to…
FocusGraph: Graph-Structured Frame Selection for Embodied Long Video Question Answering
Tatiana Zemskova, Solomon Andryushenko, Ilya Obrubov +4
The ability to understand long videos is vital for embodied intelligent agents, because their effectiveness depends on how well they can accumulate, organize, and leverage long-hor…
3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding
Tatiana Zemskova, Dmitry Yudin
A 3D scene graph represents a compact scene model by capturing both the objects present and the semantic relationships between them, making it a promising structure for robotic app…
Beyond Bare Queries: Open-Vocabulary Object Grounding with 3D Scene Graph
Sergey Linok, Tatiana Zemskova, Svetlana Ladanova +4
Locating objects described in natural language presents a significant challenge for autonomous agents. Existing CLIP-based open-vocabulary methods successfully perform 3D object gr…