Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
GESTO: Human-Centric Spatio-Temporal Memory for Reasoning in Dynamic Scenes
Ermanno Bartoli, Buwei He, Dennis Rotondi +6
Robots operating in human environments need memories that capture not only what objects exist and where, but also how people use them over time and how individual interactions comp…
cs.CV2026
Social 3D Scene Graphs: Modeling Human Actions and Relations for Interactive Service Robots
Ermanno Bartoli, Dennis Rotondi, Buwei He +3
Understanding how people interact with their surroundings and each other is essential for enabling robots to act in socially compliant and context-aware ways. While 3D Scene Graphs…
cs.CV2026
MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue
Anna Deichler, Jim O'Regan, Fethiye Irmak Dogan +4
Grounding language in the physical world requires AI systems to interpret references that emerge dynamically during conversation. While current vision-language models (VLMs) excel…