3 papers
cs.RO2025
Embodied4C: Measuring What Matters for Embodied Vision-Language Navigation
Tin Stribor Sohn, Maximilian Dillitzer, Jason J. Corso +1
Vision-language navigation requires agents to reason and act under constraints of embodiment. While vision-language models (VLMs) demonstrate strong generalization, current benchma…
cs.CV2025
SNOW: Spatio-Temporal Scene Understanding with World Knowledge for Open-World Embodied Reasoning
Tin Stribor Sohn, Maximilian Dillitzer, Jason J. Corso +1
Autonomous robotic systems require spatio-temporal understanding of dynamic environments to ensure reliable navigation and interaction. While Vision-Language Models (VLMs) provide…
cs.CV2025
R4: Retrieval-Augmented Reasoning for Vision-Language Models in 4D Spatio-Temporal Space
Tin Stribor Sohn, Maximilian Dillitzer, Jason J. Corso +1
Humans perceive and reason about their surroundings in four dimensions by building persistent, structured internal representations that encode semantic meaning, spatial layout, and…