3 papers
cs.CV2026
Learning Multi-View Spatial Reasoning from Cross-View Relations
Suchae Jeong, Jaehwi Song, Haeone Lee +9
Vision-language models (VLMs) have achieved impressive results on single-view vision tasks, but lack the multi-view spatial reasoning capabilities essential for embodied AI systems…
cs.RO2025
TRACE: A Self-Improving Framework for Robot Behavior Forecasting with Vision-Language Models
Gokul Puthumanaillam, Paulo Padrao, Jose Fuentes +6
Predicting the near-term behavior of a reactive agent is crucial in many robotic scenarios, yet remains challenging when observations of that agent are sparse or intermittent. Visi…
cs.RO2024
TAB-Fields: A Maximum Entropy Framework for Mission-Aware Adversarial Planning
Gokul Puthumanaillam, Jae Hyuk Song, Nurzhan Yesmagambet +2
Autonomous agents operating in adversarial scenarios face a fundamental challenge: while they may know their adversaries' high-level objectives, such as reaching specific destinati…