8 papers
Top-down Traffic Scenario Generation via Joint Initial-Goal Diffusion and Trajectory Infilling
Da Saem Lee, Yash Vardhan Pant, Sebastian Fischmeister
Robust traffic simulators are crucial for developing and testing autonomous vehicles to reduce the costly, labor-intensive real-world data collection process and the need for physi…
IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation
Yifan Li, Lichi Li, Anh Dao +15
While Visual Large Language Models (VLLMs) show great promise as embodied agents, they continue to face substantial challenges in spatial reasoning. Existing embodied benchmarks la…
Why Far Looks Up: Probing Spatial Representation in Vision-Language Models
Cheolhong Min, Jaeyun Jung, Daeun Lee +5
Vision-language models (VLMs) achieve strong performance on spatial reasoning benchmarks, yet it remains unclear whether this reflects structured 3D understanding or reliance on st…
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
Daeun Lee, Subhojyoti Mukherjee, Branislav Kveton +6
Streaming video understanding requires models not only to process temporally incoming frames, but also to anticipate user intention for realistic applications such as Augmented Rea…
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
Daeun Lee, Jaehong Yoon, Jaemin Cho +1
Recent text-to-video (T2V) diffusion models have made remarkable progress in generating high-quality videos. However, they often struggle to align with complex text prompts, partic…
VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting
Daeun Lee, Shoubin Yu, Yue Zhang +1
Video reasoning requires models to locate and track question-relevant evidence across frames. While reinforcement learning (RL) with verifiable rewards improves accuracy, it still…