6 papers · 1 filter
Why Far Looks Up: Probing Spatial Representation in Vision-Language Models
Cheolhong Min, Jaeyun Jung, Daeun Lee +5
Vision-language models (VLMs) achieve strong performance on spatial reasoning benchmarks, yet it remains unclear whether this reflects structured 3D understanding or reliance on st…
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
Daeun Lee, Subhojyoti Mukherjee, Branislav Kveton +6
Streaming video understanding requires models not only to process temporally incoming frames, but also to anticipate user intention for realistic applications such as Augmented Rea…
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
Daeun Lee, Jaehong Yoon, Jaemin Cho +1
Recent text-to-video (T2V) diffusion models have made remarkable progress in generating high-quality videos. However, they often struggle to align with complex text prompts, partic…
VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting
Daeun Lee, Shoubin Yu, Yue Zhang +1
Video reasoning requires models to locate and track question-relevant evidence across frames. While reinforcement learning (RL) with verifiable rewards improves accuracy, it still…
Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning
Daeun Lee, Jaehong Yoon, Jaemin Cho +1
Recent advances in Chain-of-Thought (CoT) reasoning have improved complex video understanding, but existing methods often struggle to adapt to domain-specific skills (e.g., event d…
HD Maps are Lane Detection Generalizers: A Novel Generative Framework for Single-Source Domain Generalization
Daeun Lee, Minhyeok Heo, Jiwon Kim
Lane detection is a vital task for vehicles to navigate and localize their position on the road. To ensure reliable driving, lane detection models must have robust generalization p…