collaborators

11 papers

cs.CV2026

DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior

Junjia Huang, Binbin Yang, Pengxiang Yan +6

Storyboard synthesis plays a crucial role in visual storytelling, aiming to generate coherent shot sequences that visually narrate cinematic events with consistent characters, scen…

cs.CV2026

Self-Prophetic Decoding to Unlock Visual Search in LVLMs

Zhendong He, Qiyuan Dai, Guanbin Li +2

Large Vision-Language Models (LVLMs) are rapidly evolving toward true multimodal reasoning, with visual search representing a concrete instantiation of the thinking-with-images par…

cs.CV2026

LookasideVLN: Direction-Aware Aerial Vision-and-Language Navigation

Yuwei Ning, Ganlong Zhao, Yipeng Qin +4

Aerial Vision-and-Language Navigation (Aerial VLN) enables unmanned aerial vehicles (UAVs) to follow natural language instructions and navigate complex urban environments. While re…

cs.CV2026

OptiMVMap: Offline Vectorized Map Construction via Optimal Multi-vehicle Perspectives

Zedong Dan, Zijie Wang, Wei Zhang +6

Offline vectorized maps constitute critical infrastructure for high-precision autonomous driving and mapping services. Existing approaches rely predominantly on single ego-vehicle…

cs.CV2025

Sim-DETR: Unlock DETR for Temporal Sentence Grounding

Jiajin Tang, Zhengxuan Wei, Yuchen Zhu +4

Temporal sentence grounding aims to identify exact moments in a video that correspond to a given textual query, typically addressed with detection transformer (DETR) solutions. How…

cs.CV2025

Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI

Yang Liu, Weixing Chen, Yongjie Bai +4

Embodied Artificial Intelligence (Embodied AI) is crucial for achieving Artificial General Intelligence (AGI) and serves as a foundation for various applications (e.g., intelligent…