collaborators

7 papers

cs.CV2026

A DVDrive Approach for doScenes Instructed Driving Challenge

Zijian Fu, Xiangyang Chu, Mengshi Qi +3

Instruction-conditioned trajectory prediction is an emerging problem in autonomous driving, where a model predicts the future ego trajectory not only from visual scene context and…

cs.CV2026

XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments

Kangan Qian, ChuChu Xie, Yang Zhong +13

Vision-Language-Action (VLA) models drive next-generation autonomous systems, but training them requires scalable, high-quality annotations from complex environments. Current cloud…

cs.CV2025

DSI-Bench: A Benchmark for Dynamic Spatial Intelligence

Ziang Zhang, Zehan Wang, Guanghao Zhang +5

Reasoning about dynamic spatial relationships is essential, as both observers and objects often move simultaneously. Although vision-language models (VLMs) and visual expertise mod…

cs.CV2025

CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation

Guanghao Zhang, Tao Zhong, Yan Xia +8

While previous multimodal slow-thinking methods have demonstrated remarkable success in single-image understanding scenarios, their effectiveness becomes fundamentally constrained…

cs.CV2025

Towards Enhanced Image Generation Via Multi-modal Chain of Thought in Unified Generative Models

Yi Wang, Mushui Liu, Wanggui He +13

Unified generative models have shown remarkable performance in text and image generation. For image synthesis tasks, they adopt straightforward text-to-image (T2I) generation. Howe…

cs.CV2025

Streaming Video Question-Answering with In-context Video KV-Cache Retrieval

Shangzhe Di, Zhelun Yu, Guanghao Zhang +7

We propose ReKV, a novel training-free approach that enables efficient streaming video question-answering (StreamingVQA), by seamlessly integrating with existing Video Large Langua…