9 papers
ARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual Conditioning
Zijie Wang, Wei Zhang, Weiming Zhang +4
ARDepth proposes an auto-regressive approach to monocular depth estimation that builds depth maps progressively across increasing spatial resolutions, using scale‑progressive condi…
World Models as Group Actions
Zijie Wang, Wei Zhang, Weiming Zhang +4
Video world models have achieved strong visual realism, but this does not ensure that their dynamics are truly governed by actions. In this work, we argue that action faithfulness…
HisTrackMap: Global Vectorized High-Definition Map Construction via History Map Tracking
Jing Yang, Sen Yang, Xiao Tan +1
As an essential component of autonomous driving systems, high-definition (HD) maps provide rich and precise environmental information for auto-driving scenarios; however, existing…
Thinking While Driving: A Concurrent Framework for Real-Time, LLM-Based Adaptive Routing
Xiaopei Tan, Muyang Fan
We present Thinking While Driving, a concurrent routing framework that integrates LLMs into a graph-based traffic environment. Unlike approaches that require agents to stop and del…
CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
Yu Qi, Yumeng Zhang, Chenting Gong +4
Large Vision-Language Models (LVLMs) have demonstrated remarkable success in a broad range of vision-language tasks, such as general visual question answering and optical character…
Fin3R: Fine-tuning Feed-forward 3D Reconstruction Models via Monocular Knowledge Distillation
Weining Ren, Hongjun Wang, Xiao Tan +1
We present Fin3R, a simple, effective, and general fine-tuning method for feed-forward 3D reconstruction models. The family of feed-forward reconstruction model regresses pointmap…