collaborators

9 papers

cs.CV2026

ARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual Conditioning

Zijie Wang, Wei Zhang, Weiming Zhang +4

ARDepth proposes an auto-regressive approach to monocular depth estimation that builds depth maps progressively across increasing spatial resolutions, using scale‑progressive condi…

cs.CV2026

World Models as Group Actions

Zijie Wang, Wei Zhang, Weiming Zhang +4

Video world models have achieved strong visual realism, but this does not ensure that their dynamics are truly governed by actions. In this work, we argue that action faithfulness…

cs.CV2026

HisTrackMap: Global Vectorized High-Definition Map Construction via History Map Tracking

Jing Yang, Sen Yang, Xiao Tan +1

As an essential component of autonomous driving systems, high-definition (HD) maps provide rich and precise environmental information for auto-driving scenarios; however, existing…

cs.MA2025

Thinking While Driving: A Concurrent Framework for Real-Time, LLM-Based Adaptive Routing

Xiaopei Tan, Muyang Fan

We present Thinking While Driving, a concurrent routing framework that integrates LLMs into a graph-based traffic environment. Unlike approaches that require agents to stop and del…

cs.CV2025

CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks

Yu Qi, Yumeng Zhang, Chenting Gong +4

Large Vision-Language Models (LVLMs) have demonstrated remarkable success in a broad range of vision-language tasks, such as general visual question answering and optical character…

cs.CV2025

Fin3R: Fine-tuning Feed-forward 3D Reconstruction Models via Monocular Knowledge Distillation

Weining Ren, Hongjun Wang, Xiao Tan +1

We present Fin3R, a simple, effective, and general fine-tuning method for feed-forward 3D reconstruction models. The family of feed-forward reconstruction model regresses pointmap…