collaborators

5 papers

cs.CV2026

UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding

Yueming Xu, Jiahui Zhang, Ze Huang +12

Despite the impressive progress on understanding and generating images shown by the recent unified architectures, the integration of 3D tasks remains challenging and largely unexpl…

cs.CV2026

From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D

Jiahui Zhang, Yurui Chen, Yanpeng Zhou +10

Recent advances in LVLMs have improved vision-language understanding, but they still struggle with spatial perception, limiting their ability to reason about complex 3D scenes. Unl…

cs.CV2025

4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration

Jiahui Zhang, Yurui Chen, Yueming Xu +8

Leveraging diverse robotic data for pretraining remains a critical challenge. Existing methods typically model the dataset's action distribution using simple observations as inputs…

cs.CV2025

Translating Images to Road Network: A Sequence-to-Sequence Perspective

Jiachen Lu, Ming Nie, Bozhou Zhang +6

The extraction of road network is essential for the generation of high-definition maps since it enables the precise localization of road landmarks and their interconnections. Howev…

cs.CV2025

LaneCorrect: Self-supervised Lane Detection

Ming Nie, Xinyue Cai, Hang Xu +1

Lane detection has evolved highly functional autonomous driving system to understand driving scenes even under complex environments. In this paper, we work towards developing a gen…