collaborators

8 papers

cs.CV2025

Segmenting Objectiveness and Task-awareness Unknown Region for Autonomous Driving

Mi Zheng, Guanglei Yang, Zitong Huang +3

With the emergence of transformer-based architectures and large language models (LLMs), the accuracy of road scene perception has substantially advanced. Nonetheless, current road…

cs.CV2025

MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM

Bowen Dong, Minheng Ni, Zitong Huang +3

Multimodal hallucination in multimodal large language models (MLLMs) restricts the correctness of MLLMs. However, multimodal hallucinations are multi-sourced and arise from diverse…

cs.CV2025

FILP-3D: Enhancing 3D Few-shot Class-incremental Learning with Pre-trained Vision-Language Models

Wan Xu, Tianyu Huang, Tianyu Qu +3

Few-shot class-incremental learning (FSCIL) aims to mitigate the catastrophic forgetting issue when a model is incrementally trained on limited data. However, many of these works l…

cs.CV2024

MetricDepth: Enhancing Monocular Depth Estimation with Deep Metric Learning

Chunpu Liu, Guanglei Yang, Wangmeng Zuo +1

Deep metric learning aims to learn features relying on the consistency or divergence of class labels. However, in monocular depth estimation, the absence of a natural definition of…

cs.CV2024

Geo-ConvGRU: Geographically Masked Convolutional Gated Recurrent Unit for Bird-Eye View Segmentation

Guanglei Yang, Yongqiang Zhang, Wanlong Li +5

Convolutional Neural Networks (CNNs) have significantly impacted various computer vision tasks, however, they inherently struggle to model long-range dependencies explicitly due to…

cs.CV2024

Multi-Modality Driven LoRA for Adverse Condition Depth Estimation

Guanglei Yang, Rui Tian, Yongqiang Zhang +3

The autonomous driving community is increasingly focused on addressing corner case problems, particularly those related to ensuring driving safety under adverse conditions (e.g., n…