8 papers
Segmenting Objectiveness and Task-awareness Unknown Region for Autonomous Driving
Mi Zheng, Guanglei Yang, Zitong Huang +3
With the emergence of transformer-based architectures and large language models (LLMs), the accuracy of road scene perception has substantially advanced. Nonetheless, current road…
MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM
Bowen Dong, Minheng Ni, Zitong Huang +3
Multimodal hallucination in multimodal large language models (MLLMs) restricts the correctness of MLLMs. However, multimodal hallucinations are multi-sourced and arise from diverse…
FILP-3D: Enhancing 3D Few-shot Class-incremental Learning with Pre-trained Vision-Language Models
Wan Xu, Tianyu Huang, Tianyu Qu +3
Few-shot class-incremental learning (FSCIL) aims to mitigate the catastrophic forgetting issue when a model is incrementally trained on limited data. However, many of these works l…
MetricDepth: Enhancing Monocular Depth Estimation with Deep Metric Learning
Chunpu Liu, Guanglei Yang, Wangmeng Zuo +1
Deep metric learning aims to learn features relying on the consistency or divergence of class labels. However, in monocular depth estimation, the absence of a natural definition of…
Geo-ConvGRU: Geographically Masked Convolutional Gated Recurrent Unit for Bird-Eye View Segmentation
Guanglei Yang, Yongqiang Zhang, Wanlong Li +5
Convolutional Neural Networks (CNNs) have significantly impacted various computer vision tasks, however, they inherently struggle to model long-range dependencies explicitly due to…
Multi-Modality Driven LoRA for Adverse Condition Depth Estimation
Guanglei Yang, Rui Tian, Yongqiang Zhang +3
The autonomous driving community is increasingly focused on addressing corner case problems, particularly those related to ensuring driving safety under adverse conditions (e.g., n…