9 papers
Role-Break in Attention Heads: Understanding and Detecting Hallucinations in VLMs
Mingyu Wang, Weilin Jin, Wenbo Li +5
Despite remarkable progress in vision-language generation, Vision-Language Models (VLMs) remain prone to hallucinations, producing content that is inconsistent with or unsupported…
Geo3R: Mitigating Spatial Reasoning Hallucination in Multimodal Large Language Models
Mingyu Wang, Weilin Jin, Wenbo Li +3
Despite remarkable progress in visual understanding, Multimodal Large Language Models (MLLMs) remain prone to hallucinations when reasoning about spatial relationships, often produ…
CDPR: Cross-modal Diffusion with Polarization for Reliable Monocular Depth Estimation
Rongjia Yu, Tong Jia, Hao Wang +4
Monocular depth estimation is a fundamental yet challenging task in computer vision, especially under complex conditions such as textureless surfaces, transparency, and specular re…
CSPCL: Category Semantic Prior Contrastive Learning for Deformable DETR-Based Prohibited Item Detectors
Mingyuan Li, Tong Jia, Hao Wang +5
Prohibited item detection based on X-ray images is one of the most effective security inspection methods. However, the foreground-background feature coupling caused by the overlapp…
MMCL: Correcting Content Query Distributions for Improved Anti-Overlapping X-Ray Object Detection
Mingyuan Li, Tong Jia, Hui Lu +7
Unlike natural images with occlusion-based overlap, X-ray images exhibit depth-induced superimposition and semi-transparent appearances, where objects at different depths overlap a…
FOAM: A General Frequency-Optimized Anti-Overlapping Framework for Overlapping Object Perception
Mingyuan Li, Tong Jia, Han Gu +7
Overlapping object perception aims to decouple the randomly overlapping foreground-background features, extracting foreground features while suppressing background features, which…