From the 1 of 3 linked papers with an AI index.
4 papers · 1 filter
Gaussian Mixture Modeling for Event-Aware Visual Allocation in Long Video Understanding
Yifan Lu, Ziqi Zhang, Chunfeng Yuan +3
The paper introduces GMM-EVA, a training-free framework that uses Gaussian Mixture Models to detect event-level structures in long videos and allocate visual tokens by selecting on…
IPAD-CLIP: Teaching CLIP to Detect Image Local Perceptual Artifacts
Juan Wang, Xinyu Sun, Ke Zhang +4
Current image quality assessment methods are heavily biased towards global distortions (e.g., noise, blur), neglecting local perceptual artifacts such as ghosting, lens flare, and…
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
Yifan Lu, Ziqi Zhang, Chunfeng Yuan +5
Large Vision-Language Models (LVLMs) suffer from serious hallucination problems, where the model-generated responses are inconsistent with the visual inputs. Existing hallucination…
Real-time Monocular Depth Estimation on Embedded Systems
Cheng Feng, Congxuan Zhang, Zhen Chen +2
Depth sensing is of paramount importance for unmanned aerial and autonomous vehicles. Nonetheless, contemporary monocular depth estimation methods employing complex deep neural net…