Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering
Zhiyang Li, Ao Ke, Yukun Cao +1
Multi-modal Large Language Models (MLLMs) for Visual Question Answering (VQA) often suffer from dual limitations: knowledge hallucination and insufficient fine-grained visual perce…
cs.CV2024
Prototype-based Optimal Transport for Out-of-Distribution Detection
Ao Ke, Wenlong Chen, Chuanwen Feng +4
Detecting Out-of-Distribution (OOD) inputs is crucial for improving the reliability of deep neural networks in the real-world deployment. In this paper, inspired by the inherent di…