7 papers
mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQA
Xu Yuan, Liangbo Ning, Qingqing Ye +2
Retrieval-Augmented Generation (RAG) has emerged as an effective paradigm for expanding the knowledge capacity of Multimodal Large Language Models (MLLMs) by incorporating external…
SUPERGLASSES: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glasses
Zhuohang Jiang, Xu Yuan, Haohao Qu +4
The rapid advancement of AI-powered smart glasses-one of the hottest wearable devices-has unlocked new frontiers for multimodal interaction, with Visual Question Answering (VQA) ov…
QA-Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering
Zhuohang Jiang, Pangjing Wu, Xu Yuan +2
Retrieval-Augmented Generation (RAG) has been introduced to mitigate hallucinations in Multimodal Large Language Models (MLLMs) by incorporating external knowledge into the generat…
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
Zhuohang Jiang, Pangjing Wu, Ziran Liang +7
Structure reasoning is a fundamental capability of large language models (LLMs), enabling them to reason about structured commonsense and answer multi-hop questions. However, exist…
BadCM: Invisible Backdoor Attack Against Cross-Modal Learning
Zheng Zhang, Xu Yuan, Lei Zhu +2
Despite remarkable successes in unimodal learning tasks, backdoor attacks against cross-modal learning are still underexplored due to the limited generalization and inferior stealt…
Semantic-Aware Adversarial Training for Reliable Deep Hashing Retrieval
Xu Yuan, Zheng Zhang, Xunguang Wang +1
Deep hashing has been intensively studied and successfully applied in large-scale image retrieval systems due to its efficiency and effectiveness. Recent studies have recognized th…