5 papers
MyGram: Modality-aware Graph Transformer with Global Distribution for Multi-modal Entity Alignment
Zhifei Li, Ziyue Qin, Xiangyu Luo +6
Multi-modal entity alignment aims to identify equivalent entities between two multi-modal Knowledge graphs by integrating multi-modal data, such as images and text, to enrich the s…
MacVQA: Adaptive Memory Allocation and Global Noise Filtering for Continual Visual Question Answering
Zhifei Li, Yiran Wang, Chenyi Xiong +6
Visual Question Answering (VQA) requires models to reason over multimodal information, combining visual and textual data. With the development of continual learning, significant pr…
KeenKT: Knowledge Mastery-State Disambiguation for Knowledge Tracing
Zhifei Li, Lifan Chen, Jiali Yi +6
Knowledge Tracing (KT) aims to dynamically model a student's mastery of knowledge concepts based on their historical learning interactions. Most current methods rely on single-poin…
Structurally Refined Graph Transformer for Multimodal Recommendation
Ke Shi, Yan Zhang, Miao Zhang +5
Multimodal recommendation systems utilize various types of information, including images and text, to enhance the effectiveness of recommendations. The key challenge is predicting…
SCRA-VQA: Summarized Caption-Rerank for Augmented Large Language Models in Visual Question Answering
Yan Zhang, Jiaqing Lin, Miao Zhang +4
Acquiring high-quality knowledge is a central focus in Knowledge-Based Visual Question Answering (KB-VQA). Recent methods use large language models (LLMs) as knowledge engines for…