activity
20242026
collaborators

8 papers

cs.CV2026

HyLoVQA: Dynamic Hypernetwork-Generated Low-Rank Adaptation for Continual Visual Question Answering

Yiran Wang, Chenyi Xiong, Ziyue Qin +3

Continual Visual Question Answering (VQA) requires learning from non-stationary streams of visual inputs and questions while preserving past knowledge. Most prior methods adapt by…

cs.AI2026

MyGram: Modality-aware Graph Transformer with Global Distribution for Multi-modal Entity Alignment

Zhifei Li, Ziyue Qin, Xiangyu Luo +6

Multi-modal entity alignment aims to identify equivalent entities between two multi-modal Knowledge graphs by integrating multi-modal data, such as images and text, to enrich the s…

cs.CV2026

MacVQA: Adaptive Memory Allocation and Global Noise Filtering for Continual Visual Question Answering

Zhifei Li, Yiran Wang, Chenyi Xiong +6

Visual Question Answering (VQA) requires models to reason over multimodal information, combining visual and textual data. With the development of continual learning, significant pr…

cs.AI2025

KeenKT: Knowledge Mastery-State Disambiguation for Knowledge Tracing

Zhifei Li, Lifan Chen, Jiali Yi +6

Knowledge Tracing (KT) aims to dynamically model a student's mastery of knowledge concepts based on their historical learning interactions. Most current methods rely on single-poin…

cs.IR2025

Structurally Refined Graph Transformer for Multimodal Recommendation

Ke Shi, Yan Zhang, Miao Zhang +5

Multimodal recommendation systems utilize various types of information, including images and text, to enhance the effectiveness of recommendations. The key challenge is predicting…

cs.CV2025

Integrating Object Interaction Self-Attention and GAN-Based Debiasing for Visual Question Answering

Zhifei Li, Feng Qiu, Yiran Wang +4

Visual Question Answering (VQA) presents a unique challenge by requiring models to understand and reason about visual content to answer questions accurately. Existing VQA models of…