9 papers
PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation
Yichuan Wang, Zhifei Li, Zirui Wang +5
Augmenting large language models (LLMs) with retrieved web text has become a dominant paradigm, yet the web is not natively textual: existing systems depend on complex parsing pipe…
HyLoVQA: Dynamic Hypernetwork-Generated Low-Rank Adaptation for Continual Visual Question Answering
Yiran Wang, Chenyi Xiong, Ziyue Qin +3
Continual Visual Question Answering (VQA) requires learning from non-stationary streams of visual inputs and questions while preserving past knowledge. Most prior methods adapt by…
Modality-Aware Identity Construction and Counterfactual Structure Learning for ID-Free Multimodal Recommendation
Hongjian Ma, Wenxin Huang, Yan Zhang +2
Multimodal recommendation has attracted extensive attention by leveraging heterogeneous modality information to alleviate data sparsity and improve recommendation accuracy. Existin…
MyGram: Modality-aware Graph Transformer with Global Distribution for Multi-modal Entity Alignment
Zhifei Li, Ziyue Qin, Xiangyu Luo +6
Multi-modal entity alignment aims to identify equivalent entities between two multi-modal Knowledge graphs by integrating multi-modal data, such as images and text, to enrich the s…
MacVQA: Adaptive Memory Allocation and Global Noise Filtering for Continual Visual Question Answering
Zhifei Li, Yiran Wang, Chenyi Xiong +6
Visual Question Answering (VQA) requires models to reason over multimodal information, combining visual and textual data. With the development of continual learning, significant pr…
KeenKT: Knowledge Mastery-State Disambiguation for Knowledge Tracing
Zhifei Li, Lifan Chen, Jiali Yi +6
Knowledge Tracing (KT) aims to dynamically model a student's mastery of knowledge concepts based on their historical learning interactions. Most current methods rely on single-poin…