3 papers
cs.CL2026
FASA: Frequency-aware Sparse Attention
Yifei Wang, Yueqi Wang, Zhenrui Yue +6
The deployment of Large Language Models (LLMs) faces a critical bottleneck when handling lengthy inputs: the prohibitive memory footprint of the Key Value (KV) cache. To address th…
cs.IR2025
Your Causal Self-Attentive Recommender Hosts a Lonely Neighborhood
Yueqi Wang, Zhankui He, Zhenrui Yue +2
In the context of sequential recommendation, a pivotal issue pertains to the comparative analysis between bi-directional/auto-encoding (AE) and uni-directional/auto-regressive (AR)…
cs.IR2024
Train Once, Deploy Anywhere: Matryoshka Representation Learning for Multimodal Recommendation
Yueqi Wang, Zhenrui Yue, Huimin Zeng +2
Despite recent advancements in language and vision modeling, integrating rich multimodal knowledge into recommender systems continues to pose significant challenges. This is primar…