3 papers
cs.IR2026
Learning User-Aware Recall: Personalized Retrieval in Long-Term Conversational Memory
ZhiShu Jiang, Haibo Liu, Xin Shen +6
Long-term conversational agents are expected to remember past interactions, but memory is useful only when the right evidence is recalled for the right user. Existing memory-augmen…
cs.IR2026
LASAR: Latent Adaptive Semantic Aligned Reasoning for Generative Recommendation
Yiwen Chen, Fuwei Zhang, Zehao Chen +8
Large Language Models (LLMs) have demonstrated powerful reasoning capabilities through Chain-of-Thought (CoT) in various tasks, yet the inefficiency of token-by-token generation hi…
cs.AI2026
It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches
Dayu Wang, Jiaye Yang, Weikang Li +4
Reinforcement learning with verifiable rewards has emerged as a standard approach for enhancing reasoning in large language models, which typically optimizes the policy by contrast…