activity
20242026
collaborators

18 papers

cs.CL2026

Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning

Yuting Liu, Wei Wu, Jianzhe Zhao +1

Natural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often contain information irrelevant to a part…

cs.IR2026

PCTD: Preference-Guided Counterfactual Task Decomposition for Agent Tool Retrieval

Chu Zhao, Lei Tang, Minghang Li +5

Task decomposition aims to transform ambiguous instructions into executable atomic subtasks, thereby guiding high-precision tool retrieval. However, our analysis reveals that direc…

cs.IR2026

Causal Direct Preference Optimization for Distributionally Robust Generative Recommendation

Chu Zhao, Enneng Yang, Jianzhe Zhao +1

Direct Preference Optimization (DPO) guides large language models (LLMs) to generate recommendations aligned with user historical behavior distributions by minimizing preference al…

cs.LG2026

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning

Chu Zhao, Enneng Yang, Yuting Liu +2

Test-time reinforcement learning generates multiple candidate answers via repeated rollouts and performs online updates using pseudo-labels constructed by majority voting. To reduc…

cs.CL2026

Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning

Weiyu Ma, Yongcheng Zeng, Yan Song +4

Reinforcement Learning (RL) has achieved impressive success in post-training Large Language Models (LLMs) and Vision-Language Models (VLMs), with on-policy algorithms such as PPO,…

cs.IR2026

Hard Negative Sampling via Large Language Models for Recommendation

Chu Zhao, Enneng Yang, Yuting Liu +2

Hard negative sampling improves recommendation performance by accelerating convergence and sharpening the decision boundary. However, most existing methods rely on heuristic strate…