activity
20242026
collaborators

23 papers

cs.CL2026

Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning

Yuting Liu, Wei Wu, Jianzhe Zhao +1

Natural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often contain information irrelevant to a part…

cs.IR2026

PCTD: Preference-Guided Counterfactual Task Decomposition for Agent Tool Retrieval

Chu Zhao, Lei Tang, Minghang Li +5

Task decomposition aims to transform ambiguous instructions into executable atomic subtasks, thereby guiding high-precision tool retrieval. However, our analysis reveals that direc…

cs.LG2026

Enhancing Protein Representation Learning via Manifold Restore Mixing

Yizhou Dang, Chuang Zhao, Lianbo Ma +3

Data augmentation (DA) has been proven to be an effective means for improving protein representation learning (PRL) by generating additional training samples. Although mainstream p…

cs.IR2026

Causal Direct Preference Optimization for Distributionally Robust Generative Recommendation

Chu Zhao, Enneng Yang, Jianzhe Zhao +1

Direct Preference Optimization (DPO) guides large language models (LLMs) to generate recommendations aligned with user historical behavior distributions by minimizing preference al…

cs.LG2026

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning

Chu Zhao, Enneng Yang, Yuting Liu +2

Test-time reinforcement learning generates multiple candidate answers via repeated rollouts and performs online updates using pseudo-labels constructed by majority voting. To reduc…

cs.IR2026

Pay Attention to Sequence Split: Uncovering the Impacts of Sub-Sequence Splitting on Sequential Recommendation Models

Yizhou Dang, Yifan Wu, Minhan Huang +5

Sub-sequence splitting (SSS) has been demonstrated as an effective approach to mitigate data sparsity in sequential recommendation (SR) by splitting a raw user interaction sequence…