23 papers
Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning
Yuting Liu, Wei Wu, Jianzhe Zhao +1
Natural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often contain information irrelevant to a part…
PCTD: Preference-Guided Counterfactual Task Decomposition for Agent Tool Retrieval
Chu Zhao, Lei Tang, Minghang Li +5
Task decomposition aims to transform ambiguous instructions into executable atomic subtasks, thereby guiding high-precision tool retrieval. However, our analysis reveals that direc…
Enhancing Protein Representation Learning via Manifold Restore Mixing
Yizhou Dang, Chuang Zhao, Lianbo Ma +3
Data augmentation (DA) has been proven to be an effective means for improving protein representation learning (PRL) by generating additional training samples. Although mainstream p…
Causal Direct Preference Optimization for Distributionally Robust Generative Recommendation
Chu Zhao, Enneng Yang, Jianzhe Zhao +1
Direct Preference Optimization (DPO) guides large language models (LLMs) to generate recommendations aligned with user historical behavior distributions by minimizing preference al…
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning
Chu Zhao, Enneng Yang, Yuting Liu +2
Test-time reinforcement learning generates multiple candidate answers via repeated rollouts and performs online updates using pseudo-labels constructed by majority voting. To reduc…
Pay Attention to Sequence Split: Uncovering the Impacts of Sub-Sequence Splitting on Sequential Recommendation Models
Yizhou Dang, Yifan Wu, Minhan Huang +5
Sub-sequence splitting (SSS) has been demonstrated as an effective approach to mitigate data sparsity in sequential recommendation (SR) by splitting a raw user interaction sequence…