Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Probabilistic Uncertain Reward Model
Wangtao Sun, Xiang Cheng, Xing Yu +5
Reinforcement learning from human feedback (RLHF) is a critical technique for training large language models. However, conventional reward models based on the Bradley-Terry model (…
cs.LG2025
DATA: Decomposed Attention-based Task Adaptation for Rehearsal-Free Continual Learning
Huanxuan Liao, Shizhu He, Yupu Hao +2
Continual learning (CL) is essential for Large Language Models (LLMs) to adapt to evolving real-world demands, yet they are susceptible to catastrophic forgetting (CF). While tradi…
cs.LG2025
Shuttle Between the Instructions and the Parameters of Large Language Models
Wangtao Sun, Haotian Xu, Huanxuan Liao +5
The interaction with Large Language Models (LLMs) through instructions has been extensively investigated in the research community. While instructions have been widely used as the…