5 papers
Probabilistic Uncertain Reward Model
Wangtao Sun, Xiang Cheng, Xing Yu +5
Reinforcement learning from human feedback (RLHF) is a critical technique for training large language models. However, conventional reward models based on the Bradley-Terry model (…
Shuttle Between the Instructions and the Parameters of Large Language Models
Wangtao Sun, Haotian Xu, Huanxuan Liao +5
The interaction with Large Language Models (LLMs) through instructions has been extensively investigated in the research community. While instructions have been widely used as the…
Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models
Wangtao Sun, Chenxiang Zhang, XueYou Zhang +7
Although Large Language Models (LLMs) have demonstrated strong ability, they are further supposed to be controlled and guided by in real-world scenarios to be safe, accurate, and i…
LongAgent: Scaling Language Models to 128k Context through Multi-Agent Collaboration
Jun Zhao, Can Zu, Hao Xu +6
Large language models (LLMs) have demonstrated impressive performance in understanding language and executing complex reasoning tasks. However, LLMs with long context windows have…
ItD: Large Language Models Can Teach Themselves Induction through Deduction
Wangtao Sun, Haotian Xu, Xuanqing Yu +4
Although Large Language Models (LLMs) are showing impressive performance on a wide range of Natural Language Processing tasks, researchers have found that they still have limited a…