collaborators

5 papers

cs.LG2025

Probabilistic Uncertain Reward Model

Wangtao Sun, Xiang Cheng, Xing Yu +5

Reinforcement learning from human feedback (RLHF) is a critical technique for training large language models. However, conventional reward models based on the Bradley-Terry model (…

cs.LG2025

Shuttle Between the Instructions and the Parameters of Large Language Models

Wangtao Sun, Haotian Xu, Huanxuan Liao +5

The interaction with Large Language Models (LLMs) through instructions has been extensively investigated in the research community. While instructions have been widely used as the…

cs.CL2024

Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models

Wangtao Sun, Chenxiang Zhang, XueYou Zhang +7

Although Large Language Models (LLMs) have demonstrated strong ability, they are further supposed to be controlled and guided by in real-world scenarios to be safe, accurate, and i…

cs.CL2024

LongAgent: Scaling Language Models to 128k Context through Multi-Agent Collaboration

Jun Zhao, Can Zu, Hao Xu +6

Large language models (LLMs) have demonstrated impressive performance in understanding language and executing complex reasoning tasks. However, LLMs with long context windows have…

cs.CL2024

ItD: Large Language Models Can Teach Themselves Induction through Deduction

Wangtao Sun, Haotian Xu, Xuanqing Yu +4

Although Large Language Models (LLMs) are showing impressive performance on a wide range of Natural Language Processing tasks, researchers have found that they still have limited a…