activity
20222025
most citedInvestigate-Consolidate-Exploit: A General Strategy for Inter-Task Agent Self-Evolution

7 citations · 24 across the 12 of their papers we have counts for

collaborators
Showing cs.CLShow all

12 papers · 1 filter

cs.CL20252 cited

RM-R1: Reward Modeling as Reasoning

Xiusi Chen, Gaotang Li, Ziqi Wang +9

Reward modeling is essential for aligning large language models with human preferences through reinforcement learning. To provide accurate reward signals, a reward model (RM) shoul…

cs.CL20251 cited

The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination

Yuji Zhang, Sha Li, Cheng Qian +8

Hallucination is a persistent challenge in large language models (LLMs), where even with rigorous quality control, models often generate distorted facts. This paradox, in which err…

cs.CL2024

Aligning LLMs with Individual Preferences via Interaction

Shujin Wu, May Fung, Cheng Qian +3

As large language models (LLMs) demonstrate increasingly advanced capabilities, aligning their behaviors with human values and preferences becomes crucial for their wide adoption.…

cs.CL2024

The Right Time Matters: Data Arrangement Affects Zero-Shot Generalization in Instruction Tuning

Bingxiang He, Ning Ding, Cheng Qian +10

Understanding alignment techniques begins with comprehending zero-shot generalization brought by instruction tuning, but little of the mechanism has been understood. Existing work…

cs.CL20244 cited

Tell Me More! Towards Implicit User Intention Understanding of Language Model Driven Agents

Cheng Qian, Bingxiang He, Zhong Zhuang +8

Current language model-driven agents often lack mechanisms for effective user participation, which is crucial given the vagueness commonly found in user instructions. Although adep…

cs.CL20247 cited

Investigate-Consolidate-Exploit: A General Strategy for Inter-Task Agent Self-Evolution

Cheng Qian, Shihao Liang, Yujia Qin +6

This paper introduces Investigate-Consolidate-Exploit (ICE), a novel strategy for enhancing the adaptability and flexibility of AI agents through inter-task self-evolution. Unlike…