2 papers
cs.CL2025
Aligning Large Language Models via Fully Self-Synthetic Data
Shangjian Yin, Zhepei Wei, Xinyu Zhu +2
Traditional reinforcement learning from human feedback (RLHF) for large language models (LLMs) relies on expensive human-annotated datasets, while Reinforcement Learning from AI Fe…
cs.AI2025
Beyond Outcome Reward: Decoupling Search and Answering Improves LLM Agents
Yiding Wang, Zhepei Wei, Xinyu Zhu +1
Enabling large language models (LLMs) to utilize search tools offers a promising path to overcoming fundamental limitations such as knowledge cutoffs and hallucinations. Recent wor…