Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
DeepSynth-Eval: Objectively Evaluating Information Consolidation in Deep Survey Writing
Hongzhi Zhang, Yuanze Hu, Tinghai Zhang +9
The evolution of Large Language Models (LLMs) towards autonomous agents has catalyzed progress in Deep Research. While retrieval capabilities are well-benchmarked, the post-retriev…
cs.CL2025
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
Hongzhi Zhang, Jia Fu, Jingyuan Zhang +4
Reinforcement learning (RL) for large language models is an energy-intensive endeavor: training can be unstable, and the policy may gradually drift away from its pretrained weights…