Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs
Ziliang Wang, Kang An, Xuhui Zheng +6
While search-augmented large language models (LLMs) exhibit impressive capabilities, their reliability in complex multi-hop reasoning remains limited. This limitation arises from t…
cs.CL2025
StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization
Ziliang Wang, Xuhui Zheng, Kang An +4
Efficient multi-hop reasoning requires Large Language Models (LLMs) based agents to acquire high-value external knowledge iteratively. Previous work has explored reinforcement lear…