Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search
Fang Wu, Weihao Xuan, Heli Qi +4
Although RLVR has become an essential component for developing advanced reasoning skills in language models, contemporary studies have documented training plateaus after thousands…
cs.AI2026
Multiplayer Nash Preference Optimization
Fang Wu, Xu Huang, Weihao Xuan +8
Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human preferences. However, reward-based methods grou…
cs.AI2026
Experience-Driven Multi-Agent Systems Are Training-free Context-aware Earth Observers
Pengyu Dai, Weihao Xuan, Junjue Wang +4
Recent advances have enabled large language model (LLM) agents to solve complex tasks by orchestrating external tools. However, these agents often struggle in specialized, tool-int…