Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Planner-Centric Reinforcement Learning for Deep Research with Structure-Aware Reward
Mustafa Anis Hussain, Xinle Wu, Yao Lu
Deep research tasks require LLMs to plan what to investigate, retrieve evidence, and synthesize long-form answers across multiple branches of inquiry. Existing training paradigms e…
cs.AI2026
Towards Autonomous Memory Agents
Xinle Wu, Rui Zhang, Mustafa Anis Hussain +1
Recent memory agents improve LLMs by extracting experiences and conversation history into an external storage. This enables low-overhead context assembly and online memory update w…
cs.AI2025
Reward Model Routing in Alignment
Xinle Wu, Yao Lu
Reinforcement learning from human or AI feedback (RLHF / RLAIF) has become the standard paradigm for aligning large language models (LLMs). However, most pipelines rely on a single…