5 citations · 5 across the 1 of their papers we have counts for
3 papers · 1 filter
Planner-Centric Reinforcement Learning for Deep Research with Structure-Aware Reward
Mustafa Anis Hussain, Xinle Wu, Yao Lu
Deep research tasks require LLMs to plan what to investigate, retrieve evidence, and synthesize long-form answers across multiple branches of inquiry. Existing training paradigms e…
Towards Autonomous Memory Agents
Xinle Wu, Rui Zhang, Mustafa Anis Hussain +1
Recent memory agents improve LLMs by extracting experiences and conversation history into an external storage. This enables low-overhead context assembly and online memory update w…
NovGrid: A Flexible Grid World for Evaluating Agent Response to Novelty
Jonathan Balloch, Zhiyu Lin, Mustafa Hussain +5
A robust body of reinforcement learning techniques have been developed to solve complex sequential decision making problems. However, these methods assume that train and evaluation…