8 papers
CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts
Rui Zhang, Xinle Wu, Yao Lu
Reinforcement learning (RL) with verifiable rewards has achieved strong progress in reasoning-oriented LLMs, but extending it to multi-domain RL remains challenging due to reward u…
Planner-Centric Reinforcement Learning for Deep Research with Structure-Aware Reward
Mustafa Anis Hussain, Xinle Wu, Yao Lu
Deep research tasks require LLMs to plan what to investigate, retrieve evidence, and synthesize long-form answers across multiple branches of inquiry. Existing training paradigms e…
Automatic Configuration of LLM Post-Training Pipelines
Channe Chwa, Xinle Wu, Yao Lu
LLM post-training pipelines that combine supervised fine-tuning and reinforcement learning are difficult to configure under realistic compute budgets: the configuration space is hi…
Towards Autonomous Memory Agents
Xinle Wu, Rui Zhang, Mustafa Anis Hussain +1
Recent memory agents improve LLMs by extracting experiences and conversation history into an external storage. This enables low-overhead context assembly and online memory update w…
Invariant Representation Learning for Source-Free Time Series Forecasting with LLM-Centric Proxy Denoising
Kangjia Yan, Chenxi Liu, Hao Miao +4
Effective time series forecasting enables various real-world applications, benefiting from the proliferation of mobile devices. However, the volume of time series data may vary sig…
Reward Model Routing in Alignment
Xinle Wu, Yao Lu
Reinforcement learning from human or AI feedback (RLHF / RLAIF) has become the standard paradigm for aligning large language models (LLMs). However, most pipelines rely on a single…