activity
20242026
collaborators

8 papers

cs.LG2026

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts

Rui Zhang, Xinle Wu, Yao Lu

Reinforcement learning (RL) with verifiable rewards has achieved strong progress in reasoning-oriented LLMs, but extending it to multi-domain RL remains challenging due to reward u…

cs.AI2026

Planner-Centric Reinforcement Learning for Deep Research with Structure-Aware Reward

Mustafa Anis Hussain, Xinle Wu, Yao Lu

Deep research tasks require LLMs to plan what to investigate, retrieve evidence, and synthesize long-form answers across multiple branches of inquiry. Existing training paradigms e…

cs.LG2026

Automatic Configuration of LLM Post-Training Pipelines

Channe Chwa, Xinle Wu, Yao Lu

LLM post-training pipelines that combine supervised fine-tuning and reinforcement learning are difficult to configure under realistic compute budgets: the configuration space is hi…

cs.AI2026

Towards Autonomous Memory Agents

Xinle Wu, Rui Zhang, Mustafa Anis Hussain +1

Recent memory agents improve LLMs by extracting experiences and conversation history into an external storage. This enables low-overhead context assembly and online memory update w…

cs.LG2025

Invariant Representation Learning for Source-Free Time Series Forecasting with LLM-Centric Proxy Denoising

Kangjia Yan, Chenxi Liu, Hao Miao +4

Effective time series forecasting enables various real-world applications, benefiting from the proliferation of mobile devices. However, the volume of time series data may vary sig…

cs.AI2025

Reward Model Routing in Alignment

Xinle Wu, Yao Lu

Reinforcement learning from human or AI feedback (RLHF / RLAIF) has become the standard paradigm for aligning large language models (LLMs). However, most pipelines rely on a single…