2 papers
cs.LG2026
HARTS: Efficient Agentic Reinforcement Learning for Hybrid-Attention Models over Arbitrary Rollout Trees
Boyuan Meng, Peihua Bao, Hong Liu +4
Agentic reinforcement learning (RL) often produces irregular rollout trees with shared histories. Training root-to-leaf trajectories independently recomputes these shared prefixes.…
cs.DC2026
RoutePack: Expert Placement and Attention-Aware Data Packing for MoE Reinforcement Learning
Yibo Shen, Xudong Han, Xiaowei Zhu +2
Training Mixture-of-Experts (MoE) models for reinforcement learning (RL) couples two load-balancing problems: sequence composition determines dense attention work in each data-para…