3 papers
cs.AI2026
DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models
ZhiYan Hou, Xinyu Tang, Hongyan An +9
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals…
cs.RO2026
Minimizing Worst-Case Weighted Latency for Multi-Robot Persistent Monitoring: Theory and RL-Based Solutions
Weizhen Wang, Ziheng Wang, Jianping He +2
We study multi-robot persistent monitoring on weighted graphs, where node weights encode monitoring priorities and edge weights encode travel distances. The goal is to design joint…
cs.LG2025
Analysis of On-policy Policy Gradient Methods under the Distribution Mismatch
Weizhen Wang, Jianping He, Xiaoming Duan
Policy gradient methods are one of the most successful approaches for solving challenging reinforcement learning problems. Despite their empirical successes, many state-of-the-art…