markov decision process 1multi-robot coordination 1persistent monitoring 1reinforcement learning 1weighted latency 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.AI2026
DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models
ZhiYan Hou, Xinyu Tang, Hongyan An +9
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals…
cs.RO2026
Minimizing Worst-Case Weighted Latency for Multi-Robot Persistent Monitoring: Theory and RL-Based Solutions
Weizhen Wang, Ziheng Wang, Jianping He +2
The paper proposes new tail-performance objectives for multi-robot persistent monitoring on weighted graphs and solves the resulting optimization via a reformulated Markov decision…
cs.LG2026
Analysis of On-policy Policy Gradient Methods under the Distribution Mismatch
Weizhen Wang, Jianping He, Xiaoming Duan
Policy gradient methods are one of the most successful approaches for solving challenging reinforcement learning problems. Despite their empirical successes, many state-of-the-art…