4 papers
Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor
Xiaocan Li, Shiliang Wu, Zheng Shen
MXFP4 arithmetic can dramatically accelerate reinforcement learning (RL) post-training of large language models (LLMs), yet the quantization error introduces severe accuracy degrad…
A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation
Xiaocan Li, Shiliang Wu, Zheng Shen
Decoupled PPO has been a successful reinforcement learning (RL) algorithm to deal with the high data staleness under the asynchronous RL setting. Decoupled loss used in decoupled P…
Generalized Multi-hop Traffic Pressure for Heterogeneous Traffic Perimeter Control
Xiaocan Li, Xiaoyu Wang, Ilia Smirnov +2
Perimeter control (PC) prevents loss of traffic network capacity due to congestion in urban areas. Homogeneous PC allows all access points to a protected region to have identical p…
Multi-hop Upstream Anticipatory Traffic Signal Control with Deep Reinforcement Learning
Xiaocan Li, Xiaoyu Wang, Ilia Smirnov +2
Coordination in traffic signal control is crucial for managing congestion in urban networks. Existing pressure-based control methods focus only on immediate upstream links, leading…