3 papers
cs.DC2025
Stable-MoE: Lyapunov-based Token Routing for Distributed Mixture-of-Experts Training over Edge Networks
Long Shi, Bingyan Ou, Kang Wei +3
The sparse activation mechanism of mixture of experts (MoE) model empowers edge intelligence with enhanced training efficiency and reduced computational resource consumption. Howev…
cs.LG2025
Mitigating Estimation Bias with Representation Learning in TD Error-Driven Regularization
Haohui Chen, Zhiyong Chen, Aoxiang Liu +1
Deterministic policy gradient algorithms for continuous control suffer from value estimation biases that degrade performance. While double critics reduce such biases, the explorati…
cs.LG2025
Mildly Conservative Regularized Evaluation for Offline Reinforcement Learning
Haohui Chen, Zhiyong Chen
Offline reinforcement learning (RL) seeks to learn optimal policies from static datasets without further environment interaction. A key challenge is the distribution shift between…