activity
20232026
most citedStable CDE Autoencoders with Acuity Regularization for Offline Reinforcement Learning in Sepsis Treatment

1 citations · 1 across the 17 of their papers we have counts for

collaborators
Showing cs.ROShow all

10 papers · 1 filter

cs.RO2026

ActSafeGuard: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching Policies

Jianming Ma, Rongjun Jin, Xiaxi Si +3

Vision-Language-Action (VLA) and World-Action Models (WAMs) have demonstrated strong capabilities in general-purpose robotic manipulation, yet their generated actions may violate h…

cs.RO2026

P3: Probabilistic Policy Propagation for Stable VAE-Based Robot Learning

Liyun Yan, Jianming Ma, Yang Zhang +5

Variational Autoencoders are widely used to encode high-dimensional and noisy observations in robotics. However, their stochastic latent creates a mismatch with Proximal Policy Opt…

cs.RO2026

UniLab: A Heterogeneous Architecture for Robot RL Beyond GPU-Dominant Paradigms

Yufei Jia, Zhanxiang Cao, Mingrui Yu +48

Simulation-based RL for contemporary robot control is increasingly organized around GPU-resident simulation: physics, rollout collection, and learning are placed on a single GPU-ce…

cs.RO2026

HierKick: Hierarchical Reinforcement Learning for Vision-Guided Soccer Robot Control

Yizhi Chen, Zheng Zhang, Zhanxiang Cao +7

Controlling soccer robots involves multi-time-scale decision-making, which requires balancing long-term tactical planning and short-term motion execution. Traditional end-to-end re…

cs.RO2025

Coordinated Humanoid Robot Locomotion with Symmetry Equivariant Reinforcement Learning Policy

Buqing Nie, Yang Zhang, Rongjun Jin +4

The human nervous system exhibits bilateral symmetry, enabling coordinated and balanced movements. However, existing Deep Reinforcement Learning (DRL) methods for humanoid robots n…

cs.RO2025

Keep on Going: Learning Robust Humanoid Motion Skills via Selective Adversarial Training

Yang Zhang, Zhanxiang Cao, Buqing Nie +6

Humanoid robots are expected to operate reliably over long horizons while executing versatile whole-body skills. Yet Reinforcement Learning (RL) motion policies typically lose stab…