activity
20232026
most citedBehavior Proximal Policy Optimization

8 citations · 14 across the 24 of their papers we have counts for

collaborators

24 papers

cs.RO2026

World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems

Runze Li, Hongyin Zhang, Junxi Jin +5

Vision-Language-Action (VLA) models have emerged as a promising paradigm for building embodied agents that ground perception and language into action. However, most existing approa…

cs.RO2026

NFPO: Stabilized Policy Optimization of Normalizing Flow for Robotic Policy Learning

Diyuan Shi, Yiqi Tang, Zifeng Zhuang +1

Deep Reinforcement Learning (DRL) has experienced significant advancements in recent years and has been widely used in many fields. In DRL-based robotic policy learning, however, c…

cs.RO2026

MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation

Yang Liu, Pengxiang Ding, Tengyue Jiang +10

Vision-Language-Action (VLA) models map visual observations and natural-language instructions to robot actions; however, hierarchical and autoregressive paradigms often incur archi…

cs.RO2025

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models

Minghui Lin, Pengxiang Ding, Shu Wang +7

Vision-Language-Action (VLA) models have recently enabled robotic manipulation by grounding visual and linguistic cues into actions. However, most VLAs assume the Markov property,…

cs.LG2025

Boundary-to-Region Supervision for Offline Safe Reinforcement Learning

Huikang Su, Dengyun Peng, Zifeng Zhuang +4

Offline safe reinforcement learning aims to learn policies that satisfy predefined safety constraints from static datasets. Existing sequence-model-based methods condition action g…

cs.CV2025

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL

Fengyuan Dai, Zifeng Zhuang, Yufei Huang +4

Diffusion models have emerged as powerful generative tools across various domains, yet tailoring pre-trained models to exhibit specific desirable properties remains challenging. Wh…