#offline reinforcement learning
9 papers · 1 filter
RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy
Zhengyang Yan, Junhao Li, Fangqi Zhu +6
RedFlow is an offline reinforcement learning framework that turns failure experiences into action-level corrective supervision for flow-matching vision‑language‑action policies, im…
RL-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models
Derek Ming Siang Tan, Shailesh Shailesh, Srikrishna Iyer +4
The paper presents RL², an adaptive test‑time steering framework that uses offline reinforcement learning on latent features from a frozen Vision‑Language‑Action model to compose a…
Temporal-Distance JEPA: Plan-Aware Representation Learning for Latent World Model Predictive Control
Jiaxin Bai, Jiaxuan Xiong
The paper introduces Temporal-Distance JEPA, a method that learns a directed temporal cost from offline trajectories to improve latent world model predictive control, enhancing pla…
Learning from the Unseen: Offline Reinforcement Learning with Hidden Actions
Zeyu Bian, Ying Zhou, Yifan Cui
The paper introduces a method for offline reinforcement learning when the true actions are unobserved, using next-state information to estimate policy values and providing a robust…
Evaluating covariate balance for long time horizon Markov decision processes
Joshua Spear, Rebecca Pope, Neil J Sebire
The paper investigates how covariate balance diagnostics can be used to detect hidden confounding and model miss‑specification in offline reinforcement learning for long‑horizon tr…
RAD: Retrieval High-quality Demonstrations to Enhance Decision-making
Lu Guo, Yixiang Shan, Zhengbang Zhu +5
The paper proposes RAD, a method that improves offline reinforcement learning by retrieving high-return states from the dataset and generating sub-trajectories toward these targets…