7 papers
BICPO-VLA: Behavior-Identified Continuation Preference Optimization for Smooth Asynchronous Vision-Language-Action Control
Ming Shang, Yuchen Huang, Jiaoyang Chen +8
The request-to-handoff gap has three coupled sources: ambiguity about the behavior intended at request time, physical-state drift accumulated during action generation, and residual…
PBD-AG: Persistent Baseline-Delta Active Graphs with Uncertainty-Aware Inspection for Long-Horizon Service Robots
Shuo Bao, Wei Dong, Shuyue Zhang +8
Long-horizon service robots require persistent world models that can be built autonomously in unseen environments and revised as task-relevant objects change. Existing methods rely…
Ratio-Variance Regularized Policy Optimization
Yu Luo, Shuo Han, Yihan Hu +5
Standard on-policy reinforcement learning relies on heuristic clipping to enforce trust regions, but this mechanism imposes a severe cost by indiscriminately truncating high-return…
FLAC: Maximum Entropy RL via Kinetic Energy Regularized Bridge Matching
Lei Lv, Yunfei Li, Yu Luo +2
Iterative generative policies, such as diffusion models and flow matching, offer superior expressivity for continuous control but complicate Maximum Entropy Reinforcement Learning…
Flow-Based Policy for Online Reinforcement Learning
Lei Lv, Yunfei Li, Yu Luo +4
We present \textbf{FlowRL}, a novel framework for online reinforcement learning that integrates flow-based policy representation with Wasserstein-2-regularized optimization. We arg…
Multi-segment Soft Robot Control via Deep Koopman-based Model Predictive Control
Lei Lv, Lei Liu, Lei Bao +7
Soft robots, compared to regular rigid robots, as their multiple segments with soft materials bring flexibility and compliance, have the advantages of safe interaction and dexterou…