6 papers
ReFPO: Reflow Regularization for Flow Matching Policy Gradients
Ge Wang, Yibo Peng, Fan Feng +10
We present Reflow-regularized Flow Matching Policy Gradients (ReFPO), a simple online RL method that adds explicit Reflow regularization to FPO for efficient flow-based control. We…
MimicIK: Real-Time Generative Inverse Kinematics from Teleoperation with FK Consistency
Jiahao Yang, Shenhao Yan, Fan Feng +5
Inverse kinematics (IK) remains a critical bottleneck for real-time robot manipulation. Classical numerical solvers achieve high geometric precision but often suffer from discontin…
Acting While Understanding: Asynchronous Semantic-Action Decoupling for Real-Time Vision-Language-Action Models
Shenhao Yan, Ge Wang, Qi Liu +7
Vision-Language-Action models (VLAs) have demonstrated strong task understanding and generalization in robotic manipulation, yet the high computational cost of full-model inference…
KoopmanFlow: Spectrally Decoupled Generative Control Policy via Koopman Structural Bias
Chengsi Yao, Ge Wang, Kai Kang +9
Generative Control Policies (GCPs) show immense promise in robotic manipulation but struggle to simultaneously model stable global motions and high-frequency local corrections. Whi…
SVLL: Staged Vision-Language Learning for Physically Grounded Embodied Task Planning
Yuyuan Yang, Junkun Hong, Hongrong Wang +11
Embodied task planning demands vision-language models to generate action sequences that are both visually grounded and causally coherent over time. However, existing training parad…
OS-HGAdapter: Open Semantic Hypergraph Adapter for Large Language Models Assisted Entropy-Enhanced Image-Text Alignment
Rongjun Chen, Chengsi Yao, Jinchang Ren +6
Text-image alignment constitutes a foundational challenge in multimedia content understanding, where effective modeling of cross-modal semantic correspondences critically enhances…