collaborators

6 papers

cs.RO2026

ReFPO: Reflow Regularization for Flow Matching Policy Gradients

Ge Wang, Yibo Peng, Fan Feng +10

We present Reflow-regularized Flow Matching Policy Gradients (ReFPO), a simple online RL method that adds explicit Reflow regularization to FPO for efficient flow-based control. We…

cs.RO2026

MimicIK: Real-Time Generative Inverse Kinematics from Teleoperation with FK Consistency

Jiahao Yang, Shenhao Yan, Fan Feng +5

Inverse kinematics (IK) remains a critical bottleneck for real-time robot manipulation. Classical numerical solvers achieve high geometric precision but often suffer from discontin…

cs.RO2026

Acting While Understanding: Asynchronous Semantic-Action Decoupling for Real-Time Vision-Language-Action Models

Shenhao Yan, Ge Wang, Qi Liu +7

Vision-Language-Action models (VLAs) have demonstrated strong task understanding and generalization in robotic manipulation, yet the high computational cost of full-model inference…

cs.RO2026

KoopmanFlow: Spectrally Decoupled Generative Control Policy via Koopman Structural Bias

Chengsi Yao, Ge Wang, Kai Kang +9

Generative Control Policies (GCPs) show immense promise in robotic manipulation but struggle to simultaneously model stable global motions and high-frequency local corrections. Whi…

cs.CV2026

SVLL: Staged Vision-Language Learning for Physically Grounded Embodied Task Planning

Yuyuan Yang, Junkun Hong, Hongrong Wang +11

Embodied task planning demands vision-language models to generate action sequences that are both visually grounded and causally coherent over time. However, existing training parad…

cs.CV2025

OS-HGAdapter: Open Semantic Hypergraph Adapter for Large Language Models Assisted Entropy-Enhanced Image-Text Alignment

Rongjun Chen, Chengsi Yao, Jinchang Ren +6

Text-image alignment constitutes a foundational challenge in multimedia content understanding, where effective modeling of cross-modal semantic correspondences critically enhances…