Publications (6)
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Ye Wang, Ziheng Wang, Boshen Xu +14
Temporal Video Grounding (TVG), the task of locating specific video segments based on language queries, is a core challenge in long-form video understanding. While recent Large Vis…
3D-CovDiffusion: 3D-Aware Diffusion Policy for Coverage Path Planning
Chenyuan Chen, Haoran Ding, Ran Ding +6
Diffusion models have shown strong potential for robot skill learning, yet their role in coverage path planning remains underexplored. In industrial surface processing (painting, p…
Task-Specified Compliance Bounds for Humanoids via Lipschitz-Constrained Policies
Zewen He, Yoshihiko Nakamura
Reinforcement learning (RL) has demonstrated substantial potential for humanoid bipedal locomotion and the control of complex motions. To cope with oscillations and impacts induced…
Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
Jiaze Li, Hao Yin, Haoran Xu +6
Reinforcement learning has emerged as a principled post-training paradigm for Temporal Video Grounding (TVG) due to its on-policy optimization, yet existing GRPO-based methods rema…
CoTaP: Compliant Task Pipeline and Reinforcement Learning of Its Controller with Compliance Modulation
Zewen He, Chenyuan Chen, Dilshod Azizov +1
Humanoid whole-body locomotion control is a critical approach for humanoid robots to leverage their inherent advantages. Learning-based control methods derived from retargeted huma…
Instance Scale Normalization for image understanding
Zewen He, He Huang, Yudong Wu +2
Scale variation remains a challenging problem for object detection. Common paradigms usually adopt multiscale training & testing (image pyramid) or FPN (feature pyramid network) to…