papers

Publications (6)

cs.CV2025

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Ye Wang, Ziheng Wang, Boshen Xu +14

Temporal Video Grounding (TVG), the task of locating specific video segments based on language queries, is a core challenge in long-form video understanding. While recent Large Vis…

cs.RO2025

3D-CovDiffusion: 3D-Aware Diffusion Policy for Coverage Path Planning

Chenyuan Chen, Haoran Ding, Ran Ding +6

Diffusion models have shown strong potential for robot skill learning, yet their role in coverage path planning remains underexplored. In industrial surface processing (painting, p…

cs.RO2026

Task-Specified Compliance Bounds for Humanoids via Lipschitz-Constrained Policies

Zewen He, Yoshihiko Nakamura

Reinforcement learning (RL) has demonstrated substantial potential for humanoid bipedal locomotion and the control of complex motions. To cope with oscillations and impacts induced…

cs.CV2026

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation

Jiaze Li, Hao Yin, Haoran Xu +6

Reinforcement learning has emerged as a principled post-training paradigm for Temporal Video Grounding (TVG) due to its on-policy optimization, yet existing GRPO-based methods rema…

cs.RO2025

CoTaP: Compliant Task Pipeline and Reinforcement Learning of Its Controller with Compliance Modulation

Zewen He, Chenyuan Chen, Dilshod Azizov +1

Humanoid whole-body locomotion control is a critical approach for humanoid robots to leverage their inherent advantages. Learning-based control methods derived from retargeted huma…

cs.CV2020

Instance Scale Normalization for image understanding

Zewen He, He Huang, Yudong Wu +2

Scale variation remains a challenging problem for object detection. Common paradigms usually adopt multiscale training & testing (image pyramid) or FPN (feature pyramid network) to…