activity
20242026
collaborators

12 papers

cs.CV2026

Diffusion Bridge or Flow Matching? A Unifying Framework and Comparative Analysis

Kaizhen Zhu, Mokai Pan, Zhechuan Yu +3

Diffusion Bridge and Flow Matching have both demonstrated compelling empirical performance in transformation between arbitrary distributions. However, there remains confusion about…

cs.RO2026

DreamPolicy: A Unified World-model Policy for Scalable Humanoid Locomotion

Yahao Fan, Tianxiang Gui, Kaiyang Ji +8

Achieving versatile humanoid locomotion with a single policy presents a critical scalability challenge. Prevailing methods often rely on distilling multiple terrain-specific teache…

cs.RO2026

Commanding Humanoid by Free-form Language: A Large Language Action Model with Unified Motion Vocabulary

Zhirui Liu, Kaiyang Ji, Ke Yang +4

Enabling humanoid robots to follow free-form natural language commands is a critical step toward seamless human-robot interaction and general-purpose embodied AI. However, existing…

cs.CV2026

Gaze-guided Hand-Object Interaction Synthesis: Dataset and Method

Jie Tian, Ran Ji, Lingxiao Yang +6

Gaze plays a crucial role in revealing human attention and intention, particularly in hand-object interaction scenarios, where it guides and synchronizes complex tasks that require…

cs.CV2026

NLPrompt: Noise-Label Prompt Learning for Vision-Language Models

Bikang Pan, Qun Li, Xiaoying Tang +6

The emergence of vision-language foundation models, such as CLIP, has revolutionized image-text representation, enabling a broad range of applications via prompt learning. Despite…

cs.CV2025

A Unified and Fast-Sampling Diffusion Bridge Framework via Stochastic Optimal Control

Mokai Pan, Kaizhen Zhu, Yuexin Ma +4

Recent advances in diffusion bridge models leverage Doob's -transform to establish fixed endpoints between distributions, demonstrating promising results in image translation an…