collaborators

5 papers

cs.CV2026

HumanCLAW: Can Vision-Language Models Act Through a Body?

Siyao Li, Li Siyao, Jiawei Gu +16

The paper introduces HumanCLAW, a framework that separates decision making of vision‑language models from low‑level motor execution, allowing evaluation of a model's action intelli…

cs.CV2026

Ms. Forcing: Efficient Streaming Video Generation with Multi-Scale Patchification and Attention

Zekun Li, Xiaoyan Cong, Hongyu Li +5

Streaming video diffusion models have made substantial progress toward interactive and dynamic world simulation, but the nested autoregressive and denoising loops of conventional n…

cs.CV2026

IAM: Identity-Aware Human Motion and Shape Joint Generation

Wenqi Jia, Zekun Li, Abhay Mittal +6

Recent advances in text-driven human motion generation enable models to synthesize realistic motion sequences from natural language descriptions. However, most existing approaches…

cs.CV2026

UMO: Unified In-Context Learning Unlocks Motion Foundation Model Priors

Xiaoyan Cong, Zekun Li, Zhiyang Dou +9

Large-scale foundation models (LFMs) have recently made impressive progress in text-to-motion generation by learning strong generative priors from massive 3D human motion datasets…

cs.CV2026

EgoReAct: Egocentric Video-Driven 3D Human Reaction Generation

Libo Zhang, Zekun Li, Tianyu Li +10

Humans exhibit adaptive, context-sensitive responses to egocentric visual input. However, faithfully modeling such reactions from egocentric video remains challenging due to the du…