activity
20242026
most citedSketch2PoseNet: Efficient and Generalized Sketch to 3D Human Pose Prediction

1 citations · 1 across the 11 of their papers we have counts for

collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2026

HumanCLAW: Can Vision-Language Models Act Through a Body?

Li Siyao, Jiawei Gu, Shuai Liu +15

Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task…

cs.CV2026

Ms. Forcing: Efficient Streaming Video Generation with Multi-Scale Patchification and Attention

Zekun Li, Xiaoyan Cong, Hongyu Li +5

Streaming video diffusion models have made substantial progress toward interactive and dynamic world simulation, but the nested autoregressive and denoising loops of conventional n…

cs.CV2026

IAM: Identity-Aware Human Motion and Shape Joint Generation

Wenqi Jia, Zekun Li, Abhay Mittal +6

Recent advances in text-driven human motion generation enable models to synthesize realistic motion sequences from natural language descriptions. However, most existing approaches…

cs.CV2026

UMO: Unified In-Context Learning Unlocks Motion Foundation Model Priors

Xiaoyan Cong, Zekun Li, Zhiyang Dou +9

Large-scale foundation models (LFMs) have recently made impressive progress in text-to-motion generation by learning strong generative priors from massive 3D human motion datasets…

cs.CV2026

LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens

Zekun Li, Sizhe An, Chengcheng Tang +7

Recent progress in large models has led to significant advances in unified multimodal generation and understanding. However, the development of models that unify motion-language ge…

cs.CV2025

Pressure2Motion: Hierarchical Human Motion Reconstruction from Ground Pressure with Text Guidance

Zhengxuan Li, Qinhui Yang, Yiyu Zhuang +7

We present Pressure2Motion, a novel motion capture algorithm that reconstructs human motion from a ground pressure sequence and text prompt. At inference time, Pressure2Motion requ…