From the 1 of 12 linked papers with an AI index.
9 papers · 1 filter
HumanCLAW: Can Vision-Language Models Act Through a Body?
Siyao Li, Li Siyao, Jiawei Gu +16
The paper introduces HumanCLAW, a framework that separates decision making of vision‑language models from low‑level motor execution, allowing evaluation of a model's action intelli…
Ms. Forcing: Efficient Streaming Video Generation with Multi-Scale Patchification and Attention
Zekun Li, Xiaoyan Cong, Hongyu Li +5
Streaming video diffusion models have made substantial progress toward interactive and dynamic world simulation, but the nested autoregressive and denoising loops of conventional n…
IAM: Identity-Aware Human Motion and Shape Joint Generation
Wenqi Jia, Zekun Li, Abhay Mittal +6
Recent advances in text-driven human motion generation enable models to synthesize realistic motion sequences from natural language descriptions. However, most existing approaches…
LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens
Zekun Li, Sizhe An, Chengcheng Tang +7
Recent progress in large models has led to significant advances in unified multimodal generation and understanding. However, the development of models that unify motion-language ge…
UMO: Unified In-Context Learning Unlocks Motion Foundation Model Priors
Xiaoyan Cong, Zekun Li, Zhiyang Dou +9
Large-scale foundation models (LFMs) have recently made impressive progress in text-to-motion generation by learning strong generative priors from massive 3D human motion datasets…
Pressure2Motion: Hierarchical Human Motion Reconstruction from Ground Pressure with Text Guidance
Zhengxuan Li, Qinhui Yang, Yiyu Zhuang +7
We present Pressure2Motion, a novel motion capture algorithm that reconstructs human motion from a ground pressure sequence and text prompt. At inference time, Pressure2Motion requ…