From the 1 of 7 linked papers with an AI index.
7 papers
MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models
Sihan Chen, Jiale Li, Jianghang Lin +1
The paper introduces MoHallBench, a large benchmark designed to evaluate and diagnose motion hallucination—incorrectly inferred human motions—in video large language models, coveri…
Language-Guided Transformer Tokenizer for Human Motion Generation
Sheng Yan, Yong Wang, Xin Du +2
In this paper, we focus on motion discrete tokenization, which converts raw motion into compact discrete tokens--a process proven crucial for efficient motion generation. In this p…
UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
Ziyi Wang, Xinshun Wang, Shuang Chen +2
We present UniMotion, to our knowledge the first unified framework for simultaneous understanding and generation of human motion, natural language, and RGB images within a single a…
Universal Skeleton Understanding via Differentiable Rendering and MLLMs
Ziyi Wang, Peiming Li, Xinshun Wang +3
Multimodal large language models (MLLMs) exhibit strong visual-language reasoning, yet cannot process structured, non-visual data such as human skeletons. Existing methods either c…
Prompt When the Animal is: Temporal Animal Behavior Grounding with Positional Recovery Training
Sheng Yan, Xin Du, Zongying Li +3
Temporal grounding is crucial in multimodal learning, but it poses challenges when applied to animal behavior data due to the sparsity and uniform distribution of moments. To addre…
MoSa: Motion Generation with Scalable Autoregressive Modeling
Mengyuan Liu, Sheng Yan, Yong Wang +3
We introduce MoSa, a novel hierarchical motion generation framework for text-driven 3D human motion generation that enhances the Vector Quantization-guided Generative Transformers…