5 papers
Language-Guided Transformer Tokenizer for Human Motion Generation
Sheng Yan, Yong Wang, Xin Du +2
In this paper, we focus on motion discrete tokenization, which converts raw motion into compact discrete tokens--a process proven crucial for efficient motion generation. In this p…
GLM-OCR Technical Report
Shuaiqi Duan, Yadong Xue, Weihan Wang +20
GLM-OCR is an efficient 0.9B-parameter compact multimodal model designed for real-world document understanding. It combines a 0.4B-parameter CogViT visual encoder with a 0.5B-param…
Prompt When the Animal is: Temporal Animal Behavior Grounding with Positional Recovery Training
Sheng Yan, Xin Du, Zongying Li +3
Temporal grounding is crucial in multimodal learning, but it poses challenges when applied to animal behavior data due to the sparsity and uniform distribution of moments. To addre…
Cross-Modal Retrieval for Motion and Text via DropTriple Loss
Sheng Yan, Yang Liu, Haoqiang Wang +3
Cross-modal retrieval of image-text and video-text is a prominent research area in computer vision and natural language processing. However, there has been insufficient attention g…
MoSa: Motion Generation with Scalable Autoregressive Modeling
Mengyuan Liu, Sheng Yan, Yong Wang +3
We introduce MoSa, a novel hierarchical motion generation framework for text-driven 3D human motion generation that enhances the Vector Quantization-guided Generative Transformers…