7 papers
PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion Representation
Junchuan Zhao, Qifan Liang, Ye Wang
Co-speech gesture generation aims to synthesize realistic body movements that are semantically coherent with speech and faithful to a user-specified gestural style. Existing VQ-VAE…
Learning Visual Feature-Based World Models via Residual Latent Action
Xinyu Zhang, Zhengtong Xu, Yutian Tao +3
World models predict future transitions from observations and actions. Existing works predominantly focus on image generation only. Visual feature-based world models, on the other…
TripVVT: A Large-Scale Triplet Dataset and a Coarse-Mask Baseline for In-the-Wild Video Virtual Try-On
Dingbao Shao, Song Wu, Shenyi Wang +9
Due to the scarcity of large-scale in-the-wild triplet data and the improper use of masks, the performance of video virtual try-on models remains limited. In this paper, we first i…
MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings
Zijie Li, Yichun Shi, Jingxiang Sun +8
We present MMCORE, a unified framework designed for multimodal image generation and editing. MMCORE leverages a pre-trained Vision-Language Model (VLM) to predict semantic visual e…
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model
Bin Cao, Sipeng Zheng, Ye Wang +5
Human motion generation has emerged as a critical technology with transformative potential for real-world applications. However, existing vision-language-motion models (VLMMs) face…
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
Hao Luo, Yicheng Feng, Wanpeng Zhang +7
We introduce Being-H0, a dexterous Vision-Language-Action model (VLA) trained on large-scale human videos. Existing VLAs struggle with complex manipulation tasks requiring high dex…