3 papers
cs.CV2026
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
Junliang Ye, Kenkun Liu, Guocun Wang +13
Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling…
cs.AI2026
UniMo: Unified Motion Generation and Understanding with Chain of Thought
Guocun Wang, Kenkun Liu, Jing Lin +3
Existing 3D human motion generation and understanding methods often exhibit limited interpretability, restricting effective mutual enhancement between these inherently related task…
cs.CV2025
Towards Fine-Grained Human Motion Video Captioning
Guorui Song, Guocun Wang, Zhe Huang +4
Generating accurate descriptions of human actions in videos remains a challenging task for video captioning models. Existing approaches often struggle to capture fine-grained motio…