4 papers · 1 filter
Open-UniMo: Towards Unified Motion-Language Understanding and Generation in the Open World
Guocun Wang, Kenkun Liu, Guorui Song +7
Unified motion generation and understanding is crucial for embodied AI systems that can both synthesize and interpret human actions in open-world environments. Existing motion-lang…
DIVA: Exploiting Cross-Step Conditional Propagation for Visual Jailbreaks in Discrete Diffusion Vision-Language Models
Guorui Song, Runqing Tang, Jingye Zhang +11
Large vision-language models (VLMs) are increasingly deployed in safety-critical settings, yet existing visual jailbreak research has focused almost exclusively on autoregressive a…
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
Junliang Ye, Kenkun Liu, Guocun Wang +13
Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling…
Towards Fine-Grained Human Motion Video Captioning
Guorui Song, Guocun Wang, Zhe Huang +4
Generating accurate descriptions of human actions in videos remains a challenging task for video captioning models. Existing approaches often struggle to capture fine-grained motio…