Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Open-UniMo: Towards Unified Motion-Language Understanding and Generation in the Open World
Guocun Wang, Kenkun Liu, Guorui Song +7
Unified motion generation and understanding is crucial for embodied AI systems that can both synthesize and interpret human actions in open-world environments. Existing motion-lang…
cs.CV2026
DIVA: Exploiting Cross-Step Conditional Propagation for Visual Jailbreaks in Discrete Diffusion Vision-Language Models
Guorui Song, Runqing Tang, Jingye Zhang +11
Large vision-language models (VLMs) are increasingly deployed in safety-critical settings, yet existing visual jailbreak research has focused almost exclusively on autoregressive a…