5 papers · 1 filter
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
Yuxue Yang, Shuyao Shang, Jiahe Wang +13
Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated…
FiberTune: Preserving Action-Fiber Visual Residuals in Vision-Language-Action Fine-Tuning
Haihao Lin, Xiangsheng Huang, Xiao Yang +7
Action-supervised fine-tuning of vision-language-action (VLA) policies fits demonstrations effectively but constrains only the directions that change predicted actions, leaving vis…
minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models
Min Zhao, Hongzhou Zhu, Bokai Yan +9
Recent video diffusion foundation models have achieved remarkable progress in high-quality video generation, yet turning them into real-time interactive video world models remains…
LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model
Jiachun Jin, Zetong Zhou, Xiao Yang +4
Unified models (UMs) hold promise for their ability to understand and generate content across heterogeneous modalities. Compared to merely generating visual content, the use of UMs…
Feedback-based Modal Mutual Search for Attacking Vision-Language Pre-training Models
Renhua Ding, Xinze Zhang, Xiao Yang +1
Although vision-language pre-training (VLP) models have achieved remarkable progress on cross-modal tasks, they remain vulnerable to adversarial attacks. Using data augmentation an…