7 papers
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
Yuxue Yang, Shuyao Shang, Jiahe Wang +13
Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated…
Self-Evolving Just-In-Time Memory for Proactive Embodied Safety
Bingrui Sima, Lizhong Wang, Xiaoya Lu +2
While Vision-Language Models (VLMs) have empowered embodied agents to execute complex household tasks, they struggle to proactively handle dynamically emerging hazards during close…
FiberTune: Preserving Action-Fiber Visual Residuals in Vision-Language-Action Fine-Tuning
Haihao Lin, Xiangsheng Huang, Xiao Yang +7
Action-supervised fine-tuning of vision-language-action (VLA) policies fits demonstrations effectively but constrains only the directions that change predicted actions, leaving vis…
PACT: Self-Evolving Physical Safety Alignment for Diffusion Policies in Embodied Manipulation
Lingxuan Wu, Zijian Zhu, Lizhong Wang +5
Diffusion policies have achieved remarkable success in robotic manipulation, yet they often fail to satisfy strict physical constraints required for safe deployment. Existing appro…
minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models
Min Zhao, Hongzhou Zhu, Bokai Yan +9
Recent video diffusion foundation models have achieved remarkable progress in high-quality video generation, yet turning them into real-time interactive video world models remains…
Invisible to Humans, Triggered by Agents: Stealthy Jailbreak Attacks on Mobile Vision-Language Agents
Renhua Ding, Xiao Yang, Zhengwei Fang +3
Large Vision-Language Models (LVLMs) empower autonomous mobile agents, yet their security under realistic mobile deployment constraints remains underexplored. While agents are vuln…