3 papers
cs.LG2026
Self-Evolving Just-In-Time Memory for Proactive Embodied Safety
Bingrui Sima, Lizhong Wang, Xiaoya Lu +2
While Vision-Language Models (VLMs) have empowered embodied agents to execute complex household tasks, they struggle to proactively handle dynamically emerging hazards during close…
cs.CV2026
HomeGuard: VLM-based Embodied Safeguard for Identifying Contextual Risk in Household Task
Xiaoya Lu, Yijin Zhou, Zeren Chen +6
Vision-Language Models (VLMs) empower embodied agents to execute complex instructions, yet they remain vulnerable to contextual safety risks where benign commands become hazardous…
cs.CV2025
VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models
Bingrui Sima, Linhua Cong, Wenxuan Wang +1
The emergence of Multimodal Large Language Models (MLRMs) has enabled sophisticated visual reasoning capabilities by integrating reinforcement learning and Chain-of-Thought (CoT) s…