5 papers
RoboSafe: Safeguarding Embodied Agents via Executable Safety Logic
Le Wang, Zonghao Ying, Xiao Yang +7
Embodied agents powered by vision-language models (VLMs) are increasingly capable of executing complex real-world tasks, yet they remain vulnerable to hazardous instructions that m…
AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
Zonghao Ying, Le Wang, Yisong Xiao +7
The integration of vision-language models (VLMs) is driving a new generation of embodied agents capable of operating in human-centered environments. However, as deployment expands,…
Manipulating Multimodal Agents via Cross-Modal Prompt Injection
Le Wang, Zonghao Ying, Tianyuan Zhang +5
The emergence of multimodal large language models has redefined the agent paradigm by integrating language and vision modalities with external data sources, enabling agents to bett…
BFA: Best-Feature-Aware Fusion for Multi-View Fine-grained Manipulation
Zihan Lan, Weixin Mao, Haosheng Li +4
In real-world scenarios, multi-view cameras are typically employed for fine-grained manipulation tasks. Existing approaches (e.g., ACT) tend to treat multi-view features equally an…
CogMorph: Cognitive Morphing Attacks for Text-to-Image Models
Zonglei Jing, Zonghao Ying, Le Wang +4
The development of text-to-image (T2I) generative models, that enable the creation of high-quality synthetic images from textual prompts, has opened new frontiers in creative desig…