5 papers
Visible Yet Unreadable: A Systematic Blind Spot of Vision Language Models Across Writing Systems
Jie Zhang, Ting Xu, Gelei Deng +5
Writing is a universal cultural technology that reuses vision for symbolic communication. Humans display striking resilience: we readily recognize words even when characters are fr…
Cowpox: Towards the Immunity of VLM-based Multi-Agent Systems
Yutong Wu, Jie Zhang, Yiming Li +4
Vision Language Model (VLM)-based agents are stateful, autonomous entities capable of perceiving and interacting with their environments through vision and language. Multi-agent sy…
PoseGuard: Pose-Guided Generation with Safety Guardrails
Kongxin Wang, Jie Zhang, Peigui Qi +3
Pose-guided video generation has become a powerful tool in creative industries, exemplified by frameworks like Animate Anyone. However, conditioning generation on specific poses in…
BURN: Backdoor Unlearning via Adversarial Boundary Analysis
Yanghao Su, Jie Zhang, Yiming Li +6
Backdoor unlearning aims to remove backdoor-related information while preserving the model's original functionality. However, existing unlearning methods mainly focus on recovering…
IRCopilot: Automated Incident Response with Large Language Models
Xihuan Lin, Jie Zhang, Gelei Deng +4
Incident response plays a pivotal role in mitigating the impact of cyber attacks. In recent years, the intensity and complexity of global cyber threats have grown significantly, ma…