Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Risk Awareness Injection: Calibrating Vision-Language Models for Safety without Compromising Utility
Mengxuan Wang, Yuxin Chen, Gang Xu +3
Vision language models (VLMs) extend the reasoning capabilities of large language models (LLMs) to cross-modal settings, yet remain highly vulnerable to multimodal jailbreak attack…
cs.AI2026
When Agents See Humans as the Outgroup: Belief-Dependent Bias in LLM-Powered Agents
Zongwei Wang, Bincheng Gu, Hongyu Yu +5
This paper reveals that LLM-powered agents exhibit not only demographic bias (e.g., gender, religion) but also intergroup bias under minimal "us" versus "them" cues. When such grou…