Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation
Genglin Liu, Muye Zhang, Krishnamurthy Viswanathan +5
Multimodal Large Language Models (MLLMs) are increasingly deployed for nuanced content safety and moderation tasks, yet they remain vulnerable to adversarial attacks and out-of-dis…
cs.AI2023
JAB: Joint Adversarial Prompting and Belief Augmentation
Ninareh Mehrabi, Palash Goyal, Anil Ramakrishna +6
With the recent surge of language models in different applications, attention to safety and robustness of these models has gained significant importance. Here we introduce a joint…