4 papers
Culturally-Adapted Red-Teaming Across East and Southeast Asian Contexts: A Methodological and Comparative Analysis
Hyeji Choi, Yongtaek Lim, Minwoo Kim
Multilingual safety evaluation of large language models (LLMs) has predominantly relied on direct translation (DT) of English benchmarks into target languages - an approach that co…
Reliable to Expressive: A Curriculum for Rubric-Following Safety Judges
Yongtaek Lim, Hyeji Choi, Minwoo Kim
Safety judges are increasingly deployed to evaluate model outputs against evolving criteria, yet recent meta-evaluation work shows they remain brittle under prompt and rubric varia…
STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming
MinJae Jung, YongTaek Lim, Chaeyun Kim +3
While Large Language Models (LLMs) are widely used, they remain susceptible to jailbreak prompts that can elicit harmful or inappropriate responses. This paper introduces STAR-Team…
CAGE: A Framework for Culturally Adaptive Red-Teaming Benchmark Generation
Chaeyun Kim, YongTaek Lim, Kihyun Kim +2
Existing red-teaming benchmarks, when adapted to new languages via direct translation, fail to capture socio-technical vulnerabilities rooted in local culture and law, creating a c…