4 papers
FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming
Chaeyun Kim, Daeyoung Park, Junghwan Kim +4
Existing safety benchmarks target general adversarial scenarios but miss finance-specific risks. Financial LLMs face regulatory compliance violations, fraud facilitation, and syste…
Interpreting Style Representations via Style-Eliciting Prompts
Junghwan Kim, David Jurgens
Style representation learning is a powerful tool for authorship analysis and modeling writing style, yet the latent nature of learned representations makes them difficult to interp…
STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming
MinJae Jung, YongTaek Lim, Chaeyun Kim +3
While Large Language Models (LLMs) are widely used, they remain susceptible to jailbreak prompts that can elicit harmful or inappropriate responses. This paper introduces STAR-Team…
CAGE: A Framework for Culturally Adaptive Red-Teaming Benchmark Generation
Chaeyun Kim, YongTaek Lim, Kihyun Kim +2
Existing red-teaming benchmarks, when adapted to new languages via direct translation, fail to capture socio-technical vulnerabilities rooted in local culture and law, creating a c…