Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Interpreting Style Representations via Style-Eliciting Prompts
Junghwan Kim, David Jurgens
Style representation learning is a powerful tool for authorship analysis and modeling writing style, yet the latent nature of learned representations makes them difficult to interp…
cs.CL2026
STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming
MinJae Jung, YongTaek Lim, Chaeyun Kim +3
While Large Language Models (LLMs) are widely used, they remain susceptible to jailbreak prompts that can elicit harmful or inappropriate responses. This paper introduces STAR-Team…