Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
Kai Hu, Abhinav Aggarwal, Mehran Khodabandeh +6
This paper introduces Jailbreak-Zero, a novel red teaming methodology that shifts the paradigm of Large Language Model (LLM) safety evaluation from a constrained example-based appr…
cs.CL2025
Domain Gating Ensemble Networks for AI-Generated Text Detection
Arihant Tripathi, Liam Dugan, Charis Gao +6
As state-of-the-art language models continue to improve, the need for robust detection of machine-generated text becomes increasingly critical. However, current state-of-the-art ma…