jailbreak attacks 2adversarial evaluation 1adversarial prompting 1best-of-N search 1code encoding 1model safety 1recovery decoding 1safety guards 1self-check defense 1vision-language models 1
From the 2 of 8 linked papers with an AI index.
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Temporal Flattening in LLM-Generated Text: Comparing Human and LLM Writing Trajectories
Zhanwei Cao, YeoJin Go, Yifan Hu +1
Large language models (LLMs) are increasingly used in daily applications, from content generation to code writing, where each interaction treats the model as stateless, generating…
cs.CL2025
PADBen: A Comprehensive Benchmark for Evaluating AI Text Detectors Against Paraphrase Attacks
Yiwei Zha, Rui Min, Shanu Sushmita
While AI-generated text (AIGT) detectors achieve over 90\% accuracy on direct LLM outputs, they fail catastrophically against iteratively-paraphrased content. We investigate why it…