6 papers
Show, Don't TELL: Explainable AI-Generated Text Detection
Aldan Creo, Suraj Ranganath
Research on AI-generated text detection has presented a number of approaches to discern human from AI prose, some of which achieving high in-distribution performance. However, real…
Fine-Tuning Without Forgetting via Loss-Adaptive Learning Rates
Parjanya Prajakta Prashant, Jiongli Zhu, Aldan Creo +1
Fine-tuning large language models on new data improves task performance but degrades capabilities learned during pretraining, a phenomenon known as catastrophic forgetting. Existin…
Complete Evasion, Zero Modification: PDF Attacks on AI Text Detection
Aldan Creo
AI-generated text detectors have become essential tools for maintaining content authenticity, yet their robustness against evasion attacks remains questionable. We present PDFuzz,…
Mass-Scale Analysis of In-the-Wild Conversations Reveals Complexity Bounds on LLM Jailbreaking
Aldan Creo, Raul Castro Fernandez, Manuel Cebrian
As large language models (LLMs) become increasingly deployed, understanding the complexity and evolution of jailbreaking strategies is critical for AI safety. We present a mass-sca…
Ask a Local: Detecting Hallucinations With Specialized Model Divergence
Aldan Creo, Héctor Cerezo-Costas, Pedro Alonso-Doval +1
Hallucinations in large language models (LLMs) - instances where models generate plausible but factually incorrect information - present a significant challenge for AI. We introduc…
SilverSpeak: Evading AI-Generated Text Detectors using Homoglyphs
Aldan Creo, Shushanta Pudasaini
The advent of Large Language Models (LLMs) has enabled the generation of text that increasingly exhibits human-like characteristics. As the detection of such content is of signific…