Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods
arXiv:2406.15583 · doi:10.1613/jair.1.16665
Abstract
Large language models (LLMs) have advanced to a point that even humans have difficulty discerning whether a text was generated by another human, or by a computer. However, knowing whether a text was produced by human or artificial intelligence (AI) is important to determining its trustworthiness, and has applications in many domains including detecting fraud and academic dishonesty, as well as combating the spread of misinformation and political propaganda. The task of AI-generated text (AIGT) detection is therefore both very challenging, and highly critical. In this survey, we summarize state-of-the art approaches to AIGT detection, including watermarking, statistical and stylistic analysis, and machine learning classification. We also provide information about existing datasets for this task. Synthesizing the research findings, we aim to provide insight into the salient factors that combine to determine how "detectable" AIGT text is under different scenarios, and to make practical recommendations for future work towards this significant technical and societal challenge.
References in corpus (46)
- Training language models to follow instructions with human feedback
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense
- GPT detectors are biased against non-native English writers
- Intrinsic Dimension Estimation for Robust Detection of AI-Generated Texts
- MGTBench: Benchmarking Machine-Generated Text Detection
- A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions
- RADAR: Robust AI-Text Detection via Adversarial Learning
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence Frontiers
- M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text Detection
- DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated Text
- Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature
- Robust Distortion-free Watermarks for Language Models
- CHEAT: A Large-scale Dataset for Detecting ChatGPT-writtEn AbsTracts
- Machine Generated Text: A Comprehensive Survey of Threat Models and Detection Methods
- Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
- Ghostbuster: Detecting Text Ghostwritten by Large Language Models
- DetectLLM: Leveraging Log Rank Information for Zero-Shot Detection of Machine-Generated Text
- Multiscale Positive-Unlabeled Detection of AI-Generated Texts
- On the Risk of Misinformation Pollution with Large Language Models
- Detecting LLM-Generated Text in Computing Education: A Comparative Study for ChatGPT Cases
- Disinformation Detection: An Evolving Challenge in the Age of LLMs
- Watermarks in the Sand: Impossibility of Strong Watermarking for Generative Models
- REMARK-LLM: A Robust and Efficient Watermarking Framework for Generative Large Language Models
- Overview of AuTexTification at IberLEF 2023: Detection and Attribution of Machine-Generated Text in Multiple Domains
- Calibrated Language Model Fine-Tuning for In- and Out-of-Distribution Data
- Deepfake Text Detection: Limitations and Opportunities
- HC3 Plus: A Semantic-Invariant Human ChatGPT Comparison Corpus
- Faking Fake News for Real Fake News Detection: Propaganda-loaded Training Data Generation
- SeqXGPT: Sentence-Level AI-Generated Text Detection
- Few-Shot Detection of Machine-Generated Text using Style Representations
- Robust Multi-bit Natural Language Watermarking through Invariant Features
- Fighting Fire with Fire: The Dual Role of LLMs in Crafting and Detecting Elusive Disinformation
- OUTFOX: LLM-Generated Essay Detection Through In-Context Learning with Adversarially Generated Examples
- Contra generative AI detection in higher education assessments
- Neural Deepfake Detection with Factual Structure of Text
- Red Teaming Language Model Detectors with Language Models
- GPT-who: An Information Density-based Machine-Generated Text Detector
- Attribution and Obfuscation of Neural Text Authorship: A Data Mining Perspective
- Stumbling Blocks: Stress Testing the Robustness of Machine-Generated Text Detectors Under Attacks
- FACT-GPT: Fact-Checking Augmentation via Claim Matching with LLMs
- J-Guard: Journalism Guided Adversarially Robust Detection of AI-generated News
- Advancing Beyond Identification: Multi-bit Watermark for Large Language Models
- RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors