Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods
arXiv:2406.15583 · doi:10.1613/jair.1.16665
Abstract
Large language models (LLMs) have advanced to a point that even humans have difficulty discerning whether a text was generated by another human, or by a computer. However, knowing whether a text was produced by human or artificial intelligence (AI) is important to determining its trustworthiness, and has applications in many domains including detecting fraud and academic dishonesty, as well as combating the spread of misinformation and political propaganda. The task of AI-generated text (AIGT) detection is therefore both very challenging, and highly critical. In this survey, we summarize state-of-the art approaches to AIGT detection, including watermarking, statistical and stylistic analysis, and machine learning classification. We also provide information about existing datasets for this task. Synthesizing the research findings, we aim to provide insight into the salient factors that combine to determine how "detectable" AIGT text is under different scenarios, and to make practical recommendations for future work towards this significant technical and societal challenge.
References in corpus (40)
- Training language models to follow instructions with human feedback
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense
- GPT detectors are biased against non-native English writers
- Intrinsic Dimension Estimation for Robust Detection of AI-Generated Texts
- MGTBench: Benchmarking Machine-Generated Text Detection
- A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions
- RADAR: Robust AI-Text Detection via Adversarial Learning
- M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text Detection
- DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated Text
- Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature
- Robust Distortion-free Watermarks for Language Models
- Machine Generated Text: A Comprehensive Survey of Threat Models and Detection Methods
- CHEAT: A Large-scale Dataset for Detecting ChatGPT-writtEn AbsTracts
- Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
- Ghostbuster: Detecting Text Ghostwritten by Large Language Models
- DetectLLM: Leveraging Log Rank Information for Zero-Shot Detection of Machine-Generated Text
- Multiscale Positive-Unlabeled Detection of AI-Generated Texts
- On the Risk of Misinformation Pollution with Large Language Models
- Detecting LLM-Generated Text in Computing Education: A Comparative Study for ChatGPT Cases
- Watermarks in the Sand: Impossibility of Strong Watermarking for Generative Models
- Disinformation Detection: An Evolving Challenge in the Age of LLMs
- REMARK-LLM: A Robust and Efficient Watermarking Framework for Generative Large Language Models
- HC3 Plus: A Semantic-Invariant Human ChatGPT Comparison Corpus
- Overview of AuTexTification at IberLEF 2023: Detection and Attribution of Machine-Generated Text in Multiple Domains
- Deepfake Text Detection: Limitations and Opportunities
- Few-Shot Detection of Machine-Generated Text using Style Representations
- SeqXGPT: Sentence-Level AI-Generated Text Detection
- Robust Multi-bit Natural Language Watermarking through Invariant Features
- OUTFOX: LLM-Generated Essay Detection Through In-Context Learning with Adversarially Generated Examples
- Fighting Fire with Fire: The Dual Role of LLMs in Crafting and Detecting Elusive Disinformation
- Contra generative AI detection in higher education assessments
- Red Teaming Language Model Detectors with Language Models
- Attribution and Obfuscation of Neural Text Authorship: A Data Mining Perspective
- GPT-who: An Information Density-based Machine-Generated Text Detector
- J-Guard: Journalism Guided Adversarially Robust Detection of AI-generated News
- Advancing Beyond Identification: Multi-bit Watermark for Large Language Models
- FACT-GPT: Fact-Checking Augmentation via Claim Matching with LLMs
- Stumbling Blocks: Stress Testing the Robustness of Machine-Generated Text Detectors Under Attacks
- RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors