From the 1 of 6 linked papers with an AI index.
6 papers
Pangram 4 Technical Report
Ben Glickenhaus, Katherine Thai, Jenna Russell +4
The paper introduces Pangram 4, a deep‑learning model for detecting AI‑generated text that achieves high accuracy, strong out‑of‑distribution robustness, and improved detection of…
AI use in American newspapers is widespread, uneven, and rarely disclosed
Jenna Russell, Marzena Karpinska, Destiny Akinode +4
AI is rapidly transforming journalism, but the extent of its use in published newspaper articles remains unclear. We address this gap by auditing a large-scale dataset of 186K arti…
EditLens: Quantifying the Extent of AI Editing in Text
Katherine Thai, Bradley Emi, Elyas Masrour +1
A significant proportion of queries to large language models ask them to edit user-provided text, rather than generate new text from scratch. While previous work focuses on detecti…
BEARCUBS: A benchmark for computer-using web agents
Yixiao Song, Katherine Thai, Chau Minh Pham +3
Modern web agents possess computer use abilities that allow them to interact with webpages by sending commands to a virtual keyboard and mouse. While such agents have considerable…
Literary Evidence Retrieval via Long-Context Language Models
Katherine Thai, Mohit Iyyer
How well do modern long-context language models understand literary fiction? We explore this question via the task of literary evidence retrieval, repurposing the RELiC dataset of…
One Thousand and One Pairs: A "novel" challenge for long-context language models
Marzena Karpinska, Katherine Thai, Kyle Lo +2
Synthetic long-context LLM benchmarks (e.g., "needle-in-the-haystack") test only surface-level retrieval capabilities, but how well can long-context LLMs retrieve, synthesize, and…