works on

From the 1 of 6 linked papers with an AI index.

activity
20242026
collaborators

6 papers

cs.CL2026

Pangram 4 Technical Report

Ben Glickenhaus, Katherine Thai, Jenna Russell +4

The paper introduces Pangram 4, a deep‑learning model for detecting AI‑generated text that achieves high accuracy, strong out‑of‑distribution robustness, and improved detection of…

cs.CL2026

AI use in American newspapers is widespread, uneven, and rarely disclosed

Jenna Russell, Marzena Karpinska, Destiny Akinode +4

AI is rapidly transforming journalism, but the extent of its use in published newspaper articles remains unclear. We address this gap by auditing a large-scale dataset of 186K arti…

cs.CL2025

EditLens: Quantifying the Extent of AI Editing in Text

Katherine Thai, Bradley Emi, Elyas Masrour +1

A significant proportion of queries to large language models ask them to edit user-provided text, rather than generate new text from scratch. While previous work focuses on detecti…

cs.AI2025

BEARCUBS: A benchmark for computer-using web agents

Yixiao Song, Katherine Thai, Chau Minh Pham +3

Modern web agents possess computer use abilities that allow them to interact with webpages by sending commands to a virtual keyboard and mouse. While such agents have considerable…

cs.CL2025

Literary Evidence Retrieval via Long-Context Language Models

Katherine Thai, Mohit Iyyer

How well do modern long-context language models understand literary fiction? We explore this question via the task of literary evidence retrieval, repurposing the RELiC dataset of…

cs.CL2024

One Thousand and One Pairs: A "novel" challenge for long-context language models

Marzena Karpinska, Katherine Thai, Kyle Lo +2

Synthetic long-context LLM benchmarks (e.g., "needle-in-the-haystack") test only surface-level retrieval capabilities, but how well can long-context LLMs retrieve, synthesize, and…