From the 1 of 7 linked papers with an AI index.
68 citations · 69 across the 4 of their papers we have counts for
7 papers
Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions
Zachary Izzo
The paper examines how closely a large language model's next-token predictions match the empirical next-token distribution derived from its training data, showing strong agreement…
Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
Weixin Liang, Zachary Izzo, Yaohui Zhang +9
We present an approach for estimating the fraction of text in a large corpus which is likely to be substantially modified or produced by a large language model (LLM). Our maximum l…
Subgroup Discovery with the Cox Model
Zachary Izzo, Iain Melvin
We study the problem of subgroup discovery for survival analysis, where the goal is to find an interpretable subset of the data on which a Cox model is highly accurate. Our work is…
To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis
Federico Bianchi, Yongchan Kwon, Zachary Izzo +2
How many mistakes do published AI papers contain? Peer-reviewed publications form the foundation upon which new research and knowledge are built. Errors that persist in the literat…
Quantitative Bounds for Length Generalization in Transformers
Zachary Izzo, Eshaan Nichani, Jason D. Lee
We study the problem of length generalization (LG) in transformers: the ability of a model trained on shorter sequences to maintain performance when evaluated on much longer, previ…
Group Relative Augmentation for Data Efficient Action Detection
Deep Anil Patel, Iain Melvin, Zachary Izzo +1
Adapting large Video-Language Models (VLMs) for action detection using only a few examples poses challenges like overfitting and the granularity mismatch between scene-level pre-tr…