5 papers
AI Assistance for Human Review of Default Judgments
Theodora Worledge, Othman Bensouda Koraichi, Daniel Bernal +4
Overwhelmed courts in the United States review millions of default judgments each year. Unfortunately, such manual reviews are time-consuming and prone to error. In an audit of 188…
End-to-End Test-Time Training for Long Context
Arnuv Tandon, Karan Dalal, Xinhao Li +11
We formulate long-context language modeling as a problem in continual learning rather than architecture design. Under this formulation, we only use a standard architecture -- a Tra…
One-Minute Video Generation with Test-Time Training
Karan Dalal, Daniel Koceja, Gashon Hussein +12
Transformers today still struggle to generate one-minute videos because self-attention layers are inefficient for long context. Alternatives such as Mamba layers struggle with comp…
The Extractive-Abstractive Spectrum: Uncovering Verifiability Trade-offs in LLM Generations
Theodora Worledge, Tatsunori Hashimoto, Carlos Guestrin
Across all fields of academic study, experts cite their sources when sharing information. While large language models (LLMs) excel at synthesizing information, they do not provide…
Benchmarking Distributional Alignment of Large Language Models
Nicole Meister, Carlos Guestrin, Tatsunori Hashimoto
Language models (LMs) are increasingly used as simulacra for people, yet their ability to match the distribution of views of a specific demographic group and be \textit{distributio…