activity
20242026
collaborators

8 papers

cs.CL2026

Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States

Denis Peskoff, Joe Barrow, Christopher Vu +1

Progress in legal AI increasingly depends on access to authoritative legal text at scale. Yet one of the most consequential layers of American law remains largely absent from exist…

cs.CL2026

Standardizing the Measurement of Text Diversity: A Tool and a Comparative Analysis of Scores

Chantal Shaib, Venkata S. Govindarajan, Joe Barrow +4

The diversity across outputs generated by LLMs shapes perception of their quality and utility. High lexical diversity is often desirable, but there is no standard method to measure…

cs.CL2025

SafePassage: High-Fidelity Information Extraction with Black Box LLMs

Joe Barrow, Raj Patel, Misha Kharkovski +2

Black box large language models (LLMs) make information extraction (IE) easy to configure, but hard to trust. Unlike traditional information extraction pipelines, the information "…

cs.CV2025

CommonForms: A Large, Diverse Dataset for Form Field Detection

Joe Barrow

This paper introduces CommonForms, a web-scale dataset for form field detection. It casts the problem of form field detection as object detection: given an image of a page, predict…

cs.CL2025

Personalization of Large Language Models: A Survey

Zhehao Zhang, Ryan A. Rossi, Branislav Kveton +18

Personalization of Large Language Models (LLMs) has recently become increasingly important with a wide range of applications. Despite the importance and recent progress, most exist…

cs.CV2025

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality

Mohamed Elmoghany, Ryan Rossi, Seunghyun Yoon +26

Despite the significant progress that has been made in video generative models, existing state-of-the-art methods can only produce videos lasting 5-16 seconds, often labeled "long-…