4 papers
Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners
Charlotte Li, Nick Hagar, Sachita Nishal +2
Benchmarks play a significant role in how technology companies communicate about model capabilities and how researchers and the public understand generative AI systems. However, ex…
Not Wrong, But Untrue: LLM Overconfidence in Document-Based Queries
Nick Hagar, Wilma Agustianto, Nicholas Diakopoulos
Large language models (LLMs) are increasingly used in newsroom workflows, but their tendency to hallucinate poses risks to core journalistic practices of sourcing, attribution, and…
On-Premise AI for the Newsroom: Evaluating Small Language Models for Investigative Document Search
Nick Hagar, Nicholas Diakopoulos, Jeremy Gilbert
Investigative journalists routinely confront large document collections. Large language models (LLMs) with retrieval-augmented generation (RAG) capabilities promise to accelerate t…
LLM-Assisted News Discovery in High-Volume Information Streams: A Case Study
Nick Hagar, Ethan Silver, Clare Spencer +1
Journalists face mounting challenges in monitoring ever-expanding digital information streams to identify newsworthy content. While traditional automation tools gather information…