activity
20162025
most citedBenchmarking Large Language Models for News Summarization

64 citations · 84 across the 18 of their papers we have counts for

collaborators

18 papers

cs.CL2025

Artificial Impressions: Evaluating Large Language Model Behavior Through the Lens of Trait Impressions

Nicholas Deas, Kathleen McKeown

We introduce and study artificial impressions--patterns in LLMs' internal representations of prompts that resemble human impressions and stereotypes based on language. We fit linea…

cs.CL2025

Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment

Yunfan Zhang, Kathleen McKeown, Smaranda Muresan

Large Language Models (LLMs) are typically trained to reflect a relatively uniform set of values, which limits their applicability to tasks that require understanding of nuanced hu…

cs.CL2025

Mining Contextualized Visual Associations from Images for Creativity Understanding

Ananya Sahu, Amith Ananthram, Kathleen McKeown

Understanding another person's creative output requires a shared language of association. However, when training vision-language models such as CLIP, we rely on web-scraped dataset…

cs.AI20251 cited

ModelingAgent: Bridging LLMs and Mathematical Modeling for Real-World Challenges

Cheng Qian, Hongyi Du, Hongru Wang +6

Recent progress in large language models (LLMs) has enabled substantial advances in solving mathematical problems. However, existing benchmarks often fail to reflect the complexity…

cs.LG2025

DYSTIL: Dynamic Strategy Induction with Large Language Models for Reinforcement Learning

Borui Wang, Kathleen McKeown, Rex Ying

Reinforcement learning from expert demonstrations has long remained a challenging research problem, and existing state-of-the-art methods using behavioral cloning plus further RL t…

cs.CL20242 cited

TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization

Liyan Tang, Igor Shalyminov, Amy Wing-mei Wong +11

Single document news summarization has seen substantial progress on faithfulness in recent years, driven by research on the evaluation of factual consistency, or hallucinations. We…