activity
20172026
most citedBenchmarking Cognitive Biases in Large Language Models as Evaluators

25 citations · 105 across the 54 of their papers we have counts for

collaborators
Showing 2025 · cs.CLShow all

8 papers · 2 filters

cs.CL2025

Mary, the Cheeseburger-Eating Vegetarian: Do LLMs Recognize Incoherence in Narratives?

Karin de Langis, Püren Öncel, Ryan Peters +4

Leveraging a dataset of paired narratives, we investigate the extent to which large language models (LLMs) can reliably separate incoherent and coherent stories. A probing study fi…

cs.CL2025

Tracing How Annotators Think: Augmenting Preference Judgments with Reading Processes

Karin de Langis, William Walker, Khanh Chi Le +1

We propose an annotation approach that captures not only labels but also the reading process underlying annotators' decisions, e.g., what parts of the text they focus on, re-read o…

cs.CL2025

How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMs

Karin de Langis, Jong Inn Park, Andreas Schramm +5

Large language models (LLMs) exhibit increasingly sophisticated linguistic capabilities, yet the extent to which these behaviors reflect human-like cognition versus advanced patter…

cs.CL2025

Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models

Zae Myung Kim, Chanwoo Park, Vipul Raheja +2

Reward-based alignment methods for large language models (LLMs) face two key limitations: vulnerability to reward hacking, where models exploit flaws in the reward signal; and reli…

cs.CL2025

Stealing Creator's Workflow: A Creator-Inspired Agentic Framework with Iterative Feedback Loop for Improved Scientific Short-form Generation

Jong Inn Park, Maanas Taneja, Qianwen Wang +1

Generating engaging, accurate short-form videos from scientific papers is challenging due to content complexity and the gap between expert authors and readers. Existing end-to-end…

cs.CL2025

LawFlow: Collecting and Simulating Lawyers' Thought Processes on Business Formation Case Studies

Debarati Das, Khanh Chi Le, Ritik Sachin Parkar +8

Legal practitioners, particularly those early in their careers, face complex, high-stakes tasks that require adaptive, context-sensitive reasoning. While AI holds promise in suppor…