4 papers
LLMs struggle to simulate human belief updates in controlled environments
Sebastian Pohl, Harsh Mehta, Pranav Mambayil +4
The paper tests whether large language models can replicate how humans update their beliefs after reading Reddit comments, comparing model outputs to data from 391 UK participants,…
Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025
Maria Kunilovskaya, Gagan Bhatia, Lisa Sophie Albertelli +10
Human annotation is the empirical foundation of much NLP research, from dataset construction to model evaluation, but papers often leave unclear who produced the annotations and ho…
DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain
Hsuvas Borkakoty, Sebastian Pohl, Cheng Wang +2
LLM-based agents are increasingly deployed for routine but consequential tasks in real-world domains, where their behavior is governed by inherently ambiguous domain policies that…
Towards a Principled Evaluation of Knowledge Editors
Sebastian Pohl, Max Ploner, Alan Akbik
Model editing has been gaining increasing attention over the past few years. For Knowledge Editing in particular, more challenging evaluation datasets have recently been released.…