activity
20232025
collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2025

Teaching People LLM's Errors and Getting it Right

Nathan Stringham, Fateme Hashemi Chaleshtori, Xinyuan Yan +3

People use large language models (LLMs) when they should not. This is partly because they see LLMs compose poems and answer intricate questions, so they understandably, but incorre…

cs.CL2025

BriefMe: A Legal NLP Benchmark for Assisting with Legal Briefs

Jesse Woo, Fateme Hashemi Chaleshtori, Ana Marasović +1

A core part of legal work that has been under-explored in Legal NLP is the writing and editing of legal briefs. This requires not only a thorough understanding of the law of a juri…

cs.CL2025

What Has Been Lost with Synthetic Evaluation?

Alexander Gill, Abhilasha Ravichander, Ana Marasović

Large language models (LLMs) are increasingly used for data generation. However, creating evaluation benchmarks raises the bar for this emerging paradigm. Benchmarks must target sp…

cs.CL2025

Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps

Martin Tutek, Fateme Hashemi Chaleshtori, Ana Marasović +1

When prompted to think step-by-step, language models (LMs) produce a chain of thought (CoT), a sequence of reasoning steps that the model supposedly used to produce its prediction.…

cs.CL2024

Chain-of-Thought Unfaithfulness as Disguised Accuracy

Oliver Bentham, Nathan Stringham, Ana Marasović

Understanding the extent to which Chain-of-Thought (CoT) generations align with a large language model's (LLM) internal computations is critical for deciding whether to trust an LL…

cs.CL2023

Whispers of Doubt Amidst Echoes of Triumph in NLP Robustness

Ashim Gupta, Rishanth Rajendhran, Nathan Stringham +2

Do larger and more performant models resolve NLP's longstanding robustness issues? We investigate this question using over 20 models of different sizes spanning different architect…