activity
20242026
collaborators

6 papers

cs.CL2026

Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages

Anaelia Ovalle, Candace Ross, Sebastian Ruder +4

Large language models demonstrate strong reasoning capabilities through chain-of-thought prompting, but whether this reasoning quality transfers across languages remains underexplo…

cs.CL2026

Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language Models

Chantal Shaib, Vinith M. Suriyakumar, Levent Sagun +2

For an LLM to correctly respond to an instruction it must understand both the semantics and the domain (i.e., subject area) of a given task-instruction pair. However, syntax can al…

cs.CL2025

LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance

Patrick Haller, Mark Ibrahim, Polina Kirichenko +2

For Large Language Models (LLMs) to be reliable, they must learn robust knowledge that can be generally applied in diverse settings -- often unlike those seen during training. Yet,…

cs.CL2025

On the Role of Speech Data in Reducing Toxicity Detection Bias

Samuel J. Bell, Mariano Coria Meglioli, Megan Richards +6

Text toxicity detection systems exhibit significant biases, producing disproportionate rates of false positives on samples mentioning demographic groups. But what about toxicity de…

cs.CL2025

The Root Shapes the Fruit: On the Persistence of Gender-Exclusive Harms in Aligned Language Models

Anaelia Ovalle, Krunoslav Lehman Pavasovic, Louis Martin +5

Natural-language assistants are designed to provide users with helpful responses while avoiding harmful outputs, largely achieved through alignment to human preferences. Yet there…

cs.CL2024

Chained Tuning Leads to Biased Forgetting

Megan Ung, Alicia Sun, Samuel J. Bell +3

Large language models (LLMs) are often fine-tuned for use on downstream tasks, though this can degrade capabilities learned during previous training. This phenomenon, often referre…