collaborators

20 papers

cs.CL2026

Shared Lexical Task Representations Explain Behavioral Variability In LLMs

Zhuonan Yang, Jacob Xiaochen Li, Francisco Piedrahita Velez +5

One of the most common complaints about large language models (LLMs) is their prompt sensitivity -- that is, the fact that their ability to perform a task or provide a correct answ…

cs.AI2026

Adversarial Concept Search: Predicting Compositional Errors From Feature Geometry

Jennifer Meng Lu, Ruochen Zhang, Isabelle Lee +3

Humans cannot always intuit what scenarios are most challenging to LLMs. Hoping to capture challenging edge cases, developers either design problems to be difficult for humans or c…

cs.CL2026

Lessons from the Trenches on Reproducible Evaluation of Language Models

Stella Biderman, Hailey Schoelkopf, Lintang Sutawika +27

Reliable evaluation of language models (LMs) remains an open challenge. Re- searchers and engineers face methodological issues such as the sensitivity of models to evaluation setup…

cs.CL2026

Mitigating Extrinsic Gender Bias for Bangla Classification Tasks

Sajib Kumar Saha Joy, Arman Hassan Mahy, Meherin Sultana +4

In this study, we investigate extrinsic gender bias in Bangla pretrained language models, a largely underexplored area in low-resource languages. To assess this bias, we construct…

cs.CL2026

How Do Language Models Compose Functions?

Apoorv Khandelwal, Ellie Pavlick

While large language models (LLMs) appear to be increasingly capable of solving compositional tasks, it is an open question whether they do so using compositional mechanisms. In th…

cs.LG2026

Handling and Interpreting Missing Modalities in Patient Clinical Trajectories via Autoregressive Sequence Modeling

Andrew Wang, Ellie Pavlick, Ritambhara Singh

An active challenge in developing multimodal machine learning (ML) models for healthcare is handling missing modalities during training and deployment. As clinical datasets are inh…