activity
20242026
collaborators

10 papers

cs.CL2026

What's in a Name? Morphological Shortcuts by LLMs in Pharmacology

Kaijie Mo, Thomas Yang, Chantal Shaib +6

The morphological form of a word can often give cues to its meaning, but purely relying on these mappings can lead to overgeneralization in high-stakes domains. In the medical doma…

cs.CL2026

Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence

Kaijie Mo, Siddhartha Venkatayogi, Chantal Shaib +4

In high-stakes domains like medicine, it may be generally desirable for models to faithfully adhere to the context provided. But what happens if the context does not align with mod…

cs.LG2026

Brittlebench: Quantifying LLM robustness via prompt sensitivity

Angelika Romanou, Mark Ibrahim, Candace Ross +8

Existing evaluation methods largely rely on clean, static benchmarks, which can overestimate true model performance by failing to capture the noise and variability inherent in real…

cs.CL2026

Standardizing the Measurement of Text Diversity: A Tool and a Comparative Analysis of Scores

Chantal Shaib, Venkata S. Govindarajan, Joe Barrow +4

The diversity across outputs generated by LLMs shapes perception of their quality and utility. High lexical diversity is often desirable, but there is no standard method to measure…

cs.CL2026

Measuring AI "Slop" in Text

Chantal Shaib, Tuhin Chakrabarty, Diego Garcia-Olano +1

AI "slop" is an increasingly popular term used to describe low-quality AI-generated text, but there is currently no agreed upon definition of this term nor a means to measure its o…

cs.CL2026

Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language Models

Chantal Shaib, Vinith M. Suriyakumar, Levent Sagun +2

For an LLM to correctly respond to an instruction it must understand both the semantics and the domain (i.e., subject area) of a given task-instruction pair. However, syntax can al…