collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2026

LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering

Huan Wu, Ali Emami, Muhammad Furquan Hassan +5

African American English (AAE), a rule-governed dialect spoken by over 30 million people, is routinely misinterpreted and "corrected" by large language models (LLMs). Across six in…

cs.CL2026

SAGE: A Search-AuGmented Evaluation of Large Language Models on Free-Form QA

Sher Badshah, Ali Emami, Hassan Sajjad

As Large Language Models (LLMs) become increasingly used for question-answering (QA), relying on static, pre-annotated references for evaluation poses significant challenges in cos…

cs.CL2026

SCOPE: Selective Conformal Optimized Pairwise LLM Judging

Sher Badshah, Ali Emami, Hassan Sajjad

Large language models (LLMs) are increasingly used as scalable judges in pairwise evaluation, but they remain prone to miscalibration and biases. We propose \textsc{Scope} (Selecti…

cs.CL2026

Common to Whom? Regional Cultural Commonsense and LLM Bias in India

Sangmitra Madhusudan, Trush Shashank More, Steph Buongiorno +3

Existing cultural commonsense benchmarks treat nations as monolithic, assuming uniform practices within national boundaries. But does cultural commonsense hold uniformly within a n…

cs.CL2026

The Dog the Cat Chased Stumped the Model: Measuring When Language Models Abandon Structure for Shortcuts

Sangmitra Madhusudan, Kaige Chen, Ali Emami

When language models correctly parse "The cat that the dog chased meowed," are they analyzing syntax or simply familiar with dogs chasing cats? Despite extensive benchmarking, we l…

cs.CL2025

Which Words Matter Most in Zero-Shot Prompts?

Nikta Gohari Sadr, Sangmitra Madhusudan, Hassan Sajjad +1

While zero-shot instructional prompts like "Let's think step-by-step" have revolutionized Large Language Model performance, a fundamental question remains unanswered: which specifi…