4 papers
Scaling the Queue: Reinforcement Learning for Equitable Call Classification Capacity in NYC Municipal Complaint Systems
Irene Aldridge, Ellie Bae, Siddhesh Darak +25
Municipal 311 call centers and complaint intake systems face a structural mismatch between incoming volume and classification capacity. The staff and heuristics available to triage…
Redirected, Not Removed: Task-Dependent Stereotyping Reveals the Limits of LLM Alignments
Divyanshu Kumar, Ishita Gupta, Nitin Aravind Birur +3
How biased is a language model? The answer depends on how you ask. A model that refuses to choose between castes for a leadership role will, in a fill-in-the-blank task, reliably a…
SocioEval: A Template-Based Framework for Evaluating Socioeconomic Status Bias in Foundation Models
Divyanshu Kumar, Ishita Gupta, Nitin Aravind Birur +3
As Large Language Models (LLMs) increasingly power decision-making systems across critical domains, understanding and mitigating their biases becomes essential for responsible AI d…
Beyond Western Politics: Cross-Cultural Benchmarks for Evaluating Partisan Associations in LLMs
Divyanshu Kumar, Ishita Gupta, Nitin Aravind Birur +3
Partisan bias in LLMs has been evaluated to assess political leanings, typically through a broad lens and largely in Western contexts. We move beyond identifying general leanings t…