collaborators

7 papers

cs.CL2026

Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models

Anirudh Bharadwaj, Chaitanya Malaviya, Nitish Joshi +1

Language models serve as proxies for human preference judgements in alignment and evaluation, yet they exhibit systematic miscalibration, prioritizing superficial patterns over sub…

cs.CL2026

Calibrating Large Language Models with Sample Consistency

Qing Lyu, Kumar Shridhar, Chaitanya Malaviya +6

Accurately gauging the confidence level of Large Language Models' (LLMs) predictions is pivotal for their reliable application. However, LLMs are often uncalibrated inherently and…

cs.CL2025

ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics

Li S. Yifei, Allen Chang, Chaitanya Malaviya +1

Evaluating long-form responses to research queries heavily relies on expert annotators, restricting attention to areas like AI where researchers can conveniently enlist colleagues.…

cs.CL2025

EvalAgent: Discovering Implicit Evaluation Criteria from the Web

Manya Wadhwa, Zayne Sprague, Chaitanya Malaviya +3

Evaluation of language model outputs on structured writing tasks is typically conducted with a number of desirable criteria presented to human evaluators or large language models (…

cs.IR2025

LogiCoL: Logically-Informed Contrastive Learning for Set-based Dense Retrieval

Yanzhen Shen, Sihao Chen, Xueqiang Xu +3

While significant progress has been made with dual- and bi-encoder dense retrievers, they often struggle on queries with logical connectives, a use case that is often overlooked ye…

cs.CL2025

Contextualized Evaluations: Judging Language Model Responses to Underspecified Queries

Chaitanya Malaviya, Joseph Chee Chang, Dan Roth +3

Language model users often issue queries that lack specification, where the context under which a query was issued -- such as the user's identity, the query's intent, and the crite…