activity
20242026
collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models

Anirudh Bharadwaj, Chaitanya Malaviya, Nitish Joshi +1

Language models serve as proxies for human preference judgements in alignment and evaluation, yet they exhibit systematic miscalibration, prioritizing superficial patterns over sub…

cs.CL20251 cited

ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics

Li S. Yifei, Allen Chang, Chaitanya Malaviya +1

Evaluating long-form responses to research queries heavily relies on expert annotators, restricting attention to areas like AI where researchers can conveniently enlist colleagues.…

cs.CL2025

Contextualized Evaluations: Judging Language Model Responses to Underspecified Queries

Chaitanya Malaviya, Joseph Chee Chang, Dan Roth +3

Language model users often issue queries that lack specification, where the context under which a query was issued -- such as the user's identity, the query's intent, and the crite…

cs.CL2024

DOLOMITES: Domain-Specific Long-Form Methodical Tasks

Chaitanya Malaviya, Priyanka Agrawal, Kuzman Ganchev +7

Experts in various fields routinely perform methodical writing tasks to plan, organize, and report their work. From a clinician writing a differential diagnosis for a patient, to a…

cs.CL2024

Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering

Xingyu Fu, Ben Zhou, Sihao Chen +2

Recent advances in multimodal large language models (LLMs) have shown extreme effectiveness in visual question answering (VQA). However, the design nature of these end-to-end model…