2 papers
cs.LG2026
Distribution-Calibrated Inference Time Compute for Thinking LLM-as-a-Judge
Hamid Dadkhahi, Firas Trabelsi, Parker Riley +2
Thinking Large Language Models (LLMs) used as judges for pairwise preferences remain noisy at the single-sample level, and common aggregation rules (majority vote, soft self-consis…
cs.CL2024
Learning from others' mistakes: Finetuning machine translation models with span-level error annotations
Lily H. Zhang, Hamid Dadkhahi, Mara Finkelstein +3
Despite growing interest in incorporating feedback to improve language models, most efforts focus only on sequence-level annotations. In this work, we explore the potential of util…