5 citations · 13 across the 6 of their papers we have counts for
8 papers
PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning
Afra Feyza Akyürek, Advait Gosai, Chen Bo Calvin Zhang +21
Frontier model progress is often measured by academic benchmarks, which offer a limited view of performance in real-world professional contexts. Existing evaluations often fail to…
Online Rubrics Elicitation from Pairwise Comparisons
MohammadHossein Rezaei, Robert Vacareanu, Zihao Wang +4
Rubrics provide a flexible way to train LLMs on open-ended long-form answers where verifiable rewards are not applicable and human preferences provide coarse signals. Prior work sh…
Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability
Afra Feyza Akyürek, Ekin Akyürek, Leshem Choshen +2
While language models (LMs) can sometimes generate factually correct text and estimate truth values of individual claims, these generally do not reflect a globally coherent, manipu…
DUnE: Dataset for Unified Editing
Afra Feyza Akyürek, Eric Pan, Garry Kuwanto +1
Even the most advanced language models remain susceptible to errors necessitating to modify these models without initiating a comprehensive retraining process. Model editing refers…
RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs
Afra Feyza Akyürek, Ekin Akyürek, Aman Madaan +4
Despite their unprecedented success, even the largest language models make mistakes. Similar to how humans learn and improve using feedback, previous work proposed providing langua…
On Measuring Social Biases in Prompt-Based Multi-Task Learning
Afra Feyza Akyürek, Sejin Paik, Muhammed Yusuf Kocyigit +3
Large language models trained on a mixture of NLP tasks that are converted into a text-to-text format using prompts, can generalize into novel forms of language and handle novel ta…