4 papers
Mediocrity is the key for LLM as a Judge Anchor Selection
Shachar Don-Yehiya, Asaf Yehudai, Leshem Choshen +1
The ``LLM-as-a-judge'' paradigm has become a standard method for evaluating open-ended generation. To address the quadratic scalability costs of pairwise comparisons, popular bench…
The ShareLM Collection and Plugin: Contributing Human-Model Chats for the Benefit of the Community
Shachar Don-Yehiya, Leshem Choshen, Omri Abend
Human-model conversations provide a window into users' real-world scenarios, behavior, and needs, and thus are a valuable resource for model development and research. While for-pro…
Naturally Occurring Feedback is Common, Extractable and Useful
Shachar Don-Yehiya, Leshem Choshen, Omri Abend
Human feedback data is a critical component in developing language models. However, collecting this feedback is costly and ultimately not scalable. Inspired by the way human interl…
The Future of Open Human Feedback
Shachar Don-Yehiya, Ben Burtenshaw, Ramon Fernandez Astudillo +17
Human feedback on conversations with language language models (LLMs) is central to how these systems learn about the world, improve their capabilities, and are steered toward desir…