40 citations · 58 across the 5 of their papers we have counts for
5 papers
Online Rubrics Elicitation from Pairwise Comparisons
MohammadHossein Rezaei, Robert Vacareanu, Zihao Wang +4
Rubrics provide a flexible way to train LLMs on open-ended long-form answers where verifiable rewards are not applicable and human preferences provide coarse signals. Prior work sh…
TutorBench: A Benchmark To Assess Tutoring Capabilities Of Large Language Models
Rakshith S Srinivasa, Zora Che, Chen Bo Calvin Zhang +11
As students increasingly adopt large language models (LLMs) as learning aids, it is crucial to build models that are adept at handling the nuances of tutoring: they need to identif…
Auto-GPT for Online Decision Making: Benchmarks and Additional Opinions
Hui Yang, Sifu Yue, Yunzhong He
Auto-GPT is an autonomous agent that leverages recent advancements in adapting Large Language Models (LLMs) for decision-making tasks. While there has been a growing interest in Au…
HierCat: Hierarchical Query Categorization from Weakly Supervised Data at Facebook Marketplace
Yunzhong He, Cong Zhang, Ruoyan Kong +5
Query categorization at customer-to-customer e-commerce platforms like Facebook Marketplace is challenging due to the vagueness of search intent, noise in real-world data, and imba…
Que2Engage: Embedding-based Retrieval for Relevant and Engaging Products at Facebook Marketplace
Yunzhong He, Yuxin Tian, Mengjiao Wang +7
Embedding-based Retrieval (EBR) in e-commerce search is a powerful search retrieval technique to address semantic matches between search queries and products. However, commercial s…