5 papers
FireBench: Evaluating Instruction Following in Enterprise and API-Driven LLM Applications
Yunfan Zhang, Yijie Bei, Jetashree Ravi +1
Instruction following is critical for LLMs deployed in enterprise and API-driven settings, where strict adherence to output formats, content constraints, and procedural requirement…
LiveNewsBench: Evaluating LLM Web Search Capabilities with Freshly Curated News
Yunfan Zhang, Kathleen McKeown, Smaranda Muresan
Large Language Models (LLMs) with agentic web search capabilities show strong potential for tasks requiring real-time information access and complex fact retrieval, yet evaluating…
Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment
Yunfan Zhang, Kathleen McKeown, Smaranda Muresan
Large Language Models (LLMs) are typically trained to reflect a relatively uniform set of values, which limits their applicability to tasks that require understanding of nuanced hu…
Forecasting Conversation Derailments Through Generation
Yunfan Zhang, Kathleen McKeown, Smaranda Muresan
Forecasting conversation derailment can be useful in real-world settings such as online content moderation, conflict resolution, and business negotiations. However, despite languag…
SketchFill: Sketch-Guided Code Generation for Imputing Derived Missing Values
Yunfan Zhang, Changlun Li, Yuyu Luo +1
Missing value is a critical issue in data science, significantly impacting the reliability of analyses and predictions. Missing value imputation (MVI) is a longstanding problem bec…