3 papers
cs.CL2026
A Single Rewrite Suffices: Empirical Lessons from Production Skill Description Optimization
Yangqiaoyu Zhou, Mohammad Alqudah, Kwei-Herng Lai +3
Enterprise AI agents route user queries to specialized skills by matching queries against natural language skill descriptions. When two skills share overlapping descriptions, the r…
cs.CL2026
ConflictScore: Identifying and Measuring How Language Models Handle Conflicting Evidence
Siyi Liu, Aaron Halfaker, Dan Roth +1
Existing metrics for factuality and faithfulness evaluate whether an answer is supported or contradicted by its grounding documents, but they fail to capture when both supporting a…
cs.CL2025
Characterizing Deep Research: A Benchmark and Formal Definition
Abhinav Java, Ashmit Khandelwal, Sukruta Midigeshi +6
Information tasks such as writing surveys or analytical reports require complex search and reasoning, and have recently been grouped under the umbrella of \textit{deep research} --…