2 papers
cs.CL2025
SummExecEdit: A Factual Consistency Benchmark in Summarization with Executable Edits
Onkar Thorat, Philippe Laban, Chien-Sheng Wu
Detecting factual inconsistencies in summarization is critical, yet existing benchmarks lack the necessary challenge and interpretability for robust evaluation. In this paper, we i…
cs.CL2025
CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
Kung-Hsiang Huang, Akshara Prabhakar, Onkar Thorat +6
While AI agents hold transformative potential in business, effective performance benchmarking is hindered by the scarcity of public, realistic business data on widely used platform…