2 papers
cs.CL2026
CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems
Jingbo Yang, Guanyu Yao, Bairu Hou +5
As Large Language Models (LLMs) are increasingly deployed as task-oriented agents in enterprise environments, ensuring their strict adherence to complex, domain-specific operationa…
cs.CL2025
FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts
Hagyeong Shin, Binoy Robin Dalal, Iwona Bialynicka-Birula +4
Large language models (LLMs) are known to hallucinate, producing natural language outputs that are not grounded in the input, reference materials, or real-world knowledge. In enter…