conversational agent evaluation 1governed pipelines 1llm-as-a-judge 1retail chatbots 1selective re-evaluation 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.AI2026
Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking
Niranjan Kumar M, Balaji Nagarajan, Karthik Nair +2
The paper introduces GenAI Evaluation, a configurable, governed pipeline that uses LLM-as-a-judge scoring to automatically evaluate retail conversational agents across multiple dim…
math.OC2026
OPTIMUS: Optimization Productivity Tool for Intelligent Management of Utilizable Space
Souvik Bhattacharyya, Nisha Singh, Salman Haider +4
We study department-level retail space optimization, where limited bay capacity must be allocated among planograms (POGs) under business and operational constraints. The problem is…