3 papers
cs.AI2025
Adobe Summit Concierge Evaluation with Human in the Loop
Yiru Chen, Sally Fang, Sai Sree Harsha +6
Generative AI assistants offer significant potential to enhance productivity, streamline information access, and improve user experience in enterprise contexts. In this work, we pr…
cs.IR2024
RETAIN: Interactive Tool for Regression Testing Guided LLM Migration
Tanay Dixit, Daniel Lee, Sally Fang +4
Large Language Models (LLMs) are increasingly integrated into diverse applications. The rapid evolution of LLMs presents opportunities for developers to enhance applications contin…
cs.HC2024
Evaluation and Continual Improvement for an Enterprise AI Assistant
Akash V. Maharaj, Kun Qian, Uttaran Bhattacharya +8
The development of conversational AI assistants is an iterative process with multiple components. As such, the evaluation and continual improvement of these assistants is a complex…