4 papers
Human-AI Co-reasoning for Clinical Diagnosis with Evidence-Integrated Language Agent
Zhongzhen Huang, Yan Ling, Hong Chen +7
We present PULSE, a medical reasoning agent that combines a domain-tuned large language model with scientific literature retrieval to support diagnostic decision-making in complex…
Adobe Summit Concierge Evaluation with Human in the Loop
Yiru Chen, Sally Fang, Sai Sree Harsha +6
Generative AI assistants offer significant potential to enhance productivity, streamline information access, and improve user experience in enterprise contexts. In this work, we pr…
Evaluation and Incident Prevention in an Enterprise AI Assistant
Akash V. Maharaj, David Arbour, Daniel Lee +6
Enterprise AI Assistants are increasingly deployed in domains where accuracy is paramount, making each erroneous output a potentially significant incident. This paper presents a co…
Evaluation and Continual Improvement for an Enterprise AI Assistant
Akash V. Maharaj, Kun Qian, Uttaran Bhattacharya +8
The development of conversational AI assistants is an iterative process with multiple components. As such, the evaluation and continual improvement of these assistants is a complex…