4 papers
Prioritization of Risks from Artificial Intelligence: A Delphi Study of 272 International Experts
Alexander K. Saeri, Jess Graham, Michael Noetel +185
Artificial intelligence poses many risks, ranging from familiar present-day harms to unprecedented and potentially catastrophic ones. Effective risk management requires prioritizat…
INSURE-Dial: A Phase-Aware Conversational Dataset & Benchmark for Compliance Verification and Phase Detection
Shubham Kulkarni, Alexander Lyzhov, Preetam Joshi +1
Administrative phone tasks drain roughly 1 trillion USD annually from U.S. healthcare, with over 500 million insurance-benefit verification calls manually handled in 2024. We intro…
All Required, In Order: Phase-Level Evaluation for AI-Human Dialogue in Healthcare and Beyond
Shubham Kulkarni, Alexander Lyzhov, Shiva Chaitanya +1
Conversational AI is starting to support real clinical work, but most evaluation methods miss how compliance depends on the full course of a conversation. We introduce Obligatory-I…
HalluciNot: Hallucination Detection Through Context and Common Knowledge Verification
Bibek Paudel, Alexander Lyzhov, Preetam Joshi +1
This paper introduces a comprehensive system for detecting hallucinations in large language model (LLM) outputs in enterprise settings. We present a novel taxonomy of LLM responses…