3 papers
cs.CL2026
Communication between Frozen Large Language Models via Prompt Optimization in a Referential Game
Vivek Anand, Muthu Chandrasekaran, Shiva Chaitanya
We study communication between two frozen large language models from different providers, with different tokenizers, accessed through their API endpoints. The two play a referentia…
cs.CL2026
INSURE-Dial: A Phase-Aware Conversational Dataset & Benchmark for Compliance Verification and Phase Detection
Shubham Kulkarni, Alexander Lyzhov, Preetam Joshi +1
Administrative phone tasks drain roughly 1 trillion USD annually from U.S. healthcare, with over 500 million insurance-benefit verification calls manually handled in 2024. We intro…
cs.AI2026
All Required, In Order: Phase-Level Evaluation for AI-Human Dialogue in Healthcare and Beyond
Shubham Kulkarni, Alexander Lyzhov, Shiva Chaitanya +1
Conversational AI is starting to support real clinical work, but most evaluation methods miss how compliance depends on the full course of a conversation. We introduce Obligatory-I…