4 papers
From Conversation to Query Execution: Benchmarking User and Tool Interactions for EHR Database Agents
Gyubok Lee, Woosog Chay, Heeyoung Kwak +5
Despite the impressive performance of LLM-powered agents, their adoption for Electronic Health Record (EHR) data access remains limited by the absence of benchmarks that adequately…
ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant Simulation
Jiho Kim, Junseong Choi, Woosog Chay +4
As large language models (LLMs) become increasingly integrated into daily life, there is growing demand for AI assistants that are not only reactive but also proactive and personal…
SCARE: A Benchmark for SQL Correction and Question Answerability Classification for Reliable EHR Question Answering
Gyubok Lee, Woosog Chay, Edward Choi
Recent advances in Large Language Models (LLMs) have enabled the development of text-to-SQL models that allow clinicians to query structured data stored in Electronic Health Record…
DialSim: A Dialogue Simulator for Evaluating Long-Term Multi-Party Dialogue Understanding of Conversational Agents
Jiho Kim, Woosog Chay, Hyeonji Hwang +6
Recent advancements in Large Language Models (LLMs) have significantly enhanced conversational agents, making them applicable to various fields (e.g., education, entertainment). De…