5 papers
Multi-lingual Multi-institutional Electronic Health Record based Predictive Model
Kyunghoon Hur, Heeyoung Kwak, Jinsu Jang +2
Large-scale EHR prediction across institutions is hindered by substantial heterogeneity in schemas and code systems. Although Common Data Models (CDMs) can standardize records for…
From Conversation to Query Execution: Benchmarking User and Tool Interactions for EHR Database Agents
Gyubok Lee, Woosog Chay, Heeyoung Kwak +5
Despite the impressive performance of LLM-powered agents, their adoption for Electronic Health Record (EHR) data access remains limited by the absence of benchmarks that adequately…
SCARE: A Benchmark for SQL Correction and Question Answerability Classification for Reliable EHR Question Answering
Gyubok Lee, Woosog Chay, Edward Choi
Recent advances in Large Language Models (LLMs) have enabled the development of text-to-SQL models that allow clinicians to query structured data stored in Electronic Health Record…
FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering
Gyubok Lee, Elea Bach, Eric Yang +5
The recent shift toward the Health Level Seven Fast Healthcare Interoperability Resources (HL7 FHIR) standard opens a new frontier for clinical AI, demanding LLM agents to navigate…
KorMedMCQA: Multi-Choice Question Answering Benchmark for Korean Healthcare Professional Licensing Examinations
Sunjun Kweon, Byungjin Choi, Gyouk Chu +7
We present KorMedMCQA, the first Korean Medical Multiple-Choice Question Answering benchmark, derived from professional healthcare licensing examinations conducted in Korea between…