Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
From Conversation to Query Execution: Benchmarking User and Tool Interactions for EHR Database Agents
Gyubok Lee, Woosog Chay, Heeyoung Kwak +5
Despite the impressive performance of LLM-powered agents, their adoption for Electronic Health Record (EHR) data access remains limited by the absence of benchmarks that adequately…
cs.AI2024
TrustSQL: Benchmarking Text-to-SQL Reliability with Penalty-Based Scoring
Gyubok Lee, Woosog Chay, Seonhee Cho +1
Text-to-SQL enables users to interact with databases using natural language, simplifying the retrieval and synthesis of information. Despite the remarkable success of large languag…