From the 1 of 9 linked papers with an AI index.
8 papers
TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology
Hyunjin Seo, Hyeon Hwang, Gyubok Lee +7
The paper introduces TheBioCollection, a 52.6‑billion‑token unified corpus that aggregates diverse biological resources for pre‑training large language models, and shows that train…
Leveraging Biokinetic Knowledge Priors for Data-Scarce Bioprocess Modeling
Kyunghoon Hur, Eunjung Jeon, Hyun Woo Kim +2
While deep learning has accelerated drug discovery, its impact on biomanufacturing has been considerably more limited. The reason is data scarcity. Bioreactor experiments are high-…
VibeProteinBench: An Evaluation Benchmark for Language-interfaced Vibe Protein Design
Hyunjin Seo, Hongjoon Ahn, Jimin Park +16
Protein design aims to compose amino-acid sequences that fold into stable three-dimensional structures while satisfying targeted functional properties. The field is increasingly sh…
EHRSQL: A Practical Text-to-SQL Benchmark for Electronic Health Records
Gyubok Lee, Hyeonji Hwang, Seongsu Bae +6
We present a new text-to-SQL dataset for electronic health records (EHRs). The utterances were collected from 222 hospital staff members, including physicians, nurses, and insuranc…
From Conversation to Query Execution: Benchmarking User and Tool Interactions for EHR Database Agents
Gyubok Lee, Woosog Chay, Heeyoung Kwak +5
Despite the impressive performance of LLM-powered agents, their adoption for Electronic Health Record (EHR) data access remains limited by the absence of benchmarks that adequately…
SCARE: A Benchmark for SQL Correction and Question Answerability Classification for Reliable EHR Question Answering
Gyubok Lee, Woosog Chay, Edward Choi
Recent advances in Large Language Models (LLMs) have enabled the development of text-to-SQL models that allow clinicians to query structured data stored in Electronic Health Record…