2 papers
cs.CV2026
ClinOCR-Bench: A Comprehensive Clinical Scanned Document Dataset for Optical Character Recognition Model Evaluation
Enshuo Hsu, Jin Zhou, Kirk Roberts
Extracting textual information from scanned medical documents, such as external laboratory reports and manually filled forms, has been a major challenge in modern electronic health…
cs.LG2024
LLM-IE: A Python Package for Generative Information Extraction with Large Language Models
Enshuo Hsu, Kirk Roberts
Objectives: Despite the recent adoption of large language models (LLMs) for biomedical information extraction, challenges in prompt engineering and algorithms persist, with no dedi…