5 papers
Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
Minglai Yang, Xinyan Velocity Yu, Pengyuan Li +22
Document parsing and recognition are fundamental capabilities for vision-language models (VLMs) and document processing systems. However, existing Optical Character Recognition (OC…
Refining and Reusing Annotation Guidelines for LLM Annotation
Kon Woo Kim, Jin-Dong Kim, Akiko Aizawa
While Large Language Models (LLMs) demonstrate remarkable performance on zero-shot annotation tasks, they often struggle with the specialized conventions of gold-standard benchmark…
Synthetic Mixed Training: Scaling Parametric Knowledge Acquisition Beyond RAG
Seungju Han, Konwoo Kim, Chanwoo Park +5
Synthetic data augmentation helps language models learn new knowledge in data-constrained domains. However, naively scaling existing synthetic data methods by training on more synt…
Data-efficient pre-training by scaling synthetic megadocs
Konwoo Kim, Suhas Kotha, Yejin Choi +3
Synthetic data augmentation has emerged as a promising solution when pre-training is constrained by data rather than compute. We study how to design synthetic data algorithms that…
Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
Kon Woo Kim, Rezarta Islamaj, Jin-Dong Kim +2
This study investigates how existing annotation guidelines can be repurposed to instruct large language model (LLM) annotators for text annotation tasks. Traditional guidelines are…