3 papers
cs.AI2026
KoVRE: Training an Efficient Embedding Model for Korean Visual Document Retrieval
Yongbin Choi, Gyuho Shim, Youngjoon Jang
Visual Document Retrieval (VDR) directly matches text queries against document images, preserving visual and structural information that may be lost during text extraction. However…
cs.AI2026
Revise: A Framework for Revising OCRed text in Practical Information Systems with Data Contamination Strategy
Gyuho Shim, Seongtae Hong, Heuiseok Lim
Recent advances in Large Language Models (LLMs) have significantly improved the field of Document AI, demonstrating remarkable performance on document understanding tasks such as q…
cs.CL2025
Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks
Dongjun Kim, Gyuho Shim, Yongchan Chun +3
Large Language Models are commonly judged by their scores on standard benchmarks, yet such scores often overstate real capability since they mask the mix of skills a task actually…