3 papers
cs.CL2025
Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Models
Sina J. Semnani, Jirayu Burapacheep, Arpandeep Khatua +3
Wikipedia is the largest open knowledge corpus, widely used worldwide and serving as a key resource for training large language models (LLMs) and retrieval-augmented generation (RA…
cs.CL2025
CHURRO: Making History Readable with an Open-Weight Large Vision-Language Model for High-Accuracy, Low-Cost Historical Text Recognition
Sina J. Semnani, Han Zhang, Xinyan He +2
Accurate text recognition for historical documents can greatly advance the study and preservation of cultural heritage. Existing vision-language models (VLMs), however, are designe…
cs.CL2025
LEMONADE: A Large Multilingual Expert-Annotated Abstractive Event Dataset for the Real World
Sina J. Semnani, Pingyue Zhang, Wanyue Zhai +6
This paper presents LEMONADE, a large-scale conflict event dataset comprising 39,786 events across 20 languages and 171 countries, with extensive coverage of region-specific entiti…