2 papers
cs.CL2026
Climate-ModernBERT: Revisiting Corpus Composition for Domain-Adaptive Continued Pretraining
Yongan Yu, Shantam Raj, Jingwei Ni +3
Natural Language Processing (NLP) in the climate domain requires models to process heterogeneous text sources, including scientific literature, policy disclosures, and synthetic re…
cs.CL2026
pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs
Tobias Schimanski, Imene Kolli, Yu Fan +4
PDFs are the second-most used document type on the internet (after HTML). Yet, existing QA datasets commonly start from text sources or only address specific domains. In this paper…