3 citations · 3 across the 4 of their papers we have counts for
4 papers
Pretraining Language Models on Historical Text
Xiaoxi Luo, Zachary Shinnick, Niclas Griesshaber +5
We introduce TypewriterLM, a 7.24B History language model (LM) trained exclusively on English text predating 1913. Developing History LMs requires addressing challenges in data qua…
Chronos: The AI Co-Historian
Lorenz Hufe, Niclas Griesshaber, Gavin Greif +5
AI is increasingly supporting, accelerating, and automating scientific discovery across subjects. Yet, the adoption of AI in historical research remains limited due to the lack of…
Multimodal LLMs for Historical Dataset Construction from Archival Image Scans: German Patents (1877-1918)
Niclas Griesshaber, Jochen Streb
We leverage multimodal large language models (LLMs) to construct a dataset of 306,070 German patents (1877-1918) from 9,562 archival image scans using our LLM-based pipeline powere…
Multimodal LLMs for OCR, OCR Post-Correction, and Named Entity Recognition in Historical Documents
Gavin Greif, Niclas Griesshaber, Robin Greif
We explore how multimodal Large Language Models (mLLMs) can help researchers transcribe historical documents, extract relevant historical information, and construct datasets from h…