8 papers
LitTraceQA: A Benchmark for Multi-Stage Grounding and Verification in Scientific Question Answering
Xuye Liu, Yimu Wang, Peng Shi +7
Scientific literature is increasingly used as a knowledge source for language models, retrieval-augmented generation systems, and research assistants, but answering research questi…
The Mechanistic Emergence of Symbol Grounding in Language Models
Shuyu Wu, Ziqiao Ma, Xiaoxi Luo +4
Symbol grounding (Harnad, 1990) describes how symbols such as words acquire their meanings by connecting to real-world sensorimotor experiences. Recent work has shown preliminary e…
Pretraining Language Models on Historical Text
Xiaoxi Luo, Zachary Shinnick, Niclas Griesshaber +5
We introduce TypewriterLM, a 7.24B History language model (LM) trained exclusively on English text predating 1913. Developing History LMs requires addressing challenges in data qua…
How Tokenization Limits Phonological Knowledge Representation in Language Models and How to Improve Them
Disen Liao, Freda Shi
Tokenization is the first step in every language model (LM), yet it never takes the sounds of words into account. We investigate how tokenization influences text-only LMs' ability…
Translation or Recitation? Calibrating Evaluation Scores for Machine Translation of Extremely Low-Resource Languages
Danlu Chen, Ka Sing He, Jiahe Tian +4
The landscape of extremely low-resource machine translation (MT) is characterized by perplexing variability in reported performance, often making results across different language…
LogogramNLP: Comparing Visual and Textual Representations of Ancient Logographic Writing Systems for NLP
Danlu Chen, Freda Shi, Aditi Agarwal +2
Standard natural language processing (NLP) pipelines operate on symbolic representations of language, which typically consist of sequences of discrete tokens. However, creating an…