collaborators

8 papers

cs.CL2026

LitTraceQA: A Benchmark for Multi-Stage Grounding and Verification in Scientific Question Answering

Xuye Liu, Yimu Wang, Peng Shi +7

Scientific literature is increasingly used as a knowledge source for language models, retrieval-augmented generation systems, and research assistants, but answering research questi…

cs.CL2026

The Mechanistic Emergence of Symbol Grounding in Language Models

Shuyu Wu, Ziqiao Ma, Xiaoxi Luo +4

Symbol grounding (Harnad, 1990) describes how symbols such as words acquire their meanings by connecting to real-world sensorimotor experiences. Recent work has shown preliminary e…

cs.CL2026

Pretraining Language Models on Historical Text

Xiaoxi Luo, Zachary Shinnick, Niclas Griesshaber +5

We introduce TypewriterLM, a 7.24B History language model (LM) trained exclusively on English text predating 1913. Developing History LMs requires addressing challenges in data qua…

cs.CL2026

How Tokenization Limits Phonological Knowledge Representation in Language Models and How to Improve Them

Disen Liao, Freda Shi

Tokenization is the first step in every language model (LM), yet it never takes the sounds of words into account. We investigate how tokenization influences text-only LMs' ability…

cs.CL2026

Translation or Recitation? Calibrating Evaluation Scores for Machine Translation of Extremely Low-Resource Languages

Danlu Chen, Ka Sing He, Jiahe Tian +4

The landscape of extremely low-resource machine translation (MT) is characterized by perplexing variability in reported performance, often making results across different language…

cs.CL2026

LogogramNLP: Comparing Visual and Textual Representations of Ancient Logographic Writing Systems for NLP

Danlu Chen, Freda Shi, Aditi Agarwal +2

Standard natural language processing (NLP) pipelines operate on symbolic representations of language, which typically consist of sequences of discrete tokens. However, creating an…