5 papers
On the Structural Limits of Machine Learning Decision Systems: An Information-Theoretic, Interaction-Based, and Stochastic-Dynamical Perspective
Nestor R. Barraza, Gabriel Pena
Machine learning procedures are commonly evaluated in terms of predictive accuracy and computational efficiency. However, their achievable performance is fundamentally constrained…
Revisiting Non-Verbatim Memorization in Large Language Models: The Role of Entity Surface Forms
Yuto Nishida, Naoki Shikoda, Yosuke Kishinami +4
Understanding what kinds of factual knowledge large language models (LLMs) memorize is essential for evaluating their reliability and limitations. Entity-based QA is a common frame…
Instability in Downstream Task Performance During LLM Pretraining
Yuto Nishida, Masaru Isonuma, Yusuke Oda
When training large language models (LLMs), it is common practice to track downstream task performance throughout the training process and select the checkpoint with the highest va…
Long-Tail Crisis in Nearest Neighbor Language Models
Yuto Nishida, Makoto Morishita, Hiroyuki Deguchi +2
The -nearest-neighbor language model (NN-LM), one of the retrieval-augmented language models, improves the perplexity for given text by directly accessing a large datastore b…
How to Make the Most of LLMs' Grammatical Knowledge for Acceptability Judgments
Yusuke Ide, Yuto Nishida, Justin Vasselli +4
The grammatical knowledge of language models (LMs) is often measured using a benchmark of linguistic minimal pairs, where the LMs are presented with a pair of acceptable and unacce…