7 papers
Understanding Benchmark Language Under Weakened Formal Semantics
Haoyang Chen, Kumiko Tanaka-Ishii
State-of-the-art NLP benchmarks require interpretation of natural language that specifies conditions, procedures, and exceptions, often relying on implicit assumptions and external…
Escaping Mode Collapse in LLM Generation via Geometric Regulation
Xin Du, Kumiko Tanaka-Ishii
Mode collapse is a persistent challenge in generative modeling and appears in autoregressive text generation as behaviors ranging from explicit looping to gradual loss of diversity…
Repeated Sequences Reveal Gaps between Large Language Models and Natural Language
Kumiko Tanaka-Ishii
Evaluating whether large language models (LLMs) capture the structure of natural language beyond local fluency remains an open challenge. Existing evaluation methods, largely based…
Artificial intelligence is creating a new global linguistic hierarchy
Giulia Occhini, Kumiko Tanaka-Ishii, Anna Barford +9
Artificial intelligence (AI) has the potential to transform healthcare, education, governance and socioeconomic equity, but its benefits remain concentrated in a small number of la…
Correlation Dimension of Auto-Regressive Large Language Models
Xin Du, Kumiko Tanaka-Ishii
Large language models (LLMs) have achieved remarkable progress in natural language generation, yet they continue to display puzzling behaviors -- such as repetition and incoherence…
Scale-free Characteristics of Multilingual Legal Texts and the Limitations of LLMs
Haoyang Chen, Kumiko Tanaka-Ishii
We present a comparative analysis of text complexity across domains using scale-free metrics. We quantify linguistic complexity via Heaps' exponent (vocabulary growth), Taylor…