activity
20242026
collaborators

7 papers

cs.CL2026

Understanding Benchmark Language Under Weakened Formal Semantics

Haoyang Chen, Kumiko Tanaka-Ishii

State-of-the-art NLP benchmarks require interpretation of natural language that specifies conditions, procedures, and exceptions, often relying on implicit assumptions and external…

cs.CL2026

Escaping Mode Collapse in LLM Generation via Geometric Regulation

Xin Du, Kumiko Tanaka-Ishii

Mode collapse is a persistent challenge in generative modeling and appears in autoregressive text generation as behaviors ranging from explicit looping to gradual loss of diversity…

cs.CL2026

Repeated Sequences Reveal Gaps between Large Language Models and Natural Language

Kumiko Tanaka-Ishii

Evaluating whether large language models (LLMs) capture the structure of natural language beyond local fluency remains an open challenge. Existing evaluation methods, largely based…

cs.CY2026

Artificial intelligence is creating a new global linguistic hierarchy

Giulia Occhini, Kumiko Tanaka-Ishii, Anna Barford +9

Artificial intelligence (AI) has the potential to transform healthcare, education, governance and socioeconomic equity, but its benefits remain concentrated in a small number of la…

cs.CL2025

Correlation Dimension of Auto-Regressive Large Language Models

Xin Du, Kumiko Tanaka-Ishii

Large language models (LLMs) have achieved remarkable progress in natural language generation, yet they continue to display puzzling behaviors -- such as repetition and incoherence…

cs.CL2025

Scale-free Characteristics of Multilingual Legal Texts and the Limitations of LLMs

Haoyang Chen, Kumiko Tanaka-Ishii

We present a comparative analysis of text complexity across domains using scale-free metrics. We quantify linguistic complexity via Heaps' exponent (vocabulary growth), Taylor…