2 papers
cs.CL2026
It's Not RoPE that Creates Sinks: The Role of Self-Concentration and Value-Non-Mixing in Attention
Raito Kiya, Satoki Ohashi, Kosuke Sato +6
Large Language Models (LLMs) often exhibit "Attention Sink" (AS) and the accompanying "Massive Activations" (MAs) at the initial position of a sequence. These phenomena frequently…
cs.CL2026
Language Models Compare Quantities Using Number-specific and Unit-specific Heuristics
Mutsumi Sasaki, Go kamoda, Ryosuke Takahashi +4
Quantities with measurement units, such as 110 cm and 1.2 m, require language models (LMs) to combine a numeral with a symbolic unit scale. Here, we study how LMs compare such quan…