3 papers
cs.CL2026
Measuring Affinity between Attention-Head Weight Subspaces via the Projection Kernel
Hiroaki Yamagiwa, Yusuke Takase, Hidetoshi Shimodaira
Understanding relationships between attention heads is essential for interpreting the internal structure of Transformers, yet existing metrics do not capture this structure well. W…
cs.CL2025
DeLTa: A Decoding Strategy based on Logit Trajectory Prediction Improves Factuality and Reasoning Ability
Yunzhen He, Yusuke Takase, Yoichi Ishibashi +1
Large Language Models (LLMs) are increasingly being used in real-world applications. However, concerns about the reliability of the content they generate persist, as it frequently…
cs.CL2025
Mapping 1,000+ Language Models via the Log-Likelihood Vector
Momose Oyama, Hiroaki Yamagiwa, Yusuke Takase +1
To compare autoregressive language models at scale, we propose using log-likelihood vectors computed on a predefined text set as model features. This approach has a solid theoretic…