activity
20242026
collaborators

6 papers

cs.CL2026

Measuring Affinity between Attention-Head Weight Subspaces via the Projection Kernel

Hiroaki Yamagiwa, Yusuke Takase, Hidetoshi Shimodaira

Understanding relationships between attention heads is essential for interpreting the internal structure of Transformers, yet existing metrics do not capture this structure well. W…

cs.CL2025

Likelihood Variance as Text Importance for Resampling Texts to Map Language Models

Momose Oyama, Ryo Kishino, Hiroaki Yamagiwa +1

We address the computational cost of constructing a model map, which embeds diverse language models into a common space for comparison via KL divergence. The map relies on log-like…

cs.CL2025

Beyond Chains: Bridging Large Language Models and Knowledge Bases in Complex Question Answering

Yihua Zhu, Qianying Liu, Akiko Aizawa +1

Knowledge Base Question Answering (KBQA) aims to answer natural language questions using structured knowledge from KBs. While LLM-only approaches offer generalization, they suffer…

cs.CL2025

DeLTa: A Decoding Strategy based on Logit Trajectory Prediction Improves Factuality and Reasoning Ability

Yunzhen He, Yusuke Takase, Yoichi Ishibashi +1

Large Language Models (LLMs) are increasingly being used in real-world applications. However, concerns about the reliability of the content they generate persist, as it frequently…

cs.CL2025

Mapping 1,000+ Language Models via the Log-Likelihood Vector

Momose Oyama, Hiroaki Yamagiwa, Yusuke Takase +1

To compare autoregressive language models at scale, we propose using log-likelihood vectors computed on a predefined text set as model features. This approach has a solid theoretic…

cs.CL2024

Quantifying Lexical Semantic Shift via Unbalanced Optimal Transport

Ryo Kishino, Hiroaki Yamagiwa, Ryo Nagata +2

Lexical semantic change detection aims to identify shifts in word meanings over time. While existing methods using embeddings from a diachronic corpus pair estimate the degree of c…