2 papers
cs.CL2026
Linear Representations of Hierarchical Concepts in Language Models
Masaki Sakata, Benjamin Heinzerling, Takumi Ito +2
We investigate how and to what extent hierarchical relations (e.g., Japan Eastern Asia Asia) are encoded in the internal representations of language models. Bui…
cs.CL2025
STEP: Staged Parameter-Efficient Pre-training for Large Language Models
Kazuki Yano, Takumi Ito, Jun Suzuki
Pre-training large language models (LLMs) faces significant memory challenges due to the large size of model parameters. We introduce STaged parameter-Efficient Pre-training (STEP)…