◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Shun Kiyono

RIKEN AIP, Tohoku University

4 papers hereh-index 13977 citations32 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author4

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.CL4
affiliations
  • RIKEN AIP, Tohoku University
Homepage

identity via Semantic Scholar / OpenAlex

collaborators

4 papers

cs.CL2026

Pre-training LLM without Learning Rate Decay Enhances Supervised Fine-Tuning

Kazuki Yano, Shun Kiyono, Sosuke Kobayashi +2

We investigate the role of learning rate scheduling in the large-scale pre-training of large language models, focusing on its influence on downstream performance after supervised f…

cs.CL2026

Efficient Construction of Model Family through Progressive Training Using Model Expansion

Kazuki Yano, Sho Takase, Sosuke Kobayashi +2

As Large Language Models (LLMs) gain widespread practical application, offering model families with varying parameter sizes has become standard practice to accommodate diverse comp…

cs.CL2025

Spike No More: Stabilizing the Pre-training of Large Language Models

Sho Takase, Shun Kiyono, Sosuke Kobayashi +1

Loss spikes often occur during pre-training of large language models. The spikes degrade the performance of large language models and sometimes ruin the pre-training. Since the pre…

cs.CL2025

Large Vocabulary Size Improves Large Language Models

Sho Takase, Ryokan Ri, Shun Kiyono +1

This paper empirically investigates the relationship between subword vocabulary size and the performance of large language models (LLMs) to provide insights on how to define the vo…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.