◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Kaiyue Wen

3 papers hereh-index 290 citations3 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author2

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.LG2
  • cs.CL1
same name
  • Kaiyue Wen — 6 papers, h 5
  • Kaiyue Wen — 6 papers, h 5
  • Kaiyue Wen — 4 papers, h 8
  • Kaiyue Wen — 1 paper, h 4

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.LG2026

Fantastic Pretraining Optimizers and Where to Find Them II: Hyperball Optimization

Kaiyue Wen, Xingyu Dang, Kaifeng Lyu +2

Matrix based optimizers such as Muon can substantially speed up language model pretraining, but their gains over AdamW are observed to shrink as model size and data scale grow when…

cs.LG2025

Weight Ensembling Improves Reasoning in Language Models

Xingyu Dang, Christina Baek, Kaiyue Wen +2

We investigate a failure mode that arises during the training of reasoning models, where the diversity of generations begins to collapse, leading to suboptimal test-time scaling. N…

cs.CL2025

Overtrained Language Models Are Harder to Fine-Tune

Jacob Mitchell Springer, Sachin Goyal, Kaiyue Wen +5

Large language models are pre-trained on ever-growing token budgets under the assumption that better pre-training performance translates to improved downstream models. In this work…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.