◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Fu-Ming Guo

4 papers hereh-index 5397 citations12 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author3

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.CL2
  • cs.LG2

identity via Semantic Scholar / OpenAlex

collaborators

4 papers

cs.CL2026

Smooth Scaling Laws Hide Stepwise Token Learning

Pingjie Wang, Zechen Hu, Peiru Yang +2

Language model loss follows remarkably regular scaling laws over model and data size, yet it remains unclear why the aggregate loss should exhibit a power-law form. Existing explan…

cs.LG2026

JTok: On Token Embedding as another Axis of Scaling Law via Joint Token Self-modulation

Yebin Yang, Huaijin Wu, Fu Guo +5

LLMs have traditionally scaled along dense dimensions, where performance is coupled with near-linear increases in computational cost. While MoE decouples capacity from compute, it…

cs.LG2025

AdamHD: Decoupled Huber Decay Regularization for Language Model Pre-Training

Fu-Ming Guo, Yingfang Fan

Adaptive optimizers with decoupled weight decay, such as AdamW, are the de facto standard for pre-training large transformer-based generative models. Yet the quadratic nature of th…

cs.CL2025

dots.llm1 Technical Report

Bi Huo, Bin Tu, Cheng Qin +24

Mixture of Experts (MoE) models have emerged as a promising paradigm for scaling language models efficiently by activating only a subset of parameters for each input token. In this…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.