◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Marco Chen

7 papers hereh-index 356 citations7 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author6

Across the 7 of 7 papers where every author was matched, so the position is known.

fields
  • cs.LG4
  • cs.CV3

identity via Semantic Scholar / OpenAlex

collaborators
Showing cs.LGShow all

4 papers · 1 filter

cs.LG2026

SageBwd: A Trainable Low-bit Attention

Jintao Zhang, Marco Chen, Haoxu Wang +5

Low-bit attention, such as SageAttention, has emerged as an effective approach for accelerating model inference, but its applicability to training remains poorly understood. In pri…

cs.LG2026

Delving into Muon and Beyond: Deep Analysis and Extensions

Xianbiao Qi, Marco Chen, Jiaquan Ye +2

The Muon optimizer has recently attracted considerable attention for its strong empirical performance and use of orthogonalized updates on matrix-shaped parameters, yet its underly…

cs.LG2026

SimpleGPT: Improving GPT via A Simple Normalization Strategy

Marco Chen, Xianbiao Qi, Yelin He +2

In this work, we revisit Transformer optimization through the lens of second-order geometry and establish a direct connection between architectural design, activation scale, the He…

cs.LG2025

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD

Xianbiao Qi, Marco Chen, Wenjie Xiao +4

Transformers have become the de facto backbone of modern deep learning, yet their training typically demands an advanced optimizer with adaptive learning rate like AdamW, rather th…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.