◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Yifeng Liu

5 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author3

Across the 5 of 5 papers where every author was matched, so the position is known.

fields
  • cs.LG3
  • cs.CL1
  • cs.CR1
same name
  • Yifeng Liu — 8 papers, h 16
  • Yifeng Liu — 2 papers
  • Yifeng Liu — 1 paper, h 4
  • Yifeng Liu — 1 paper
  • Yifeng Liu — 1 paper
  • Yifeng Liu — 1 paper

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

collaborators

5 papers

cs.LG2025

Robust Layerwise Scaling Rules by Proper Weight Decay Tuning

Zhiyuan Fan, Yifeng Liu, Qingyue Zhao +2

Empirical scaling laws prescribe how to allocate parameters, data, and compute, while maximal-update parameterization (μP) enables learning-rate transfer across widths by equaliz…

cs.LG2025

MARS-M: When Variance Reduction Meets Matrices

Yifeng Liu, Angela Yuan, Quanquan Gu

Matrix-based preconditioned optimizers, such as Muon, have recently been shown to be more efficient than scalar-based optimizers for training large-scale neural networks, including…

cs.CR2025

A Novel Approach to Differential Privacy with Alpha Divergence

Yifeng Liu, Zehua Wang

As data-driven technologies advance swiftly, maintaining strong privacy measures becomes progressively difficult. Conventional (ε,δ)-differential privacy, while prevalent, exhib…

cs.CL2025

Tensor Product Attention Is All You Need

Yifan Zhang, Yifeng Liu, Huizhuo Yuan +4

Scaling language models to handle longer input sequences typically necessitates large key-value (KV) caches, resulting in substantial memory overhead during inference. In this pape…

cs.LG2024

MARS: Unleashing the Power of Variance Reduction for Training Large Models

Huizhuo Yuan, Yifeng Liu, Shuang Wu +2

Training deep neural networks--and more recently, large models demands efficient and scalable optimizers. Adaptive gradient algorithms like Adam, AdamW, and their variants have bee…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.