◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Pengle Zhang

4 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author3

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.LG4

identity via Semantic Scholar / OpenAlex

collaborators

4 papers

cs.LG2025

SageAttention2++: A More Efficient Implementation of SageAttention2

Jintao Zhang, Xiaoming Xu, Jia Wei +5

The efficiency of attention is critical because its time complexity grows quadratically with sequence length. SageAttention2 addresses this by utilizing quantization to accelerate…

cs.LG2025

SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training

Jintao Zhang, Jia Wei, Pengle Zhang +6

The efficiency of attention is important due to its quadratic time complexity. We enhance the efficiency of attention through two key contributions: First, we leverage the new FP4…

cs.LG2025

Accurate INT8 Training Through Dynamic Block-Level Fallback

Pengle Zhang, Jia Wei, Jintao Zhang +2

Transformer models have achieved remarkable success across various AI applications but face significant training costs. Low-bit training, such as INT8 training, can leverage comput…

cs.LG2024

SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization

Jintao Zhang, Haofeng Huang, Pengle Zhang +3

Although quantization for linear layers has been widely used, its application to accelerate the attention process remains limited. To further enhance the efficiency of attention co…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.