◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Zhao Song

4 papers hereh-index 4129 citations5 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • last author4

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.LG3
  • cs.DS1
same name
  • Zhao Song — 45 papers, h 19
  • Zhao Song — 28 papers, h 37
  • Zhao Song — 27 papers
  • Zhao Song — 18 papers, h 17
  • Zhao Song — 12 papers, h 6
  • Zhao Song — 11 papers, h 8

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20232025
most citedHow to Capture Higher-order Correlations? Generalizing Matrix Softmax Attention to Kronecker Computation

1 citations · 2 across the 4 of their papers we have counts for

collaborators

4 papers

cs.LG2025

Only Large Weights (And Not Skip Connections) Can Prevent the Perils of Rank Collapse

Josh Alman, Zhao Song

Attention mechanisms lie at the heart of modern large language models (LLMs). Straightforward algorithms for forward and backward (gradient) computation take quadratic time, and a…

cs.LG2025

Fast RoPE Attention: Combining the Polynomial Method and Fast Fourier Transform

Josh Alman, Zhao Song

The transformer architecture has been widely applied to many machine learning tasks. A main bottleneck in the time to perform transformer computations is a task called attention co…

cs.LG2024★ 1 cited

The Fine-Grained Complexity of Gradient Computation for Training Large Language Models

Josh Alman, Zhao Song

Large language models (LLMs) have made fundamental contributions over the last a few years. To train an LLM, one needs to alternatingly run `forward' computations and `backward' co…

cs.DS2023★ 1 cited

How to Capture Higher-order Correlations? Generalizing Matrix Softmax Attention to Kronecker Computation

Josh Alman, Zhao Song

In the classical transformer attention scheme, we are given three n×d size matrices Q,K,V (the query, key, and value tokens), and the goal is to compute a new $n \time…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.