◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Ao Sun

10 papers hereh-index 493 citations11 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author6

Across the 8 of 10 papers where every author was matched, so the position is known.

fields
  • cs.DC4
  • cs.CL3
  • cs.LG2
  • cs.CV1
same name
  • Ao Sun — 7 papers, h 1
  • Ao Sun — 5 papers, h 8
  • Ao Sun — 4 papers, h 2
  • Ao Sun — 4 papers, h 2
  • Ao Sun — 3 papers, h 2
  • Ao Sun — 3 papers, h 2

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20242026
collaborators
Showing cs.DCShow all

4 papers · 1 filter

cs.DC2026

StateFlow: Sequence Pipeline Parallelism for Long-Context Modeling with Linear Recurrence

Wenxuan Zhao, Yingfa Chen, Xu Han +7

Long-context training is increasingly important for large language models, and linear attention and state space models have become popular for improving long-context efficiency. Ho…

cs.DC2026

InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training

Shiju Wang, Yujie Wang, Ao Sun +5

Long context training is crucial for LLM's context extension. Existing schemes, such as sequence parallelism, incur substantial communication overhead. Pipeline parallelism (PP) re…

cs.DC2025

BurstEngine: an Efficient Distributed Framework for Training Transformers on Extremely Long Sequences of over 1M Tokens

Ao Sun, Weilin Zhao, Xu Han +4

Existing methods for training LLMs on long-sequence data, such as Tensor Parallelism and Context Parallelism, exhibit low Model FLOPs Utilization as sequence lengths and number of…

cs.DC2024

Seq1F1B: Efficient Sequence-Level Pipeline Parallelism for Large Language Model Training

Ao Sun, Weilin Zhao, Xu Han +5

The emergence of large language models (LLMs) relies heavily on distributed training strategies, among which pipeline parallelism plays a crucial role. As LLMs' training sequence l…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.