◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Fan Lai

4 papers hereh-index 322 citations5 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author1
  • last author3

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.DC2
  • cs.LG1
  • cs.SD1
same name
  • Fan Lai — 7 papers, h 3
  • Fan Lai — 5 papers, h 4
  • Fan Lai — 4 papers, h 3
  • Fan Lai — 4 papers, h 13
  • Fan Lai — 2 papers, h 2
  • Fan Lai — 1 paper, h 2

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

works on
diffusion models 1efficient inference 1GPU scheduling 1model serving 1regional execution 1

From the 1 of 4 linked papers with an AI index.

collaborators

4 papers

cs.DC2026

FlashDiff: Efficient Regional Execution and Scheduling for Diffusion Model Serving

Yaqi Qiao, Ping He, Songrun Xie +4

FlashDiff is a system that speeds up diffusion model inference by dynamically selecting which latent regions need further processing and efficiently scheduling those regions across…

cs.SD2026

SoundWeaver: Semantic Warm-Starting for Text-to-Audio Diffusion Serving

Ayush Barik, Sofia Stoica, Nikhil Sarda +4

Text-to-audio diffusion models produce high-fidelity audio but require tens of function evaluations (NFEs), incurring multi-second latency and limited throughput. We present SoundW…

cs.DC2025

HyGen: Efficient LLM Serving via Elastic Online-Offline Request Co-location

Ting Sun, Penghan Wang, Fan Lai

Large language models (LLMs) have facilitated a wide range of applications with distinct service-level objectives (SLOs), from latency-sensitive online tasks like interactive chatb…

cs.LG2025

IC-Cache: Efficient Large Language Model Serving via In-context Caching

Yifan Yu, Yu Gan, Nikhil Sarda +7

Large language models (LLMs) have excelled in various applications, yet serving them at scale is challenging due to their substantial resource demands and high latency. Our real-wo…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.