◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Xingqi Cui

5 papers hereh-index 218 citations4 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author3

Across the 5 of 5 papers where every author was matched, so the position is known.

fields
  • cs.DC3
  • cs.LG1
  • cs.OS1

identity via Semantic Scholar / OpenAlex

collaborators
Showing cs.DCShow all

3 papers · 1 filter

cs.DC2026

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving

Xingqi Cui, Chieh-Jan Mike Liang, Ziang Tang +2

Achieving cost efficiency while meeting strict user-facing SLOs (e.g., time-to-first-token) remains a fundamental challenge for cloud GPU clusters serving large language models (LL…

cs.DC2026

Characterization-Guided GPU Fault Resilience in NVIDIA MPS

Rixin Liu, Xingqi Cui, Kaijian Wang +4

NVIDIA Multi-Process Service (MPS) enables fine-grained GPU sharing by allowing multiple processes to execute concurrently on the same GPU, making it an important mechanism for imp…

cs.DC2025

From Models to Operators: Rethinking Autoscaling Granularity for Large Generative Models

Xingqi Cui, Chieh-Jan Mike Liang, Jiarong Xing +1

Serving large generative models such as LLMs and multi- modal transformers requires balancing user-facing SLOs (e.g., time-to-first-token, time-between-tokens) with provider goals…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.